2025:企业编码之年

2025: The Year of Enterprise Coding

奥利维耶·戈德芒 Olivier Godement · Unsupervised Learning · 2025-12-10 · 约 58 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Olivier Godement 探讨 AI 模型的演进、企业应用以及对科学研究的影响。

Olivier Godement discusses the evolution of AI models, enterprise adoption, and the impact on scientific research.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 30)

全文 · Full transcript(中英对照)

引言与2025预测 Introduction and 2025 predictions

Host

2024 年是编程之年。我认为 2025 年是企业级编程之年。我开始看到采用的意义。是的。那 2026 年呢?

2024 is the year of coding. I think 2025 is the year of coding in the enterprise. Like I'm starting to see like the meaning for adoption. Yeah. What's 2026?

Olivier Godement

我不知道,兄弟。比如,我刚刚发布了一篇论文。我需要深入处理,大概需要 30 分钟才能完成。论文作者告诉我们,这相当于几周的工作量。新模型比如 5 或 5.1 出来了。这个过程实际上是什么样的?你想象它会如何随时间变化或演变?那种基本上只需热切换一个 API 参数就能从一个模型换到另一个模型的日子已经基本结束了。

I don't know man. Like uh here's a paper I just released. I need to deeply pro uh I think something like 30 minutes like to achieve it. The author of the paper was telling us like it was like weeks of work. A new model like five or 51 come out. What does that like process actually look like and how do you imagine that changing or evolving over time? The days where you could just like hot swap essentially like one API parameter from like you know the one model to the next are basically gone.

Host

那些真正优秀的顶尖团队在试验新模型时都做些什么?

What do the really good top teams do when like they're experimenting with a new model?

Olivier Godement

这对我来说可能是过去两年最大的收获,比如为开发者企业构建产品。嘿,我是 OpenAI 企业产品负责人 Olivier Godement。我是 Jacob Efron,今天在《无监督学习》中,我向 Olivier 提出了我最关心的问题。我们讨论了他所看到的 AI 在科学领域的进展,以及 OpenAI 模型如何助力。我们讨论了围绕模型出现的脚手架和框架,他在不同领域看到的模式,以及他认为这些模式会在多大程度上趋同,还有基础模型提供商是否会为初创公司提供这些。我们讨论了他对 Andrej Karpathy 关于智能体仍需十年才能实现的评论的看法。我们还讨论了模型的下一个前沿,以及 Olivier 认为哪些因素对解锁企业进一步价值最为重要。Olivier 还给出了很多非常棒的建议,关于最优秀的企业如何使用 OpenAI 来采用新模型并充分利用这些产品。这是一次与一位真正聪明的人的精彩对话。我想大家会喜欢的。闲话少说,有请 Olivier。

That's probably to me the biggest earnings like of the past 2 years of like you know building for like developers businesses is hey Olivia Goodman is the head of products for enterprise at OpenAI. I'm Jacob Efron and today on unsupervised learning I got to ask Olivier all my top questions. We talked about the progress he's seen in AI for science and how open AI models are helping there. We talked about the scaffolding and harnesses that are emerging around models, the patterns he's seen in different spaces, and the extent to that he thinks it will converge, as well as have foundation model providers provide this for startups. We talked about his reaction to Andre Carpathy and his comments about agents still being a decade away. And we also talked about the next frontiers for models and what Olivier thinks will be most important to unlock further value in enterprises. Olivier also gave a lot of really great tips on what the best enterprises using OpenAI do to adopt new models as well as fully get the most out of these products. Just an awesome conversation with a really brilliant mind. I think folks will really enjoy it. Without further ado, here's Olivia.

Host

非常感谢你来做客播客。真的很感激。

Well, thanks so much for coming on the podcast. Really appreciate it.

Olivier Godement

谢谢你邀请我。

Thank you for having me.

Host

而且能在 OpenAI 总部做这件事,格外有趣。

And it's extra fun to get to do it in, you know, opening IH HQ.

Olivier Godement

我知道。

I know.

Host

我可能是最没资格坐在演示日发布席上的人。所以这很有趣。

I'm probably the the least qualified person to ever sit in like the demo day launch seat. And so it's uh it's it's fun.

Olivier Godement

你是在模型船上。

You're in the in the model ship.

Host

你知道,今天有很多事情要讨论,但我想也许可以从 5.1 开始。显然,你们有了一个新模型。我认为在个性方面和整体推理上有一些非常有趣的改进。也许谈谈这些模型改进实际上是如何产生的,以及你到目前为止看到哪些让你兴奋的东西,人们是如何使用它的。

You know, lots of things to discuss today, but I figured maybe one place to start would be with, you know, 5.1. Obviously, you've got a new model. I think you know some really interesting improvements on the personality side on reasoning as a whole. Maybe just talk about how how do these things actually come about these sets of model improvements and what do you you've seen so far that you're excited about and how people are using it.

Olivier Godement

是的,完全同意。所以,我们上周发布了一批新模型,DP 5.1 和用于编程的 DP 5.1 Codex。这些模型基本上是基于 DP 5 的反馈训练的。当 DP 5 推出时,人们喜欢它的智能。人们喜欢它遵循指令的能力和流式传输能力。但人们不喜欢它的速度。这个模型在长时间思考时非常好,但对于更基础的查询,GP 5 太慢了,所以这是 DP 5.1 的主要设计目标之一:保持智能,但尽量压缩思考 token,以便更快地响应。我认为在这方面,目标已经实现了。我们看到人们可以相当无缝地在低思考努力(用于相当基础的查询)和更高思考努力(用于更复杂的查询)之间切换。Codex 模型也获得了相当多的采用。我认为 Codex 现在正迎来它的时刻。

Yeah, totally. So, we shipped a bunch of new models last week, DP 5.1 and DP 5.1 Codex for coding. Um those models basically were um trained based on the feedback of DP5. When DP5 came around, people love the intelligence. People love like the ability to follow instructions, the streamability. People didn't like the speed. Like the model was really good like you know when it was like thinking for a long time but like for more basic queries GP5 was too slow and so that was one of the main sort of design goals of DP5.1 like keep that intelligence but try to you know compress like the thinking tokens like you know as much as we can in order to respond like you know way faster. Um I think with that regard like the goal has been achieved. uh we're seeing like people like you know switch like pretty seamlessly between like you know low thinking effort for like you know fairly basic queries and like larger thinking efforts for you know more complex queries. The codex model has been like gaining quite a bit of adoption as well. Uh I think Codex is having like a moment at the moment.

Host

它确实看起来正迎来它的时刻。开发者们很喜欢它。我们开始看到越来越频繁的使用,我们也看到公司也在采用 Codex。

It definitely seems like it's having a moment. Like you know developers are loving it. We're starting to see like you know more and more like frequent use like you know we're seeing companies as well adopt Codex uh in.

Olivier Godement

这让你感到惊讶吗?还是你们内部已经充分试用过,知道它会这样?这可能是我们内部使用最多的模型,因为每个软件工程师、研究员在 OpenAI 都使用或曾使用过 Codex。我认为最新的统计数据是,团队中的工程师因为 Codex 能够多交付大约 70% 的工作。所以,是的,这已经变得很明显,它已经成为工具栈的一部分。

Did that surprise you or like had you played around with it enough internally to know this was coming? It's probably the model that we dog food the most internally just by the virtue of like you know every single like software engineer like researcher open AI like use or has used codeex. Um I think the latest stat I think something like you know engineers like on the team like are able to push something like 70% like more um because of codex. So yeah it's became it's become clear like you know and sort of in part of the tool stack at.

Host

是的。是的。嗯,不,我的意思是,看到延迟的改善和整体 Codex 模型的改进令人印象深刻。你看到解锁了哪些新的用例,或者人们开始以哪些有趣的方式使用这些模型?我认为大多数用例都和我们预期的一样,但更好的编程生产力,任何知识查询,比如检索信息、客户支持、客户体验,都一样,但更本质的是,坦率地说,过去几个月让我相当惊讶的一个领域是深度 5.1 在科学界的应用。我们看到越来越多的报告说科学家和研究人员使用 LLM 来完成他们的工作。这项工作意味着压缩、聚合科学文献和知识,更快地测试假设。这真的很令人兴奋,因为有趣的是,当我几年前加入时,公司的目标之一就是:嘿,如果能加速科学研究,那该多酷。我想说,在过去的几个月里,我第一次真正感觉到这正在发生。当然,这还很早期,我们还有很长的路要走,但第一次我们看到相当出色的科学家告诉我们:是的,我更快地完成了工作,我更快地实现了发现,因为有了它。

Yeah. Yeah. Um, no, I mean it's been uh it's been impressive to see and I guess any like you know that kind of improvement in latency and know the improvement in the overall codeex models. Any like new use cases you've seen unlocked or like fun ways that you've seen uh folks start to use these models? I think most of the use case were pretty much what we expected but better coding productivity any like knowledge like query like you know retrieving information customer support customer experience same but more essentially I think the one domain which has surprised me quite a bit over the past few months frankly is usage of deep and 5.1 among the scientific community we're seeing more and more reports of scientists and researchers use um LLMs in order to perform like their job. Um that job meaning like you know condensating aggregating like scientific like literature and knowledge just get faster like test hypothesis. It's just been like a really exciting because it's funny like when I join a couple years ago like that was always like one of the goal of the company is like hey how cool would it be if you could like accelerate scientific research. I would say for the first time in the past few months I'm actually feeling that it's happening. Of course it's early like you know and you know we have way to go but for the first time we're seeing like pretty stellar like you know scientists tell us like yes I have done my job faster you know I have done like you know um achieve the discovery that prove like you know faster because of.

Host

从你与这些人的交流中,你觉得速度提高了多少?你认为我们现在达到了什么样的加速?这是否是你在模型改进时考虑攀登的一个评估指标?

How much faster like do you think this is you know anecotally as you talk to these folks like you know what kind of speed up do you think we are at now and is that like an eval that you that you you know think about hill climbing on as these models get better.

科学加速的迹象 Glimmers of scientific acceleration

Host

我们的首席研究官 Mark Chen 曾与一位研究黑洞的物理学家合作。那位物理学家给模型出了一个非常难的任务:一篇他刚发表的论文,所以不在训练集里。他要求模型基本上复现其中的数学推导。Deeply 5 Pro 花了大约 30 分钟就完成了。我不是物理学家,无法评判,但论文作者告诉我们,一位专业物理学家需要几周的工作才能做到。所以我认为我们开始看到一些明显的加速迹象。

Mark Chen, our chief research officer, was working with a physicist who studies black holes. The physicist gave the model a really hard task: a paper he just released, so it wasn't in the training set. He asked it to essentially reproduce the math. It took Deeply 5 Pro about 30 minutes to achieve it. I'm not a physicist, so I can't judge, but the author of the paper told us it was weeks of work for a professional physicist to achieve that. So I think we're starting to see some clear glimmers of acceleration in that work.

Host

一方面,编程和科学支持方面,模型及其能力取得了令人兴奋的进展。另一方面,公众也在讨论这些模型到底处于什么水平。有一个重要时刻,Andrej Karpathy 在 Dwarkesh 的播客上说,行业跳跃太大,假装有些东西很了不起,其实有些是垃圾。他说至少需要十年,AI 才能有意义地自动化整个工作。你每天与企业打交道,看到这些模型能做什么不能做什么,你对 Andrej 的这个观点有什么反应?

On the one hand, coding and science support feel like there's been really exciting progress in models and their abilities. On the other hand, there's been a public conversation about where we are with these models. There was a big moment when Andrej Karpathy went on Dwarkesh's podcast and said the industry is making too big of a jump, pretending some stuff is amazing when some of it is slop. He said it will take at least a decade until AI can meaningfully automate entire jobs. I'm curious, working with enterprises every day, what was your reaction to that take from Andrej?

Olivier Godement

构建一个真正优秀的智能体很难,这已经不是秘密。我们还没有达到只需一天就能自动化任何白领工作的阶段。但我们在一些特定领域开始看到相当强大的自动化用例。编程就是其中之一。我认为我们已经到了这样一个阶段:如果我从软件工程师那里拿走 AI 编程工具,可能会引发骚乱。人们基本上会粉饰他们的工作。所以这些事情正在发生。我认为自动化可能还没有达到完全自动化软件工程师工作的水平,但我们有实现这一目标的路线图。

It's no secret that building a really good agent is hard. We haven't reached a stage where we can just take a day to automate any white-collar job. But we are starting to see some quite strong automation use cases in a few specific fields. Coding is the one that comes to mind. I think we've reached the point where if I were to take away AI coding tools from software engineers, there would probably be a riot. People would polish their jobs essentially. So that stuff is happening. I would say the automation is probably not yet at the level of completely automating the job of a software engineer, but I think we have a line of sight to get there.

Olivier Godement

我们谈到了客户体验,比如销售和客户支持。我们开始看到相当强的采用案例。我一直在与美国电信公司 T-Mobile 的人员合作,为他们的客户提供更好的体验。我们在有意义的规模上取得了相当好的质量成果。但这很难。除了模型本身,你还需要构建一个非常好的工具链:如何将模型连接到工具,如何提示模型,你必须构建一个非常好的评估框架,然后你还需要构建某种带有人类反馈的飞轮,以不断改进那个模型工具链。这工作量很大。但我的感觉是,在未来一两年内,我们可能会在可可靠自动化的任务数量上带来惊喜。

We talked about customer experience, like sales and customer support. We're starting to see fairly strong cases of adoption. I've been working with the folks at T-Mobile, the telecom company in the US, to provide a better experience for their customers. We're starting to achieve fairly good results in terms of quality at a meaningful scale. But that stuff is hard. On top of the model, you have to build a really good harness: how do you connect the model to tools, how do you prompt the model, you have to build a really good evaluation framework, and then you have to build some sort of a flywheel with human-in-the-loop to constantly improve that model harness. It's a lot of work. But my sense is we'll probably surprise in the next year or two on the amount of tasks that can be automated reliably.

Host

有没有哪些行业或用例你觉得正处于临界点,如果一两年内没有更多活动你会感到惊讶?

Are there any industries or use cases that you feel are on the cusp, where you'd be surprised if there wasn't a lot more activity in a year or two?

Olivier Godement

我通常押注生命科学。我一直在与 Amgen 和其他几家公司合作。这非常有趣。当你问 Amgen 他们作为一家公司存在的意义时,他们的目标是设计新药。大致有两大部分工作:实际的研发实验和科学家验证,以及更多的行政工作。行政工作量巨大。从锁定药物配方到上市,需要数月甚至数年。这涉及生成非常复杂的监管文件,让委员会审查,让监管机构审查。大量的信息共享、转换和验证,而模型在这方面很擅长。它们擅长聚合、整合大量结构化和非结构化数据,发现文档中的差异和变化。所以我们与他们合作了很多。我的直觉是,一旦我们找到一种方法在内部进行适当的版本控制、审计和发布模型,我们将在生命科学领域看到更多的采用。结果可能是为人们提供更多的药物和新药,这非常酷。

My bet is often on life sciences. For my companies, I've been working a bunch with Amgen and a few others. It's really interesting. When you ask Amgen why they exist as a company, their goal is to design new drugs. There are roughly two big chunks of work: the actual R&D experiments and validation by scientists, and then more admin work. The admin work is huge. The time it takes from locking the recipe of a drug to having it on the market is months, sometimes years. That's about generating really complex regulated documents, having committees review them, having regulators review them. A lot of information sharing, transformation, validation, which turns out models are pretty good at. They're good at aggregating, consolidating tons of structured and unstructured data, spotting diffs and changes in documents. So we've been working quite a bit with them. My hunch is, once we figure out a way to properly version, audit, and release the model with the right permissions internally, we're going to see a bunch more adoption in life sciences. The outcome will probably be more medications and new drugs for people, which is pretty cool.

Host

我注意到,任何监管严格、需要大量文书工作和提交材料的行业,都是 LLM 的完美用例。

I'm struck by how any industry that's heavily regulated and requires lots of paperwork and submissions is a pretty perfect LLM use case.

Olivier Godement

是的,而且到处都有。前几天我见了一家相当大的投资公司。类似地,他们的工作是实时聚合和分析海量数据,并试图理解这些数据以做出判断。同样,模型的超能力是在瞬间完成这种大规模分析。所以我们开始看到在对冲基金、投资基金和银行方面有不少采用。

Yeah, and there are some all over the place. I was meeting with a pretty big investment firm the other day. Similarly, their job is to aggregate and analyze mountains of data in real time and try to make sense of it to make judgment calls. Again, the model's superpower is to do that sort of massive analysis in a split second. So we're starting to see quite a bit of adoption in hedge funds, investment funds, and banks on that front.

Host

这很有趣。我注意到的一点是,如今的应用层很大程度上是那些深入某个行业、向前部署并询问你问题的公司。他们构建这些工作流程。即使对于你提供的所有例子——T-Mobile、Amgen、那家投资公司——也有像 Sierra 这样的初创公司可能想服务 T-Mobile,或者像 Colate 这样的公司可能想服务 Amgen 的用例。你怎么看什么时候适合你们与客户紧密合作,什么时候适合更广泛的生态系统?

It's interesting. One thing I'm struck by is how much of the application layer today feels like companies that go deep in some industry, forward deploy, and ask what your problems are. They build those workflows. Even for all those examples you provided—T-Mobile, Amgen, the investment firm—there are also startups like Sierra that might want to serve T-Mobile, or a company like Colate that might want to serve the Amgen use case. How do you think about when it makes sense for you guys to be the ones working super closely with customers, and when it makes sense for the broader ecosystem?

听众问答介绍 Introduction to listener Q&A

Host

今天我非常激动地介绍一个节目新环节,我们把话语权交给你们。如果你对 AI 生态有任何迫切问题想问 Jacob 和即将到来的嘉宾,请通过下方节目说明中的 Google 表单提交。我们非常好奇你们的想法,也非常感谢你们对节目的支持。

Today I'm super excited to introduce a new concept to the show where we open up the floor to you. If you have any pressing questions about the AI ecosystem that you'd want to ask Jacob and an upcoming guest, please submit them using the Google form that's listed in the show notes below. We're super curious to hear about what you have to say and very appreciative of your support of the show.

OpenAI角色与生态伙伴 OpenAI's role vs ecosystem partners

Host

你怎么看待什么时候适合你们自己与客户紧密合作,什么时候适合更广泛的生态?

How do you think about when it makes sense for you guys to be the ones working super closely with the customers and when it makes sense for the broader ecosystem?

Olivier Godement

嗯,这是个好问题。过去几年我在推动企业采用 AI 的过程中学到的是,问题的规模和深度都很大。比如你随便选一家公司,像 Amgen、T-Mobile、BNY 等,内部为了以那样的规模和质量运营,其复杂性和用例数量都极其庞大。所以我认为 OpenAI 没有幻想我们会是唯一构建优秀智能体或创造产品的公司。坦率地说,使命的一部分是赋能生态,比如第三方,让他们能更好地与我们合作。我们正在产品描述层面思考实现方式。今年 Dev Day 我们宣布了 ChatGPT 中的 Apps,我认为这是与生态合作的首批重要方向之一。

Yeah, it's a good question. What I've learned over the past couple of years working on AI adoption among businesses and the enterprise is frankly the size and the depth of the problem. Like once you pick any of them, like Amgen, T-Mobile, BNY, any company, the amount of complexity and use cases that you have internally in order to operate at that scale with that level of quality is absolutely huge. So I think we at OpenAI have no illusions that we're going to be the only ones to build really good agents or create products. Frankly, part of the mission is to enable the ecosystem, like third parties, to work better with us. We're thinking on ways to achieve that in the product description wise. I think we announced at Dev Day this year Apps in ChatGPT, which is one of the first big vectors of partnership with the ecosystem.

企业反馈与ChatGPT通用界面 Enterprise feedback and ChatGPT as universal interface

Olivier Godement

我们从企业那里听到的反馈基本上是,员工喜欢 ChatGPT 作为一种通用界面。他们想在 ChatGPT 中做更多事,而目前它基本上只限于 OpenAI 的功能。有大量初创公司想为特定市场构建特定功能,并从 ChatGPT 的采用、记忆和连接器中受益。所以这是一个方向,我认为未来我们会做更多类似的事情。

The feedback we've heard essentially from enterprises is that employees love ChatGPT as a sort of universal interface. They want to do more in ChatGPT, and at the moment it's pretty capped on just essentially OpenAI features. There are tons of startups that want to build that specific feature for that specific market and just benefit from the adoption, memory, and connectors of ChatGPT on day one. So that's one vector, and I think we'll do more of those in the future.

Host

未来是不是几乎每个企业员工一上班就会打开 ChatGPT?这是你的设想吗?

Is pretty much every enterprise worker in the future going to open ChatGPT right when they start their day? Is that kind of how you envision?

Olivier Godement

我觉得是的。当然很难准确预测。我不认为 ChatGPT 会取代所有工具。比如,如果我是金融分析师,我一天还是会花在 Excel 或某个为我的特定用例专门设计的工具上。但我确实期望 ChatGPT 会越来越成为你早上第一个查看的地方。

I think so. It's hard to predict for sure. I don't think ChatGPT is going to replace every tool. For sure, if I'm a financial analyst, I'm going to spend my day in Excel or some tool which was purpose-designed for my specific use case. But I do expect ChatGPT to become more and more like the first place that you check in the morning.

Host

我不知道你用了多少 Pulse。Pulse 对我来说是一个非常重要的功能。现在它非常精准,帮我准备一天。比如,嘿,昨晚那封邮件来了,那个会议要开了,顺便那篇论文出来了。它正在成为一个非常准确且有洞察力的高效信息源。所以我认为我们会看到更多这样的功能。我认为我们会看到更多人在 ChatGPT 中执行操作,也包括简单操作。所以我的直觉是,ChatGPT 可能会成为早上的第一个网站,然后当然一些深度工作流你会双击并在其他地方完成,比如在 IDE 中编码或在电子表格中处理数据。

I don't know how much you've been using Pulse. Pulse has been a pretty monumental feature for me. It's very on point now, preparing my day. Like, hey, last night that email came in, that meeting is coming up, by the way that paper came up. It's becoming a really accurate and insightful source of productive information. So I think we'll see more and more of that. I think we'll see more people taking actions in ChatGPT, simple actions as well. So my hunch is that ChatGPT could become the first website in the morning, and then of course some deep workflows you would double-click and do somewhere else, like coding in an IDE or data manipulation in a spreadsheet.

当前能力与未来飞跃 Current capabilities vs future leaps

Host

这个领域的一个迷人之处在于,当我们发现这些模型的能力时,它们也在变得更好。我想问的是,如果我们冻结今天的模型能力,是用当前这套能力还有无穷无尽的东西可以构建,还是我们只是在不断发现更多,或者你还在等待研究团队的下一步飞跃来解锁更多可能性?

One thing that's fascinating about the space is just as we're discovering the capability of these models, they also get way better. I think there's a question if we froze model capabilities today. Is there an endless amount of stuff to go build with this current set of capabilities, or are we just discovering more and more, or are you still kind of waiting for the next leap from the research team to unlock even further things?

Olivier Godement

哦,我觉得我总是在两三个不同的时间维度上运作。第一个时间维度就是当前模型能力,解锁尚未实现的用例。这可能很简单,比如有合适的框架、合适的用例、合适的数据,但用当前能力。然后我会提前几个月思考,比如对于 GPT-5.1,对于某些用例,需要一点智能的后训练来真正让体验更好、更快、更对齐、更可控。这是第二个时间维度。第三个是根本性突破,比如 o1 的形式。比如,如果模型能可靠地思考 30 分钟,你能实现什么?这个更难预测,因为你无法预测研究何时落地。但可以肯定的是,我们深入思考我们发布的产品和为客户启用的用例。它们会带来有意义的收益吗?通常当这发生时,就是魔法出现的地方。

Oh man, I feel like I always operate on two or three different time horizons. Time horizon number one is just with the current model capabilities, like unlock use cases which haven't been enabled yet. That could be as simple as having the right harness, right use case, right data, but with current capabilities. Then I'm trying to think a couple of months in advance, like for GPT-5.1, for some of those use cases it takes a little bit of smart post-training on the radar to really make the experience better, faster, more aligned, more steerable. So that would be the second time horizon. And then the third one is fundamental breakthroughs, like in the shape of o1 essentially. Like, what could you achieve if the model could reliably think for 30 minutes? That one is harder to predict because you don't predict when research is going to land. But for sure, we think pretty deeply about the products we release and the use cases we enable with customers. Are they going to meaningfully benefit? And usually when that happens, that's where the magic shows up.

下一代模型前沿 Next model frontiers

Host

你思考的下一个模型前沿是什么,或者你会觉得,哦,当 X 或 Y 实现时会很棒?

What are the next model frontiers that you think about, or you're like, oh it'd be amazing when X or Y?

Olivier Godement

哦,有很多。从客户业务用例的角度来看,我认为一旦我们攻克持续学习,那可能会是一个非常重大的转变。

Oh man, there are many of them. From a customer business use case perspective, I think once we crack continuous learning, that probably will be a very meaningful exchange.

智能体实习生类比 Agent as Intern Analogy

Olivier Godement

我的想法很简单:我把智能体看作入职第一天的实习生。他们有很多学术知识,但缺乏大量实践经验,因为我从没把自己做的每件事都记录下来。本质上,你必须边干边学。

I think of it very simply: I think of an agent as like hiring an intern on day one. They have a bunch of academic knowledge, but they don't have a ton of practical knowledge on the job because I have never been documenting everything I do. You have to learn on the job essentially.

Host

所以我认为与智能体的关系很大程度上将基于反馈和标注,比如“嘿,你那么做了”或“你那么回应了,这次你应该稍微不同地做”。模型会随着时间吸收这些反馈。目前,我们通过提示词等方式实现自动化,但一旦模型能根据人类反馈(比如实时信号反馈)实际更新权重,我认为这将解锁大量用例:编程、客户体验、金融。没错,它将无处不在。

And so I think the relationship with an agent is very much going to be based on feedback, annotations, and like, "Hey, you did that," or "You responded that way, you should do it slightly differently this time." And the model incorporates that over time. At the moment, we automate it through prompting and that sort of stuff, but once the model is able to actually update its weights based on human feedback, like in-time sign feedback, I think that's going to unlock quite a few use cases: coding, customer experience, finance. Yeah, it's going to be all over the place.

Olivier Godement

基本上,你的智能体每天早上都像实习生一样出现,但带着越来越好的指令,所以如果它能学习就好了。

Basically your agent shows up as an intern each morning, but with better and better instructions, and so it would be nice if it learned.

Host

它们消化吸收了,结果变得更聪明了。

They slept on it and they're smarter as a result.

产品市场契合类别 Product-Market Fit Categories

Host

从我这边看,我很好奇。如果一年前你问我哪些类别在 AI 领域真正实现了产品-市场匹配,我会说编程、客户支持,可能还有医疗和法律作为应用。如果你今天问我,我仍然觉得这四类是主导。跨领域也有一些有趣的语音应用。但你在企业一线总能见到这种情况。我猜你会加上生命科学。还有别的吗?你归类为“疯狂的产品-市场匹配”、“一般匹配”、“仍早期”的?你觉得还有什么能放进“疯狂的产品-市场匹配”这个桶里?

I'm curious from my seat. It feels if you asked me a year ago about the categories that really had product-market fit in AI, I would have said coding, customer support, maybe healthcare and legal too as applications. And if you ask me today, I still feel like those four are kind of the dominant ones. There have been some interesting voice applications as well across domains. But you see this all the time on the ground with enterprises. I guess you'd add life sciences. Anything else that you categorize as insane product-market fit, kind of product-market fit, still early? Anything else that you feel goes in insane product-market fit bucket?

Olivier Godement

是的。有时我提醒团队,当我们谈论编程、客户体验、金融时,这些都是巨大的市场。

Yeah. Sometimes I remind my team when we talk about coding, customer experience, finance, those are gigantic markets.

Host

哦,100%。

Oh, 100%.

Olivier Godement

编程市场,软件市场。我是说,我和团队聊过:我们不知道它有多大,对吧?我们从未有过这么便宜的……软件的时间有多大?我们确定的是软件工程师短缺,所以可能比当前软件工程师的薪酬还要多,可能多得多。所以坦白说,我们的思考方式是,GPT-3 之后的头几年很像是在四处撒网,看什么能奏效。现在我认为我们对行业、用例、领域有了更清晰的认识,我们认为第一,存在巨大的客户问题;第二,模型会持续改进,使市场更高效、更有效。所以我目前的理念是,尝试在这些市场上加倍投入。当然我们一直在扩展。我们谈过科学。我不知道我们是否会做,但鉴于收到的反馈,我们可能应该做点科学相关的事。但在编程上,我们可以推得更远。客户支持:好吧,你自动化了一级工单。走得更远意味着什么?把客户支持变成公司真正的收入来源意味着什么?让它更加个性化,诸如此类。所以是的,我会说类似的领域,但越来越深入。

The market for coding, the market for software. I mean, I was chatting with the team: we have no idea how big it is, right? We've never had such cheap... How big is the time of software? What we know for sure is that there's a shortage of software engineers, and so probably more than the current pay of software engineers, probably way more than that. And so frankly, the way we think about it, I would say the first few years post-GPT-3 were very much like spraying and trying to see what sticks. Now I think we have a much better picture of the industries, use cases, domains on which we think one, there is a massive customer problem, and number two, the models are going to keep improving and make it more efficient, essentially effective in that market. And so my current philosophy frankly is to try to double down more on those markets. Of course we expand all the time. We talked about science. I don't know if we'll do it, but we should probably do something on science given the feedback we're receiving. But on coding, we can push it much further. Customer support: okay, you automated tier one tickets. What does it mean to go much further? What does it mean to turn customer support into an actual revenue maker for the company? Have it be way more personalized, that sort of stuff. So yeah, I would say similar domains but going deeper and deeper.

未来领域与企业采用 Future Domains and Enterprise Adoption

Host

是的。所以你认为一年后我们会有类似的清单。

Yeah. So you think a year from now we'll kind of have a similar list of the stuff.

Olivier Godement

我确信会有新领域。如果一年前你问我,我不确定生命科学、制药、医疗是否在其中。现在我能看到了。而且其中一部分不仅仅是模型能力;还有纯粹的软件和变革管理。

I'm sure there'll be new domains. If you had asked me a year ago, I'm not sure life sciences, pharma, healthcare were in there. Now I can see it. And some of it is not just model capabilities; it's pure software and change management.

Host

完全同意。

Totally.

Olivier Godement

那些企业有庞大的系统、大量员工,所以让 AI 适应并调整以很好地服务于这些案例需要大量工作。所以回答你的问题,假设我们把研究冻结几年,还会有很多年的企业采用,仍然非常有价值。

Those enterprises have massive systems, lots of employees, and so having AI be adapted and adjusted essentially to work really well for those cases is a lot of work. And so to your question, let's say we froze research for a few years, there would be many years of enterprise adoption that would still be quite valuable.

Host

感觉在某些行业,你几乎达到了一个临界点,足够多的人在使用它,以至于不使用它开始变得像是不负责任。我觉得即使是客户支持:人们一开始只是试水,然后有些人取得了巨大成功,现在它变成了这样一件事:如果你不尝试这些解决方案,你几乎在干什么?

And it feels like in some of these industries you kind of just reach a tipping point where enough people are using it that it starts to become like irresponsible not to. I feel like even customer support: people were kind of dipping their toes in, and then some people had huge success with it, and now it's become one of these things like if you're not trying one of these solutions, what are you almost doing?

Olivier Godement

当然。我是说我们在每个市场都看到这一点。坦白说,有一些先驱、企业初创公司,本质上就是愿意承担风险的公司。一开始感觉摇摇欲坠;需要大量脚手架,东西几乎站不住。但到了某个点,你能看到它起作用了,你加固它,然后基本上每个人都跟着做。所以是的,我们在整个行业都看到同样的动态。

Of course. I mean that's what we see with every market. And frankly there are a few pioneers, enterprise startups, companies essentially who are willing to take the risks. It feels shaky at first; it requires a lot of scaffolding and the thing barely stands. But at some point you can see the thing working, you harden it, and then everyone follows essentially. So yeah, we see the same motion across basically industry.

跨用例的脚手架模式 Scaffolding Patterns Across Use Cases

Host

人们构建的脚手架和框架在不同用例中是否相当相似?比如在生命科学、支持和编程中有一套通用的脚手架模式,还是脚手架针对最终问题有多定制化?

Does it feel like the scaffolding and the harnesses that people are building are pretty similar across use cases? Like a common set of patterns for scaffolding in life sciences and support and coding, or how bespoke to the end problem is the scaffolding?

Olivier Godement

我会说目前相当定制化,坦白说这是一个挑战。我的说法是,坦白说人们正在不惜一切代价让它工作。

I would say it's fairly bespoke at the moment, which is a challenge frankly. The way I would put it is frankly people are trying to make it work whatever it takes.

Host

是的。一个智能体、多个智能体、中间有一些确定性门控。人们正在尝试许多不同的东西。

Yeah. One agent, multiple agents, some deterministic gates in between. People are trying many different things.

智能体架构与标准化 Agent architecture and standardization

Host

我觉得还没有——我是说我们开始看到了,但还没有一个真正标准的智能体架构或运行时被跨行业采用。这是我们正在积极推动的事情。

I think there hasn't been, I mean we're starting to see it, but like there hasn't been a really standard agent architecture or runtime, you know, which has been adopted across industries. That's something that we are actively working on.

Olivier Godement

你为什么觉得还没有?

Why don't you think there has been?

Host

我的意思是,这才三年而已,你明白吗?我们花了——我在 OpenAI 干了大概两年半。第一年我光顾着跟上增长,心里想‘这到底是怎么回事?’第二年就是‘好吧,坐下来看看我们能实现什么,什么行得通,什么行不通。’现在我觉得我们又开始收敛了:‘好了,我们对能力、客户问题、正在做的事情有了清晰的认识。栈里哪些部分如果标准化了,就能显著加速采用?’所以我的无聊答案其实就是时间问题。

I mean, it's only been three years, you see what I mean? It took us—I mean, I've been at OpenAI for like two and a half years. The first year I was just trying to keep up with the growth and be like, 'Okay, what the heck is going on?' Year two was like, 'Okay, let's sit down, what can we achieve, what is working, what is not working?' And now I think we're starting to converge again on like, 'Okay, we have a good idea of capabilities, customer problems, what is being done. What are the pieces in the stack that, if they were standardized, would meaningfully accelerate adoption?' So my boring answer is frankly just time.

Host

那你觉得那套标准的脚手架可能是什么样的?

And how are you thinking about what that standard set of scaffolding might look like?

Olivier Godement

我觉得我们看到的其中一点是,代码和编程是一种比纯软件工程更通用的能力。模型在生成和编写代码方面非常出色,而且进步速度快得惊人。所以这似乎是一个值得押注的好方向。

I think something that we're seeing is that code and coding is a much more general-purpose capability than just software engineering. The models are really, really good at generating and writing code, on an insane progress path too. So it seems like a great thing to bet on.

Host

没错。比如写个脚本,在 shell 这样的工具里执行它,然后拿回结果。所以我认为,让智能体通过计算机访问资源的整个努力,很可能会成为行业标准。在数据和 API 连接方面,我觉得 MCP 是一个很好的标准。我预计行业会继续围绕它规范化。在智能体之间的通信方面,目前还没有真正的突破或真正被使用的标准。评估方面——我觉得我们在生成轨迹、评估轨迹、尝试从中推断改进方面做得越来越好了。所以这其实是一场寸步之争。

Exactly. Like writing a script, executing that thing in a tool like a shell, having it back. So I think the whole effort towards giving access to agents over computer is probably going to become a standard in the industry. On the whole data and API connections, I think MCP was a really good standard. I expect the industry to keep normalizing around it. On agent-to-agent communication, there hasn't been yet a true breakthrough or a true standard that we see being actually used. Evaluation—I think we're getting much better at generating traces, evaluating them, trying to infer some improvements on the traces. So it's a bit of a game of inches, actually.

Host

但你觉得——听起来你话里的意思是,把更多这些东西迁移到代码上非常有帮助,因为这是这些模型进步最快的方向。所以你能做的——任何你能迁移到代码上的脚手架部分——可能都相当有用。

But you think—it sounds like behind what you're saying is that moving more of this stuff to code feels very helpful because it's the vector along which these models are improving at the most rapid pace. So anything you can—any part of your scaffolding that you can move to that—is probably quite helpful.

Olivier Godement

完全正确。坦白说,这有点像人类。如果你让我不带笔记本电脑做点工作,我不是没用,但也离没用不远了。而如果你给我一台带网络、shell、IDE 的笔记本电脑——好东西——你的能力就会大得多。所以,如果我们能用智能体和模型复制这种模式,我认为我们会看到大量进展。而且,正如你所说,我们对模型核心编码能力所做的每一次进步,都会因此产生数量级更大的影响。

Exactly. It's a bit like a human, frankly. If you ask me to do some work without a laptop, I'm not useless, but I'm not far away from useless. Versus you give me a laptop with internet, a shell, an IDE—good stuff—your capabilities are way, way larger. So to the extent we can replicate essentially that model with an agent with a model, I think we are going to see a bunch of progress. And to your point, every progress that we do to the core coding capabilities of the models is going to have orders of magnitude more impact as a result.

成本作为限制因素 Cost as a limiting factor

Host

我觉得似乎当这些模型推出时,每个人都处于用例探索模式,只是寻找模型能做的所有事情。我想知道,成本在今天是不是某些事情的限制因素?我听到你在 BG2 上提到过,我很好奇,因为感觉在某个价格点上存在有趣的用例,但在另一个价格点上可能就不存在了?

I guess it seems like when these models come out, everyone is in use case exploration mode, just finding all the things the models can do. I'm wondering, does it feel like cost is a limiting factor today on some of these things? Where I heard you say that on BG2 and I was curious, because where does it feel like there are interesting use cases at one price point but maybe not at another?

Olivier Godement

退一步看,过去两三年我们看到的成本和价格下降,在科技领域基本上是闻所未闻的。我们在三年内将 GPT-4 级别查询的成本降低了一到两个数量级,而且没有牺牲利润率之类的东西。这纯粹是压缩模型大小、拥有更好的硬件、更好地将 GPU 联网——栈的每一层。我与从事栈各层工作的团队交流时看到的是,仍有大量优化和改进的空间,所以我们不会止步于此。在商业方面,对于一些高风险的用例,比如编程,经济账是划算的——将软件工程师的生产力翻倍杠杆率极高,因此每月支付几十或几百美元可能是值得的。但还有很多其他用例目前被阻塞了。例如,模型在个性化、理解意图和调整内容方面相当不错。那么为什么世界上每个内容网站的主页都没有融入这些功能?我想我知道答案:成本和延迟。我非常认为不断降低成本是 OpenAI 使命的一部分。过去两年我看到的是,我们进行了几十次成本削减。每次,量的增长都超过了价格效应。这告诉我,仍然存在巨大的未开发需求,基本上受限于成本。

If I step back, the reduction of prices and cost that we've seen over the past two or three years is basically unheard of in technology. We reduced the cost of GPT-4-level queries by one to two orders of magnitude in three years, without sacrificing margin or something. It's pure compressing the model size, having better hardware, being better at networking together GPUs—every layer of the stack. What I see talking with teams who are working on each layer of the stack is that there is still a ton of room for optimization and improvement, so we are not going to stop there. On the business side, for some very high-stakes use cases, say coding, the economics work—it's so high leverage to multiply by two the productivity of your software engineer that paying dozens or hundreds of dollars every month could be worth it. But then there are many other use cases which are currently being blocked. For example, models are pretty good at personalization and understanding intent and adjusting content. So why isn't every content website in the world's homepage infused with that? I think I know the answer: cost and probably latency. I very much see it as part of the OpenAI mission to keep driving the cost down. What I've seen over the past two years is that we've done dozens and dozens of cost cuts. Every time, the increase in volume is larger than the price effect. That tells me that there is still a wild untapped demand which is basically limited by cost.

Host

是的,我也很期待看到它在 Sora 模型上的表现。就像你有 API,很明显随着时间的推移,这些模型的成本会大幅下降,我认为人们会找到各种各样的用途。

Yeah, I'm really excited to see it play out on the Sora models too. Like you have the API and it's clear it's going to be a massive cost reduction over time in those models, and I think people will find all sorts of uses.

Olivier Godement

完全正确。所以如果你看看我们目前一些最成功的智能体,它们运行几个小时或几十分钟——这很快就会变得相当昂贵。如果你真的想进入一种模式,让我们每个人都拥有 1000 个智能体几乎一直异步并行运行在后台,我们整个行业需要降低栈的每一层成本。我有信心我们会实现这一点。

Exactly. So if you look at some of our most successful agents at the moment, they run for hours or tens of minutes—that can get pretty expensive quickly. If you truly want to move to a mode where each of us has 1000 agents running in parallel asynchronously in the background pretty much all the time, we need as an industry to bring down the cost at every layer of the stack. I have good conviction we'll get there.

Host

是的。那 RFT 呢?我觉得它最近非常流行。

Yeah. What about RFT? I feel like it's very in the zeitgeist these days.

模型调优的阶跃变化 Step change in model tuning efficacy

Host

你知道吗,我觉得很多人似乎认为,与 SFT 范式相比,调整这些模型的效率有了真正的阶跃式提升。你实际上与这些企业密切合作吗?你看到人们开始使用这个了吗,以及我们在越来越多的人为他们的用例调整模型的旅程中处于什么位置?

You know, I feel like a lot of people seem to think there's been a real step change in improvement in efficacy of tweaking these models versus maybe the SFT paradigm. Are you actually working closely with these enterprises? Are you seeing folks start to use this, and where are we in the journey of more and more people actually tweaking these models for their use cases?

Olivier Godement

我们开始看到了,但还没有广泛采用。目前我对企业机会的看法是,大部分市场正在追赶,真正在追赶前沿。就像我们之前讨论的,他们还没有充分利用 GPT-4o 基础模型的能力,所以这就是市场的要点。然后有一些创新者明显被前沿所阻碍,他们确切知道需要什么才能达到那里。几个例子:我曾与一家会计软件公司合作,进行极其准确的税务会计分析,模型开箱即用不够好,太慢了。大概需要几十个高质量环境和导航器的样本,来更新和改进模型,在他们的黄金标准评估上提升大约 20% 到 30%,这本质上就是让他们从不太可行变成实际可行的差异。所以我们开始看到一些创新者识别出这些问题,尝试调优之外的一切方法让模型工作,但没有成功,因此有了突破。话虽如此,这仍然需要大量工作。你必须构建一个真正高质量的环境,耗时数月。你必须确保你的评分器很好。一旦你启动强化学习任务,可能需要几个小时,有时甚至几天。所以这更重量级。我认为我们发布了一个 API,我相信这可能是市场上第一个这样的 API,所以我们看到了相当多的兴奋,但这显然还不是大众市场,坦率地说可能永远不会。但这没关系。只要我们在为一些客户推动前沿,我就很高兴。

We are starting to see it, but it's not widely adopted yet. The way I think about the enterprise opportunity at the moment is that most of the market is catching up, really catching up to the frontier. Like our discussion earlier, they haven't yet fully leveraged the capabilities of GPT-4o base, so that's the gist of the market. Then you have a few innovators who are clearly blocked at the frontier, know exactly what they need to get there. A couple of examples: I was working with an accounting software firm to make extremely accurate tax accounting analysis, and the model was just not quite good enough, too slow essentially out of the box. It took maybe a few dozen samples of really high quality environments and navigators to update and improve the model by maybe 20 or 30% on their gold standard eval, which was essentially the diff that allowed them to get from not really viable to actually viable. So we're starting to see some innovators identifying those things, trying everything they can outside of tuning to make it work, and not succeeding, and therefore having a knock with that. With that said, it's still a lot of work. You have to build a really high quality environment, months. You have to really make sure your graders are good. Once you kick off a reinforcement learning job, it can take hours, can take days sometimes. So it's more heavy-handed. I think we released an API which I believe was probably the first one on the market in that space, so we're seeing quite a bit of excitement, but it's clearly not yet a mass market, maybe never will frankly. But that's fine. As long as we're pushing the frontier for some customers, I'm happy about it.

多数企业会用RL吗? Will most enterprises use RL?

Host

你认为随着时间的推移,大多数人最终会使用它,还是说你可以只是依赖通用模型的改进?对于大多数企业来说,谁需要将前沿推进 6 个月或 12 个月?

Do you think over time most people end up using it, or is it like you can kind of just tie yourself to the general model improvements? Who needs to push the frontiers 6 months ahead or 12 months ahead for most enterprises?

Olivier Godement

我推测大多数企业会有一些强化学习用例,但我推测这些强化学习用例不会是他们业务的核心。他们会在模型的某些方面进行创新,但对于大规模自动化运营、构建知识库、维护和更新,我预计基础模型开箱即用就会相当不错。

I suspect that most enterprises will have some RL use cases, but I suspect those RL use cases are not going to be the gist of their business. They will want to innovate on some aspect of the model, but for the massive automating your operation, making your knowledge base, maintaining it, updating it, I expect the base model will be pretty good out of the box.

对初创公司的影响 Impact on startups

Host

那在初创公司方面呢?我觉得在很长一段时间里,训练自己的模型或者花太多时间在微调上似乎很傻,对吧?它帮助不大。现在强化学习似乎确实能提升性能。有一段时间我会说,嘿,大多数 AI 初创公司需要知道如何很好使用这些模型的人,但可能不需要那些在这些模型上做所有事情非常深入的人。你认为在这个范式下情况改变了吗?

What about on the startup side? So I feel like for the longest time it seemed really silly to train your own model or spend too much time on even fine-tuning, right? It didn't help that much. Now it does seem like RL can push performance. For a while I would have said hey, most AI startups they need people that know how to use these models well, but maybe not people that are super deep in doing all on them. Do you think that's changed in this paradigm?

Olivier Godement

这是个好问题。我喜欢提醒初创公司和人们我们在 OpenAI 做的事情。我们大量后训练模型,大量微调模型。所以当然我认为我们做得很好,模型针对大多数用例进行了后训练。但当然有些用例中模型训练得不够好,它的行为不会完全正确。风格、格式、语气、简洁性都不会完全到位。所以我预计那些达到一定规模或想要真正提升到下一级能力的初创公司会继续微调。这工作量很大。也许在某个时候我们能够实现那种持续学习或某种自动微调。我们还没有完全做到。但是的,我确实预计一部分初创公司会继续使用微调来突破极限。

It's a good question. I like to remind startups and people of what we do at OpenAI. We post-train a lot the models, we fine-tune a lot the models. So of course I like to think of us doing a really good job and the model is post-trained for most use cases. But of course there are some use cases where the model will be less well trained, it will behave not exactly the right way. The style, the formatting, the tone, the conciseness will not be quite it. And so I expect that startups who are achieving a certain scale or want to get really to the next level of capabilities will continue to fine-tune. It is a lot of work. Maybe at some point we'll be able to achieve that continuous learning or some sort of automated fine-tuning. We're not quite there yet. But yeah, I do expect that a fraction of startups will continue to use fine-tuning to push the envelope.

模型选择的关键因素 Key factors for model choice

Host

当你思考开发者在选择模型时所做的决定时,显然有这些模型的整体质量。你认为还有什么最终会驱动初创公司和开发者选择在哪个平台上构建?

As you think about the choices developers make in the models they use, obviously there's the overall quality of these models. What else do you think will ultimately drive where startups and developers choose to build on?

Olivier Godement

是的,我喜欢把它想成三大类。第一类显然是模型能力和行为,我们会分解它。第二类是成本和延迟。第三类是氛围、趋势,基本上就像 Twitter,这变得越来越重要,因为……

Yeah, I like to think of it as three big buckets. One is clearly model capabilities and behaviors, and we'll decompose that. Number two is cost latency. Number three is vibes, trends, like Twitter essentially, and that's becoming more and more important because...

Host

这就像现在的评估标准,对吧?Gemini 3 今天发布了,我想,是啊,所有百分比看起来都不错,但让我们看看 Twitter 怎么说。坦率地说,跟上所有东西真的很难,所以我开始看到越来越多严肃的初创公司和企业在很大程度上依赖特定的影响者、媒体,我不知道。但无论如何,能力和行为肯定是最大的类别。我几乎一直看到的趋势是,人们不会过早地优化成本和延迟。首要目标是让它工作,并希望让它工作得便宜且足够快。我在这里看到的趋势是,在某个时候,你可以把大量信息塞进学术基准测试。如果我在构建一个客户支持智能体,当然,SWE-bench、MMLU 可能会给我一些指示,但坦率地说,可能没有那么多。我们开始看到一些有趣的行业级基准测试。Tow-Bench 在服务业是一个很好的例子。我们正在努力构建更多这样的基准测试。几个月前我们发布了 GDP,它本质上是为了分析模型在现实世界经济任务上的能力。

That is like the eval these days, right? Gemini 3 comes out today and I'm like, yeah, all the percents seem good, but let's see what Twitter says. It's really hard to keep up frankly with everything, and so I'm starting to see more and more serious startups and enterprises rely way more on specific influencers, media, I don't know. But anyway, capabilities behavior for sure the biggest bucket. The trend I'm seeing pretty much all the time is people are not prematurely optimizing for cost and latency. The first goal is to make it work, and hopefully to make it work cheaply and fast enough. The trend I'm seeing here is that at some point there is so much information you can cram into academic benchmarks. If I'm building a customer support agent, sure, SWE-bench, MMLU probably give me some indication, but frankly probably not that much. We're starting to see some interesting industry level benchmarks. Tow-Bench is a good one in the services industry. We're trying to build more and more those. We released a few months ago GDP, which was essentially meant to analyze the model capabilities on real world economic tasks.

模型疲劳与评估挑战 Model Fatigue and Evaluation Challenges

Olivier Godement

但我想说的是,行业目前可能处于追赶模式,初创公司没有奢侈的时间等我们慢慢来。所以我看到很多基本上是定性测试,这挺有意思的。我开始在这些初创公司中看到一些人,他们对模型的细微差别有很好的品味,就像那些非常擅长写作或绘画的人。他们不一定能详细阐述框架,但他们有那种感觉。我开始看到同样的事情发生在模型上,这很酷。成本和延迟非常重要。我的预期是成本将在未来一年左右继续降低数倍。然后还有 Twitter 上的氛围。我觉得我们在某种程度上重新发明了 Gartner 等公司存在的理由,那就是在某个时候你无法比较世界上所有的会计软件;你必须信任某人。所以我认为我们正在达到那个阶段,而且我们会停留在那里。

But I would say here the industry is probably in a catch-up mode, and startups do not have the luxury to wait for us to come up. So I'm seeing a lot of frankly qualitative testing, which is pretty interesting. I'm starting to see among these startups some people who have such a good taste for nuances of models, just like people who are really good at writing or painting. They're not necessarily able to elaborate on the framework, but they have that sense. I'm starting to see the same happen for models, which is pretty cool. Cost and latency are pretty important. My expectation is cost will continue to be reduced by multiple times over the next year or so. And then you have Twitter vibes. I feel in a way we are reinventing why Gartner and others exist, which is at some point you cannot compare all the accounting software in the world; you have to trust someone. So I think we're getting at that stage, and I think we'll stay there.

Host

我喜欢 Twitter 就像中心化的把关人。是的,说得好。有一件事我觉得非常有趣,请原谅我提起这个,我想是在 Anthropic 模型的背景下,当 4.5 Sonnet 发布时,Cognition 被提到了,他们不得不把所有东西迁移过去,这实际上需要他们做大量全新的工作。我想知道,你与所有这些企业合作,然后有新模型如 5 或 5.1 发布。这个过程实际上是什么样的,你想象它会如何随时间变化或演变?

I like that Twitter's like the centralized guard. Yeah, that's a good way to put it. One thing I thought was really interesting, and you'll have to forgive me for referring to it, I think it was in the context of Anthropic models, but Cognition was talked about when 4.5 Sonnet came out, I think it was, and they had to move everything over and it actually required a ton of net new work from them. I'm wondering, you work with all these enterprises and then you have a new model like 5 or 5.1 come out. What does that process actually look like and how do you imagine that changing or evolving over time?

Olivier Godement

这就是模型疲劳的一部分。我认为对于非平凡用例来说,那种只需切换一个 API 参数就能从一个模型热替换到另一个模型的日子基本上已经过去了。模型的特异性,尤其是在不同提供商之间,变得越来越明显。有些模型对某些类型的指令响应更好。有些模型针对特定的工具签名、特定的工具名称进行了预训练或后训练。有些模型能处理更多或更少的长上下文召回。所以你必须从根本上理解模型的所有怪癖,然后调整你的提示词和框架来适应它。这工作量很大。即使在最顶尖、最成熟的初创公司中,我们看到每次都能准确完成这项工作很难,需要大量工作,所以他们宁愿不做,除非有重大变化。所以你可以想象在企业中,主要工作不是实施;主要工作是设计药物或销售手机。他们非常渴望转向一个更规律的节奏,带有清晰的变更日志。我觉得我们在某种程度上重新发现了如何部署软件。如果我每天给你一个新二进制文件,你会说‘祝你好运,酷,但我该怎么办?’而如果是‘这是版本 1.1,是一个主要版本,这是变更日志,你可以在 3 个月内期待更多。’人们只想要可预测性和透明度。

That's part of the model fatigue. I think the days where you could just hot swap essentially one API parameter from one model to the next are basically gone for non-trivial use cases. The idiosyncrasies of the models are becoming, especially among different providers, more and more distinct. Some models respond better to certain types of instructions. Some models have been pre-trained or post-trained for specific tool signatures, specific tool names. Some models handle more or less very long context recall. So you have to essentially understand all the quirks of models and then adjust your prompt, your harness, to adapt to it. It's a lot of work. Even among the top, most sophisticated startups, what we see is that doing it every time and doing it accurately is hard, takes a lot of work, so they would rather not do it unless there's a meaningful change. So you can imagine in the enterprise, the primary job is not implementation; the primary job is to design drugs or sell cell phones. They are very much eager to move to a much more regular cadence with clear changelogs. I feel we are sort of rediscovering how to deploy software. If I drop a new binary every day, you're like, 'Good luck, cool, but what should I do about it?' versus 'It's version 1.1, it's a major version, here's a changelog, and you can expect more in 3 months.' People just want predictability and transparency.

Host

你认为随着时间的推移,脚手架会演变得更通用,从而能更好地吸收这些模型的变化,还是说永远都会是‘嘿,你得花一周时间搞清楚什么变了以及你需要适应什么’?

Do you think that over time the scaffolding evolves to be more generalizable such that it can better absorb changes in these models, or is it always going to be like, 'Hey, you're going to have to take that week sprint to figure out what changed and what you need to adapt'?

Olivier Godement

我认为随着我们标准化智能体架构,它必须变得更加通用。我认为目前不同模型之间存在如此大差异的一个原因是,每个实验室基本上都在针对不同的框架、不同的目的、不同的用例进行训练。而且还没有一个统一的通用架构或框架可以遵循。所以我希望行业在某个时候会收敛到一个更通用的智能体框架,本质上类似于 MCP,但涵盖智能体的每一个维度,这将使客户更容易比较不同模型、采用新模型以及为不同用例采用多个模型。

I think as we standardize the agent architecture, it has to become more generalized. I think one reason why at the moment there is so much discrepancy across different models is that each lab is training for different harnesses, different purposes, different use cases essentially. And there is not yet a single common architecture or framework to follow. So my hope is that at some point the industry will converge to a much more universal agent framework, something like MCP essentially, but across every single dimension of the agent, and that will make it easier for customers to compare different models, to adopt new ones, and to adopt multiple models for different use cases.

Host

你对构建者有什么建议吗?那些真正优秀的顶尖团队在尝试新模型并准备将所有东西迁移到新模型时是怎么做的?对其他人有什么经验教训吗?

Do you have any advice for builders? What do the really good top teams do when they're experimenting with a new model and trying to get ready to shift everything to that new model? And any lessons for the rest of the world?

Olivier Godement

坦率地说,他们会花时间。我看到的那些过于急躁的团队,只想热替换模型名称,然后说‘哦,糟糕,那个模型很烂,效果不好。’所以最好的团队,我认为,有很强的品味但保持开放的心态。他们花时间对模型进行实战测试,也花时间与我们合作。坦率地说,我们控制后训练,所以如果有些团队特别希望某个工具以特定方式表现,我们可以影响这一点。我们收到的带有具体例子的反馈越具体,我们实际上可以在下一个模型快照中为此进行调整。所以通常就是这样。

Frankly, they take the time. What I've seen is teams that are way too impatient and just want to hot swap the model name and then say, 'Oh shoot, that model sucks, doesn't work as well.' So the best teams, I would say, have strong taste but come with an open mind. They take the time to battle test the model, take the time to work with us as well. Frankly, we control the post-training, so if some teams are really ganged up on having a specific tool behave in a specific way, we can influence that. The more specific feedback we receive with specific examples, we can actually tweak for next snapshots of models for that. So usually that's what happens.

Host

我们还没有谈到语音,我想聊聊这个。感觉这是过去一年变化很大的一个方面。有很多人在用实时 API 进行构建。你如何看待那里的下一个前沿?你在哪里看到很多契合点,以及你认为我们还需要做哪些事情来解锁下一阶段?

We haven't hit on voice yet, and I want to. It feels like that's one thing that's just changed so much in the last year. There's so many people building with the real-time API. How do you think about the next frontiers there? Where are you seeing a bunch of fit and where do you think we still need to do X, Y, or Z to unlock the next set?

语音AI与图灵测试 Voice AI and the Turing Test

Host

这很有意思。GPT-4o 是去年五六月发布的,到现在也就一年多。对我来说,它可能是继 ChatGPT 之后,第二个让我真正感受到 AGI 的突破。就像,世界将从此不同。一个模型能表达如此丰富的情感和语调,还能理解人类的情感和语调,这真的让我震撼。话虽如此,我觉得我们在语音方面显然还没有通过图灵测试。在文本方面,我敢说我已经无法区分人类和机器人了。但在语音上,打断的时机、节奏还是不太对。所以我认为语音的下一个前沿就是:让世界更智能。但其次,要达到自然和表现力的程度,基本上在同等智能水平下,你愿意接受 AI 服务就像接受人类服务一样。我觉得我们有希望实现这一点,所以我认为要真正跨越鸿沟,我们开始看到部署了。我回到客户支持,因为这是世界上主要的语音通话用例之一。我们开始看到一线客户支持电话中有意义的部署,这很好,有几个原因。第一,模型有无限的耐心。所以对于需要更多时间得到答案的人,模型最多可以花五分钟。第二,我当初没意识到的是,多语言能力对客户支持有多关键。比如你在美国,是一家大型零售商。大多数客户说英语,其他人说西班牙语、中文,还有长尾语言。如果你要为每种语言配备客服人员,基本上行不通。所以你不得不做出艰难的取舍,结果就是流失客户。因此,我们看到了多语言能力带来的强劲效果和客户反馈。

So it's interesting. GPT-4o, which came I think in May or June last year, so a little more than a year ago, to me was probably the second breakthrough in terms of feeling the AGI after ChatGPT. Like, man, the world will not be the same. Essentially, having a model being able to express that range of tone and emotions, and to understand as well, as a result, from the human such tone and emotion, was like, put in my brain. With that said, I think we clearly haven't crossed the Turing test yet for voice. For text at that point, I'm pretty sure I wouldn't be able, frankly, to distinguish between a human and a bot. On voice, it still feels like the interruptions, the cadence, still not quite exactly it. And so I think that's probably the next frontier on voice: make the world more intelligent. But second, get to a point of naturalness and expressiveness where basically, provided the same level of intelligence, you're okay to be served by AI just as you'd be okay to be served by a human. I think we have a line of sight to get there, and so I expect that's going to be through to really cross the chasm. We are starting to see deployment. I come back to customer support because that's one of the main voice calls use cases in the world. We're starting to see meaningful deployments among tier one customer support calls, which is good for a couple of reasons. Number one is that the model is infinitely patient. So for people who need more time to have their answer, the model can take five minutes at most. The second thing which I did not quite realize is how critical multilingual capabilities were for customer support. Because let's say you're in the US, you're a big retailer in the US. Most of your customers speak English, others will speak Spanish, Chinese. There's a long tail of languages. And if you have to staff customer support agents for every language, you're not going to make it essentially. So you have to make really hard trade-offs and literally lose customers as a result. And so we've been seeing pretty strong results and reception from customers on those multilingual capabilities.

Codex与软件工程领域 Codex and the Software Engineering Space

Host

是的,我们之前聊到,过去几个月 Codex 几乎成了生态系统的主角,显然进步巨大。你怎么看这个领域的发展?除了模型本身的能力,你觉得最终是什么因素决定人们会选择 Codex、Claude Code 还是其他工具?

Yeah, we were talking before about how Codex has kind of become the main character of the ecosystem for the past months, and obviously there's tremendous improvement there. How do you see that space playing out? And besides the underlying capabilities of the models, what do you think will ultimately determine whether folks reach for Codex or Claude Code or one of these other tools going forward?

Olivier Godement

好问题。Codex 团队非常出色。在我看来,他们是一个小而精的团队,专注于一个用例,全力以赴:模型、工具链、集成、数据,无所不包。所以他们真的非常优秀。我对软件工程领域如何演变的看法是:目前,模型在生成和理解代码方面已经非常出色。它们还可以更好,但已经取得了重大进展。但当我想到软件工程师时,那只是工作的一部分,对吧?另一部分工作是待命值班,与队友沟通,确定变更范围,做出艰难的架构决策,弃用一些 API。还有很多其他事情。所以我认为,在提升模型编写和理解代码的能力之上,协作这第二个维度可能是一个重大突破,能让 AI 的好处更广泛地传播。所以这可能是其一。第二点,可能有点琐碎,就是让它在企业中获得采用。我和很多企业聊过,他们还没有升级,仍然停留在 GitHub Copilot v1,因为他们从未完成所有的安全和配置流程,以确保这些智能体能在代码库中正确使用。所以,是的。

It's a good question. The Codex team is incredible. To me, they are the epitome of a really small, talented team singularly focused on a use case and cranking at it at whatever it takes: model, harness, integration, data. So they're really, really good. My read on how the software engineering space is going to evolve: at the moment, the models are really good at generating code and understanding code. They could be better, but they've made meaningful progress. But when I think about software engineers, that's only part of the job, right? The other part of the job is to be on call, another part is to communicate with your teammate, to scope some changes, to make some tough architectural decisions, to deprecate some APIs. There is much more to it. And so I think that on top of improving the model capabilities on writing and understanding code, that second axis of collaboration essentially is probably a major unlock to spread the benefits of AI more broadly. So that's probably one. A second one, maybe trivial, is just for that to gain adoption in the enterprise. When I talk to so many enterprises who have not yet essentially are still stuck on GitHub Copilot v1, because they've never gone through all the security and provisioning process to really make sure that these agents can be used properly in code bases. So yeah.

企业采用智能体编码工具 Enterprise Adoption of Agentic Coding Tools

Host

你觉得企业是否已经准备好接受这些智能体式编码工具了?还是说在安全或合规方面还有三年的障碍?

Do you think we're close to enterprises being like, 'All right, time for these agentic coding tools'? Or does it feel like there are three years of hurdles on security or compliance?

Olivier Godement

我开始看到一批关键数量的企业真正投入进来,配置了成千上万的许可证,让工程师在特定用例上实验和迭代。我的直觉是,2025 年是企业编码之年。你提到 2024 年是编码之年,我认为 2025 年是企业编码之年。我开始看到有意义的采用。

I think I'm starting to see a critical mass of enterprises who are truly leaning in and provisioning thousands, hundreds of thousands of licenses, letting engineers experiment and iterate on specific use cases. My gut is that 2025 is the year of coding in the enterprise. You mentioned 2024 is the year of coding. I think 2025 is the year of coding in the enterprise. I'm starting to see meaningful adoption.

Host

是啊。那 2026 年呢?

Yeah. What's 2026?

Olivier Godement

我不知道啊。也许是 APM,自动项目经理?不知道。我的猜测是,一方面模型更可靠了,能更好地编写和理解代码。另一方面,模型在协作方面变得非常出色,所以你开始拥有一个更多维度的 AI 软件工程师,可以与之合作。

I don't know, man. Maybe an APM, automated PM? I don't know. My bet would be that one, the models are more reliable, essentially to write and understand code better. And second, the models are getting really good at collaboration, so you start to have a more multi-dimensional AI software engineer that you can work with.

Host

协作方面的改进有多少是模型本身变好了,又有多少是围绕它们构建的工具链的功劳?

How much of the improvements in collaboration is just models getting better versus the harnesses you put around them?

Olivier Godement

两者都有。我的评估是,越来越难分清哪些是模型的能力,哪些是工具链的贡献。但我看到,一些最好的智能体是经过训练的,模型是为特定工具链训练的。我认为这就是 Codex 如此出色的原因。所以我越来越把它们看作一种共生关系。

The two. My assessment is that it's getting harder and harder to disentangle what is the model versus the harness. But I see that some of the best agents out there are trained, models are being trained for a specific harness. I think that's why Codex is so good, frankly. And so I think of them more and more as a sort of symbiosis.

Host

我想这又回到了我之前问的问题:初创公司是否会训练自己的模型。我想知道,如果有一套标准的工具链,他们可以利用你们的模型,还是说如果他们有自己的工具链,最终需要针对那个工具链进行非常具体的训练。

I guess it goes back to this question I was asking earlier about whether startups will train their own models. I wonder whether if there's a standard set of harnesses they can leverage your models, or if they have their own harness, they'll ultimately need to train very specifically on that harness.

开源框架与行业演进 Open-sourcing the harness and industry evolution

Olivier Godement

所以我们尽可能做的就是开源 Codex 上的内容。我们在 GitHub 上开源实际代码,开源工具定义,让人们能够充分利用 Codex 的能力。如今,如果你想在 Cursor 或任何其他 IDE 中使用 Codex,基本上都可以做到。所以我认为这就是行业的发展方向。模型提供商将从单纯的模型推理 API 转变为同时提供模型、工具链,甚至一些用户界面。

So that's what we try to do as much as we can, which is we open source the codex on us essentially. We open source the actual code on GitHub, we open source the tool definition for people to be able to fully utilize Codex abilities. And today, if you wanted to use Codex in Cursor or any other IDE, you could essentially. So I think that's how the industry is going to evolve. Model providers are going to move from being just model inference APIs to providing both the model, the harness, maybe some UI.

Host

基本上就是工具链的参考设计,然后其他人可以……

The reference design for the harness basically, and then everyone else can...

Olivier Godement

没错,就像一种标准架构,一个标准蓝图,用来充分发挥模型的能力。回顾过去两年为开发者业务构建的经验,对我来说最大的教训是:你不能只是把新模型丢进 API 里。除非你提供更多的蓝图、文档或特定的工具链,否则人们很难最大化地利用模型的多重能力。模型在某种程度上非常美妙又奇特,除非你有很好的使用指南或大量与模型交互的经验,否则很难大规模地利用它们。

Exactly, like a sort of standard architecture, a standard blueprint essentially to use the model to the best of its capabilities. If I step back, that's probably to me the biggest learning of the past two years of building for developer businesses: hey, you can't just drop new models in an API. It's going to be really hard for people to maximally utilize the multi capabilities unless you give them more of a blueprint, more documentation, or a specific harness. It's hard to discover. Models are so beautiful and weird in some way. Unless you have a really good recipe or a ton of experience interacting with models, it's hard to massively leverage them.

Host

你认为未来普通企业能够直接与模型和这些工具链交互吗?还是说,存在一整类应用,它们本质上是模型能力与终端行业之间的翻译层,而我们处于中间位置?

And do you think that in the future, your average enterprise will be able to directly interact with the models and these harnesses themselves? Or it seems like there's this whole set of applications that are basically translation layers between the capabilities of the models and your end industry, and we sit in between.

Olivier Godement

我的直觉是,企业大多会购买工具链。他们大多会购买解决方案。我认为会有一些例外。如果你谈论的是企业的核心业务用例,我认为确实有理由构建自己的工具链。但如果你是一家零售商,需要运营销售、财务、IT,坦白说,你直接买软件就行了。在某个时候,你为什么要……我认为人们通常低估了以高质量完成任何事情所需的工作量。对于智能体来说,这一点同样成立,甚至更甚。所以,我的判断是,对于大多数用例,购买而非自建。

My gut is enterprises are mostly going to buy harnesses. They're mostly going to buy solutions. I think there will be some exceptions. If you're talking about a use case which is the core business of the enterprise, I think there is some reason frankly to build your own harness. But if you're a retailer and you have to operate sales, finance, IT, frankly you just buy software. At some point, why would you... I think people usually tend to underestimate the amount of effort it takes to do anything at a great level of quality. The same will be true, or even more true, for agents. So yeah, my bet would be buy versus build for most use cases.

Host

是的,看看各个实验室的工具链是会趋同,还是最终看起来截然不同,这将非常有趣。结果就是,你不得不弄清楚你想参与哪个工具链生态系统。

Yeah, it'll be fascinating to see if the harnesses from the labs converge or they actually end up looking quite different, and as a result you kind of have to figure out which harness ecosystem you want to play in.

Olivier Godement

我不知道。这是一个有趣的发散与收敛的游戏。

I don't know. It's a fun game of divergent convergence.

Host

是的。

Yeah.

Olivier Godement

我不知道,也许我太天真了,但我确实认为长期来看收敛会胜出。但在科研中,你必须让百花齐放,然后找出最有前途的那一个,并聚集在它身后。

I don't know, maybe I'm being naive, but I do expect convergence will win in the long term. But in science research, you have to let a thousand flowers bloom to then figure out which one's the most promising and gather behind it.

AI新手企业速查表 Cheat sheet for enterprises new to AI

Host

是的。你显然与许多不同类型的企业合作过。当你接触一个可能对 AI 领域还比较陌生的全新企业时,你是否有一个速查清单,列出他们真正应该做或知道的事情,这些经验来自你最复杂的客户?或者,如果你能从最成熟的企业中提炼一些经验,会是什么?

Yeah. I mean, you obviously work with so many different kinds of enterprises. When you go into a net new enterprise that's maybe a little newer to the game, do you have a cheat sheet of the few things they really should do or know, that you kind of take from your most complex customers? Or if you could distill some lessons from the most sophisticated ones, what would they be?

Olivier Godement

哦,到那时我想我已经与 200 多家企业合作过。我们谈到过 T-Mobile、MGEN、Salesforce、BNY,还有很多数据库公司和其他企业。我的速查清单。有很多很多技巧和窍门。第一个是经典的企业软件问题:如果你的数据一团糟,你将一事无成。我可以给你最强大的编码智能体。如果你不能将该智能体接入正确的代码库、正确的身份权限和数据库,那么这个智能体就很难发挥作用。所以坦白说,很多工作只是解释如何组织你的数据。如果没有 API、没有服务,如何建立正确的服务?如何使用或不使用 MCP?如何验证这些请求?如何记录这些请求?所以有很多数据工程工作要做。这是第一步。第二步是真正向团队解释如何评估模型。我之前谈到过人们进行“感觉评估”,这在某种程度上是必要的,但如果你有一个生产级用例,你必须进行评估。严格的评估对大多数团队来说并非自然而然,除非他们以前做过。我经常告诉他们,我们在 OpenAI 经常谈论训练模型的团队,他们极其重要。我们有一个全职评估模型的精英团队,坦白说他们可能同样重要。所以花大量时间记录黄金数据集、记录流程和标准操作程序,然后构建严格的评估,以确保你在正确的方向上爬山。这是第二步。

Oh, at that point I think I worked with, I don't know, 200 plus enterprises. We talk about T-Mobile, MGEN, Salesforce, BNY, many of them database and many others. My cheat sheet. There are many, many tips and tricks. The first one is a classic like enterprise software, but if your data is a mess, you'll be able to achieve nothing. I could give you the most powerful coding agent. If you're not able to plug that coding agent into the right code bases, the right identity permissions, database, it's going to be really hard for that agent to be useful. So frankly, a lot of the work is just explaining how to structure your data. If there is no API, no services, how do you stand up the right services? How to use MCP or not use MCP? How do you authenticate those requests? How do you log those requests? So there's a lot of data engineering work to do. That would be step one. Step two would be to really explain to teams how to evaluate models. I was talking earlier about people vibe evaluating, which is necessary to some degree, but at some point if you have a production grade use case, you have to evaluate it. Rigorous evaluation doesn't come naturally to most teams unless they've done it before. And I always tell them, we often talk at OpenAI about the teams that train models, and they're extremely important. We have a crack team that evaluates models full-time, and they're probably equally important frankly. So spending a lot of time to document golden sets, document the procedures, the SOPs, and then build rigorous evaluations in order to make sure you are hill climbing in the right direction. That would be number two.

Host

你在评估方面看到的最常见错误是什么?

What's the most common mistake you see on the eval side?

Olivier Godement

评估太少或太多。

Too little or too much evals.

企业AI挑战:隐性知识与评估 Challenges in Enterprise AI: Tacit Knowledge and Evals

Host

就像太少,比如你只做了五个,然后第一个客户来了,完全超出分布范围,你会想,好吧,我学到了什么?我不知道。我的意思是,与企业合作时最大的发现之一是,大部分知识都在人们的脑子里。

Like too little, like you know, you've done only like five, like you set, and you know, first customers comes in and like it's completely out of distribution, you're like okay, like you know, what did I learn? I don't know. I mean, one of my biggest findings like working with enterprises is, um, most of the knowledge is in people's brains.

Host

是啊,通常你进来做客户支持,你肯定以为某个地方有个 JIRA 或者 Confluence 页面,上面写着所有流程,但这种事从来不会发生。你知道,如果你有 20%-30% 的流程被写下来就算幸运了,剩下的就是,哦,你知道,Sarah 和 Mark 对此非常了解,你应该跟他们聊聊。

Yeah, like you usually come in and you're like customer support, I'm sure there is like whatever a J or confence somewhere with like every procedures being written, that never happens. Like you know, you're lucky if you have like 20-30% of the proc being written, the rest is like, oh you know, Sarah and Mark like know really well about it and you know you should talk to them.

Host

所以,构建真正好的评估与其说是把文本转换成评估。

And so as a result, like building really good evals is like not so much about, you know, converting like text like to eval.

Host

是啊,是找到。

Yeah, it's finding.

Host

但关键是找到对的人,找到那些 Sarah 和 John,也就是了解情况的人,所以这是一个相当迭代的过程。你不是一天做完就完事了。你开始做,发布东西,然后发现不太对劲,就去问 John,为什么不行?然后再构建。

But finding the right people, finding the Sarahs and the Johns essentially, well you know, the ones who know about it, and so that's like a fairly iterative process. Like you know, you don't do it like on one day then you're done. Like you start to do it, you ship the thing, you're like okay that thing is not quite working, let me talk to John, you know why is that not? And then you build.

Host

我知道我打断你了。还有第三个吗?关于企业中最常见的问题?

I know I cut you off. Was there a third one, uh, on like the things that are most common on the enterprise?

Olivier Godement

变更管理。

Change management.

Host

是啊。

Yeah.

Olivier Godement

这项技术对我们所有人来说都是新的。也许有时候我自己也会被它的强大和有时的不确定性吓到,所以花时间向团队和客户解释它是如何工作的,这一点我再怎么强调也不为过。

Like that technology is new to all of us. Maybe for me sometimes, like you know, I'm like weirded out by you know how powerful it is and sometimes you know how like unequive it is, and so yeah, taking the time like to explain to teams, to customers, you know how does it work? I cannot emphasize enough, you know, how critical it is.

Sora API的企业用例 Enterprise Use Cases for Sora API

Host

你有没有看到企业用 Sora API 做了一些很酷的事情?我知道它才刚推出不久。

Have you seen enterprises do anything cool with Sora API yet? I know it's only been out for a little bit.

Olivier Godement

是的,实际上。我们在两个行业看到了不少活力。一个是广告/内容生成,人们在创建一些非常个性化的内容,这很有趣。第二个更像是制片厂、制作公司。这个其实很有意思。我学到了很多关于制作好电影需要什么。现在,能够在 30 秒内向别人展示你脑海中的画面和图像,显然对团队头脑风暴非常有用。所以这是一个有趣的用例。但视频生成,坦率地说,我认为我们还处于早期阶段。它很贵,很慢,但我们已经开始看到它将如何彻底改变这些用例和工作的执行方式。所以,很期待接下来会发生什么。

Yes, actually. Um, we're seeing quite a bit of energy in particular in two industries. Um, one is like ads/content generation where people are creating some crazy personalized, like you know, content which is really fun. The second one is like more like the studios, the productions companies. Um, that one is really interesting actually. I learned a lot about what it takes like to build like really good movies. And now, like being able like to just show to someone in like 30 seconds what you have in mind and you know the sort of the picture, the imagery that you have in mind is apparently, like you know, being like quite useful, like you know, for teams like to brainstorm like way more. So yeah, that's a fun use case. But yeah, video generation I think we're still like in the early innings, frankly, of video generation. It's expensive, it's slow, and you know, but we can start to see how it's going to totally, like you know, transform how that sort of use cases, like jobs are being performed. And so, yeah, quite excited to see what's happen next.

快问快答:AI中过度炒作与低估 Quickfire: Overhyped and Underhyped in AI

Host

太棒了。我总喜欢在采访结束时问一组标准的快速问答,听听你对一些宽泛问题的看法。那么首先,你认为当今 AI 世界中什么被过度炒作,什么被低估了?

That's awesome. Well, I always like to end my interviews with a standard set of quickfire questions where we get your thoughts on some overly broad questions that we stuff into the end. Uh, and so maybe to start, what's one thing that you think's overhyped and underhyped in the AI world today?

Olivier Godement

哦,糟糕。我没准备这个问题。嗯,过度炒作,低估,低估。我老是回到科学、药物设计、药物发现。我觉得我们谈论得比较少,因为对我们这些在软件行业工作了一段时间的人来说,这不是一个很自然的用例。但归根结底,如果你看历史轨迹,那是进步的基础。完全正确。所以如果我们能够哪怕只加速 5% 的发现速度,对经济和技术其他部分的影响都是巨大的。

Oh, shoot. I did not prepare for that one. Uh, overhyped, underhyped, underhyped. I keep coming back to science, drug design, drug discovery. I think we tend to talk less about it because for us who've been working in software like, you know, for a while, doesn't come super naturally as a use case. But at the end of the day, like you know, if you look at like the arc of history, like that's like the substrate of like progress. Totally. And so if we're able to accelerate if only by like 5% like the rate of discovery, like the sort of the implications on like the rest of the economy and like technology are just like you know humongous.

Host

所以,是的,让模型非常擅长科学领域。

And so yes, making models um harnesses extremely good for like you know scientific.

Host

你必须建一个实验室来做这个,我的意思是,要真正构建一个生物模型,你需要某种实验反馈循环,对吧?

You have to build a lab to do that, I mean it feels like, uh, to really build a biomodel you need some sort of feedback loop of experiments, right?

Olivier Godement

是的,你需要实验室、数据,还需要处于 LLM 和各自领域交叉点的人才,所以这很难。但这可能是复利最高的收益。

Yes, you need a lab, you need data, like you need people, like you know, who are like at the intersection of like you know LLMs and like you know their field, so you know it's hard. But probably like you know the highest like compounding, uh, like you know benefits.

观念转变:模型非万能,框架与数据重要 Changed Mind: Model is Not Everything, Harnesses and Data Matter

Host

在过去一年里,你对 AI 的什么看法发生了改变?

What's one thing you've changed your mind on in AI in the last year?

Olivier Godement

模型就是一切。有几个原因。第一,就像我们讨论过的,工具层也变得极其重要,而且越来越难将两者分开。但我确实期望那些极其强大的智能体将拥有极其优秀的工具层,其进化速度甚至比模型本身还快。

The model is everything. A couple of reasons. Number one, just like we discussed, the harnesses are getting extremely important as well, and you know it's getting harder and harder like to detangle the two. But I do expect like those you know extremely powerful agents will have like you know extremely good harnesses which are evolving like you know even faster than the models are. Yeah.

Olivier Godement

嗯,这很重要。第三件事是,在企业中,模型的目标是获取越来越多的数据来产生输出,所以如果没有高质量的数据,输出就不会那么好。因此,我认为对于 AI 在企业中的广泛采用,同样重要的是有一套标准的基础设施框架,能在正确的时间向模型提供正确的数据。我认为一旦我们行业解决了这个问题,企业采用率可能会大幅提升。

Um, that's important. The third thing is like in the enterprise, like the whole like goal of like the model is like to get more and more data like to perform outputs, and so if you don't get like high quality data, like you know the more output will not be that good. And so I think equally important for like AI to be adopted in the enterprise widely is you know a set of like standard infrastructure framework to present the right data at the right time to the model. And so I think once we crack that as industry, yeah, like we're probably going to see enterprise adoption like you know scale lock it.

Host

是啊。你认为工具层的大部分进步会发生在模型公司内部,因为它们离模型最近,还是会在创业公司中发生?

Yeah. Do you think that most of the advances in harnesses will happen within the model companies because they're as close to the models that are posted on these things, or do you think will happen in the startup world?

Olivier Godement

我认为在创业公司中也会有很多采用。由于工具层的开源以及提供了最佳利用模型的配方,人们会进行创新。他们会发现模型的一些特性,从而调整工具层定义以获得更多结果。所以,我确实期待相当多的创新。

I think we'll see a bunch of adoption in startup world as well. I think as a result of like open sourcing the harnesses and sort of giving like the recipe essentially to best utilize the models, like people are going to innovate. They're going to see some quirks of the model, like some way essentially to tweak the harness definition to get like more results out of it. And so yeah, I know I do expect a fair bit of innovation.

反思OpenAI的错误 Reflecting on OpenAI's Mistakes

Host

从外部看,感觉 OpenAI 似乎总是一帆风顺。回顾过去两年半,有没有什么事情让你觉得,“哦,那件事我们搞错了”?

From the outside, it feels like everything just always goes right at OpenAI. Anything looking back on your last two and a half years that you're like, "Oh, we got that thing pretty wrong."

Olivier Godement

哦,天哪,从何说起呢?我认为世界对 OpenAI 相当宽容,因为我们在很多事情上都是先行者,所以人们对先驱者有更高的容忍度或灵活性。我们做错了什么?嗯,我们发布了很多产品功能和工具,它们要么没有找到产品市场契合点,要么没有充分利用模型的能力。

Oh man, where do I start? Uh, I think the world has been pretty forgiving to OpenAI because we are first on many things, and so I think people like you know have like a higher like sort of a tolerance or you know flexibility, like you know, to um, to sort of pioneers. What have we gotten wrong? Uh, I mean there are plenty of like product features, like tools that we ship that you know did not like you know find product-market fit, or like you know did not like you know utilize like the models as well as we could.

反思过往失败与实验 Reflections on past failures and experiments

Olivier Godement

我们在 2023 年对智能体有过一些定义,可能太早了,没能被广泛采用,因此我们不得不淘汰了一些 API。我们在各种音频技术上投入了很多,但老实说,效果并不理想。所以,我认为归根结底,历史记住的是成功,但与此同时也有相当多的失败或实验没有成功。

We had some definitions of what an agent was like in 2023, which was probably too early and didn't catch adoption, so we had to sunset some APIs as a result. We invested a bunch in different kinds of audio technologies that did not really pan out as well, frankly. So yeah, I think at the end of the day, history remembers the successes, but there's a fair amount of failures or experiments that didn't pan out in the meantime.

Host

嗯。

Yeah.

对Gemini 3的看法 Thoughts on Gemini 3

Host

你怎么看 Gemini 3?

What do you think of Gemini 3?

Olivier Godement

我还没用过。基准测试看起来很好。我看到一些视觉示例,效果非常强。所以谷歌团队似乎真的打造了一个很棒的模型。但我等不及要亲自测试了。

I haven't played with it yet. The benchmarks look really good. I saw a couple example visual examples that look really strong. So it seems like the Google team really cooked a great model. But yeah, I cannot wait to actually test it myself.

Host

嗯。你通常怎么测试它?你会做什么?

Yeah. What do you do to test it? Like what are you going to do?

Olivier Godement

我自己会做两三种不同的测试。第一种通常是风格、语气,本质上就是个性。我有一些私人问题,喜欢通过向模型输入大量上下文来提问。这是第一种。第二种是经典的从零开始生成一个应用,然后看看前端,稍微测试一下。第三种是我喜欢测试非常长程的能力,比如输入几十万个 token,然后尝试问一个关于倒数第二个 token 的非常难的问题。

I have two or three different kinds of tests I do myself. Number one is usually on style, tone, personality essentially. I have some personal questions which I like to ask by dumping a lot of context to models. That's one. Number two is the classic generating an app from scratch and looking at the front end, testing it a bit. Number three is I love to test very long horizon capabilities, like dumping literally hundreds of thousands of tokens in there and trying to ask a very hard question on the second to last token.

Host

太棒了。嗯,这真是一次精彩的对话。我想把最后的话留给你。话筒交给你。大家可以去哪里了解更多关于你,或者关于 OpenAI 推出的任何你想推荐的东西?交给你了。

That's awesome. Well, this has been a fascinating conversation. I want to make sure to leave the last word to you. The mic is yours. Where can folks go to learn more about you, about anything that OpenAI has shipped that you want to point folks to? I'll leave it to you.

Olivier Godement

我们经常发推文。我们发推文。也许不够多,但我们应该多发。不过,是的,在 Twitter 上,谢谢。Twitter OpenAI。我也在 Twitter 上。我会回复每个人。任何反馈、任何功能请求,直接 @ 我,我会很高兴联系。

We tweet a lot. We tweet. Maybe not enough, but we should tweet more. But yeah, on Twitter, thank you. Twitter OpenAI. I'm also on Twitter. I respond to everyone. Any feedback, any feature request, just tweet at me and I'll be happy to connect.

Host

太棒了。嗯,非常感谢。这非常有趣。

Amazing. Well, thanks so much. This was a ton of fun.

Olivier Godement

非常感谢。

Thanks so much.

互动版:逐字朗读 + 针对本期提问 →