从超级应用到无形界面:AI 代理的未来

From Super App to Invisible Interface: The Future of AI Agents

格雷格·布罗克曼 Greg Brockman · Alex Kantrowitz · 2026-07-01 · 约 45 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

OpenAI 的 Greg 讨论从对话式 AI 到主动代理的演变,这些代理能代表你完成任务,最终界面将消失。

OpenAI's Greg discusses the evolution from conversational AI to proactive agents that accomplish tasks on your behalf, with the interface eventually disappearing.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 15)

全文 · Full transcript(中英对照)

超级应用 vs AGI 愿景 Super app vs AGI vision

Host

你知道吗,Greg,这是我们第四次对话,每次我们都聊 OpenAI 的产品方向,我想我开始理解了。之前有人说,超级应用这个词不适合你们正在做的产品——把 Codex(OpenAI 的编程产品)、浏览器和 ChatGPT 整合在一起。当你用超级应用这个词时,人们会说:‘不,超级应用是那种你可以在里面使用其他所有应用的东西。’但现在我们看到这些产品整合在一起,实际上超级应用可能是个正确的词。至少对我们外人来说,我们开始看到,当你需要做任何事情时,都会从 ChatGPT 的提示开始,然后 OpenAI 的技术会用你的浏览器或电脑帮你完成。这样理解对吗?

You know, Greg, this is our fourth time speaking and we've spoken every time about OpenAI's product direction and I think I'm starting to get it. There was this conversation that a super app was the wrong term for what you were doing with the app you're building, bringing Codex, which is the coding side of OpenAI's product, browser, and ChatGPT together. And when you used the word super app, people would say, 'No, a super app is actually something that you can just use every other app within.' And now as we've seen these products come together, actually super app might be the correct term. At least for us on the outside, we're starting to see that when you need to do anything, it will start with a prompt in ChatGPT and then OpenAI's technology will use either your browser or your computer to get that done for you. Is that the right way to think about it?

Greg

我认为这是一个很好的视角。真正放大来看,我们实际上要构建的是 AGI。想想人们自从 ChatGPT 以来一直在用的,那是一个语言模型。这两者之间有很大的差距。你能和它对话,它也能回应你,这很神奇,很棒。但我们在 2022 年推出时,它没有记忆,没有连接任何工具,没有上下文。所以对话智能只是人们完成工作、实现目标所需的一部分。我们的方向是拥有一个真正为你着想的 AI,你可以提供目标和方向,它会不断思考今天能为你做什么。它可以解决超级困难的问题,也能处理非常平凡的问题。你醒来时收件箱已经整理好,但如果你在考虑一个健康计划,它也能帮你实现,找出医疗方案,或者至少提供这类信息。我们花了很多时间思考你想要什么样的界面,答案是:你几乎不想要界面,不想要产品。你希望它像你我之间的界面一样,只需与一个持久的实体对话,它就能为你完成目标。构建这样的东西很难,需要时间,但我们有很多组件。我们正在越来越多地整合产品层,努力让模型更好,让整个系统减少点击按钮、切换开关和改变模式。不是说过程中不会有这些,但长期趋势是简化和统一。

I think that's a pretty good perspective. And to really zoom out, the thing we're actually trying to build is AGI. If you think about what people have been using since ChatGPT, it's a language model. There's a big gap between these. It's amazing you can talk to it and it talks back. Great. Wonderful. But when we launched in 2022, there was no memory, it's not hooked up to any tools, has no context. So conversational intelligence is only one part of what people really need to get work done and achieve their goals. Where we're going is to have an AI that's really looking out for you, that you can provide goals and directions, and it's constantly thinking about what it can do for you today. It can solve super hard problems and very mundane problems. You wake up and your inbox is organized, but also if there's a health plan you're thinking about, it can help you achieve that, figure out medical treatments, or provide you with that kind of information. The question of what interface you want is what we spend a lot of time thinking about, and the answer is you want almost no interface, no product. You want this to be like the interface between you and me, just being able to talk to a persistent entity that can accomplish goals for you. Building that is hard and will take time, but we have a lot of the pieces. We're increasingly bringing together the product layer, trying to make the models better, trying to make the whole system so there's less clicking buttons and toggles and changing modes. Not to say there won't be some of those along the way, but the long-term trajectory is towards simplification and unification.

从对话到行动 From conversation to action

Host

你说界面会消失,这很有意思。再深入一点,我们很多使用 ChatGPT 这类产品的人会看到,机器人在最后会提出建议。你问营养问题,它会说‘要不要我帮你制定一个健康计划?’你问旅行,它会给你一个日程。我理解得对吗?在 ChatGPT 中,你谈论健康决策,它可能会说‘你可能需要去看这个专家,让我帮你预约’,然后它真的会代表你采取行动?所以它从单纯的对话界面,变成了真正理解你的意图,然后出去为你完成它。

It's very interesting that you say the interface will melt away. To go a little deeper, many of us who use products like ChatGPT today will see that the bot makes a suggestion at the end. You ask about nutrition and it says, 'Should I make a health plan for you?' You ask about travel and it gives you an agenda. Am I hearing you right that what's going to happen within ChatGPT is you talk to it about your health decisions and it might say, 'You probably need to go to this specialist, let me make an appointment for you,' and then it will actually take that action on your behalf? So it goes from simply a conversation interface to actually understanding your intent and then going out and accomplishing that for you.

Greg

完全正确。如果你用过 Codex——顺便问一下,在座有多少人用过 Codex?(观众回应)不少。我们的目标是把 Codex 的力量带给每个人,把智能体带给每个人。这项技术现在已经存在。你可以把 Codex 连接到 Slack、我的 Gmail、我的日历。OpenAI 内部有很多非技术用户——名字里有‘代码’,但它其实不是关于代码的。它是关于拥有一个使用智能体的通用工具。例如,我们通讯团队的一位同事在组织活动时,它就会询问所有活动参与者的饮食偏好,设置整个座位表,完成所有工作,这样她就可以专注于她想做的部分。我们会在各个领域看到这种情况。认为 AI 连接到这些工具不再是科幻。我记得我们第一次在 ChatGPT 中尝试工具使用是在 2023 年,大概是三月或四月,我们发布了插件。有人记得插件吗?(观众回应)它完全没成功,因为模型还没准备好。形式是对的——显然你会有一个能和你 Gmail 对话的 AI。但我们一次只能向模型暴露三个连接器,否则它就会开始遗忘。我们只有大约 2K 或 4K token 的上下文。根本没有记忆。这就像 60 年代或 70 年代的早期计算机,内存很小。今天你的手机比那个时代的任何超级计算机都强。这就是这些模型的发展方向。改进的速度非常快。现在你可以访问数百种不同的工具,我们有能力将它们连接到整个文件系统。你几乎可以让模型拥有互联网和几乎所有应用的完整能力。而且它很聪明——它有 5200 万 token 的上下文,取决于你怎么看。能力水平也变得越来越强大。这些模型现在正在解决未解的数学问题和物理问题,真正帮助人们实现他们原本无法做到的事情。我们正处于智能体时代的边缘,它将彻底改变我们的运作方式,无论是在软件工程、金融、法律、销售,还是在我们的个人生活中。

That's exactly right. And if you've used Codex—by the way, how many people in the room have used Codex? (audience responds) A decent number. Our goal is to really bring the power of Codex to everyone, to bring agents to everyone. That technology exists right now. You can hook up Codex to Slack, to my Gmail, to my calendar. There are many people within OpenAI, non-technical users—it's got 'code' in the name but it's not really about code. It's about having this general-purpose tool using an agent. For example, someone on our comms team was organizing an event and it would just ask all the event attendees for their dietary preferences, set up a whole seating chart, did all that work so she could focus on the parts she wanted to. We're going to see this across the board. It's not sci-fi anymore to think about an AI hooked up to these tools. I remember our very first attempt at tool use in ChatGPT was 2023, around March or April, we released plugins. Do people remember plugins? (audience responds) It didn't work at all because the models weren't ready. The form factor is correct—obviously you're going to have an AI that can talk to your Gmail. But we could only have three different connectors exposed to the model at a time or it would start forgetting. We had like 2K maybe 4K token context. There's just no memory. It's like early computers in the 60s or 70s with tiny memory banks. Today your phone is better than any supercomputer from that era. That's where we're going with these models. The rate of improvement has been so steep. Now you can have hundreds of different tools accessible, we have the ability to hook them up to whole file systems. You can almost have the full power of the internet and almost any application at the model's fingertips. And it's smart—it's got 52 million token context, depends how you look at it. The capability level is also getting so powerful. These models are now solving unsolved math problems and physics problems, really helping people achieve things they couldn't otherwise. We are on the edge of this era of agents really transforming how we all operate, whether in software engineering, finance, legal, sales, and in our personal lives too.

智能体 AI 与信任 Agentic AI and Trust

Host

那么,我们来详细说说你刚才举的例子:你的一个同事正在和 ChatGPT 聊一个活动,然后它建议说,‘嘿,我们该怎么联系活动参与者?’ 然后,它不会说‘好的,我得去做这件事’然后打开活动程序,而是界面会从那里接管,一旦它说这是个好主意并且你同意,它就会接入你正在使用的任何工具,然后为你完成这件事。

So just to unpack that example that you were giving there: one of your colleagues is chatting with ChatGPT about an event and then suggests, 'Hey, how should we contact event attendees about something?' And instead of saying, 'Okay, I have to do that' and going into an event program, basically what happens is the interface will take over from there once it says it's a good idea and you agree, and then hook into whatever tools you're using and do it for you.

Greg

完全正确。所以它使用 Gmail 连接器,搜索你的收件箱,找到所有参加活动的人,然后如果你在查看饮食限制之类的,它会发现,‘哦,这些人的饮食限制我已经有了,这些人的我还没有。’ 它会根据你的具体设置起草一封邮件。它可能会说,‘嘿,我起草了这些邮件,我能发送吗?’ 如果你有一个甚至不允许它发送邮件的连接器,它会说,‘我已经起草好了,你需要发送它们。’ 而在另一种情况下,你也可以想象,你已经与系统建立了足够的信任,它会说,‘我已经起草了邮件,并且实际上已经发送了。’ 我认为这实际上指出了智能体时代一个非常重要的方面,那就是信任,对吧?我们需要真正学会如何与这些系统建立信任,了解它们擅长什么,不擅长什么。弄清楚你想把什么委托给它们,以及你希望如何赋予它们责任。而我们认为这是需要赢得的,对吧?这不是我们可以直接给予的,而是通过向操作者——也就是这个 AI 所代表的人——提供大量的工具、控制、监督和监管,我们认为这将是一件非常重要的事情,所以这是一个关键的产品特性和差异化因素。

Exactly. So it uses its Gmail connector, searches through your inbox to find all the people who are attending, and then if you're on the whatever, dietary restrictions, it sees, 'Oh, these people I already have their dietary restrictions, these people I do not.' It drafts an email depending on exactly how you have things set up. It might say, 'Hey, I drafted these emails, can I send them?' If you have a connector that doesn't even let it send emails, it says, 'I drafted it, you need to send them.' And in a different world, you could also imagine that you've built enough trust with the system where it says, 'I drafted the emails and I actually sent them.' And I think that this actually points to a really important aspect of the agentic era, which is trust, right? That we need to really learn how to build trust with these systems, where they're good, where they're not. Figure out what you want to delegate to them and how you want to entrust them with responsibility. And that's something we view as earned, right? It's not something that we can grant, but by providing lots of tools and control and oversight and supervision to the operator, to the person who this AI is operating on behalf of, we think that is going to be such an important thing, and so that's a key product feature and differentiator.

早期尝试与 Codex 应用 Early Attempts and the Codex App

Host

是的。当你回顾早期的一些尝试时,OpenAI 曾有过让你在 ChatGPT 内叫 Uber 的功能。这延续了一长串试图让你在聊天中采取行动的公司的做法,但从未真正流行起来。而这里的区别可能在于,聊天机器人可以控制你的浏览器或控制你的电脑,然后你就不必担心‘这个插件能用吗?’它通过接管你的机器来为你完成这件事。所以我想知道,你是否预期会与当今的用户界面(也就是所有其他应用、所有软件)发生冲突?因为要真正有用,ChatGPT 必须不被阻止,才能代表用户去执行这些操作。

Yeah. And when you go back to some of the early attempts at this, there was this move that OpenAI had to let you call an Uber within ChatGPT. And it followed a long line of companies that have tried to get you to take action within chat, but it never really took off. And the difference here might be that the chatbot can take control of your browser or take control of your computer, and then you don't necessarily have to worry about, 'Is this plug-in going to work?' It goes and accomplishes that for you by taking over your machine. So I wonder if you expect a fight from the user interfaces that we have today, aka all the other apps, all the software, where to be truly useful, ChatGPT will have to not be blocked to be able to go out and execute these actions on behalf of a user.

Greg

嗯,首先,我要说的是,这已经不是理论上的了,对吧?人们已经在使用 Codex 了。所以,它是一个独立的产品,独立的应用程序。你需要单独安装它。它确实开始专注于软件工程。但在 Codex 上发生的非软件工作数量正在绝对爆炸式增长,对吧?它是一条令人难以置信的指数曲线,正是你所期望的那样。在 OpenAI 内部,我们现在在使用率上基本上达到了与 Slack 相同的渗透率,对吧?就像 OpenAI 的每个人都在一个完全基于 Slack 的公司里。我们基本上不使用电子邮件。真的,如果你不在 Slack 上,你就无法做任何工作。现在 Codex 应用也有这种感觉。而且每个人的 Codex 都连接到了所有这些工具。生态系统如何演变,我认为这将是一个非常微妙的事情,因为我认为非常重要的一点是,我们相信应该有一个充满活力和繁荣的生态系统,人们可以真正构建并看到好处。所以我们实际上从合作伙伴公司那里看到了这一点,你知道,我记得有几个不同的合作伙伴,我们说,‘嘿,我们真的想训练我们的 AI 非常擅长使用你们的软件,’我们不知道他们会说什么,实际上我们得到的回应是,‘这是我们收到过的最友好的合作伙伴外联,’对吧?这个想法是,你会让你的 AI 特别擅长使用我们的工具,他们只是看到了机会,因为他们的工具会因此被使用得更多,而且每个人都在思考如何不仅作为一家公司在 AI 时代生存,而且如何蓬勃发展?比如,你如何真正利用即将有更多活动这一事实的优势,如果你不让 AI 参与,如果你把它拒之门外,那么你实际上会衰落,而不是繁荣。

Well, look, first of all, I'd say that this is not theoretical at this point, right? That people have been using Codex. So, it's a separate product, separate app. You have to install it separately. Really starting to focus on software engineering. But the amount of non-software work that has been happening on Codex has been absolutely exploding, right? It's been this incredible exponential curve, exactly the thing that you would expect. And within OpenAI, we basically have the same level of penetration now in usage as Slack, right? It's like everyone at OpenAI is in an entirely Slack-based company. We do not use email for the most part. It's like really, if you're not on Slack, you're not going to do any work. And it's kind of feeling that way now with the Codex app as well. And that everyone's Codex is hooked up to all of these tools. How the ecosystem evolves, I think it's going to be a very nuanced thing because I think one thing that is very important is that we believe that there should be an ecosystem that gets to be vibrant and thriving and that people can really build and see the benefits. And so we've actually seen this from partner companies where, you know, I remember there's a couple different partners where we said, 'Hey, we really want to train our AI to be really good at using your software,' and we didn't know what they would say, and actually the response we got is, 'This is the most partner-friendly outreach we've ever had,' right? That the idea that you will make your AI specifically good at using our tool, and they just see the opportunity because their tool will be used just so much more as a result, and that everyone is trying to think about how do they not just survive as a company into the AI era, but thrive? Like how do you really get the advantages of the fact there's going to be so much more activity, and if you don't have AI in there, if you shut it out, then you're actually going to be declining, not thriving.

ChatGPT 作为操作系统 ChatGPT as an Operating System

Host

没错。这实际上让 OpenAI……首先,你谈到人们使用 Codex。所以你的一个同事分享过,我想你也谈到过,你们已经把 ChatGPT 带入了 Codex,这样你就可以把 Codex 带入 ChatGPT,这基本上意味着,如果我们作为 ChatGPT 的用户,我们刚才谈到的体验——ChatGPT 不仅建议你下一步可能想做什么,而且会为你去做——将会发生。所以这实际上让它成为了一个操作系统,你不觉得吗?但不是像 iOS 那样的操作系统,你会打开手机然后点击不同的应用。这几乎就像是所有与应用的交互都将通过这个界面发生。这是你们的雄心吗?

Right. This kind of makes OpenAI... So first of all, you talk about people using Codex. So one of your colleagues shared, and I think you've talked about this too, that you've brought ChatGPT into Codex, so you can bring Codex into ChatGPT, which is basically like if we're users of ChatGPT, this experience that we talked about of ChatGPT not only suggesting what you might want to do next but going to do it for you, that's going to happen. And so it makes you effectively an operating system, don't you think? But not the operating system like an iOS where you would go open up your phone and then tap different apps. It's almost as if all interaction with all apps will happen through this interface. Is that the ambition?

Greg

我认为你可以这样描述,但我的想法略有不同。我思考这个问题的方式是,什么是 AGI 的理想界面,或者我们称之为个人 AGI。我认为它再次是我们现在正在使用的同一个界面。你只是想和一个助手交谈,对吧?你想和一个能够代表你去工作、去操作的东西交谈。所以是的,那个智能体,那个 AGI,那个 AI 将拥有自己的计算机,对吧?它将有自己的访问权限。也许它可以,你知道,就像一个理想的同事,他们也可以过来在你的电脑上打字。所以一些访问权限,一些对你的系统的委托访问权限。而且你知道,也许你有时会委托访问你的收件箱,也许它有自己的收件箱,通过某种窗口了解它需要的东西,你把邮件转发给它。这些实际上,如果你仔细想想,并不是没有先例的,对吧?就像你与一个人类助手合作的方式,我们实际上,或者任何同事,我们已经花了很多时间思考如何建立这些信任边界,并确保你们能够一起工作。所以我把它看作是一个不同的东西。它不是,你可以把它看作是一个操作系统,但操作系统几乎是来自不同时代的东西,对吧?它是堆栈的不同层。这实际上更多的是关于你如何与技术进行广泛的交互。

I think that you could describe it that way, but I think of it a little differently. Like the way that I think about this is that what is the ideal interface to an AGI, or we call it kind of a personal AGI. And I think that it's again the same interface that you and I are using right now. You just want to talk to an assistant, right? You want to talk to something that can go and work and operate on your behalf. And so that yes, like that agent, that AGI, that AI will have its own computer, right? It'll have its own access to things. Maybe it can, you know, like an ideal coworker would be they can come over and type things on your computer too. So some access, some delegated access to your own system. And you know, maybe you delegate access to your inbox sometimes, maybe it has its own inbox with some sort of window into the things that it needs, you forward emails to it. These are not actually, if you think about this, it's not unprecedented, right? It's like the way that you work with an assistant who's a person, that we've actually, or any coworker really, we've spent a lot of time really thinking about how do you build these trust boundaries and make sure that you're able to operate together. And so I think of it as just a different thing. It's not, you could think of it as an operating system, but an operating system is almost something from a different time, right? It's a different layer of the stack. This is really more about how do you interface with technology broadly.

AI 作为个人智能与智能体时代 AI as personal intelligence and agentic era

Greg

我认为人工智能的美妙之处在于,它真正让机器更贴近人类,而不是让我们扭曲自己去适应文件、文件夹这些不自然的细节,对吧?这些细节更多是关于机器如何运作,而不是我们如何运作。

And I think that the beautiful thing about AI is it's really about bringing the machine closer to the human rather than us having to contort ourselves into like files and folders and like all these details that somehow are not natural, right? That are more about how the machine operates rather than how we operate.

Greg

是的。说到个人智能,呃,你上周看 WWDC 了吗?

Yeah. Talking about a personal intelligence, it sort of um I don't did you watch WWDC last week?

Host

呃,不,不,我错过了。我被禁了,但呃我在电视上看了。呃,拜托,苹果。总之,看起来你和 Siri,新的 Siri,会有点竞争,对吧?因为他们是一个应用,或者说一个智能,会位于你所有应用之上,让你采取行动。而 ChatGPT 将是 iPhone 上的一个应用。那么,谈谈这种定位对 OpenAI 来说是否会困难,以及你如何从战略上考虑这个问题。

Uh no, no, I missed it. I was I was banned, but um I watched it on TV. Um come on, Apple. Anyway, it does look like you and and and Siri, the new Siri are going to come kind of into competition, right? Because they're an app that's going to sit or an intelligence that will sit on top of all of your apps and let you take action. And ChatGPT will be an app on the iPhone. So then talk a little bit about whether that positioning is going to be difficult for OpenAI and how you're thinking about that strategically.

Greg

嗯,我只是再次以不同的方式思考。我认为我们正处于这个新的智能体式时代的开端,而 AI 的发展方式一直是:当你拥有新水平的能力时,就意味着你有机会重新思考一切,对吧?重新思考人们如何交互,技术能做什么。我认为这也不例外,对吧?在我看来,比如我在地平线上看到的东西,例如 AI 解决科学问题,对吧?我认为我们开始看到端倪。比如,今天我们在同行评审文献中宣布,有医生正在使用 o3。还记得 o3 吗?

Well, I just think again think of it a little differently. Like I think that we're in the beginning of this new agentic era and the way that this has always gone in AI is that when you have a new level of capability, it means you have an opportunity to rethink everything, right? Rethink how people interface, how like what the tech is capable of. And I think that this is no different, right? In my mind, like the kinds of things that I see on the horizon, for example, AI for solving scientific problems, right? And I think we're starting to see the inklings of this. Like for example, today we announced we have in uh peer-reviewed literature people doctors who are using o3. Remember o3?

Host

是的。

Yep.

Greg

那好像是好久以前的事了,对吧?那是我们最早的推理模型之一,用来为那些多年从医生那里得不到答案的人找到诊断。你知道,有一个例子,有人被神秘疾病困扰了 20 年。最终,通过这项技术得到了诊断。如果你说,好吧,你有能做这个的模型,它们确实能做。然后真正的问题在于,你知道,同样的分发,以及,你知道,你能访问一个应用吗?你知道,对我来说,这不对。就像我们有了根本性的新东西。所以这并不是说不会有竞争。我实际上认为会有竞争,而且对每个人都有好处,但我只是认为你使用这项技术的方式、它能做的事情以及它让你能够做的事情,完全不同于我们以前见过的任何东西。

That was like forever ago now, right? That was like one of our earliest reasoning models using that to find diagnosis for people who had no answers from doctors for many many years. You know, there's an example of someone who had spent 20 years with a mysterious ailment. Finally, it's been diagnosed through the use of this technology. And if you're like, okay, you've got models that can do that, they can do that. And then it's really about like, you know, the same like distribution and like, you know, can you get access to an app? You know, to me it's it doesn't type check. It's like we have something fundamentally new. And so that's not to say that there won't be competition. I actually think that there will be and it's going to be great for everyone, but I just think that the ways in which you're going to use this technology, the things it will be capable of and what it'll make you capable of doing are just totally different from anything we've seen before.

Host

你知道,我本来要问你,嗯,这是否意味着你必须,呃,创造你自己的设备,假设我的概念是,你知道,你必须通过苹果才能接触到用户。嗯,假设这有点道理。但答案是你已经在做了。所以 OpenAI 现在正在开发一个设备。

You know, I was going to ask you um well, does it mean that you'll have to um you know, create your own device assuming that like my concept is is you know that you're going to have to go through Apple to get to the user. Um assuming that's somewhat valid. But the answer is you already are right. You're so OpenAI is working on a device right now.

Greg

这确实被公开报道过。

It certainly has been publicly reported.

Host

我十二月在你的办公室,Sam 告诉我这正在发生。是多个设备。嗯,所以如果你再次考虑你将如何与这些 AI 交互,那个设备或一系列设备如何发挥作用?

I I was in your office in December and Sam told me that this is happening. It's it's multiple devices. Um so if you think about the way that again you're going to interface with um with these AIs, how does that device play in or series of devices?

Greg

嗯,看,我再次退一步说,我认为这是非常新事物的开始,我认为关于界面,最大的转变甚至不是关于设备之类的东西。它实际上是从对话智能(比如聊天范式,你有一个足够个性化的 AI,值得阅读它的输出,对吧?你问一个问题,得到一个答案,这对你有用)转向智能体,它们有足够的能力为你实际做事。这是一个巨大的转变,意味着你交互方式的不同。所以你只会想要一个能访问你上下文的单一智能体。这在个人生活中如此,在商业环境中也是如此,对吧?你想象一下,比如,你有一个在每个领域都有博士学位的同事,你知道,多个诺贝尔奖得主,你雇了一个,雇了一百个,但你不邀请他们参加任何会议。他们不会很有用。所以问题在于如何将上下文输入 AI,不仅是静态的,而且是动态的,随着上下文和业务流程的演变,如何有一个上下文层让 AI 访问,使 AI 能够发挥其原始智能的程度。因此,找到让 AI 在会议中可访问、符合人体工程学、易于使用的方法。嗯,我认为所有这些都需要重新思考,但对我来说,核心是从智能体形态开始,然后倒推如何让它拥有所需的上下文,而信任将是使整个等式成立的核心部分。

Well, look, I I think again I I would just step back and say that I think this is the beginning of something very new and that I think about the way that I think I want to say like I think the biggest shift that has happened in terms of interface again it's not even about devices and and and things like that. It's really about the shift from conversational intelligence like kind of the chat paradigm where it's like kind of you have an AI that's personalized enough to you that it's worth reading its output, right? You ask it a question, you get an answer, it's something that's useful to you to agents where they're capable enough to actually do things for you. Like that is a big shift and that that implies a difference in how you want to interact. And so you kind of are just going to want a single agent that has access to your context. And this will be true in personal life. This will be true in a business context, right? You imagine, for example, having a uh you know, imagine you have a PhD in every field co-worker, you know, Nobel prizes, multiple of them, and you hire one of these, you hire a hundred of them, and you don't invite them to any meetings. They're not going to be very useful. And so there's something about how do you get context into the AI and not just statically but dynamically right as context evolves as your business processes evolve how do you have a context layer that is accessible to an AI that lets the AI operate to the extent of that raw intelligence and so finding ways to make that AI be accessible so available in your meetings to make it very ergonomic it's very easy to get access to Um, I think all of that's going to require a rethink, but I think it again it's just the core for me starts from thinking about the agentic form factor and then working backwards to how do you just make this have the context it needs and again the trust is going to be such a core part of making this whole equation work.

Host

所以有点像随时带着这个设备,然后说“我需要完成这个”,它就替你去做了。我认为这将是其中的一部分。但我甚至认为,如果你没有这样的设备,你也不会出局,对吧?因为这是 AI。并不是说只有一种版本,你认为设备就是 AI,你想要你的手机成为 AI。你想要任何你想到的定制设备成为 AI,但不会是这样。它更像是一个界面。就像你的手机不是你,对吧?它是你的一个界面。它是一种方式,让我可以在需要时随时联系你。呃,每当我想问你一个问题,有不同的访问方式,比如同步的电话通话,我可以给你发短信,发邮件。我认为我们与智能体的交互也会非常相似。呃,有报道说 OpenAI 正在开发双向语音模型,我想我们过去讨论过,目标是拥有一个你可以与之交谈的 AI,它能处理并更自然地回应你。你能分享一些相关信息吗?

So kind of like having this this device with you at all times and being like I need to get that done and it goes and does it for you. And I think that that will be part of it. But I almost even think if you don't have a device like that, it's not like you're going to be out of the game, right? Because it's this a this AI. It's not because there's one thing there's one version of it where you think of it where it's like the device is the AI and you want your phone to be the AI. You want, you know, whatever whatever you know custom device you're you're thinking about to be the AI, but it's not going to be like that. It's going to be more like an interface. Like no more than your phone is you, right? It's an interface to you. It's a way that I can sort of, you know, call you up whenever I need you. uh whenever I want to ask you a question and there's different ways of accessing right there's like synchronous phone call I can text you I can email you and I think that we're going to be much the same with how we interact with our agents uh there's been some reports that open is working on these like birectional voice models I think we've talked about that in the past like the goal is to have like an AI that you can you can speak with and we'll be able to process that and speak back with you in a much more natural way can you share anything about that

Greg

不,但说真的,呃,我的意思是,看,我认为这项技术的总体形态,就像我们已经有语音模型,呃,你知道,一种非常酷的语音体验,已经有一年半、两年了。

No, but no, more seriously, um I mean, look, I think that the that the general shape of the of the technology, like the way that like we had we had we've had voice models, um you know, kind of a really cool voice experience for, you know, year and a half, two years now.

语音交互与自然对话 Voice Interaction and Natural Conversation

Greg

嗯,你知道,我们最早在 2024 年 3 月、4 月演示了它。嗯,大概那年晚些时候推向市场。它的工作方式,以及所有模型的工作方式,基本上是把几个东西串联起来——嗯,最初的做法是串联一个语音转文本模型,然后一个文本转文本模型,再一个文本转语音模型。很糟糕,对吧?这三个东西串在一起。嗯,即使你有一个统一的模型,能够接收输入并输出响应,仍然存在轮流说话的问题,对吧?想象一下,我们不能重叠,不能打断。就像你跟我说完一轮,就得等我完全说完。这不是人类对话的方式。所以我们基本上用了一个取巧的方法,用模型来判断“嗯,这一轮好像结束了”和“嗯,这一轮好像开始了”。我们为什么非要谈“轮次”呢?对吧?轮次太不自然了。这是人类在迁就机器和它的局限。所以显然你想要的是一个更像你我之间交流的 AI 模型,对吧?能够同时处理输入和输出。当然,这个领域的很多人都在朝这个方向努力。嗯,我认为当我们转向这些自然、流畅、像人类一样的对话界面时,会非常令人兴奋。没人见过这样的东西。比如,我想起现在和 ChatGPT 语音的交互。很多方面都很神奇,对吧?很多人在通勤时用它提问,但也很令人沮丧,对吧?每当它打破魔幻感,比如你想补充一点,它却一直打断你,这说不通。所以我认为,我们需要的一部分,AI 的全部意义,就是让你能流畅自然地与之交互。顺便说一句,我认为这不仅关乎个人使用场景,也关乎工作场景。我觉得我用 Codex 最神奇的体验就是通过语音操作。像很多人一样,我们有内置语音,有些人用第三方应用。嗯,当你意识到发一条简短消息给反馈很容易,但写一整段你想要的文字却很糟糕时,你会得到完全不同的体验。没人想那样做,对吧?你只想说出来,想要实时反馈循环,这一切都会发生,而且会很棒。

Um you know, we first demoed it back in March, April of 2024. Um brought it to market, um you know, maybe late that late that year. And the way that it works and the way that everyone's models work is that you basically chain together um well the original way that these things worked was that you would chain together a text to or a speechtoext model then you do a texttoext model and then you would do a texttospech model. Horribleness, right? Like these three things chained together. Um it still has been the case that even if you have one unified model that's able to kind of take in input and then you know able to output a response, you still have this problem of turn taking, right? Imagine that like we have this like you cannot overlap, you cannot interrupt. It's just like once you you you speak to me in a turn and then you got to wait for me to finish my whole response. That is not how human conversation works. And so that we we basically have like a hack where we have these models that determine oh it seems like the turn has ended and oh it seems like the turn has started and we're like why are we talking about turns, right? Like turns again are so unnatural. This is the humans contorting ourselves to the machine and its limitations. And so the obvious thing that you want to accomplish is a model in AI that works much more like you and I do, right? That's able to process input at the same time it's processing output. And all of that is of course something that many people in this field are trying to to run towards. Um I think it's going to be very very exciting as you move to these natural very human fluid like conversational interfaces. No one's seen anything like it. Like one thing that that I think about is the the current interaction with you know chat GBT voice. many ways it's magical, right? So many people use it on their commute, able to ask all these questions, but it also is so frustrating, right? Whenever it breaks the magic because it's like you realize, oh, I want to like add some followup and it keeps talking over you and it didn't it's just like that is just it doesn't make sense. And so I think that part of what we need part of like the whole point of this AI is to be something that you can interact with interact with fluidly and naturally. And by the way, I think it's not just going to be about the sort of use case, like we kind of think about the the personal use case, but it's also really the work use case. And I think some of the most magical experiences that that I've had with Codex have been when operating it through voice. Like many people, we have a voice built in. Some people use third party apps for it. Um, and that you just get a very different experience when you start to realize that like typing a quick message to give some feedback, easy, but like writing out a whole paragraph and everything you want, horrible. No one wants to do that, right? You just want to be like saying things and you want the real-time feedback loop and all of that is going to happen and it's going to be amazing.

Host

那么,我们简单谈谈模型改进。嗯,几年前有讨论说大语言模型即将碰壁。嗯,那是错的。嗯,我在想,我想我们都在想,这些模型能变得多好,改进何时会停止?有什么想法吗?

So, let's talk about model improvement briefly. Um, so there was a discussion a couple years ago that large language models were about to hit a wall. um that was wrong. And um something you know that I'm thinking about is I think we're all thinking about it is how much better can these models get and when will the improvement stop? Any thoughts?

Greg

嗯,我认为在构建这些模型时,你会获得一种从外部难以获得的直觉,因为我们看到了所有数据点,也看到了改进背后的工作。所以答案有两部分。一是,我认为基础科学是我所知、所能想象的最神秘、最重要的科学发现和实证观察之一——我们能够真正构建这些模型,而且缩放定律持续成立,你可以继续用更多数据、更多算力、更好的架构训练这些模型,并且有很多改进。但每当我们遇到“哦,这没有按预期缩放”的情况,其实是我们有问题、有 bug、数学不对、实现与数学不匹配等等。我认为这是需要内化的非常重要的一点。实际上,如果你做过研究,回顾这个领域的开端——神经网络本身是在 20 世纪 40 年代设计的,在计算机之前,作为大脑处理信息的一种模型;第一个硬件实现是 1959 年的感知机。如果你看看这个领域的里程碑式成果,它们遵循一条极其平滑、确定性的路径,即投入更多算力。所以 70 年,也许 80 年来,一直有人说这东西永远不会成功、不会扩展、会碰壁。但还没碰壁。仍然看不到墙。所以我认为基础是允许的。但实践很难,对吧?实际构建这些巨型超级计算机很难、很贵、不容易。我们有团队拼命解决这些极其困难的技术问题。我们不得不设计自己的网络协议。嗯,我们有人检查堆栈的每一层,图表上有奇怪的波动。看待这些神经网络的方式是,没有抽象层,对吧?几乎任何一个小错误都可能产生连锁反应,只在后面显现。所以你需要人深入理解所有这一切。但如果你组建了合适的团队,给人们正确的使命,他们努力奋斗,结果是值得的,对吧?而且是可以实现、有可能的。所以我认为基于这些原因,进步会继续。

Well, I think that this is a place where when you're kind of building these models, you get kind of a sense and an intuition that I think is harder to get from the outside because we see all the data points and we see also the work that goes into these improvements. And so that there's two parts to the answer. One is I think that the fundamental science is one of the most mysterious and important just scientific discoveries and empirical observations that that I can that I that I'm aware of that I can imagine right that we are able to actually build these models and that the scaling laws continue right that it just is the case that you can just keep training these models more data more compute better architectures and there's a lot of improvements that go in. But every time we've kind of run into a like, oh, this isn't quite scaling the way we expect, it's we have a problem, we have a bug, that our math wasn't quite right, that oh, our implementation isn't isn't isn't quite matching the math, whatever the thing is. And that is, I think, a very important thing to to sort of internalize. And actually if you we've done studies where you go back to the beginning of the field right that neural nets themselves were designed in like 1940s right before computers right as a model of maybe this is maybe this how the brain processes information first hardware implementation was 1959 with the perceptron and if you look at landmark results in the field that the landmark results follow this incredibly smooth deterministic path of more compute being poured into them and so 70 years of people maybe 80 years now of people saying this stuff is never going to work, never going to scale, going to hit the wall. Hasn't hit the wall yet. There's still no wall in sight. And so I think that the fundamentals allow it. Now the practicality is hard, right? Actually building these massive supercomputers. It's hard. It's expensive. It's not easy, right? That we have teams that just like work so hard to solve these incredibly hard technical problems. We have our own network protocol that we've had to design. um that we have people who look at every single layer of the stack that there's weird wiggles in the graph and you the way to think about these neural nets is that there's like there's no abstractions, right? It's almost like any little piece that's wrong can have a ripple effect that only shows up down there. And so you need people to deeply understand all of it. And yet if you get the right team together, put the right mission in front of people and people do that grind, the outcome, it's worth it, right? And it's achievable and it's possible. And so I think that for those reasons the progress will continue.

Host

那么,我想听听你的看法,模型能否从今天的位置大幅进步?嗯,假设 OpenAI 构建了最好的模型,相当于拥有 15 个博士学位、极佳情商、不抱怨、能为你做事的东西。嗯,然后下一个模型制造商会构建一个稍差的,但有 13 个博士学位,情商也不错,仍然能为你做事。那么,当达到这种智能水平时,差异化体现在哪里?因为我们看到模型制造商步调一致地前进。一个取得进展,下一个也取得进展。所以它们都变得那么聪明。

So then I'd love to hear your perspective if if models can basically progress much further from where they are today. Um let's say let's say open AI builds the best model and it the equivalent of like something with like 15 PhDs with excellent emotional intelligence that doesn't complain and goes out and does stuff for you. Um, and then the the next model maker will build a a less good, but it has 13 PhDs and it's like pretty good uh, you know, EQ and we'll still go and do things for you. So, where does the differentiation come in when you get to that level of intelligence? Because we've seen the model makers kind of move in lock step. One makes an advance, the next one comes in and makes the advance. So, they all become that smart.

差异化与算力稀缺 Differentiation and Compute Scarcity

Host

有可能实现差异化吗?

Is it possible to differentiate?

Greg

嗯,我认为答案有几个维度。第一,我确实认为存在一种吸引子状态,从商业模式角度看,每个供应商都会把算力卖光。我认为这就是我们正在走向的世界,算力根本不足以满足所有需求。我们正在走向一个算力驱动的经济,每个人都会一直使用这些模型来完成感兴趣的任务。我们就能看到这一点。现在我们在谈论算力限制,使用这些智能体的人数大约在一两千万左右。我们还没有达到全球规模。ChatGPT 有十亿用户,但我们还没有把智能体能力带到那个水平。所以你看这些因素,使用深度与我们未来相比也微不足道。所以我认为我们将处于这样一个世界:即使有不同的供应商、不同的能力水平、开源模型等等,这些新云服务,我认为算力将成为稀缺资源,而且会被充分利用。所以从某种程度上说,这对新进入者来说是个好生意吗?我的回答实际上是肯定的。我认为有一个巨大的市场我们根本无法满足,我们需要更多的能量和动力。但第二点,这也忽略了智能不是一维的事实。如果你仔细看,擅长不同领域是这样的:即使你有很高的原始智能,如果你从未练习过——比如你从未做过演讲——你第一次也不会做得好。还有很多不同的事情:你从未操作过电子表格,你就无法成功完成一些复杂的建模。所以我认为我们一直在内化一点:我们审视不同的行业和领域,必须优先考虑。我们不可能同时在每个领域都出色。当然有很多说法是‘你只要提高通用智能,它就会经历很多这些事情’,但要真正成为领域专家,真正成为那种博士,真正能够推动一个领域的雄心,那是很难的。另外,我还想说,我认为理解成功做到这一点后会发生什么很重要。看看 AlphaGo 的例子。还记得第 37 手吗?那一手改变了人们对围棋的理解,现在下围棋的人比以往任何时候都多。它实际上激励了人们做得更多。我认为我们就会看到这种情况。所以我认为深度永远不会停止。你能在科学上走多深?人们有时会想,‘嘿,我们已经发现了所有物理学,一切都好,我们完成了。’我不认为那是我们未来的样子。我认为我们的未来是每次解开一个谜团,就会解开十个更多。所以我认为未来还有更多事情要做,不同公司之间有巨大的差异化空间。

Well, I think there are several dimensions to the answer. Number one, I do think there's a bit of an attractor state where, just from a business model perspective, every provider sells out all their compute. I think that is just the world we're heading towards, where there just isn't going to be enough compute to serve all the demand. We're heading to this compute-powered economy where everyone's going to be using these models all the time to accomplish tasks of interest. And we just see it. Right now we're talking about compute constraints, and the number of people using these agents is on the order of 10 or 20 million, maybe. We're not at planet scale. ChatGPT has a billion users, but we haven't brought the agentic power there yet. So you're looking at these factors, and the depth of usage is also tiny compared to where we're going. So I think we're just going to be in a world where, even if you have different vendors, different capability levels, open-source models, all these things, these neoclouds, I think compute is just going to be the scarce resource, and it's going to go to use. So to some extent, is this a good business to be in for new entrants? My answer is actually yes. I think there is a huge market that we are just not going to be able to address, and we need much more energy and momentum there. But a second thing is that it also misses the fact that intelligence is not a unidimensional thing. If you really zoom in, being good at different domains is something where, even if you have a lot of raw intelligence, getting good if you've never practiced—like you've never actually done a pitch—you're not going to be good at it your first time. And there's lots of different things: you've never operated a spreadsheet, you're not going to be able to succeed at doing some complex modeling. So I think there is something we have been internalizing: we look across different industries and domains and we have to prioritize. We can't possibly be great at every single area at once. There is definitely a lot of 'hey, you just get the general intelligence up and it'll experience a lot of these things,' but to really become a domain expert, to really be that PhD, and to really be something that can help push forward the ambition of a field, that's hard. And by the way, one thing I also want to say is that I think understanding what happens when you successfully do that is important. You look at something like AlphaGo. Remember move 37, that move that changed people's understanding of the game, and now more people play Go than ever. It actually inspired people to do even more. I think we're just going to see that. So I think the depth is never going to stop. How deep can you go on science? People have sometimes thought, 'Hey, we found out all the physics, it's all good, we're all done.' I don't think that's the future we're signed up for. I think we're signed up for one where every time you unlock one mystery, it unlocks ten more. So I think there's just going to be so much more to do and tons of room for differentiation across different companies.

算力投资与商业可行性 Compute Investment and Business Viability

Host

所以我想我理解你的意思,你的信念是也许每个人都能扩大这些模型的规模,但最终拥有最多算力的公司会赢。而且,几个月前我们聊过,你提到内部有人问你‘我们应该买多少算力?’你说‘全部。’他们说‘不,真的,我们应该买多少?’你说‘不,全部买下。’而 OpenAI 绝对是购买算力的领导者。我是说,我们看到钱在流出。显然,很多钱通过投资进来,现在你已经建立了有客户的业务,但也有很多钱在流出。你有没有想过,‘嘿,也许我们无法偿还所有这些钱,因为这是一个全新的类别?’

So I think I'm reading you right, that your belief is maybe there's a way that everybody can scale up these models, but ultimately the company with the most compute is going to win. And you know, we spoke a couple months ago and you had mentioned that you were asked internally, 'How much compute should we buy?' And you said, 'All of it.' And they said, 'No, really, how much should we buy?' And you said, 'No, buy all of it.' And OpenAI is definitely the leader in buying compute. I mean, we see the money going out. Obviously, a lot of money coming in through investment, and now you've built a business with customers, but there's a lot of money going out. Do you ever wonder, 'Hey, maybe we're not going to be able to pay all this money back because it's a brand new category?'

Greg

嗯,我看待这个问题的方式是基于基本面。你需要真正看到这样一个事实:算力的交付需要多年时间,具体取决于你在做什么。例如,我们已经在自己的芯片项目上投资了好几年。进展非常令人兴奋——我们很快就会有更多消息宣布。但我们能够做到这一点是非常独特的。真正考虑供应链的完全垂直整合。我认为我们正在走向的世界是,再次,世界上没有足够的算力来满足所有需求。我们非常具体地看到了这一点。你看看 ChatGPT 的指数增长,看看我们现在的指数增长,想想我们能够解决的问题。实际上很有趣,我们昨天——其实是两天前——宣布了化学方面的新成果,能够合成新的改进反应。而这一切都没有引起太多关注。我刚才说的:如果你深入一个领域,你真的可以改变它,而我们甚至还没有触及表面。所以思考方式是经济如此巨大。我们非常具体地看到了这一点,从我们自己的增长、人们愿意支付的价格以及整个行业的规模和增长来看。所以我认为我最关心的是如何满足需求,如何真正拥有能够支持人们在经济中想要做的所有工作的东西。我认为这是一件非常庞大的事情。我认为我们任何人都还没有完全内化这一点。

Well, the way I look at it is on the fundamentals. You need to really look at the fact that the way compute goes is that it's multiple years out before compute actually arrives, depending on exactly what you're doing. For example, we've been investing in our own chip program now for multiple years. And super exciting progress—we'll have more to announce actually pretty soon. But the fact that we're able to do that is something very unique. Really think about the full vertical integration of the supply chain. And I think the world we're heading towards is one where, again, there's just not going to be enough compute in the world to satisfy all the demand. And we see this very concretely. You look at the exponential of ChatGPT, look at the exponentials we're on now, think about the problems that we are able to solve. It's actually kind of interesting that we just yesterday announced—it was two days ago—announced a new result in chemistry, being able to synthesize new improved reactions. And all of this is without much attention. The thing I just said: if you go deep in a domain, you can really transform it, and we're not even scratching the surface yet. So the way to think about it is the economy is so massive. And we see it very concretely in terms of our own growth, in terms of what people are willing to pay, and the size and growth of this whole industry. So I think the thing I think about the most is how do we meet the demand and how do you actually have something that can help support all the work that people want to do in the economy? And I think that is such a vast thing. I don't think any of us have internalized it yet.

价格战与市场动态 Price War and Market Dynamics

Host

你可以在节目说明中的链接观看完整纪录片。是的,但恕我直言,一场价格战正在酝酿。至少报道是这么说的。很高兴你能来谈谈。华尔街日报最近有报道称,OpenAI 即将推出的模型可能会有大幅降价。那么,在需求增长且可能需要大量资源来满足的情况下,如果还有降价压力,你怎么算得过来这笔账?

You can watch the full documentary at the link in the show notes. Yeah, but if I may, there is a price war brewing. I mean, at least that's according to the reports. It's great to have you here to talk about it. The Wall Street Journal recently had a report that an upcoming OpenAI model might have significant price cuts. And so again, like how can you know if it requires so much resources to serve this demand and it is growing demand in an environment where there might be price cuts, how do you make that math work?

Greg

嗯,我再次从不同角度看这个问题。回顾我们整个历史,我们实际上一直在提高智能的同时降低价格——对于固定水平的智能而言。而人们不知怎么地,就像日本悖论一样,这种情况一直在发生。所以我认为前沿智能永远会是价格最高的东西,但一年后,那个水平的智能会变得相当普通,而且更容易获得。我认为我们正处在一个人们开始真正思考价值的时代。过去一个季度,也许到现在,人们一直觉得 AI 智能体之类的东西很新奇,需要引入企业,不想落后,想参与未来。而现在人们开始说,好吧,让我们确保这真的能带来投资回报和价值。实际上我认为这是一个很好的状态,因为人们在问正确的问题。我今天参加了一些客户会议,他们正是这么说的。他们问:“我们怎么能有好的支出控制?我们怎么能有可观测性?”而我们今天正好发布了支出控制功能。所以,你看,就是这样。

Well, again, I look at it from a different angle. So if you look at the whole history of what we've done, we actually have been increasing the intelligence cutting price right for a fixed amount of intelligence and people somehow just like the Japanese paradox just keeps happening. And so I think frontier intelligence will always be something that is going to be, you know, it's always going to be the priciest thing but I think that a year from now that level of intelligence is going to feel pretty mundane and like you know going to be much more available. And I think that the world that we're in is one where people are starting to really think about value. And it's actually been a very interesting shift where over the past, you know, first quarter, maybe up until now, people have just been like this AI agent stuff, it's all new. We need to bring it into our enterprise. Like, we don't want to be left behind. How do we be part of this future? And now people are like, okay, like let's make sure this actually delivering ROI and value. And I actually think that's a great place to be, right? because people are asking the right questions. And I hear this, I had some customer meetings today where people were saying exactly this. They were like, "How can we have even just like good spend controls? How can we have observability?" And I think we literally today just released spend controls. So, you know, it's like

Host

好的。

Okay.

Greg

没错。我们确实在大力投资企业就绪性和客户告诉我们他们需要的工具。我认为这是我们公司经历的一个转变:不再只是想着发布模型,而是真正思考业务的端到端。如何将其用于解决真实客户的实际问题,这正在每个行业迅速发生。许多公司仍在思考如何最好地利用这些模型,我们也在同时学习。对我来说,整个游戏还处于早期阶段,市场规模增长如此之快,我们自己的收入增长也如此之快。我认为我们谁都没有预料到这一切会变得多么陡峭。

Exactly. We are really investing hard in enterprise readiness and the tools that our customers are telling us that they need. And I think that for me is the shift that we've also been going through as a company is really not just thinking about hey we're just going to release models and have a model, really thinking about the end to end of the business. How do we bring this into solving real problems for real customers and that is happening so quickly across every single industry and the number of different companies that still feel like they're wrapping their mind around how to best make use of these models we're learning at the same time. I think it's just so early in this whole game to me, the absolute size of the market growing so quickly, our own revenue ramp growing so quickly. I think it's still just like none of us are anticipating how steep that's all going to go.

Host

你们会降价吗?

Are you going to cut prices?

Greg

所以答案总是肯定的,对吧?但关键在于,我认为将会持续发生的是我们会有前沿模型。我不认为短期内会有巨大转变。我认为那种事情不会发生。但你应该预期的是,在一年时间范围内,达到今天感觉非常高级的智能水平会便宜得多。但会有新的东西好得多,你会想,我为什么还要用这个旧的?事情总是这样。

So again, the answer is always yes, right? But it's about like I think that what's going to keep happening is that we're going to have frontier models. I don't think there's going to be like a massive shift in the short term. I don't think that that is the kind of thing that's going to happen. But I think the thing you should anticipate is that over a year-long time horizon to get to today's level of intelligence that feels very premier, it's going to be much cheaper. But there's going to be a new thing that is going to be so much better and you're going to be like, why would I ever use this other one, right? It's just how it's always going to be.

模型商品化与微软竞争 Model Commoditization and Competition with Microsoft

Host

萨提亚·纳德拉最近发了一些有趣的推文并接受了采访。他最近说,“模型正在变成商品,有价值的资产是公司”,或者我可能转述一下,“有价值的资产是一个持续从你的数据中学习的公司特定 AI 系统”。你怎么看?现在和微软竞争是不是很奇怪?

So Satya Nadella has had some interesting tweets and interviews recently. He recently said, "The model is becoming a commodity and the valuable asset is company," or this might be a paraphrase, "The valuable asset is a company specific AI system that continually learns from your data." What do you think about that? And is it weird to be competing with Microsoft now?

Greg

嗯,听着,我不认为堆栈中的任何一层会从价值链中消失。我认为这些东西是相乘的。想想最底层的算力,没有算力就没有 AI。某种程度上你可以说算力已经商品化了,只是浮点运算,谁在乎呢?但实际上,看看今天的芯片股,看看那些卖算力的人,市场对他们的估值,他们看到这里有一个根本性的资产,非常关键。我认为那是因为它是一个收入中心,任何构建 AI 的人都必须依赖它,而且在效率提升和利润率等方面有很多有趣的动态。但根本上,即使你眯着眼说它商品化了,它并没有。价值不会消失,利润率不会消失。市场会奖励它,因为它有根本价值,而且其重要性会随时间增加。你可以从人们为 H100 支付的价格中看到这一点。Hoppers 并没有过时,它们是上一代芯片,在正常情况下,如果不是完全供应受限,没人会买它们。但相反,市场价格相对于之前上涨了。所以这种倒挂正在发生,而且我认为会持续发生,因为每个人都面临海啸般的需求,你会看到价格和利润率在堆栈的各个层面持续增长。

Well, look, I don't think that there's any layer of the stack here that is going to just kind of be removed from the value chain. I think that these things multiplied together. And if you think about the most base layer of compute, that is something where it's just like no compute no AI and to some extent you could say oh compute is commoditized it's just flops who cares about it but in reality like you look at today's chip stocks you look at the people who are selling compute, what the market is valuing people at and they see that there's a fundamental asset here that is just so critical and I think that is because it is a revenue center it is something that anyone who's building AI has to rely on and that there's a bunch of very interesting dynamics in terms of the efficiencies that you can squeeze out and the margins, all these things. But fundamentally, even though it's like you can kind of squint it and say it's commoditized, it's not. It's not that the value goes away. It's not that the margins go away. It's like something that the market will reward because it has fundamental value and that the importance of it's going to go up over time. You can see that with some of the prices that people are paying for H100s, right? Hoppers are, you know, kind of, not obsolete, right? They're a previous gen chip and in any normal situation where you're not totally supply constrained, no one would be buying them. But instead, the market prices are up relative to where they were before. So there's this inversion that's happening and again I think it's going to keep happening where because everyone has this avalanche of demand that you're going to see prices and margins and all of these things continuing to increase at various levels of the stack.

Greg

我认为同样适用于模型本身,它们也不是商品化的,那里有很多竞争,我认为这很好。对企业、客户、消费者都有好处。但我认为有很多领域,比如我们的模型一直是最聪明的,能够解决极其困难的问题。我认为我们刚刚开始进入一个阶段,你会看到由此带来的变革性影响。如果我们真的能通过模型加速科学进步,模型越聪明,速度就越快。这与具有对话界面、能帮你订旅行或管理日历的模型非常不同。所以这也是一个维度,我认为我们会做得很好。但我要说的是,这是一个不同的领域。

I think the same kind of applies for models where the models themselves are also again they're not, there's a lot of competition there and I think that's very good. I think it's good for the enterprise. I think it's good for customers, consumers. But I think that there's a lot of areas where for example our models have always been the sort of smartest ones, right? The ones that are able to solve these incredibly hard problems. I think we're just starting to reach a phase where you're going to see the transformative impact from that, right? It's like if we're really able to speed up science through models, the smarter the model, the faster it's going to go. And it's very very different from a model that has a conversational interface that you're able to, you know, is able to book your travel, right? Or organize your calendar. So that's also a dimension I think we're going to do a very good job in. But I'm just saying it's a different area.

生态协作与竞争 Ecosystem Collaboration and Competition

Host

嗯,然后我认为问题是,如何真正将智能与你的客户连接起来,与真实价值连接起来。你有所有这些在不同领域建立了了不起业务的企业,这是一件大事,而且如果你没有领域专业知识,这不是你就能做到的事情。部分原因在于,你需要考虑受监管的行业,考虑任何有类似情况的领域,比如教育,那里有家长、老师、学生,这些不同的群体需要以非常周到的方式互动。对于所有这些领域,所有这些领域,通过深入该领域并思考工作流程应该如何运作、这些模型应该如何编排,可以创造大量价值。所以,我真的认为有足够多的机会可以分享,我认为我们必须作为一个完整的生态系统共同努力,才能实现我认为这些系统能够带来的那种价值。

Um and then I think that the question of well how do you actually connect the intelligence to your own customers right to real value to you have all these enterprises that have built incredible businesses in different domains and it's a huge thing and it's not something where if you don't have domain expertise that you're just going to be able to do right and part of it is that you need you think about regulated industries you think about any area where there's like you know think about education where you have a parent you have a teacher you have student you have these different parties that need to interact in very thoughtful ways for all of these areas, all of these domains that there's a lot of value to be built by being in that area and thinking about how the workflow should work, how these models should be orchestrated. And so I I really think that there's more than enough to go around and I think that we have to work together as a whole ecosystem in order to deliver the kind of value that I think is possible from these systems.

Host

好的,再回到 SA 的观点一次,呃,他把模型称为商品。他试图建立自己的前沿智能。他告诉你的潜在客户,嘿,你们必须来和我们合作,因为我们将帮助构建这些从你们的数据中学习的循环。嗯,他有权访问你的知识产权,我想,直到 2032 年。那么,听到来自 Saudia 的这番话,你感觉如何?

Okay, just to go back to the SA point one more time, uh he's called models a commodity. He's trying to build his own frontier intelligence. He's telling potentially your customers, hey, you got to come work with us because we're going to help build these loops that will learn from your data. Um, he's got access to your IP, I think, till 2032. So, how does it make you feel to hear this coming from Saudia?

Greg

听着,我认为现在最重要的事情是 AI 在经济中的应用,真正地转变经济并提升每个人。所以,我认为这是我真正关注的事情。越多的人努力实现这一点,我认为对每个人都越好。

Look, I think that the most important thing that is happening right now is the usage of AI in the economy to really transform the economy and to uplift everyone. And so, I think that that is something that I'm really focused on. And the more that people are trying to make that happen, I think that that's better for everyone.

GPT-5.6 传闻 GPT-5.6 Rumors

Host

嗯,有传言说 GPT 5.6 即将问世。据说这只是一个推特上的谣言,但我还是念给你听。嗯,总是最好的谣言。比 Fable 便宜三倍,上下文窗口高达 150 万 token。呃,更强的智能体式编码工作流。嗯,其中有多少是真的?我们应该对 GPT 5.6 有什么期待?

Um, GPT 5.6 is rumored to be on its way. Supposed to be this just a Twitter rumor, but I'm going to read it to you. Um, always the best rumors. Three times cheaper than Fable up to 1.5 million token context. Uh, stronger agentic coding workflows. Um, how much of that is true? What should we expect for GPT 5.6?

Greg

我的意思是,听着,你应该总是期待更好、更快、更智能,全套的。

I mean, look, you should always expect better, faster, smarter, the whole thing.

Host

所以,一切都确认了。

So, everything confirmed.

Greg

绝对相信你在推特上读到的一切。是的,也许不是。

Definitely believe everything you read on Twitter. Yeah, maybe not.

Host

这实际上一直是我个人生活中问题的来源。嗯,好吧。

That has actually been a source of problems in my personal life. Um, okay.

AI 健康:个性化医疗 AI in Health: Personalized Medicine

Host

所以,呃,我想以健康话题结束。嗯,你提过几次。你之前在观众中确实有一个关于这个的问题。嗯,你知道,有时候有一个故事,你读到它,然后对自己说,我知道这个人正在对媒体说话,我知道他们说的听起来可能是真的。嗯,但故事有些不对劲,我们不会看到更多了。嗯,我最近读到了几个这样的故事。嗯,一个是,我想是你的朋友 GitLab CEO Sid Sberage?他得了癌症,然后他做了所有他能做的诊断测试,就是疯狂地测试,然后把数据输入到 ChatGPT 中,在一些人为它构建了专用应用程序的帮助下,他能够——我不知道治愈这个词是否合适——但一定程度上击退了癌症。还有一个故事是关于澳大利亚的一只狗 Rosie。你们听说过 Rosie 吗?最疯狂的故事,这个人——我可能会说错一些细节——但一个人给他的狗做了活检,狗得了癌症,然后他把突变数据通过 AlphaFold 运行,然后设计了一种 mRNA 疫苗,他注射给了狗,在聊天机器人的帮助下构建了这东西,结果狗又能跳过桌子了,肿瘤缩小了。嗯,当我们思考 AI 和健康的未来时,嗯,帮我们理清这个问题的真相。这些是几个制造了好头条的异常案例,但故事中有一些我们没有听到的东西,还是这将成为未来的常态?

So, uh, I want to end on health. Um, you brought it up a couple times. You actually had a question in the audience about it earlier. Um, you know, sometimes there's a story and you read it and you say to yourself, I know this person is speaking to the media and I know that what they're saying sounds like maybe it's true. Um, but there's something wrong with the story and we're not going to see more of it. Um, and I've read a couple of those recently. Um, one is I think is it your friend the GitLab CEO um, Sid Sberage? He had um he he got cancer and used uh he got all the diagnostic testing uh he could have so just went out and tested like crazy and fed that data into chatt with the assistance of some people who had built purpose-built application for it and was able I don't know if cure is the right word but to beat back the cancer to a degree. There was also this dog Rosie the dog in Australia. You guys heard of Rosie? like the craziest story where this guy I'm gonna get some detail wrong but a guy um biopsied his dog which had cancer um ran the mutations across of across Alphafold and then was able to design an mRNA vaccine that he injected into the dog um with the assistance of chatbots to build this thing which ended up being able to jump over tables again and the tumor shrunk. Um, when we think about the future of AI and health, um, help us sort out the truth with this question. Are these a couple of outliers that made good headlines, but there was something about the story we weren't hearing, or is this going to become standard in the future?

Greg

绝对会成为常态。绝对会。而且我个人有很多朋友做过非常类似的事情,获取正确的数据,你的健康诊断,然后使用这些模型从中获取见解。我认为有很多人,比如我认为每周大约有 2.3 亿人使用 ChatGPT 进行健康查询,对吧?这是一个惊人的规模。这些人有时会上传扫描结果。有时医生会告诉你相互矛盾的信息。我认为我们一直处于一个患者没有被赋权的世界,对吧?患者必须成为医生,对吧?你是决策者。你要负责,对吧?你知道,医生犯了错,你将为此付出余生代价。这是一种非常不同的激励机制。这对我来说非常个人化。你知道,我妻子有多种健康状况,我认为我们——我甚至不知道如果没有 ChatGPT,我们现在如何管理她的许多状况。我认为我们只是处于这个旅程的开始,对吧?我认为,即使你拥有最好的医疗团队、最好的资源、最好的专家,能做的也只有这么多,对吧?你想想那些人类能力之外的事情,或者有时就像有人甚至没有看图表,对吧?错过了一个细节。所有这些我们都应该能够通过这些工具大幅改进。所以我认为个性化医疗,有时是关于药物和药物发现,针对大众市场,但有时甚至针对像 NF1 这样的疾病诊断,就像我今天早些时候提到的。有时只是为了理解病情并尝试提出新的潜在疗法。所有这些我们现在都亲眼目睹。这不是理论上的。它真的在发生。所以我认为 AI 最令人惊叹的可能性之一就是它能极大地改善我们的健康。你想想这个系统的连锁反应,对吧?现在医疗系统上花费了那么多钱。那是经济的一大部分,如果你能真正帮助人们预防问题,对吧,提前应对潜在的健康问题,那实际上会减轻很多负担和压力。我们生活在一个医生精疲力尽、护士精疲力尽的世界,就像我们面前正在发生一场真正的危机。我认为 AI 将能够帮助解决所有这些问题。我们有这种潜力,只要我们明智而妥善地部署和使用它。

Absolutely going to become standard. Absolutely. And it's I I personally have a number of friends who have done very similar things of get the data right your health diagnostics and use Codex right use these models to get insights from them and I think that there are many people like I think that there's about 230 million people each week who use chat GBT for health queries right and that's been that's this is like a staggering scale Right. And this is pe these are people sometimes you upload a scan. Sometimes you have doctors who are telling you conflicting information. And I think that we've been in a world where patients are not empowered, right? Patients have to be the doctor, right? You're the decider. You are accountable, right? You know, doctor makes a mistake and you're going to be paying the price for the rest of your life. Like it's just it's a very different kind of incentive. And this is very personal for me. you know, my wife has a number of health conditions and I think that we've just been we've not like I don't even know how we'd be able to manage many of her conditions right now without the use of of chat. And I think we're just at the beginning of this journey, right? That I think that the degree to which even if you have the best medical team, the best access, the best best experts, there's only so much that can be done, right? That you think about the things that are just outside of the reach of humanity or even just sometimes it's like someone didn't even read the chart, right? and kind of missed a detail. All of that we should be able to improve massively through these tools. And so I think that the personalized medicine and sometimes it's going to be about drugs and drug discovery that are of you know for for mass market but sometimes it'll be even for the kind of NF1 things like the the disease um diagnoses that I mentioned earlier today. Sometimes it will be for just like trying to understand conditions and trying to come up with with new potential therapeutics. All of that we're seeing it happening right now in front of our eyes. It's not theoretical. It's really happening. And so one of the most I think it's like one of the most astounding possibilities of AI is how much it can improve our health. And you think about the ripple effects of the system, right? Where so much spending on the health care system happens right now. that's a massive part of the economy and that if you're actually able to help people prevent issues, right, to to get ahead of potential, you know, health health problems, that's something that actually then alleviates a lot of burden, a lot of strain. And we're in a world where doctors are burned out, nurses are burned out, like there's like a real crisis that's happening in front of us. And I think AI will be able to help with all of that. Like we have that potential if we deploy it and use it wisely and well.

结语 Closing Remarks

Greg

所以我认为,将人工智能应用于医学,这对我来说是一个个人动力,让我思考我们正在构建的整个旅程,以及我们试图通过 OpenAI 实现的目标。我希望我们作为一个世界和社区能够充分利用这一点。

And so I think that applying AI to medicine, like that's something that is really a personal motivation for me in thinking about this whole journey of what we're building, what we're trying to do with OpenAI. And I'm hopeful that we as a world and a community can make the most of that.

Host

希望如此。我认为我们会的。我非常有信心。

Let's hope. I think we will. I'm very very confident.

Host

Greg,非常感谢你。

Greg, thank you so much.

Host

谢谢。太好了。谢谢。非常感谢。天哪。谢谢大家。今天过得愉快。

Thank you. Great. Thank you. Thank you so much. Oh my god. Thank you everyone. You have a good time today.

Host

谢谢。我们明年再办一次吗?

Thank you. Should we do it again next year?

Host

你会来吗?是的。

You going to come? Yeah.

互动版:逐字朗读 + 针对本期提问 →