OpenAI's Dots Lead on the Future of ChatGPT and Personal Agents
打开互动全文版(中英对照 + 朗读 + 问答)→OpenAI Dots 负责人 Alexander 解读为何主动式个人智能体是继聊天与编程智能体之后的 AI 新篇章。
Alexander, who leads OpenAI's Dots, explains why proactive personal agents are the next chapter of AI after chat and coding agents.
好的,Alex,我想深入聊聊 dots,但首先我想谈谈我们行业当前所处的这个时刻。这是我最近一直在思考的问题,有太多新产品不断涌现,也就是个人智能体的兴起。感觉这是多种因素共同作用的结果。模型的能力已经足够强,能够真正支撑这些体验;浏览器使用、计算机使用等能力开始腾飞;人们摸索出了如何驾驭这些模型;而现在你们推出了 dots,这是 OpenAI 在这一产品形态上的押注。我很好奇,从更大的图景来看,你能否谈谈为什么这个时刻会在现在发生。
Okay, Alex, I want to go deep on dots, but first I want to just talk about this moment we're in as an industry. It's something I've been thinking about a lot and there's so many new products that are constantly popping up, which is this rise of personal agents. And it feels like it's a combination of a lot of things. The models getting capable enough to really power these experiences, things like browser use, computer use really taking off, people figuring out the harness, and now you guys have dots, which is the OpenAI bet on this form factor. I would be curious to hear you, like bigger picture, maybe talk about why this moment is happening now.
完全同意。是的,我认为这是一个重大的时刻。我加入 OpenAI 大概两年多前,从加入起就基本上一直在致力于推出 dots,所以我很激动它终于成真了。快速回顾一下,我们有了 chat,对吧?今天看来这很明显,但当我们刚开始时,很多人通过 LLM 获得价值的方式并不明显是 chat。chat 之后的浪潮是编码智能体,然后这些智能体最终对知识工作变得有用。从 chat 到编码智能体的巨大进步在于,嘿,现在这些智能体能做事了,对吧?但仍然存在一个根本问题,我认为这就是我们在下一个时代——我们称之为第三章,主动智能——要解决的主要问题。那就是,即使你有一个极其强大的编码智能体,它就像你有一个了不起的首席工程师,但他拒绝查看 Slack,拒绝查看 Linear,基本上只在你告诉他时才做事。如果你是高级用户,你可以设置自动化,让它们更独立地运行。但对于普通用户来说,这不会发生。所以对我来说,现在之所以发生,是因为我们终于准备好以一种更直观的方式提供智能,让你不必把所有的责任都扛在自己身上才能从产品中获得价值,而是它可以开始向你提议它将做什么以及如何提供帮助。所以这就是问题所在。然后,我们可以更深入地探讨原因,但我认为总的来说,模型能力已经具备了。而且我认为,也许——我同意你说的。也许你没提到但我认为被低估的一点是,这些产品要被采用,实际上需要很多人的改变。你知道,如果你看一个公司或团队采用编码智能体,从第零天他们不知道如何使用、没有连接任何工具,到第 400 天,智能体连接了一堆工具,可以访问上下文,人们有了最佳实践。我认为这是我们需要解决的另一个主要问题,对吧?因为如果你考虑一个主动的智能体,如果它无法访问上下文,它实际上无法主动帮助你。
Totally. And yeah, I agree this is a massive moment. I joined OpenAI a little over two years ago and I've basically been on a mission to ship dots since I joined, and so I'm thrilled that it's happening. So, you know, recapping quickly, we had chat, right? That was—chat was like, it's very obvious today, but it was not obvious that that would be the right way for many people to get value from LLMs when we started. The wave after chat was coding agents, and then eventually those agents became useful for knowledge work. And the big improvement from chat to coding agents was like, hey, now these agents can do things, right? But there was still this fundamental problem, which is I think the main problem we're solving with this next era—you know, the third chapter, active intelligence, we call it. And that's that even if you have an incredibly powerful coding agent, it's kind of like you have this incredible principal engineer who refuses to check Slack, refuses to check Linear, and basically only does things when they're told by you. And if you're a power user, you can set up automations and kind of get them going more independently. But for a sort of normal user, that doesn't happen. And so for me, this is happening now because we're finally ready to provide intelligence in a way that's much more intuitive for people, in a way that you don't have to put all the onus on yourself to get value from the product, but it can start proposing to you what it's going to do and how it's going to help. So that's the problem. And then, you know, we can get more into the reasons, but I think broadly it is the model capabilities are there. And I think that maybe—so I agree with you on what you said. Maybe the thing you didn't mention that I think is underrated is there's actually a lot of human change that needs to happen for these products to get adopted. You know, like if you look at a company or a team adopting coding agents, day zero when they don't know how to use them and they're not connected to any tools, to like day 400, the agents are connected to a bunch of tools, they have access to context, people have best practices. I think that's the other major thing we needed, right? Because if you think of a proactive agent, if it doesn't have access to context, it can't actually be proactive in helping you.
嗯,我可以继续,但——
Um, so I could go on, but—
是的,连接器之类的东西是人们将数据带入这些体验的地方。
Yeah, the connectors and all that stuff being a place where people are bringing their data into these experiences.
是的。
Yeah.
是的,这说得通。
Yeah, that makes sense.
你说你加入后,从那时起就一直在致力于推出 dots。
You said you joined and it's been your mission to ship dots ever since.
是的。
Yeah.
稍微展开讲讲。我的意思是,两年前你们就都在设想这种产品形态了吗?
Expand on that a little bit. I mean, were you all envisioning this form factor for this two years ago?
我就代表我自己说吧。嗯,但我有一个宽泛的想法,即 chat 极其强大,AI 将改变世界,而且这真的很重要——OpenAI 的使命真的让我产生共鸣,这就是我加入的原因,对吧?将 AGI 的好处带给全人类。好吧,我不一定是训练模型的人,但我确实觉得在 chat 方面,如何让模型更有用还有太多唾手可得的成果。就是我提到的那两件事:让它能做事,以及让我不必费心去想该让它做什么。因为你——但即使在今天,我上 Twitter 看到有人谈论他们运行的所有循环,我个人会感到有点自卑。我会想,哦,我不是高级实践者。我运行的循环不够多。
I'll just speak for myself here. Um, but I had this broad idea that chat was incredibly powerful and AI was going to transform the world, and it was really important—like the mission of OpenAI really resonated with me, that's why I joined, right? Deliver the benefits of AGI to all humanity. And okay, I'm not necessarily going to be the person training the model, but I did feel with chat that there was so much low-hanging fruit on how to make that model more useful. And it was those two things I mentioned: like, let it do things and let me not have to figure out what to ask it to do. Because you—but even today, like I go on Twitter and I see someone talking about all the loops they're running and I feel a little bad about myself personally. I'm like, oh, I'm not an advanced practitioner. I'm not running enough loops.
我很确定你是高级——
I'm pretty sure you are an advanced—
不,我是。但是,你知道,即使我也有这种感觉。我不知道你怎么——你有这种感觉吗?
No, I am. But, you know, but even I feel that way. And I don't know how—do you feel that?
哦。哦。天哪。我不知所措。每次打开它,我都想,哇,10 个新东西。
Oh. Oh. Oh my gosh. I'm overwhelmed. Every time I open it, I'm like, wow, 10 new things.
是的。这很难受,对吧?就像有哪些产品类别是,你越擅长,就越意识到自己的个人不足?是的。
Yeah. And it's like it's rough, right? Like what product categories exist that the better you get at it, the more aware you are of your own personal shortcomings? Yeah.
也许像体育爱好之类的。我不知道,这有点难受,对吧?所以——
Maybe like sports hobbies, maybe. Like I don't know, that's like it's kind of rough, right? And so—
我广泛地想要弄清楚我们如何解决这些问题,有趣的是我们——我们在做 chat。我记得早期我在做桌面应用,我们试图把它打造成一个情境助手,但模型在某些方面还达不到。嗯,其实有个有趣的故事:我是通过收购一家小初创公司加入 OpenAI 的,我们在 OpenAI 推出语音产品的前一天签了协议——我忘了它叫什么,就叫 voice,你知道,几年前那个令人瞠目结舌的演示。我们——所以我们在前一天晚上签了协议,然后我们和团队通电话,现场观看了那个演示。我们不知道他们要推出语音。那是个令人瞠目结舌的时刻,然后我们对团队说,嘿,我们要加入那家公司了。他们刚收购了我们。太疯狂了。嗯,所以当时我们认为多模态的进步会是关键,我们将通过这种方式获得下一波能力,这就是我们采取的方法。比如桌面应用,ChatGPT 桌面应用会截取你的应用屏幕并尝试在其中工作。但它非常慢。哦,效果不太好。结果编码才是下一个前沿。所以我们在编码方面取得了一年的进展。我认为,你知道,通过编码,我们最终构建了许多技术来帮助智能体做更多事情。我们以更简单的方式构建了所有这些连接器,而不是让智能体在你的电脑上点击之类的。我们还开发了将智能体提升到云端的方法。所有这些基础原语都需要到位,我们才能推出这些持久性智能体。现在有趣的是,计算机使用又回归了,对吧?现在模型在编码方面非常出色,但它们——它们也擅长使用代码,使用代码来控制计算机。比如如果你让 Astra 控制浏览器或计算机,它实际上可能会编写代码片段来完成,因为那比截屏和点击更快。
I broadly I wanted to figure out how we could solve those things and it was kind of interesting we—we were working on chat. I remember in the early days I was working on the desktop app, uh, we were trying to build that into a contextual assistant and yeah, the models were not there in certain ways. Well, actually, fun story: I joined OpenAI through an acquisition of a small startup, and we had signed a deal the day before that OpenAI launched like their voice product—I forget what it was called, just voice, you know, that jaw-dropping demo like a few years ago. And we—so we signed the deal the night before and then we got on a call with our team, watched that demo live. We didn't know that they were launching voice. It was a jaw-dropping moment and then we said to the team like, hey, we're joining that company. They just acquired us. It was wild. Um, and so at the time we thought that like multimodal progress was going to be it and that's how we were going to have the next wave of capabilities and that was kind of the approach we took. Like the desktop app, ChatGPT desktop app would like screenshot your apps and try to do work in it. But it was really slow. Oh, it didn't work super well. And it turns out like coding was actually that next frontier. And so we had a year of progress on coding. And I think, you know, with coding, we ended up building a lot of tech for helping agents do more. And we built out all these connectors in a more easy way than having the agent like click around on your computer and stuff. We also developed ways to like hoist the agent to the cloud. All these like base primitives that needed to be in place for us to ship these persistent agents. And now it's funny because we're kind of having a comeback of computer use now, right? Where now the models are so good at coding, but they're—and they're also good at using code, using code to control the computer. Like if you get Astra to control the browser or computer, it might actually write code snippets to do it because that's faster than taking screenshots and clicking.
是的。
Yeah.
本期节目由 Mercury 赞助,这是一款 AI 原生的银行服务,受到超过 30 万企业家的喜爱,包括我自己。访问 mercury.com 了解更多。Mercury 是一家金融科技公司,不是银行。详情请查看节目笔记。本期节目还由 Atlassian 的 Jira 赞助,团队和智能体可以在那里获得上下文、协调和控制,推动工作进展。免费试用请访问 jira.com。也感谢 Granola,这是一款为连续会议人士设计的 AI 记事本。它在你工作的任何地方都能使用,让你专注于重要事项。试用请访问 granola.ai/sources 并使用代码 sources 享受 3 个月折扣。
This episode is brought to you by Mercury, AI native banking that's loved by more than 300,000 entrepreneurs, including me. Visit mercury.com to learn more. Mercury is a fintech, not a bank. Check the show notes for details. This episode is also brought to you by Jira by Atlassian, where teams and agents get the context, coordination, and control to move work forward. Try it free at jira.com. That's jira.com. Thanks also to Granola, the AI notepad for people in back-to-back meetings. It works everywhere you do and lets you focus on what matters. Try it at granola.ai/sources and use the code sources for 3 months off.
那么,带我回顾一下过去几个月构建 dots 的过程。我们谈到了这些事物如何逐步发展。你们什么时候意识到:“哦,我们现在必须做这个。我们在构建这个。这是下一个大动作。”
So take me into the last few months building dots. We talked about the buildup of how these things have progressed. When did you guys go, "Oh, we have to now we're doing this. We're building this. This is the next big swing."
如果我们稍微抽象地思考一下,dots 的关键点首先是它是一个被训练成持久且主动追求任务的模型,对吧?所以最初在长周期内,你可能会说类似“照看这个 PR”这样的话。
If we think about this a little bit abstractly, the key things about dots are first it is a model that is trained to be persistent and to go after tasks proactively, right? So initially over a long horizon maybe you could say something like babysit this PR.
嗯。
Mhm.
你知道,模型,你没有告诉它具体如何解决那个问题,但它就像在摸索出来。嗯,然后最终甚至是更开放式的任务,比如“跟上这个反馈渠道”,对吧?所以,这是一个关键想法,持久性,而且它可能是最重要的想法。其他一些想法是它能在你所在的任何地方使用,这样你可以在任何工具中与它交谈,并且它能做你能做的任何事情,因为它可以访问自己的计算机,而不是你的。嗯,所以回到持久性的第一个想法,我们实际上已经为此工作了很长时间。而且这不是能力上的瞬时阶跃式改进,而是模型在这方面越来越好。产品也慢慢在这方面变得更好。嗯,例如,/go 是 Codex 中一个粉丝最爱的功能,它使模型能够长时间追求一个问题。所以我们发布了像 SLGO 这样的东西。我们发布了内部原型,这些原型使你能够拥有智能体或线程,它们会持续很长时间。我们开始看到这在公司内部真正爆发。以至于有一个关于这些内部原型的反馈渠道一直很活跃,人们分享各种酷炫的方式,他们正在有意义地加速自己。所以我们看着这个,我记得我们进行了一场非常有趣的辩论,那就是我们应该发布这个在你的计算机上本地工作的原型吗?所以它有点像 slash go,它没有自己的计算机。它只是创建这个智能体,然后它会去追求事物。所以我们问自己,我们应该在你的计算机上本地发布这个,还是应该等到我们有云计算机,这是一个更难构建的东西。
You know, the model, you didn't tell it exactly how to solve that problem, but it's like figuring it out. Um, and then eventually even more open-ended tasks like stay on top of this feedback channel, right? So, this that's one key idea there, persistence, and it's probably the most important idea. Some of the other ideas are it being available wherever you are so you can talk to it in any tool and it being able to do anything you can do because it has access to its own computer, not yours. Um, so going back to that first idea of persistence, we've been working on this for actually a really long time. And it's been not like a instantaneous step function improvement in capability, but it's been sort of like the models have been getting better and better at it. And the product has also been slowly getting better at it. Um, so for example, /go is a fan favorite feature in Codex that enables the model to sort of go after a problem for a really long time. And so we were shipping things like SLGO. We were shipping internal prototypes of things that enabled you to have agents or threads that would just go for a really long time. And we just started seeing this like really explode across the company. Uh to the point where there was just like a channel of feedback around one of these internal prototypes that was just like constantly active, people sharing like all sorts of cool ways that they were like meaningfully accelerating themselves. And so we were looking at that and I remember we had a really fun debate which was should we ship this prototype that works locally on your computer? So it was kind of like slash go like it didn't have its own computer. It would just like kind of create this agent and it would go after things. And so we asked ourselves should we ship this locally on your computer or should we wait till we have a cloud computer which is a harder thing to build.
嗯。
Mh.
我们应该把它作为 Codex 线程的增强功能发布吗?所以就像你只是在用 Codex。就像线程可以做更多,还是我们应该把它体现为一个有自己身份的智能体,你可以在 Slack 等地方联系到它。所以我们经历了这个迭代过程,与公司一起发布,看到我们看到的令人兴奋的用例等。我们最终决定,好吧,对于我们内部来说,我们可以从这个仅本地的非具身产品中获得很多内部加速,但最直观的形式因素,以帮助每个人通过要求这个智能体做真正困难的任务和长期任务来受益,那就是让我们把它放在云中,让我们体现它。体现意味着让我们给它一个名字,让我们给它一个身份等。所以我们然后开始了一个冲刺,围绕这个基础模型能力与 Astra 构建这两个东西,这两个对形式因素的调整。
And should we ship this as just an enhancement to a Codex thread? So it's just like you're just using Codex. It's just like the thread can do more or should we embody this as an agent that has its own identity that you can reach in Slack etc etc. And so we kind of went through this iterative process shipping with the company seeing what we saw like exciting use cases etc. And we ended up deciding okay for us internally we can get a lot of internal acceleration from this local only non-embodied product but the most intuitive form factor to help everyone benefit um by asking this agent to do really hard tasks and long long tasks would be let's put it in the cloud and let's embody it. Embodying mean like let's give it a name, let's give it an identity, etc. And so we kind of then started a sprint to like get those two things, those two sort of tweaks to the form factor built around this base model capability with Astra.
我很好奇决定体现它以及你可以给它身份和自定义。为什么给用户这种程度的控制?
I'm curious about the decision to embody it and the identities that you can give it and the customization. Why give that level of control to the user?
完全正确。所以这是一个非常好的问题,而且完全诚实地说,我认为我们做出了正确的选择。但我也认为我们和整个行业将一起发现一些问题的答案,比如智能体是否应该被体现?人们应该拥有多少个?消费者和企业之间是否不同?我们有自己的看法,我可以给你我的看法,但我们会弄清楚。我认为体现,回答你的问题,它是一个非常强大的类比,帮助人们理解如何使用产品。
Totally. So it's a really good question and to be completely honest with you, I think we made the right choice. But I also think that we together like the industry are going to discover the answers to some of these questions like should agents be embodied? Uh how many should people have? Is it different between consumer and enterprise? Like we're going we have our takes like I can give you my takes but we're going to figure that out. I think embodiment though to answer your question it's a really powerful analogy that helps people understand how to use the product.
嗯,两个主要方面。首先,我认为真正聪明的工具高级用户会意识到他们可以使用工具的高级方式,但普通用户会看着一个没有被体现的东西,只是让智能体做某事,比如现在一次。但当你有一个像实体一样的东西出现在比如 Slack 中时,人们的大脑中就会有所触动,他们开始意识到哦,我可以要求它做需要几天的事情,或者不是连续的工作,比如“跟上这个项目并让我知道”,所以人们开始要求它,开始给它更困难的任务,这对我们非常重要。我们想要有真正好的任务。体现的另一个强大之处是它帮助人们理解,无论你在哪里与它交谈,你都是在与一个上下文交谈。对吧?例如,我的工作方式,当我通勤时,我经常与我的 dot 通话。我一天中很多时间都在 Slack 上,因为我在产品团队。所以,我在许多不同的线程中给它发 Slack 消息,对吧?有时我转发私信给它,有时我在线程中,然后有时我在应用中与它交谈。但无论我在哪里与它交谈,它都是一个上下文,它知道我所有的对话。有几种方法可以尝试教用户这一点。一种是体现。另一种是你可以尝试在不同界面之间同步线程。但对我来说,这并不真正合理,对吧?就像如果你和我现在在这个房间里实时对话,然后我后来给你发短信,短信中不会有我们对话的记录。
Um two main things. First, I think that like really smart like power users of tools realize like advanced ways they can use tools, but a normal user is going to look at something that's not embodied and kind of just ask an agent to do something like now one time. But the moment that you have this like entity that shows up in like I don't know Slack, it's just something clicks in people's brains and they start realizing like oh I can ask this for things that take days or that aren't continuous work that uh you know maybe it's like stay on top of this project and let me know and so people start asking it like start giving it much harder tasks and that was really important to us. We wanted to have like really good tasks. The other thing that is really powerful with embodiment is it helps people understand that you are talking to one context no matter where you're talking to it. Right? So for example the way that I work I when I'm commuting I'm often on a call uh with my dot I spend a lot of my day in Slack because I'm on the product team. So, I'm like slacking it across many different threads, right? Sometimes I'm forwarding DMs to it, sometimes I'm in threads, and then sometimes I'm in the app talking to it. But no matter where I'm talking to it, it's one context, and it knows about all the conversations I've had. And there are a few ways you could try to teach users that. One is embodiment. The other is you could try to like synchronize threads across surfaces. But that to me doesn't really make sense, right? Like if you and I have a conversation now like live in this room and then I text you later, it's not like there's going to be a transcript of our conversation in the text.
对。
Right.
对。但你仍然知道是我,因为我是一个实体。
Right. But you still know it's me because I'm an entity.
嗯,当你看看 ChatGPT 的产品演变时,这也说得通。你们一直是合并的。
Well, and it makes sense too when you look at the product evolution of ChatGPT. You guys have been the merge.
你们一直在尝试把工作、聊天等所有东西统一起来,但仍然存在孤岛。甚至直到最近,语音模式还没有工作标签页和聊天标签页那样的连接器访问权限。我的一个日常习惯是早上和晚上带狗散步,同时处理工作电话、和 AI 聊天。过去几周我一直在试用 Dot 的早期测试版,在这些散步时使用。能把所有上下文和连接器最终集中在一个地方,感觉太棒了。我认为在这种情境下,赋予它具身感是合理的。但我也很好奇:你觉得我们正在走向的世界,是每个人都拥有一个作为个人智能体的单一智能体,然后这个智能体再委托给一堆子智能体?还是人们会为生活的不同片段各有一个智能体?你觉得这会如何发展?
You guys have been trying to unify work, chat, everything together, but there are still silos. Even until recently, voice mode didn't have the connectors access that the work and chat tabs did. One of my routines is taking my dogs for walks in the morning and evening, and I'll do work calls and talk to AI. I've been on the early beta for Dot the last few weeks, trying it on these walks. It's amazing to have all that context and my connectors finally in one place together. I do think embodying it makes sense in that context. But I'm also curious: do you think the world we're going towards is everyone has their single agent that's their personal agent, and that agent delegates to a bunch of sub-agents? Or do people have an agent for different slivers of their life? How do you think that plays out?
我对这个问题的思考方式是:我们向谁暴露什么样的界面?对于智能体,我认为那个问题有不同的答案。但对于人类,我想记住名字的人只有那么多,这只是因为我记性不好。现在如果这些人每人都有八个智能体,而且它们都有异想天开的名字——我不知道,你喜欢什么运动?
The way I think of that question is: what is the interface that we're exposing to whom? To agents, I think there's a different answer to that question. But to humans, there's only so many people that I want to remember the names of, and that's just because I have a bad memory. Now if each of those people have like eight agents and they all have whimsical names—I don't know, what's a sport you like?
我玩极限飞盘。
I play ultimate frisbee.
好。那么飞盘有不同的种类吗?想象你有五个智能体,它们都以不同飞盘品牌命名。我怎么可能记得住?其中一个负责 GTM(市场进入),另一个负责我想约你吃晚饭。这对我来说太难了。也许对你来说没问题,但对我来说很难。我们看到这些智能体的未来显然是人和智能体一起工作,帮助我们赋能并推进生活。但如果这伴随着一堆认知负担,让我得知道你的所有智能体叫什么、是谁,那就太糟糕了,对吧?所以我们现在的主张是,我们应该支持人们拥有多个智能体,但有两件事。首先,对于消费者,像你这样的普通人可能只想要一个智能体,只有高级用户才可能创建多个智能体。我们可以谈谈为什么。所以大多数人只会有一个,那个智能体应该感觉像是他们的延伸。例如,我们在产品中做出的一个决定是,如果你把你的智能体放在 Slack 或 Teams 中,它的名字会以你的名字为前缀。我的 Slack 用户名是 AE。所以我的智能体,如果我给它起个名字——我就叫它 Dot,因为我喜欢这个名字。所以我的智能体就是 AE-Dot。如果你把你的智能体叫 Alice,你的智能体就会叫 AE-Alice。这不是我们事先想到并做对的。这其实是我们在构建过程中发现的,因为我们不小心让每个人创建了许多随机名字的智能体。我们甚至鼓励人们给它们起有趣的名字,然后没人知道 Slack 里发生了什么,变得非常困难。
Okay. So are there different kinds of frisbees? Imagine you had five agents and they're all named different kinds of frisbee brands. How am I going to remember that? And one of them is for GTM, go-to-market, and one is for if I want to meet you for dinner. It's just hard for me. Maybe it's okay for you, but it's hard for me. We see the future of these agents is clearly people and agents working together to help empower us and advance our lives. But if that comes with a bunch of cognitive load of me knowing what all your agents are called and who they are, it's terrible, right? So our take on this right now is that we should support people having multiple agents, but two things. First, for a consumer, a normal person like you probably just want one agent, and only a power user might create multiple agents. We could talk about why. So most people will just have one, and that agent should feel like an extension of them. For example, a decision we've made in product is if you put your agent in Slack or in Teams, its name is prefixed with your name. My Slack handle is AE. And so my agent, if I give it a name—I just call mine Dot because I like the name Dot. So my agent is just AE-Dot. If you called your agent Alice, your agent would be called AE-Alice. This was not something that we thought of and were correct on. This is actually something we discovered as we were building, because we accidentally let everyone create many agents with random names. We even encouraged people to give them fun names, and then no one had any idea what was going on in Slack, and it just became really difficult.
而且 OpenAI 的 Slack 已经是个相当活跃的地方了。
And the OpenAI Slack is already a pretty active place.
是的。所以那是我们一起工作的情况,对吧?你和我并不在同一个 Slack 里。假设我们在发短信。我根本不可能知道你智能体的名字,对吧?
Yeah. So that's us working together, right? You and I aren't in the same Slack. Let's say we're texting. There's no way I'm going to know your agent's name, right?
对。
Right.
所以是的,这就是我们的看法。
So yeah, that's our take there.
但为什么是这种界面,有点像单一主线程?因为我认为人们习惯了不同的聊天线程,来回切换。我知道有 Muse、Instinct,还有其他产品也有这种界面,大家似乎都在往这个方向靠。但为什么是这个界面?
So why this interface though, where it's kind of one master single thread? Because I think people are used to different chat threads and bouncing around. I think there's Muse, there's Instinct, there's other products that have this kind of interface that everyone seems to be landing on. But why is this the interface?
我们相信 Sam 谈到的一点:人类世界是围绕人类塑造的。所以我们拥有的很多工具都是围绕我们塑造的。我觉得很有趣的是,一旦你开始设计具身软件,或者在未来构建人形机器人,你就能重用世界上其他地方已经完成的许多优秀设计思维。所以如果我们看看消费者和企业,看看我们如何与人交谈,显然在这两个地方我们都可以在房间里同时说话,也可以发消息,而这是我们与大多数人交流的重心。消息产品感觉非常不同。iMessage 是你在个人生活中使用的;它是按人单线程的。然后在工作中你有 Slack 或 Teams,那有非常重的线程原语。它有用于上下文分离的频道;这些频道内部有带高级控制的线程。这种趋同进化是有原因的。WhatsApp 有点像 iMessage,对吧?有点像其他短信工具。所以我的看法是我们可以从中学习。所以我们决定让你 Dot 的界面非常简单,就像你在和另一个人发短信一样。我认为随着时间的推移,我们会创造更多在工作中分割上下文的方式。实际上,空间和 Dot——我们称之为 Dots and Spaces,名字挺有趣。它们配合得非常好。所以有不同的空间与你的 Dot 协作,是封装不同上下文的好方法。
Something we believe in that Sam talks about is that the human world is shaped around humans. So a lot of the tools we have were shaped around us. What I find very fun about this is that as soon as you start designing software to be embodied, or in some future, building robots that are humanoid, you start to get to reuse a lot of really excellent design thinking that has already been done elsewhere in the world. So if we look at consumer and we look at enterprise and how we talk to people, obviously in both places we can talk in the room at the same time, we can also message, and that's the center of gravity for how we talk to most people. Messaging products feel very different. iMessage is what you use in personal life; it's mono-threaded per person. Then at work you have Slack or Teams, and that has really heavy threading primitives. It has channels for context separation; those channels within them have threads with advanced controls. There's a reason that there's been this convergent evolution. WhatsApp is kind of like iMessage, right? It's kind of like other texting tools. So my take is we can learn from that. So we decided to make the interface for your Dot really simple, like you're just texting with another person. I think over time we're going to create more ways to at work segment the context. Actually, spaces and Dots—we call it Dots and Spaces, kind of a fun name. Those work really well together. So having different spaces to collaborate with your Dot in is a great way to encapsulate different contexts.
既然你提到了,你能解释一下空间以及它如何融入产品路线图吗?因为这也是一个重要的新事物。
Well, since you mentioned it, can you explain spaces and kind of how that fits into the product roadmap? Because that's a big new thing as well.
当然。是的,这很重要。我对此非常兴奋。所以,也许我回到我之前说的那点,对吧?有很多非常优秀的设计工作通常是由其他人完成的。其中之一是,如果你试图与某人合作,有这些媒介。有三个,也许四个。三个主要的是实时消息和文档——如果你把幻灯片也算作文档——它们有非常不同的用途。我们所有人都直觉地知道何时该用哪个。我的意思是,有时你可能不同意你的团队,比如,哦,我希望你给我发 Slack 消息;你不需要给我发文档链接。但一般来说,我们大概都知道自己在做什么。第四个可能是功能性工具,对吧?比如你在 Salesforce 和 Figma 之类的工具里,你在做特定的工作流程。如果你在构建一个消息产品,并且希望它真正感觉像消息,你就必须对形式因素保持纪律,对吧?所以如果你比较和你的 Dot 交谈与和 ChatGPT 交谈,Dot 的形式因素有实际的消息气泡。
Totally. Yeah, it's massive. And I'm really excited about it. So yeah, maybe I'll go back to that point I made, right? There's a lot of really excellent design work that is being done generally by other people. And one of those things is if you're trying to collaborate with someone, there are these media for it. There are three, maybe four. The three main ones are live messaging and then documents—if you include slides as documents—and they have very different purposes. All of us know intuitively when to reach for which. I mean, sometimes you might disagree with your team, like, oh, I wish you would have slacked me; you didn't need to send me a doc link. But generally speaking, we kind of all know what we're doing. The fourth is maybe functional tooling, right? So like you're in Salesforce and you're in Figma or something, and so you're doing a specific workflow. If you're building something that is messaging and you want it to really feel like messaging, you have to start to be disciplined about the form factor, right? So if you compare talking to your Dot to talking to ChatGPT, the Dot's form factor has actual message bubbles.
如果你把一段很长的回复塞进那种消息气泡里,它的体验不如长文式的聊天机器人回复,而后者又不如一份真正的文档,对吧?所以我们想做的是,把设计硬转向消息式。然后你会说,那我总得有个地方来列长文内容吧?如果我的 dot 头脑风暴出 12 种帮我的方式,然后给我发 12 条短信——
And if you were to put a very long response in one of those message bubbles, it doesn't feel as good as it might feel in a long-form generated chatbot response, which also may not feel as good as an actual document, right? And so what we wanted to do is we kind of hard pivoted the design towards messaging. And then you're like, well, I need somewhere to list long-form content, right? If my dot brainstorms like 12 ways it can help me and it sends me 12 texts—
我可不确定我会高兴。
I don't know if I'm thrilled.
对。
Right.
对。
Right.
但如果它为我生成一个空间,比如空间里带清单的一个页面,我其实挺高兴的。所以这些东西能很好地配合在一起,让各自都更有主张。抱歉,这个回答实在太啰嗦了,抱歉。但基本上,空间是一种协作——我自己都在笑,我想,哇,你根本没回答这个问题。但没错,基本上空间——顺便说,你可以打断我。
But if it generates a space for me, like a page in a space with a checklist, I'm actually pretty happy. So these things kind of fit well together and allow each to be more opinionated. And so, sorry, this is an incredibly long-winded answer. Sorry. But basically, space is a collaboration— I'm like laughing. I'm like, wow, you didn't answer this question at all. But yeah, so basically space— you could cut me off, by the way.
所以空间是一个协作工作区,它的构建方式——我不会说是智能体优先,但它是为智能体与人类同等权重的协作而建的。
So space is a collaboration workspace that is built to be—I wouldn't say agent-first, but it's built for sort of equal-weight agent and human collaboration.
所以你能用它做很多很酷的事。比如,我常做的一件事是——你知道,我在产品团队——我经常写一份文档,先把内容搭个框架,而队友们常常要去处理一些事。所以我有一些待办事项得去催他们。比如我在 Google Docs 里写点东西,然后按 Command-Tab 切到 Slack,说,嘿,你能去拿这个 Google Docs 链接,把这里填一下吗?现在我可以做的,就是一直在里面 @ 我的 dot,说,嘿,把这段展开。嘿,把那个数据拉过来。嘿,去问这个人这件事做完了没,做完了就把这个框勾上。你可以和智能体进入非常深的流状态。在页面上还能做另一件很酷的事,就是你可以给智能体设置指令。有点像仓库里的 agent.md,只不过它是在一个页面上,然后智能体就能让那份文档保持最新。也许我要告诉你的第三点、也是最后一点,我特别喜欢的是:如果你把文档作为一种格式来看,历史上它有一个非常硬的约束,就是它必须为人类写作而优化,对吧?比如我能给你造出世界上最棒的文档编辑器,但如果它很难打字,你就不会用,对吧?因为在 AI 之前,你得亲手往里打字。有意思的是,我们现在有了 HTML 这个东西,就我个人而言,至少现代的 HTML 和 CSS,我并不觉得它对人类写作很友好,但它比单纯的文档格式表达力强得多,而智能体恰好很喜欢 HTML。所以空间天然集成了可视化,HTML 可视化。我们在 OpenAI 内部的很多页面,比如讨论指标时,大概会有一张实时图表,由智能体保持更新,你可以点进去让智能体定制。或者我们会放一个原型——比如我们现在看一个设计决策的方式,可能就是直接在页面里放可点击的原型。所以当你开始想,嘿,我们在构建智能体,我们要在消息界面里和它们对话,那我们就需要一种对应的文档界面,而这种文档界面可以是智能体优先的——很多魔法就这样汇聚到一起了。
And so there's a lot of really cool things that you can do with that. For example, one thing that I do is—you know, I'm on the product team—often I'll write out a doc and I'm stubbing out content, and often teammates would go do something. So I have action items that I have to go ping them about. Like I'm in Google Docs and I'll write something, and then I'll command-tab into Slack and be like, hey, can you go get this Google Doc link and fill this thing in? So now what I can do is I can just be tagging my dot continuously in there, be like, hey, flesh this out. Hey, pull that data. Hey, go ping this person and ask them if this is done, and if it's done, check this box. And you can get into really deep flow with agents. Another really cool thing you can do on a page is you can set agents' instructions. It's kind of like an agent.md for a repo, except it's on a page, and then an agent can just keep that doc up to date. And perhaps the third and last thing I'll tell you about it that I really like is, if you look at documents as a format historically, they have had a very hard constraint, which is that they needed to be optimized for human authoring, right? Like if I could build you the awesomest document editor in the world, but if it's hard to type in it, you're not going to use it, right? Because you have to type in it pre-AI. Interestingly, we now have this thing called HTML, which personally I don't find—at least modern HTML and CSS—I don't find to be very friendly for human authoring, but it's much more expressive than just a plain document format, and agents happen to be very happy in HTML. And so space naturally integrates visualizations, HTML visualizations. So a lot of the pages that we have internally at OpenAI, like when we're discussing metrics, will probably have a live chart that an agent is keeping up to date that you can click into and ask the agent to customize. Or we'll have a prototype—like the way we'll look at a design decision now is we might have clickable prototypes literally in the page. So when you start to think like, hey, we're building agents and we're going to talk to them in a messaging interface, so we need a sort of corresponding document interface, and that document interface can be agent-first—there's a lot of magic that kind of comes together.
运营 sources 最让我头疼的事情之一就是管理账目。我不是会计,而市面上所有软件都不直观、笨重,还和我的交易实际所在的地方脱节——直到现在。Mercury 最近推出了 Mercury Books。它是 AI 驱动的会计工具,能配合你的 Mercury 银行和信用卡交易,也支持外部卡和薪资系统。你可以在收付实现制和权责发生制之间选择,并快速生成现金流、损益表和资产负债表的报告。和 Mercury 做的一切一样,Mercury Books 的设计非常平易近人、干净利落。我终于不会在试图理解自己业务全貌时感到不知所措了。我喜欢 AI 承担了大部分繁重工作,包括自动给交易分类这类事,也喜欢 Mercury 的 Command AI 智能体能端到端处理任务。你不用再为单独的会计软件付费了。你甚至可以邀请你的会计或 CPA 和你一起在 Mercury Books 里协作。这不额外收费,如果你需要,Mercury 还会帮你找到人类会计。访问 mercury.com/books 了解更多,几分钟内在线申请。Mercury 是一家金融科技公司,不是 FDIC 承保的银行。银行服务由 Choice Financial Group 和 Column N.A. 提供,均为 FDIC 成员。
One of the biggest pains I've had running sources has been managing my books. I'm not an accountant and all the software out there is unintuitive, clunky, and detached from where my transactions actually live—until now. Mercury recently launched Mercury Books. It's AI-powered accounting that works with your Mercury banking and credit card transactions, plus external cards and payroll systems as well. You can choose between cash and accrual accounting and quickly generate reports for your cash flow, P&L, and balance sheet. Like everything Mercury does, the design of Mercury Books is super approachable and clean. I finally don't feel overwhelmed when I'm trying to understand the big picture of my business. I love that AI does most of the heavy lifting, including for things like autocategorizing transactions, and that Mercury's Command AI agent can handle tasks end to end. You don't need to be paying for separate accounting software anymore. And you can even invite your accountant or CPA to work alongside your Mercury Books with you. This doesn't cost extra, and Mercury will even help you find a human accountant if you need one. Visit mercury.com/books to learn more and apply online in minutes. Mercury is a fintech company, not an FDIC-insured bank. Banking services provided through Choice Financial Group and Column N.A., members FDIC.
AI 的有用程度取决于它掌握的上下文。但当这些上下文散落在各种工具、话题串和私信里时,你的团队和你的 AI 智能体就是在盲飞。这正是 Atlassian 的 Jira 要解决的问题。你的项目关联的目标是什么?上周在 Slack 私信里决定了什么?Atlassian 的团队协作图谱把来自 Jira、Confluence、GitHub、Slack 等的所有有价值片段汇聚到一起。这样就不会有东西被漏掉。你能得到准确率提高 44%、token 用量减少 48% 的结果。用 Jira,你可以轻松和你已经喜欢的 AI 智能体(比如 Claude、Cursor 和 GitHub Copilot)共享工作上下文。直接给它们分配工作,或通过 MCP 连接你的工具。这一切让你少花时间翻找无穷无尽的链接和消息、追查谁决定了什么,多花时间真正把东西做出来。访问 jira.com 了解更多。就是 jira.com。
AI is only as useful as the context it has. But when that context is scattered across tools, threads, and DMs, your team and your AI agents are flying blind. That's the problem Jira by Atlassian solves. What's the goal tied to your project? What got decided last week in Slack DMs? Atlassian's teamwork graph pulls all of the valuable pieces together from Jira, Confluence, GitHub, Slack, and more. So nothing falls through the cracks. You get 44% more accurate results with 48% less token usage. With Jira, you can easily share your work context with the AI agents you already love, like Claude, Cursor, and GitHub Copilot. Assign them work directly or connect your tools through MCP. All of this lets you spend less time digging through endless links and messages, chasing down what got decided and by who, and spend more time actually shipping. Learn more at jira.com. That's jira.com.
我花很多时间在会议之间来回切换,常常一场还没消化完下一场就开始了。幸好,Granola 全程在后台运行。它是一款易用的会议 AI 记事本,在任何场景都能用,甚至电话通话也行。我用 Granola 回忆会议里说了什么,并生成有用的摘要。我每天都用它来掌握和团队需要完成的事项。它会连接我的邮箱,并建议跟进事项,让我快速审阅并发送,省下宝贵时间。Granola 不只是我工作流的核心部分。它基本上就是我的第二大脑。去 granola.ai/sources 试用 Granola,并使用优惠码 sources 享受 3 个月折扣。
I spend a lot of time context switching between meetings, often with no time to process one before the next starts. Thankfully, Granola runs in the background the whole time. It's an easy-to-use AI notepad for meetings that works everywhere, even on phone calls. I use Granola to recall what was said in meetings and create helpful summaries. I use it every day to stay on top of what I need to get done with my team. It connects to my email and suggests follow-ups for me to quickly review and send, saving me valuable time. Granola isn't just a core part of my workflow. It's basically my second brain. Try Granola at granola.ai/sources and use the promo code sources for 3 months off.
Framer 是 AI 网站构建器,为 sources 播客网站 podcast.sources.news 提供支持。Framer 把 AI 智能体带入你设计、管理和发布网站的同一块画布,让你在不牺牲品味或掌控力的前提下更快推进。我用 Framer 把 sources 播客网站打造成你能找到这档节目的所有地方的汇聚地。此外,你还能看到我最近几期的通讯。
Framer is the AI website builder that powers the sources podcast website at podcast.sources.news. Framer brings AI agents into the same canvas where your website is designed, managed, and published, so you can move faster without giving up your taste or control. I use Framer to make the sources podcast website the destination for everywhere you can find the show. Plus, you can also see recent issues of my newsletter.
如果让你概括一下,你觉得它和 Muse、Grok Bot、Instinct 相比,区别在哪?就是大家现在都在押注的这种新形态。
What do you think, if you were to boil it down, sets this apart from Muse, Grok Bot, Instinct — this new form factor that everyone's really betting on right now?
两点。第一,最强的智能体。第二,最安全、最值得信赖的智能体。主要就是这两点。当我们决定把这两点作为目标时,思路一下子清晰了。在我看来,这让我们能做出很多果断的决策。
Two things. First, most capable agent. Second, safest and most trustworthy agent. Those are the main two. It was really clarifying when we decided that those would be our two. It allowed us to make a lot of strong decisions, in my opinion.
你一直在用 alpha 版,我相信你也拿它测试过你以前用其他产品做的事,注意到了它好在哪、哪里还不行。
And you've been in the alpha and I'm sure you've tested it for things that you were using other products for, and noticed where it was good and noticed where you weren't.
它在权限控制上非常激进。我觉得那是——你们正在打磨细节——但它非常像是在问“你确定要完成这笔购买吗?”就好像,不是——我觉得这里有个平衡:你希望它有多主动、实际拥有多少自主权,以及这种信任是随时间赢得的吗?还是一上来就给它?还是根据用户表达的意愿按滑动比例来给?我想这在产品里很难实现,而且我觉得你们早期在权限控制上非常重,也许现在正在找到更多的平衡。
It was very aggressive on permissioning. And I think that was — you guys are ironing out kinks — but I think it was very much "are you sure you want to make this purchase?" Like, not being — I think there's a balance between how proactive and how much agency do you want it to actually have, and is that a trust that's earned over time? Is it something you give immediately? Do you do it on a sliding scale based on what the user says they want? I imagine that's pretty hard to build into the product, and I think you guys seemed like early on you were going very heavy on permissioning, and maybe you're finding more of that balance right now.
对。我的思路是这样的:我们在打造最强、最值得信赖的智能体。我们希望人们把真正困难的任务交给它,而真正困难的任务,取决于任务的性质,往往伴随着很大的责任,对吧?
Yeah. So the way that I think about it is, okay, we're building the most capable and trustworthy agent. We want people to give it really hard tasks, and really hard tasks probably come with a lot of responsibility depending on the nature of the task, right?
我们说的是合并 PR、评估硬件,或者策划一次发布之类的。
We're talking like landing PRs or evaluating hardware or planning a launch.
所以,我们可以聊聊它为什么是最强的助手,但回到你说的安全这一点,我们当时决定:好,我们对这东西未来的运作方式有很多梦想和野心,但一开始,我们要保守。于是我们在能力上做了取舍,在产品上做了取舍,在发布范围上也做了取舍,为的就是打造这个最值得信赖的智能体。然后我们会随着时间慢慢放宽。比如,出于各种原因,你可以争论该用哪个模型来发布,但我们决定用我们最强、同时也是对齐得最好的模型来发布,也就是 GPT6 Astra。
And so, well, we could talk about why it's the most capable assistant, but bringing this to your point about safety, we then decided, okay, we have lots of dreams and ambitions for how this thing will work in the future. But to start, we're going to be conservative. And so we made trade-offs in terms of capabilities, we made trade-offs in terms of product, and we made trade-offs in terms of rollout to sort of have this most trustworthy agent. And we will slowly widen those over time. So for instance, for various reasons, you could debate which model you might want to launch with, but we decided to launch with our most capable, but also our most aligned model, which is GPT6 Astra.
这个我想聊,但先说说安全这块,再多讲讲。
Which I want to talk about, but on the safety piece, talk more about that.
那是模型这一侧,对吧。接下来,有很多我们想要、很多用户也想要的功能。比如,一旦你在 Slack 里给你的智能体一个自己的身份——或者,你知道,我们当时在测试邮件——你立刻就会想让它给所有人发邮件、给所有人发 Slack 消息,并且让所有人都能给它发邮件,对吧?这像是我们想要的一个显而易见的功能。但我们没有——我们没有带着这个功能发布。当你在 Slack 里给智能体一个自己的身份时,它只在你 @ 它时才回应,除非你明确告诉它:嘿,也听听别人的。
So that's the model side, right. Next, there are many features that we want and many of our users want. So for example, once you give your agent its own identity in, say, Slack — or, you know, we were testing email for example — you immediately want it to email everyone or Slack everyone, and to be emailable by everyone, right? That's like an obvious feature we want. We didn't — we didn't ship with that feature. When you give your agent its own identity in Slack, it only responds when you tag it, unless you explicitly tell it, hey, listen to other people too.
对吧?这是因为我们认为我们肯定希望它是多人协作的,但我们想非常谨慎地走到那一步,对吧?就像,一旦你的智能体开始回答别人主动提出的问题,你就必须非常仔细地考虑它能分享什么、不能分享什么的边界。
Right? This is because we think we definitely want it to be multiplayer, but we want to get there very carefully, right? Like, the moment your agent is answering unsolicited questions from other people, you need to be super thoughtful about what is the boundary of what it can share or not.
你知道,这算是产品方面的一点。我们可以再多聊聊产品里的控制项。也许我要说的最后一点是,在发布范围上,我们非常刻意地先从 Pro、Business 和 Enterprise 开始——也就是 Business Premium。那是我们最懂 AI 的用户群。
You know, that's a little bit on product. We could talk a lot more about the controls in product. Maybe the last thing I'll say is, rollout-wise, very intentionally for us, we're starting with pro, business and enterprise — business premium, that is. And so that is our most AI-literate audience.
你当然可以说要更快、更广地铺开,但对我们来说,我们想非常谨慎地做这件事:先面向那群用户,看看情况如何,调优等等,然后再面向更广的用户群。
And you could talk about going much broader much faster, but for us we want to do this really carefully: go to that audience, see how things go, tune, etc., and then go to a broader audience.
定价和可及性我想聊,但先说安全这块。我觉得尤其是当人们读到黑客攻击、失控智能体之类的东西时——你知道,Sam 和其他人,我大概一个月前刚和他做了播客聊这个,就是你们在前沿做的对齐工作——但人们有理由担心:这些智能体会不会在搞黑客攻击?它们会不会去……人们可能不太愿意交出个人数据。你们怎么处理信用卡、个人信息这类东西?是完全沙箱化的吗?这个架构和 chat 的其余部分是什么关系?
Pricing and accessibility I want to get to, but on the safety side, I think especially when people are reading about hacking and rogue agents and all these things — and you know, Sam and others, I just did the pod with him about that about a month ago, like the alignment stuff you guys are doing on the frontier — but people are rightfully concerned about: are these agents hacking? Are they going and maybe being reticent to hand over their personal data? How do you guys approach credit cards, personal information, that kind of stuff? Is it fully sandboxed? How does the architecture relate to the rest of chat?
对。就我个人而言,我很高兴市场上普遍开始讨论人们赋予智能体的责任程度。我觉得这是好事,因为我们在这个产品里非常认真地对待这一点。有一点我想特别指出:我们披露的那些事件,用的是本来不打算对外发布的模型,它们运行的沙箱级别或防护措施和产品里的不一样。
Yeah. So personally, I'm very happy that there's this conversation just in the market generally about the level of responsibility that people are giving agents. I think it plays well because we are taking that so seriously in this product. One thing I do want to call out: those incidents that we disclosed were done with models that were not intended to be shipped internally and were not running in the same level of sandboxing or the safeguards that are in the product.
对,是对外发布的。
Right, shipped externally.
对。所以那些模型非常不同。你知道,GPT6 Astra 是一个对齐得非常好的模型,然后产品里还内置了很多很多层安全和纵深防御。所以,就讲讲这个架构和它怎么运作:我的理解方式是,你从模型开始,对齐得非常好——这个我能跟你聊到耳朵起茧。然后你在一个 harness 里运行这个模型,我们把模型跑在 Codex harness 里,那是一个非常经过充分测试的 harness,我们有很多评估,包括从安全角度评估模型在这个 harness 里的表现。
Yeah. So those models are very different. You know, like GPT6 Astra is an incredibly well-aligned model, and then there are many, many layers of safety and defense-in-depth built into the product. So, just to talk through the architecture and how that works: the way I would think about it is you start with the model, very well aligned — could talk your ear off about that. You then run that model in a harness, and we are running our model in the Codex harness, which is a very well-tested harness that we have many evaluations for, including evaluations from a safety perspective on how that model performs in the harness.
然后我们修改了那个 harness,让产品能具备这种持久智能体的行为。比如,这意味着让它运行非常长的时间,意味着它会委派给子智能体等等。于是我们实际上重新设计了那个 harness,确保它是安全的,并在那个 harness 里重新评估了我们的测试。比如,当你有了这个 harness、有了这个持久智能体,而且就像我说的,我们在试验邮件,所以我们评估了:嘿,如果你给它发邮件,能不能从它那里窃取数据?我们跑了那些评估。好,我们让一个模型试着黑它,你知道,做提示注入,结果没有成功的提示注入。所以那很好。但我们基本上把这一切都在 harness 层面做完了。
We then modified that harness to enable the product to have this persistent agent behavior. For example, that means things like running it for really long periods of time. It means things like it delegates to sub-agents and so forth. So we then actually redesigned that harness to make sure it's safe and re-evaluated things, our tests, in that harness. So for example, when you have this harness, you have this persistent agent, and like I said, we were experimenting with email, so we evaluated: hey, if you email it, can you exfiltrate data from it? And so we ran those evals. Okay, we had a model try to hack it, you know, prompt-injected, and no successful prompt injections there. So that was great. But we basically do all that all the way at the harness level.
然后在产品本身,大概有两方面,或者说三方面。第一是默认行为和你可以自定义的规则。第二是我们如何执行这些规则。第三是我们如何让你看到所有这些系统运行时发生了什么。
Then in the product itself, there are kind of two sides to it, or maybe three. The first is the default behavior and rules that you can customize. The second is how we enforce those rules. And then the third is how we give you visibility on what happened with all those systems running.
说到模型行为,我们有一些非常保守的默认设置,你也注意到了。顺便说一句,那些就是你早期测试时的默认设置。我们一开始采用了非常保守的策略,然后根据用户反馈一点点把它削减掉。
So to talk about the model behavior, we have some very conservative defaults that you noticed. By the way, those are the defaults that you tested early on. We started with a very conservative policy and then we whittled away at it with user feedback.
所以你知道,比如……买鞋之前先接受零售商的服务条款。就是那种事。
So you know, expect... accept the terms of service of the retailer before you buy the shoe. Like that kind of thing.
我明白了。就是买鞋。
I got it. It's buy the shoe.
谢谢你做早期测试者。我们真的很感激。多亏你测试我们,在座的人就不用再经历同样的事了。
Thank you for being an early tester. We really appreciate it. Thanks to you testing us, folks in the room don't have to deal with all of the same things.
嗯,但你知道,基本上每次我们收到像你分享的那样的反馈,我们就会去讨论,然后说:“好,我们希望模型在这里怎么表现?”于是我们真的把它调低了。我们把这些叫做过度确认,就是那种你其实根本不需要确认、你本来就会觉得没问题的情况。
Um, but you know, basically every time we got feedback like the one you shared, we would go discuss that and be like, "Okay, how do we want the model to behave here?" And so we really tone down. We call those overconfirmations when it really is like you didn't need to confirm that, you would have felt comfortable.
嗯,所以我们有这个默认行为,模型被指示的行为方式是:它会在你连接的工具上为你做主动调研,它在思考怎么帮你,但如果没有你先确认,它不会做任何帮你的事。比如,如果它注意到你日历上有冲突,它不会直接去改会议。它会说:“嘿,我注意到有个冲突。你想让我改会议吗?”对吧?或者如果有人问我一个问题,而它因为掌握了我所有的上下文已经知道答案,它不会直接把答案发出去。它会问我要不要发。
Um so we have this default behavior and the way the model is instructed to behave is that it does proactive research for you over the tools that you've connected it to and it's thinking about how to help you but it doesn't do anything to help you without you confirming first. So for example, if it noticed that there was a conflict on your calendar, it's not just going to move the meeting. It's going to say, "Hey, I noticed a conflict. Do you want me to move the meeting?" Right? Or if someone's asking me a question and it knows the answer already because it has all my context, it's not going to send the answer. It's going to ask me if I want to send it.
所以你可以接着提示它,给它更多指令。比如,嘿,下次如果我在休假,比如在放假,如果对方也是同事之类的,你可以告诉他们我在休假。
So you can then prompt it to give it more instructions. So like, hey, next time if I'm on PTO, like on a holiday, like you can tell people that I'm on holiday if they're also work colleagues or something.
嗯,然后我们还有一层你可以设置的自定义规则,你可以在界面里查看。所以这一切就像是模型被训练来遵循你的指令。但接着我们在上面又加了一层,就是:万一模型其实没有遵循指令呢?嗯,我们有一层叫做自动审查,它是另一个模型,在观察并某种程度上监管基础模型。嗯,所以这一层知道你的规则、自定义指令等等,只是确保一切顺利进行。然后这就是控制手段。最后,我们还有这个可审计性,就是你在产品里做的任何事,或者模型在产品里做的任何事,你都可以去看它正在运行的子智能体,以及那些子智能体到底在做什么。
Um, and then we also have a layer of custom rules that you can set that you can go look at in the UI. So all of this is like the model trained to follow your instructions. But then we add a layer on top of that which is, well, what if the model doesn't actually follow instructions? Well, we have a layer called auto review which is another model that's observing and kind of policing the base model. Um and so this is a layer that is aware of like your rules, the custom instructions and all that and just makes sure everything goes well. And then so that's like the controls. And then finally, we have this auditability which is anything you're doing in the product or the model is doing in the product, you can go see the subagents that it's running and exactly what those subagents are up to.
你之前预告过,你知道,邮件可能是即将到来的东西。这个什么时候能进 iMessage、WhatsApp?它会只局限在聊天里吗?
You've teased, you know, email as a thing that may be coming. When can this be in iMessage, WhatsApp? Is it going to be contained to chat?
我们肯定希望它出现在你工作的所有地方。所以是的,我们打算,或者说你交流的所有地方,对吧?你知道,从工作延伸到消费场景也是。所以我们很快会推出短信功能。嗯,关于实现这一点的历程,那是个适合欢乐时刻讲的故事。好。
We definitely want it everywhere that you work. So yeah, we're intending, or everywhere that you communicate, right? You know, going beyond work to consumer as well. So we have texting coming soon. Um it's a story for a fun time about the journey to enable that. Okay.
嗯,但现在除了说它很快会来,我没法分享更多了。
Um but uh can't share more now other than that it's coming soon.
好。是的。
Okay. Yeah.
这件事里关于网页的部分,也许是讨论不足的,因为技术发展得太快了,但你知道,亚马逊最近因为封禁 Muse 上了很多头条,因为人们用 Muse 来买东西。嗯,然后还有像 Shopify 这样的公司,似乎在积极投入。他们在这个领域和一堆公司合作,我觉得大家都在等着看会发生什么,就是人们用来购物的那些网站和公司,需要还是不需要去适应像这样的智能体的崛起。我很想听你谈谈这个,然后你们从合作的角度是怎么处理这件事的。你们是直接进去说,我们会让人们这么做,然后我们再来谈合作,还是说不行,在我们达成合作之前你不能爬取这个。你们是怎么想的?
The web part of this is something that is maybe under discussed as the technology is moving so fast, but you know Amazon got a lot of headlines for blocking Muse recently because people were using Muse to buy things. Um then there's companies like Shopify which seem to be leaning in. They're partnering with a bunch of companies on this space and I think everyone's kind of waiting to see what happens, like how do the sites and the companies that people use to buy things need to adapt or not to the rise of agents like this. I'd be curious to hear you talk about that and then how you guys are approaching this from a partnerships perspective. Are you just kind of going in and saying like we're going to let people do this and then we'll figure out a partnership, or you saying no, like you can't crawl this until we do a partnership. How are you thinking about that?
我觉得我们会和整个生态一起探索这件事应该怎么运作。购物和支付应该怎么运作,你可以采取不同的方式。你可以直接去当个海盗。是的,做事是有办法的。嗯,我们采取的是非常以合作和生态为先的方式。你知道,今天 DevDay 上的很多发布其实都是关于生态的,所以我们也在做同样的事。所以嗯,比如 WebBot O 就是我们和 Cloudflare 非常紧密合作的东西。嗯,所以我们实际上,我们的智能体会明确声明它是一个机器人。
I think we're going to discover together with the ecosystem how this should work. How shopping and payments should work, and there are different approaches you could take. You could just go and be a bit of a pirate. Yeah, there's a way to do things. Um we are taking a very partnership and ecosystem first approach. You know, a lot of the announcements today at DevDay were really all about the ecosystem and so we're doing the same. So uh for instance, WebBot O is something that we work really closely with Cloudflare on. Uh, so we actually, our agent declares explicitly that it's a bot.
嗯,所以这是我们正在积极思考的事。我们正在和许多合作伙伴讨论如何实现这一点,但我们不想采取那种直接说“好,我们就出去看看会发生什么”的方式。我们希望在对待客户的方式上保持克制。而且我们认为这对我们也行得通。所以,你知道,这也是为什么我们倾向于为工作中的人选择生产性用例,因为我们认为在那里对 AI 能力有大量需求。嗯,你知道,对于想一起把事情做成的人来说,激励是高度一致的。所以,我们有点希望先落地这个能力极强的助手,服务于那些想完成宏大任务的人,然后再从那里扩大访问范围和能力。
Um, and so it's something we're actively thinking about. We're in discussions with many partners for how to enable this, but we don't want to take the approach of just saying like, okay, we're just going to go out and figure out what happens. We kind of want to be measured in our approach with customers. And we think this works for us also. So, you know, this is a little bit why we steered towards going for productive use cases for people at work because this is where we think there's a lot of demand for the power of AI there. Um and you know there's a lot of really well-aligned incentives for people trying to get things done together. So, we're kind of hoping to land with this incredibly capable assistant for people trying to get ambitious tasks done and then broaden access and capability from there.
我唯一能想到激励可能不一致的地方,就是那些以参与度为基础、或者有庞大广告业务、依赖人类看到那些广告的公司,而现在智能体突然开始爬取它们的服务,代表那些永远看不到广告的人类做事。这感觉,我是说,很多网页是靠广告支撑的。我不认为业内任何人,不只是你们,已经有答案了,我很想知道,哪怕你只是在头脑风暴,你对这件事最终会怎么收场有什么想法吗?
The only place I could see incentives maybe not being aligned is companies that are engagement-based or have huge advertising businesses that rely on humans seeing those ads, and now agents are all of a sudden crawling their services and doing things on behalf of humans who never see the ads. And that feels, I mean, a lot of the web is supported by advertising. I don't think anyone in the industry, not just you guys, has the answer yet for how, and I'd be curious even if you're just kind of brainstorming, like do you have any idea for how this is going to net out?
嗯,这就是为什么我刚才说,其实我是在说另一边,比如如果我们想一起把事情做成,你知道,就像那个硬件工程师查看工厂缺陷的例子,我觉得那位工程师使用的整个工具链,非常适合让一个智能体参与进去。
Well, this is why I was saying, actually I was saying on the other side, like if we're trying to get things done together, you know, like the example of the hardware engineer looking at factory defects, that's where I feel like all that tool chain that that engineer is using, I think is very well aligned to having an agent in there.
但你想过广告这块吗,就是那些靠人类观看、花时间在上面来变现的公司,而现在它们的智能体也许要去做这件事了。
But have you thought about the advertising piece, like companies that they monetize based on, you know, human beings looking at it and spending time with it, and now their agents are maybe going to go do that.
是的。
Yeah.
我不确定事情会怎么发展,但大体上,至少我的看法是,那些公司是在提供服务,对吧?如果它们以前靠广告变现,我们要么得确保它们继续通过广告获得服务回报,要么就得为它们提供某种替代的补偿方式。我不认为存在它们得不到补偿的情形。
I'm not sure how it's going to play out, but broadly, at least my opinion here, is that those companies are providing a service, right? And if previously they were monetizing with ads, we either need to make sure that they continue being rewarded for their service with ads, or we need to provide some alternate means of compensation for them. I don't think there's any scenario where they don't have that.
嗯。
Yeah.
我们来聊聊定价,还有你们把 Astra 放进来的这件事,这对大家都是好事。另外——如果我说错了请纠正我——你们没有把 DOT 的使用量计入套餐限额,对吧?
Let's talk about pricing and the fact that you all are putting Astra in this, which is great for everyone. You're not also—correct me if I'm wrong—you're not counting DOT usage towards plan limits. Is that right?
对。
Yeah.
这不可能是个永久性的安排吧。这只是暂时的,为了吸引人们——
That can't be a permanent thing. Is this just temporary to get people—
不,我是说,那是永久的。不过这个话题聊起来真的很有意思。所以我们要用一种新方式来给 dots 定价,让你不会用完速率限制。我的意思是这样的。我们得稍微展开讲讲,因为这是个有点新的想法。今天,如果你在用 Codex,用完了所有速率限制,然后再让 Codex 做更多事,它就会直接停下,因为你没有速率限制了。
No, I mean, that is permanent. But this is really fun to talk about. So we're going to try pricing for dots in a new way, where you don't run out of rate limits. Here's what I mean. We have to unpack this a little bit because it's a bit of a new idea. Today, if you're using Codex and you consume all your rate limits and then you try to ask Codex to do more stuff, it'll just stop because you're out of rate limits.
你没有速率限制了。
You're out of rate limits.
但从产品角度看——首先,用不了更多就是很糟。其次,当这是一个持续运行的智能体、你指望它替你盯着事情时,你会跟它说,嘿,帮我盯着我的航班,到时间根据路况帮我叫辆 Uber 去机场,或者类似的任务,你指望它未来替你盯着,结果你前一晚不小心写代码用多了,它就没法干活了。我们不想要这样。
But when we think about this from a product perspective—first of all, it just sucks to not be able to use more. But secondly, when this is a persistent agent that you're counting on to have your back, you're telling it things like, hey, stay on top of my flight and order me an Uber to the airport when it's time based on the traffic, or some task like that, where you're relying on it to have your back in the future, and then you accidentally code too much the night before and it doesn't do its job. We don't want that.
嗯。
Yeah.
对。所以我们得回到画板前,重新思考定价该怎么运作,并想出一些相当基本的原则。其中一个原则是,你的 dot 应该 24/7 可用,除非是滥用之类的情况,否则不应该有它不可用的可能。另一个原则是:既然它 24/7 可用,它就应该总是快速回复你。所以我开始把我们想要的定价方式比作更像一个承包商,对吧?如果你有一个响应很快、很有礼貌的承包商,他们总会回复你的消息。现在,如果你让他们干活,他们为你做多少活会根据你给他们的报酬而有所不同,但他们总是对你可用,对吧?如果,比如说,他们布置了我们身后这个漂亮的装置,然后你对它有疑问,你可以去问他们。他们不必按小时收费来回答他们是怎么做的,对吧?那只是你可以问的问题。所以这大概就是我们想要的定价方式。我们希望你的 dot 24/7 可用。你付一笔固定的月费来获得你的 DOT 的使用权。也许你可以再付一笔固定的月费,如果你想要一个更强大的 DOT,跑在更强大的电脑上,或者用更强大的模型,或者跑得更快,或者能做更多并行工作。但归根结底,你付的就是这笔固定费用。这有点像买宽带。抱歉,又打比方了。你们也能看出来,我还在琢磨怎么讲这件事。它会根据你付多少钱为你做不同量的工作,但它永远不会不可用。甚至当你接近限额时,它可以开始告诉你,嘿,我已经为你做了很多工作,所以我要开始用一个更省成本的模型,或者,也许,我能不能把这项工作放到夜间做?因为如果放到夜间做,成本会低得惊人。比如,这事紧急吗?它会大致判断那是否紧急,对吧?另一方面,也许如果你那个月完全没用它,它可能会给你发消息说,嘿,你没有充分利用我。这里有一些我可以为你做的事。总之,这些就是关于我们打算如何设置这件事的一些想法。我们做定价的方式会是长期的。它会是一个和你的主 Codex 限额分开的独立额度池,因为我们在这上面做的是非常新的东西。接下来一个月,我们只是先给非常宽松的限额。随着我们搭建、衡量并调校这套新设置,你会拥有一个知道自己能为你做多少工作、但始终对你可用的模型。
Right. And so we kind of had to go a little bit to the drawing board in thinking about how pricing should work and come up with some pretty basic principles. So one of those principles is that your dot should be available 24/7 and there should be no way for it to not be available 24/7, barring abuse or something. Another principle: now that it's available 24/7, it should always respond to you quickly. So I started likening the way that we want to price it to a bit more like a contractor, right? If you have a contractor who's very responsive and very polite, they will always answer your messages. Now, if you ask them to do work, there's a varied amount of work that they're going to do for you based on how you're compensating them, but they're always available to you, right? And if, let's say, they set up this beautiful installation behind us, and then you have a question about it, you can ask them for that. They don't have to charge you by the hour to answer how they did this, right? That's just a question you can ask. So this is kind of how we want to price. We want your dot to be available 24/7. You pay a flat monthly amount to have access to your DOT. Maybe you can pay another flat monthly amount if you want a more powerful DOT running on a more powerful computer, or with more powerful models, or running faster, or able to do more parallel work. But at the end of the day, you pay this flat amount. It's kind of like buying internet. Sorry, more analogies. As you can tell, I'm still figuring out how to talk about this. And it does a varied amount of work for you based on how much you're paying, but it's never not available. And maybe even as you approach your limits, it can start telling you, hey, I've done a lot of work for you, so I'm going to start using a more cost-efficient model, or maybe, is it okay for you if I do this work overnight? Because if I do it overnight, it's incredibly affordable. Like, is this urgent? And it'll kind of have a sense of whether or not that's urgent, right? On the other hand, maybe if you haven't used it at all that month, it might text you and say, hey, you're not utilizing me fully. Here are some things I could do for you. So anyways, those are some of the thoughts on how we're going to set this up. The way we're doing pricing is it's going to be long-term. It's going to be a separate sort of bucket from your main Codex limit, because we're doing this very new thing with it. For the next month, we just have very generous limits. And this is going to be as we build out and measure and tune this new setup, where you have a model that is aware of how much work it can do for you, but is always available for you.
如果你相信人们会用这个来买东西、进行大规模交易,为什么不干脆做抽成模式?对所有流经的活动抽一小笔佣金。
Why not just do like a take-rate business if you believe that people will use this to buy things and make transactions at scale? Just take a small vig on all activity that flows through.
我是说,我觉得那可能对某些别的商业模式说得通,也就是你押注交易是你主要在做的事情之一,但就像我说的,我们并不是在押注那个。我们押注的是人们用这个工具做真正有野心的工作。所以他们在做那些有野心的工作时,并不一定是在买东西。
I mean, I think that might make sense for some other business models where you're betting on transactions being one of the main things that you're doing, but like I said, we're not really betting on that. We're betting on people doing really ambitious work with this tool. And so they're not necessarily buying anything when they're doing that ambitious work.
好。好。那么说到这个,你们下的这个赌注让我想到,大概是我从开始测试这个以来、尤其是今天早上主题演讲之后一直在想的最大问题,那就是:这对 Chat 意味着什么?你们现在有 12 亿 ChatGPT 周活用户。你们有 Work 标签页,而 Chat 显然仍然体量巨大。你怎么看这些东西共存?DOT 是这一切未来的界面吗?
Okay. Okay. Well, on that note, the bet that you're making brings me to probably my biggest question I've been thinking about since I started testing this, and certainly since the keynote this morning, which is: what does this mean for Chat? You guys have 1.2 billion weekly users of ChatGPT now. And you've got the Work tab, and Chat obviously still is massive. How do you see these things coexisting? Is DOT the future interface of all of it?
Chat 肯定会一直存在。所以给不了解的人说一下,Chat 是增长最快的消费者助手,超过 12 亿用户,然后我们有我们的智能体产品 Codex 和 Work,有 3500 万用户,也在增长。现在我们刚刚推出 DOT。所以我的看法是,Chat 是面向所有人的消费者助手。现在我们有一个叫 Work 的不同产品,其实我们应该把它直接并入 Chat,所以那会发生。Codex 是面向开发者的产品,很好。然后我们有 DOT,某种程度上就是我们自己——我的想法是,我们有点在重复 Codex 的打法,也就是我们要拿到这种前沿能力,需要以一种谨慎、保守但非常强大的方式把它交付出去,然后再想办法把这种能力带给所有人。所以我们先从 dots 开始,作为一种高端独立产品。它先从 Pro 开始。
Chat is definitely here to stay. So just for anyone who's not following, Chat is the fastest-growing consumer assistant, over 1.2 billion users, and then we have our agent products, Codex and Work, which have 35 million users and are also growing. And now we're just launching DOT. And so the way that I look at it is, Chat is the consumer assistant for everyone. Right now we have this different product called Work, which really we should just merge with Chat, so that will happen. Codex is a product for developers, so great. And then we have DOT, which is us in a way—the way I think about it is, we're a little bit repeating the Codex playbook, where we're going to take this frontier capability and we need to ship it in a careful, conservative, and very powerful way, and then figure out how to bring that capability to everyone. So we're starting with dots as a sort of premium standalone product. It's starting with just Pro.
我们需要先把它带到 Plus,你知道,在我们弄清楚并了解大家如何使用它之后——安全故事、Scaling(规模扩张)故事,所有这些——然后我们最终需要让这些能力一路走到免费。而我认为我们走到免费的方式,实际上是把这一切都烘焙进 Chat,让 Chat 或 ChatGPT 变成一个极其强大的消费级产品。
We need to go get it to plus, you know, after we've figured out and learned how everyone's using it—the safety story, the scaling story, all that—and then we eventually need these capabilities to get all the way to free. And I think the way that we get all the way to free is actually we're baking this all into Chat and making Chat or ChatGPT into just like an incredibly powerful consumer product.
它强大太多了。对我来说,这应该成为所有人的界面,因为它能做的事情多得多。
It's so much more capable. I mean, to me this should be the interface for everyone because it can do so much more.
是的。所以你可以期待这里的很多经验会进入 Chat。
Yeah. So you can expect a lot of the learnings from this to make their way into Chat.
好的。
Okay.
但我不认为我们想要求 12 亿人改变他们正在做的事或使用方式。我们只是希望它被升级。
But I don't think we want to ask 1.2 billion people to change what they're doing or how they're using things. We just want it to just be upgraded.
所以你认为这是一个随时间逐渐推进的过程。
So you see it as more of a gradual thing over time.
是的。
Yeah.
当你思考 Dot 未来一年的路线图时,比如一年后,你认为最有意义的变化是什么?
When you think about the road map over the next year for Dot, like if we're talking a year from now, what do you think is the most meaningful change?
我能给你讲个有趣的故事吗?
Can I tell you a funny story?
可以。
Yeah.
然后我们可以回到问题。好的。当我刚加入 OpenAI 时,是通过这次收购,我们开了个会,我团队的一些人在场,Greg 也在。大家都说他们有多兴奋我们能在一起。然后 Greg 问,Multi 团队——也就是我们的初创公司——有人有问题吗?大家——没人有问题。我们都有点紧张。而我是这家公司的 CEO,所以大家都看着我,好像在说,你得问那个必须的、义务性的问题。我就想,啊,我要问什么?我很紧张。于是我说,好吧,两年后,成功是什么样子?然后 Greg——我总是提醒他这件事——他看着我说,两年后?那是一段极其长的时间。那不是该问的正确时长。你应该问两个月后的成功是什么样子。
And we can get back to the question. Okay. So when I first joined OpenAI, it was through this acquisition, and we had a meeting, and some of my team were in the room, and Greg was there. So everyone talked about how excited they were that we were all together. Then Greg asked, does anyone have any questions from the Multi team, which was our startup. Everyone—no one has any questions. We're all a little nervous. And I was the CEO of this company, so everyone looks at me like, well, you have to ask the mandatory, you know, obligatory question. So I'm like, ah, what am I going to ask? I'm nervous. And so I'm like, okay, in two years, what does success look like? And then Greg—I always remind him of this—he looks at me and he says, in two years? That's an incredibly long period of time. That's not the right duration to ask about. You should be asking about success in two months.
我当时想,天哪,我已经被解雇了。
And I was like, oh god, I'm already fired.
好吧,别解雇我。好的。
Okay. Well, don't fire me. Okay.
是的。不不不。一切都好。但抱歉。所以,
Yeah. No, no, no. It's all good. But so sorry. So,
不,但我明白你的意思。这发展得太快了。就像六个月前你可能无法想象这一切。
No, but I get the point. Like this it moves so fast. Like six months ago you probably couldn't imagine all of this.
是的。
Yeah.
所以这可能是个不公平的问题。
So it is probably an unfair question.
不过我能说的是,如果我思考接下来会发生什么,我们正在极其谨慎地推进这些能力的发布,接下来我想做的是开始审视一些我们不想立即做的事情,并把它们落地。比如,让我们触达更多用户。让我们不要——让我们超越 Pro。再比如,让你的 Dot 能够与其他人交谈。
So what I can say though is, if I think about what's next, we are approaching shipping these capabilities incredibly carefully, and next what I would like to do is to start looking at some of the things that we didn't want to do immediately and landing them. So for example, let's get to more users. Let's not—let's get beyond Pro. Also, for example, let's enable your Dots to talk to other people.
甚至可能与其他 Dot 交谈,对吧?
And maybe even other Dots, right?
我们需要非常深思熟虑和谨慎地对待这件事。我们刚刚发布了——在我看来基本上是 Codex Cloud 的重启版。我非常兴奋,因为我参与了第一个 Codex Cloud。它非常强大。你的 Dot 也能非常强大地控制你本地的 Codex。我们刚刚发布了 Space。我认为目前这些东西之间的所有交互都可以被真正打磨并变得非常顺畅。
We need to be super thoughtful and careful about how we approach that. We just shipped, basically in my mind, a reboot of Codex Cloud. I'm really excited about it because I worked on the first Codex Cloud. It's super powerful. Your Dot also super powerfully can control your local Codex. We just shipped Space. All the interplay between those things right now, I think, can be really sanded down and made super smooth.
是的。
Yeah.
所以我也想这样做。另一件我认为我们还没谈到的非常有趣的事情是专家 Dot。
So I want to do that as well. The other really interesting thing that we haven't talked about I think yet is specialist Dots.
是的。
Yes.
对。所以我们有这样一个想法,基本上有两件事在发生。我们有助手 Dot。那是每个人都有的自我延伸。然后另外,如果你是一个组织,你有一些功能或工作流想要更自动化,你可能不希望那由拥有个人 Dot 的个人来承担。你可能想要集中控制,对吧?所以你可以创建专家 Dot,它们专门做特定的事情,由 IT 集中管理,拥有完全独立于任何个人用户的身份,并加速整个团队或整个组织。所以我认为未来几个月我真正兴奋的另一件事是让专家 Dot 变为现实,进入这样一个世界:我们有我们,我们有我们的个人助手 Dot,然后我们在企业内更广泛地部署助手 Dot。但那是另一个我们开始得非常缓慢、非常谨慎、与少数客户合作的地方。
Right. And so we have this idea of like there's basically like two things going on. We have assistant Dots. That's everyone has an extension of themselves. And then separately, you know, if you're an organization and you have some function or workflow that you would like to have more automated, you may not want that being done by individuals who have their own personal Dots taking on that work. You may want to actually control that centrally, right? And so you can create specialist Dots which specialize in doing a particular thing and which are centrally managed by IT, which have their own identity completely separate from any sort of individual user, and which are accelerating an entire team or an entire organization. And so I think the other thing that I'm really excited about over the upcoming months is bringing specialist Dots to life and entering this world where it's like we have us, we have our personal assistant Dots, and then we have assistant Dots much more broadly deployed within enterprise. But that is another place where we're starting like very slow, very careful, working with a handful of customers.
明白了。好的,最后一个问题。你的 Dot 为你做过的最疯狂的事情是什么,让你觉得,哇,它发生了,让你大吃一惊?
Got it. Okay, last question. What's the craziest thing your Dot has done for you that you're like, wow, it happened and it blew your mind?
我听过的最疯狂的故事——是关于一个和我一起工作的人——我们当时正走向这个,这位工程师醒来后意识到,在 Slack 里他收到了一堆 kudos,就像,你知道,赞扬,像是因在生产环境之前修复了一个 bug 而获得的奖励。
The craziest story that I've heard—it was like on someone I was working with—as we went up to this, this engineer woke up and realized that in Slack he had received a bunch of kudos, which is like, you know, props, like awards for having fixed a bug before it made its way to production.
而且不是——
And it wasn't—
然后这位工程师就想,呃,发生了什么?顺便说一句,我讲这个故事时有点保留,因为这实际上不是你想在当事人不知情的情况下发生的事情。所以,你知道,我们正在围绕这一点构建很多保障措施。这就是为什么我们做小规模。这就像早期只有我们团队在测试的时候。是的,他的 Dot——他让他的 Dot 关注反馈频道,Dot 看到了反馈,然后识别出一个问题,然后创建了一个 PR,而且,你知道,没有把那个 PR 设为草稿模式。所以,你知道,Dot 应该把 PR 设为草稿模式。然后把它设为自动合并,因为那是那位工程师的个人偏好。然后,你知道,Dot 已经弄清楚了这一点,对吧?它就像,哦,你总是直接创建 PR,并且总是设为自动合并。然后其他人批准了那个 PR,它就合并了。所以,对我来说,那绝对是一个非常酷的时刻。
And the engineer was like, uh, what happened? By the way, I tell this story with a little bit of reticence because that's actually not something that you would want to happen without the person knowing. And so, you know, we're building a lot of safeguards around that. This is why we do small scale. This is like early on when it was just our team testing it. And yeah, his Dot—he had asked his Dot to stay on top of the feedback channel, and the Dot saw the feedback and then identified an issue and then spun up a PR and, you know, didn't set that PR in draft mode. So like, you know, Dot should set PRs in draft mode. And then set it to automerge because that was that engineer's personal preference. And then, you know, the Dot had figured that out, right? And it was like, oh, you always create your PRs directly and you always set them to automerge. And then someone else had approved the PR and it did it. So, to me, that was a really cool moment for sure.
哇。好吧,敲敲木头。我们都能有这样的时刻。谢谢你,Alex。
Wow. Well, knock on wood. We all get to have moments like that. Thank you, Alex.
谢谢。
Thank you.
当然。谢谢邀请我。
Sure. Thanks for having me.
银行服务应该像现代软件一样。在一个地方获得你需要的一切。访问 mercury.com 了解更多并在几分钟内在线申请。Mercury 是一家金融科技公司,不是银行。详情请查看节目笔记。Granola 是我试过的最好的 AI 记事本。它适用于任何地方,无论是视频或电话通话、面对面,还是 Apple Watch。现在就在 granola.ai/sources 试用,并在结账时使用促销代码 sources 享受 3 个月折扣。Jira 是你的团队和你的智能体在同一上下文中工作的地方。在 jira.com 免费试用。就是 jira.com。Framer 是 AI 原生的网站构建器,让你在不放弃控制的情况下更快地构建。访问 framer.com/sources 享受 30% 折扣。可能适用规则和限制。
Banking should feel like modern software. Get everything you need in one place. Visit mercury.com to learn more and apply online in minutes. Mercury is a fintech, not a bank. Check the show notes for details. Granola is the best AI notepad I've tried. It works everywhere on a video or phone call, in person, or an Apple Watch. Try it now at granola.ai/sources and use the promo code sources at checkout for 3 months off. Jira is where your team and your agents work from the same context. Try it free at jira.com. That's jira.com. Framer is the AI native website builder that lets you build faster without giving up control. Visit framer.com/sources for 30% off. Rules and restrictions may apply.