为什么我无法在 OpenAI 打造 Jev —— TypeSafe 联合创始人兼 CEO Diogo Almeida

Why I Couldn't Build Jev at OpenAI — Diogo Almeida, TypeSafe Co-founder & CEO

迪奥戈·阿尔梅达 Diogo Almeida · Latent Space · 2026-09-21 · 约 142 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

TypeSafe 联合创始人兼 CEO Diogo Almeida 讲述为何无法在 OpenAI 打造 Jev,并介绍一类全新的机器原生、可编程的系统一模型,其目标是让代码成为消费者。

Diogo Almeida, TypeSafe co-founder and CEO, explains why he couldn't build Jev at OpenAI and introduces a new class of machine-native, programmable system-one models where code is the consumer.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 58)

全文 · Full transcript(中英对照)

开场思考AI悖论 Opening Thoughts on AI's Paradox

Diogo

AI 怎么能聪明到这种不可思议的程度?我们能在数学上解决千禧年大奖难题,却连最基础的工作都无法自动化,比如非常基础的写作类工作?但我们有这台超强的自动化引擎,只是没有合适的插头之类的东西,插不进所有这些有经济价值的工作。你知道,如果整个 Type Safe 公司消失了,也许人们要花一两年才能真正赶上。我其实不知道要多久。如果模型质量重要,那我们会在很长一段时间里处于非常有利的位置,但它已经完成了,对吧?就像这已经改变了技术历史的路径。

Like how can AI be so unbelievably smart? How can we solve millennium prize problems in math but still not automate even the most basic of work, like really basic wrote stuff? But we have this supercharged engine of automation that just does not have the right plugs and stuff to plug into all of this economically valuable work. And you know, like if the whole company of Type Safe disappears, maybe it'll take a year or two for people to truly catch up. I actually don't know how long it'll take. If model quality matters, then we are going to be in a very good position for a long time, but it's done, right? Like this has changed the path of technological history.

赞助商信息与介绍 Sponsor Message and Introduction

Host

在我们进入今天的节目之前,我有一小段话要对听众说。谢谢你们。如果你们没有选择点击并收听我们的内容,我们就无法为你们带来你们如此明确想要的 AI 工程、科学和娱乐内容。几乎每天都有人找我们赞助。但幸运的是,你们中有足够多的人订阅了我们,让这一切在没有广告的情况下也能持续下去,我们想保持这种方式。但我只想请你们帮个忙。你能做的最强大、完全免费的一件事就是点击那个订阅按钮。这是我唯一会要求你做的事。这对我以及我的团队来说意味着一切,他们每周都如此努力地为你带来这个空间。如果你这样做了,我向你们保证,我们会一直努力让节目变得更好。现在,让我们开始吧。好的,我们在演播室里,这是一个特殊的场合,因为本周,呃,Diogo,我的好朋友,呃,发布了 Jev,它已经占据了整个时间线。呃,你感觉如何?你现在是什么状态?

Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us to keep all this sustainable without ads and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you. And it means absolutely everything to me and my team that work so hard to bring in space to you each and every week. If you do it, I promise you we'll never stop working to make the show even better. Now, let's get into it. Okay, we're in the studio, a special occasion because this week, uh, Diogo, my good buddy, uh, launched Jev and it's been taking over the complete timeline. Uh, how do you feel? What's it like to be you right now?

情绪状态与AI革命 Emotional State and AI Revolution

Diogo

情绪上,从未如此糟糕。我现在就像一具疲惫不堪的行尸走肉,因为事情太多了,而我是技术型 CEO,有很多火要救。但精神上,我一直这么说,多年来在我的 overunder 活动中也一直这么说:我觉得整个 AI 领域就像那种嘉年华的镜屋,每个人都疯了,痛苦不堪,说着最奇怪、毫无意义的话。而感觉就在这一周,我更与现实同步了,就像,哦,人们现在看到了。AI 可以比曾经想象的强大得多,是的,我们将要实现的基于 AI 的经济革命又回到了桌面上,这太棒了。我太兴奋了,开发者们懂了。是的。我想向开发者表达我内心的感激,我对社区和一切都感到非常兴奋。太棒了。

Emotionally, never been worse. I'm a ragged corpse of a person right now because there's so much going on and I'm a technical CEO so I have a lot of fires to fight. But mentally, I say this all the time and I've been saying this for years in my overunder events: I feel like the entire AI field is like one of those carnival house of mirrors and everyone is just insane, in pain and saying the weirdest stuff that doesn't make sense. And it feels like for just this week, I'm more in sync with reality and like, oh people see it now. AI can be so much more than what was once thought and yes we are going to make an AI based economic revolution is back on the table and this is awesome. I'm so jazzed the developers get it. Yeah. And I want to show my internal gratitude to developers and I'm so jammed about the community and everything. It's so great.

社区优先于VIP Prioritizing Community Over VIPs

Host

是的。你昨天说,你决定优先考虑市政厅会议,而不是一群你知道的 VIP 投资者之类的人,因为你想确保他们是你最关注的人,对吧?工程师、开发者。

Yeah. You were saying yesterday that you decided to prioritize the town hall and not a bunch of like you know VIP investor type people because you wanted to make sure that they are the people that you get your most attention, right? The engineers, the developers.

Diogo

是的。感觉有点像,哦天哪,我现在在和非常重要的人说话。我可能不该透露是谁,但对我来说感觉有点肮脏,我可能过于真诚了。就像如果在我巨大的日历事件中,要交谈的人里,你知道,社区不是其中之一,那感觉就很肮脏。实际上,在我的理想世界里,应该一直都是社区。我在想,我应该在走到你演播室的路上主持一个市政厅会议吗?然后我想,不,那太疯狂了。当然。

Yeah. It felt a little like, oh man, I'm talking to really important people right now. I probably shouldn't reveal who, but it feels a little bit dirty for me to, I'm like perhaps overly genuine in things. Like it feels dirty if in my gigantic calendar events of things to people to talk to, you know, the community isn't one of those, you know, and actually in my ideal world, it would be like community all the time. I was thinking, should I host a town hall while walking to your studio? And I'm like, nah, that's too crazy. Sure.

社区增长与社交媒体 Community Growth and Social Media

Host

是的。嗯,你们一直在 Discord 上主持市政厅会议。Discord 现在有 10 万人了。嗯,你的 Twitter 关注这些数据。所以,天哪

Yeah. Well, you guys have been hosting town halls on the Discord. Discord is now 100,000 people. Um your Twitter follow these stats. So, holy

Host

你的 Twitter 爆了。呃,你知道,真的很有趣,因为在 AIE 上,你说,“请关注我。”然后你甚至没有提供你的账号。

Your Twitter's blown up. Uh you know, it was really funny cuz like at AIE, you were like, "Follow me, please." And then you didn't like provide even your handle.

Diogo

我是个新手。我是个新手。

I'm a noob. I'm a noob.

Host

但不,但就像那是一种积极的气场,就像你不知道如何推销自己。

But no, but like that's like positive aura that like you don't know how to promote yourself.

Diogo

有人在我发帖说“天哪,我们三个都是热门话题”时指出了我。然后他们说,“那是个人推送。”当然。当然。对你来说。是的。因为那是你点击的内容。嗯,所以,好吧,让我们呃,是的。所以,恭喜一切。我们会随着你有更多细节再聊,但让我们为那些像生活在岩石下的人,或者只是再一次,像确定的事情。什么是 Jev?

Someone like called me out when I posted like, "Holy we're all three trending trending topics." And then they're like, "That's a personal feed." Of course. Of course. To you. Yes. Because it's what you clicked on. Um, so, okay, let's uh Yeah. So, congrats on everything. We'll talk about more uh details as you have them, but let's for people who are like living it under a rock or just just once like the definitive thing. What is Jev?

定义Jev与System One模型 Defining Jev and System One Models

Diogo

让我想想,那是个难题。嗯,好的。如果你愿意,我很乐意重新问。我很乐意,我很乐意就随便聊聊。我会说,对于这个问题,我第一件感到欣慰的事是,现在我不必再向我的父母回答那个问题了,因为 ChatGPT 可以解释。嗯,所以我的看法是,我们需要一类新的模型。嗯,我们并不执着于给这类模型命名。嗯,我们想出的最准确的名字是系统一模型。会有原因,但有一个原因我们不叫它们决策模型,因为它们会像系统一,但超越那个。我只能说这么多。嗯,我们没想到这会是我们的大发布。所以我们还有存货。嗯

Let me think of a That's a hard one. Um, okay. And I'm happy to like reask if you want to. I'm happy to I'm happy to like just jam on it. I will say like the first thing that I'm relieved about with this question is now I don't have to answer that question to my parents anymore cuz Chad GPT can just explain it. Um so the way I see it is we need a new class of models. Um we're not attached to naming that class of models. Um are the most accurate name we've come up with is system one models. There will be reasons but it's there's a reason why we don't call them decision models because like they will be like system one is beyond that. That's all I can say. Um we didn't expect this to be our big launch. So we have stuff in the tank. Um

Host

你应该说低调的研究预览。

You should have said low-key research preview.

Diogo

嗯,有点像,没错。有点像。但嗯,我们有一类模型,我们将其描述为机器原生、系统一、大型、可编程。我认为这些是模型类别,其目标是让代码成为消费者。所以,嗯,相对于预训练的大型语言模型,它们是为了互联网的自动补全,或者像基于人类反馈的强化学习(RLHF)模型,如聊天机器人指令遵循模型,它们是为了回复文本,嗯,或者 RLVR,它与 RLHF 处于一个奇怪的灰色地带,这些模型是为了让事物直接被代码消费,因此得名 Type Safe。所以我们真正真正真正想要的是让 AI 尽可能强大,我们认为实现这一目标的方式是将其与软件集成,我们正在设计一切,你知道,不仅仅是模型的外部,还有深层内部,都针对软件进行优化。所以第一,Jev 是我们的第一个大型可编程模型,嗯,或者系统一模型,随你怎么叫。嗯,Jev 旨在优化每美元智能。呃,因此得名 Jev,你知道,Jevad。吉文斯悖论。是的。嗯,它针对每美元智能进行了优化。我喜欢和人们辩论可靠性、成本、校准和速度之间什么最重要。嗯,Jev 旨在成为,Jev 将成为处于每美元智能前沿的模型名称。

Um it's kind of was right. It kind of was. But um we so there's a class of models that we describe them as like machine native system one large programmable. I think these are is the class of models where the goal is for code to be the consumer. So um as opposed to um um pre-trained large language models which are meant for like autocomplete of the internet or RHF models like chatbot instruction following models which are meant to like reply to text um or RLVR it's in a weird gray area with RHF like these are meant to have things that directly are consumed by code hence the name type safe. So the thing we really really really want is to have like AI like be as powerful as possible and we think the way to do that is to integrate it with software and we are designing everything you know beyond just the outside the deep internals of the model to be optimized for software. So number one Jev is our first large programmable model um or system one model whatever you want to call it. Um, Jev is meant to be optimized for intelligence per dollar. Uh, hence the name Jev, you know, Jevad. Jeans Jevans paradox. Yeah. Um, and it's optimized for intelligence per dollar. I love this debate with people about what is the most important between reliability, cost, calibration, and speed. Um, and Jev is meant to be Jev will be the name of models that will be on the frontier of intelligence per dollar.

权衡与校准 Tradeoffs and Calibration

Diogo

还有其他优化方式,比如机器学习,或者至少如果你擅长机器学习,一切都在于权衡。而我们正全力以赴。

There's other ways to optimize it, like ML, or at least if you're good at ML, it's all about tradeoffs. And we are just going all out on that.

Host

是的,对我来说,校准是人们之前不太谈论的新事物之一。我们过去和 Hugging Face 的 Clementine Fia 做过一期节目,他们当时说:‘是的,实际上,他们只是,你知道,这就是你关于基于人类反馈的强化学习(RLHF)的整个论点。他们的模式是朝着你最想听到的或最可能的方向坍缩,而不是他们自己对某事的内部信心?’

Yeah, and to me, calibration is one of the new things that people weren't talking about as much. We've done an episode in the past with Clementine Fia of Hugging Face, where they were like, 'Yeah, actually, they're just, you know, and this is your whole argument about RLHF. Is their mode collapsing towards what you want to hear the most or what is most likely, instead of their own internal confidence about a thing?'

Diogo

我能就此说几句吗?酷。我听说你的听众是最技术型的。所以我实际上想深入探讨一下。如果我极其精确地确保我们发布视频中的一切都是准确和真实的,显然这很不寻常。没有人关注的一件事是基于人类反馈的强化学习(RLHF)的缺点,特别是模式丢失。

Can I soap box on that for a second? Cool. I've heard that your audience is the most technical. So I actually want to get into that. And if I went through extreme precision to make sure everything in our launch video is accurate and real, apparently that's very unusual. One of the things that no one paid attention to was the downsides of RLHF, in particular, mode dropping.

Host

模式丢失坍缩。

Mode dropping collapse.

Diogo

呃,是一回事。是一回事。我最终想写一篇博客,但我想尽可能告诉更多人,因为我觉得这很有趣。所以辛辣观点。我非常相信 Yan Lun。我认为 Yan Lun 的观点实际上是最接近——

Uh, it's the same thing. It's the same thing. And I want to have a blog on this eventually, but I like want to tell as many people this as possible because I think it's a very interesting thing. So the spicy take. I believe in Yan Lun a lot. I think Yan Lun's takes are actually among the closest to—

Host

这个呢?

What about this?

Diogo

好吧,我应该现在处理这个,还是你稍后再处理?实际上我认为在观点中,Yan Lun 是最准确的之一。但他有一个非常著名/臭名昭著的幻灯片,关于语言模型注定失败。好的,

Well, should I address this now or should you go and no later go? I actually think that among takes, Yan Lun is among the most accurate. But he has this very famous/infamous slide about LM are doomed. Okay,

Host

你知道,就像那个,他有一个饼图,其中有一小部分,一小点,并说随着序列长度增加,它出错的概率会上升。

You know, like that one where he like has like a pie chart with like a tiny part, tiny little thing, and says that as you increase sequence length, the probability of it making an error goes in.

Diogo

是的,这个。这个我喜欢这个,因为它看起来在数学上显而易见,但显然是错误的,对吧?就像它在数学上显而易见,但经验上不成立。这是我最喜欢教人们的事情,比如哪里——

Yes, this one. This one I love this one because it's one of these things that seems mathematically obvious but is obviously wrong, right? Like it's mathematically obvious but it doesn't empirically hold. And this is my favorite thing to teach people about, like where—

Host

脱节在哪里?

What's the disconnect?

Diogo

正是。我可以还是你想告诉我关于模式坍缩?

Exactly. And may I or you want to tell me about mode collapse?

Host

哦不不不。所以模式坍缩与此相关。

Oh no no no. So mode collapse is related to this.

Diogo

是的。脱节发生是因为如果你处于覆盖模式或校准分布中,你不会因为拥有异常值而受到过度惩罚。你会预期,你知道,某些时候你会处于分布外,某些时候你会处于分布内。这就是覆盖分布时发生的情况。这就像 GAN 之前的模型。它们生成模糊图像,对吧?相反,GAN 模式丢失。它们丢弃少数类,只做非常常见的类。这就是为什么这种效应不发生,对吧?就像为了生成非常长的字符串而不出错,它们需要非常保守,因为错误发生时很容易看到。而当一个看起来正确的微妙事情发生时,很难看到。这种校准就像对字符串的概率分布是彻底的毒药。

Yeah. The disconnect happens because if you are in a mode covering or a calibrated distribution, you are not overly punished about having outliers. You'd expect like, you know, something some amount of the time you'd be out of distribution, some amount of time you'd be in distribution. That's what happens when you cover the distribution. This was like models before GANs. They made blurry images, right? Instead GANs mode drop. They like drop the minority classes and just do the really common ones. And this is why this effect doesn't happen, right? Like instead of in order to generate really long strings without making errors, they need to like be extremely conservative because it's really easy to see when an error happens. It's very hard to see when like a subtle thing that looks correct happens. And that calibration is like total poison into like the probability distributions of strings.

Host

是的。

Yeah.

Diogo

这是一个微妙的观点,我认为这就是为什么这不会发生,这就是为什么字符串在决策方面如此糟糕,或者你知道,为决策过度加载字符串模型就像糟糕的时光。当我们在讨论 Han 的话题时,你同意他的修复是像世界模型,像 JEPA 类型的嵌入东西,是正确的解决方案吗?所以基本上,它可能失败的原因之一是因为你试图在 token 输出上推理,然后再次循环,继续直到你到达句子结束。

And it's a nuanced take, and I think that this is why this doesn't happen, and this is why strings are so bad at decision-making, or you know, overloading the string models for decision-making is like a bad time. And while we're on the topic of Han, do you agree that his fix is with which is like a world model like a JEPA type embedding thing is the right solve? So basically like one of the reasons that it could fail is because you're trying to reason over token outputs and then just looping back again and going keep continuing going until you reach like an end of sentence.

Host

嗯,就像,他的解决方案是 JEPA,对吧,就像联合嵌入预测。所以这是解决方案吗,或者你知道,你对此有看法吗?

Um, like is that, and his solve is JEPA, right, which is like joint embedding prediction. So like is that the solve, or you know, do you have a take on that?

Diogo

哦天,我可能不应该过多谈论机器学习的内部。但我会说,我的品牌除了不羁之外是实用的,你知道,就像我这里的观点也是实用的,就像我是缩放定律的粉丝吗?嗯,取决于。这取决于,你知道,就像缩放定律告诉你对于投入量,你在某件事上变得多好。缩放定律确实意味着指数级更多的资源换取通常次线性的收益,这看起来是糟糕的投资,除非那些线性收益真的非常有价值。但对我来说,一切都在于我们能用我们所拥有的做什么来产生最大的可能影响。我可以骂人。

Oh man, I probably shouldn't talk too much about the insides of ML. But I will say that my brand other than unhinged is practical, you know, like even my take here is practical, and like am I a scaling law fan? Um, depends. It depends, you know, it's like scaling laws tell you how much better you get at a thing for amount in. A scaling law does mean exponentially more resources for normally sublinear gains, which looks to be a bad investment unless those like linear gains are like really really valuable. But it's all to me, it's all about like what can we do with what we have to make the biggest possible difference. I can curse.

Host

是的。

Yeah.

Diogo

是的。是的。我们被批准为成人内容。

Yeah. Yeah. We're approved for adults.

Host

当然。

Hell yeah.

Diogo

而且我们还有一个缩放定律的东西,如果你想稍后深入。哦,我可以,如果我们——那部分现在不太相关。实际上,如果你想深入我最痛苦的教训,我认为那更相关。但对我来说,我一切都关于实用性,我认为 JEPA 的东西是非常酷的早期研究。我真的很喜欢很棒的研究。它实用了吗?嗯,可能不应该说。嗯,但就像有很多——我只是认为现在研究界到处都是未打磨的钻石,因为人们不知道如何做正确的任务,我认为我们的发布所做的,它确实启动了我们作为一家公司,就像是的,它会对我们作为一家公司很好吗?是的,我认为对于程序化 AI 这个方向会更好。会有像淘金热在我们之上,因为软件被超级充电,但我认为也会有与我们平行的淘金热,关于我们可以暴露事物的所有不同方式,使软件更强大,这样人们可以做出更酷的东西,然后我们回到像早期互联网的能量,你知道,我认为这就是为什么 Twitter 就像 Jeff Jeff,你知道,嗯,就像——

And also we have a scaling law thing if you want to go into that later. Oh, I could if we— That part is not super relevant right now. I actually if you want to go into my bitterest lesson, I think that that's more relevant. But like to me, I'm all about like pragmatics and I think that the JEPA stuff is really cool early research. I really love awesome research. Is it practical yet? Um, probably shouldn't say. Um, but like there's just a lot of— I just think there's like so many diamonds in the rough all over the research world right now that haven't been polished because people don't know how to like do the right task and I think that what our launch did, it does kickstart us as a company, like yes, will it be great for us as a company? Yes, I think it's going to be like even greater for this direction of like programmatic AI. There was going to be like a gold rush on top of us for because like software is super charged, but I think there's going to be a gold rush parallel to us as well on like all the different ways we can expose things to make software more powerful so people can make even cooler stuff and then we are back to like early internet energy, you know, and I think that's why like you know the Twitter is just like Jeff Jeff, you know, um, it's like—

Host

这很鼓舞人心,因为它与我们习惯的如此不同,那就是‘对不起你不能这样做,但我们做缩放定律,只有大实验室才能做。’对。

It's inspiring because it's so different than what we're used to, which is 'I'm sorry you can't do this, but we do scaling laws and only the big labs can do it.' Right.

Diogo

实际上,如果我可以岔开话题,如果你不介意。我想你可能会喜欢这个。

That actually, if I'm going to tangent if that's okay. I think you might enjoy this.

Host

我们已经岔开五个话题了。

We're like five tangents in this.

Diogo

是的,我在所有岔开的话题中迷失了。

Yeah, I get lost on all my tangents.

Host

这对听众来说会很难搞清楚,但他们会搞清楚的。

This is going to be horrible for the listeners to figure it out, but they're going to figure it out.

Diogo

是的,我们可以编辑这个。所以,Discord 上人们一直问我,我还没时间解释的热门事情是,为什么我反对安全对齐,为什么我们不拒绝?我不反对安全作为原则,但我认为安全对齐通常与用户不一致。而拒绝就像显然是类型错误。

Yeah, we could edit this. So, popular thing on Discord that people keep asking me, I haven't had the time to explain it yet, is why am I opposed to safety alignment and why do we not refuse? I'm not opposed to safety as a principle, but I think that safety alignment is generally misaligned with users. And refusal is just like obviously a type error.

软件依赖中的拒绝 Refusals in software dependencies

Diogo

就像你作为一个人,在跟一个瓶子什么的聊天,你在云端写代码,然后出现了一次拒绝——比如“抱歉,我无法读取 DNA.py”——那真是烦人的时刻。很烦,对吧?但你可以应付,而且因为斯德哥尔摩综合征,你被迫应付。我也有这方面的故事。我需要在这里再深入跑个题。但如果你把这个东西放进一个后台运行的依赖里,如果它拒绝了呢?如果别人在用那个依赖呢?他们不知道那个系统是什么。你想让软件因为用户发了一条奇怪的消息就随机崩溃吗?那简直是疯了。这来自那些不懂软件、不懂编程的人,他们痴迷于这个“无马马车”式的 AI 同事,而不是挖掘 AI 的全部力量。

Like if you're a human being and you're chatting with a bottle or whatever, you're Claude Coding and a refusal happens — like 'I'm sorry, I can't read DNA.py' — that's an annoying time. It's annoying, right? But you can work with it, and you're forced to work with it because of Stockholm syndrome. I have stories about that too. I need another tangent deep in here. But like if you ever want this in a dependency running in the background, what happens if that refuses? What if someone else is using that dependency? They don't know what that system is. You want the software to just stochastically break because a user sent a weird message in there? That is straight up insanity. It's coming from a place of people who do not understand software, do not understand programming, and they are obsessed with this horseless carriage of an AI co-worker instead of unearthing the full power of AI.

Host

有道理。你想要的是那种核心内核,到处都能用。

Fair enough. You want something that is the core kernel that is usable everywhere.

Diogo

对,没错。就像认知内核,对吧?你需要这个东西非常通用,针对它的用例高度优化。你希望它能适用于所有未来的用例,人们正在做的所有奇怪的事情。我们显然没有训练过那些东西。它居然能用,奇怪吗?不奇怪,因为我们训练过更奇怪的东西,朋友。所以,但再跑个题,关于安全对齐——在我看来,安全对齐对产品是有意义的,比如 ChatGPT 和 Claude。安全对齐和能力对齐的区别在于,能力对齐是关于做用户想要的事。这对软件工程师来说太棒了——他们希望自己的东西做那件事,而且它越可预测,他们就越不需要测试和摆弄。现在还远没到那一步。它可以做到,但我们需要多得多的可靠性,才能让它好到像数据库查询一样,你甚至不用想它。当你需要智能时,它就在那里。但安全对齐就像指令遵循的反面。它是当你想遵循别人的指令时,比如 OpenAI。没错。这对产品来说也很有意义——比如 ChatGPT 应该做你。如果他们不想和 ChatGPT 做一些不适合工作的角色扮演,那是他们的事,因为也许他们的用户有父母和孩子,想要那样。那没问题。但在 API 里,那太疯狂了,对吧?那完全不可接受,因为人们需要围绕这个编程,那太反用户了——我可以是个愤怒的人。所以我应该试着冷静下来。

Yes. Exactly. Like the cognitive core, right? And you need this thing to be so general, so optimized for its use cases. You want it to work on all the future use cases, all the weird stuff that people are doing. We obviously didn't train on any of that stuff. Is it surprising that it works? No, because we trained on weirder stuff, my friend. So but one tangent up about safety alignment — safety alignment makes sense for a product in my opinion, for like ChatGPT and Claude. What makes safety and capability alignment different is capability alignment is about doing what the user wants. That is sick for software engineers — they want their thing to do the thing, and the more predictable it is, the less they have to test and play around with it. It is not anywhere close to that yet. It could be, but there are so many more nines of reliability that we want in order to make it so good — like a database query that you don't even have to think about. It is just there when you need intelligence. But safety alignment is like the opposite of instruction following. It's when you want to follow someone else's instructions, like OpenAI. Exactly. And this makes a lot of sense for a product again — like ChatGPT should do you. If they don't want to do some not safe for work roleplay with ChatGPT, that's on them, because maybe that's what their users who have parents and kids want. That's fine. But in an API, that's nuts, right? That's completely unacceptable because people need to program around this, and that is so anti-user that it — I can be an angry person. So I should try to calm down.

安全对齐与API使用 Safety alignment and API use

Host

人们理解你的热情,我觉得这很好。我要给你的一个反驳是,如果我们用它来杀人呢,对吧?那才是真正的不适合工作的东西。私人的、个人的,随便。但没错,我们会在战争中使用它,这是公司可以合理希望他们的 API 不被用于的事情。

It's people get your passion and I think it's really good. The one push back I'll give you is like what if we use it to kill people, right? Like that is the actual not safe for work thing. It's private, personal, whatever. But like yes, we will use it in war and that is something that companies can reasonably prefer their APIs not be used for.

Diogo

我明白。我认为在某些务实的场合可以持有这种观点。我不认为通用技术的基础是那种地方。就我个人而言,我是否更希望我们的东西不被用来杀人?显然。我是否更希望它被用于世界上各种伟大的事情?显然。我会为此施加影响吗?会。但我会在技术层面这么做吗?绝对不会。因为那会分裂智能。每次你需要它过度拟合某些奇怪的东西,你就在越来越多地分裂它的智能。而这些东西是分裂的——它们现在分裂得太厉害了。而且,对我来说,我认为智能更像数据库而不是同事。我不认为数据库有责任检查它们是否被用于不太好的事情。你知道,比如 CIA——其实我不知道 CIA 到底做什么。你可以想象,杀死甚至不是坏人的人之类的。

I get that. I think that there are pragmatic places where that opinion can be held. I don't think the foundation of a general purpose technology is that place. Personally, would I prefer that our stuff is not used to kill people? Obviously. Would I prefer it's used for all sorts of great stuff in the world? Obviously. Will I put my thumb on the scale for that? Yes. But will I do it at the technological layer? Absolutely not. Because that will fracture the intelligence. Every single time you need it to overfit to some weird stuff, you're fracturing its intelligence more and more. And these things are fractured — they're so darn fractured right now. And as a furthermore thing, to me it's like I think intelligence will be more like a database than a coworker. I don't think it's up to databases to add checks on whether or not they're used for something that's not great. You know, like the CIA — actually I don't know what the CIA does really. You can imagine killing people who are not even bad or whatever.

Host

嗯。

Mhm.

Diogo

我不认为数据库有责任为此负责。而且,对我来说很奇怪的一件事是,当人们在 Slack 上注册我们的东西,他们说:“嘿,我们要部署这个。我们能部署这个东西吗?”我就想,兄弟,我们是 API,你是开发者。这不关我的事,对吧?你甚至不应该知道整个任务是什么,因为它应该被分解成小事情。我们不应该能够知道下游用户在做什么,这是一个很好的边界,给软件工程师最大的权力。理想情况下,他们用它做好事,理想情况下我们可以帮助他们,我们谈过做开源和慈善之类的。我们现在绝对没有时间做其他任何事情。但只要我负责,他们就不会从技术层面得到任何那种偏见。

And I don't think it's the database's responsibility for that. And furthermore, a thing that has been weird to me is when people sign up for our thing on Slack and they're like, 'Hey, we're going to deploy this. Can we deploy this thing?' I am just like, my brother, we are an API, you are a developer. It's none of my business, right? You shouldn't know what the whole task even is because it should be decomposed into small things. We shouldn't be able to know what the downstream users are doing, and that is a good boundary to give software engineers maximum power. Ideally, they use it for the good stuff, and ideally we can help them, and we've talked about doing open source and charity and all of that. We have absolutely no time for anything else right now. But they will get any of that bias out of the technological layer as long as I'm in charge.

隐私与使用条款 Privacy and terms of use

Host

是的,那很好。既然谈到这个话题,我们也简要谈谈你的隐私政策、使用条款,这引起了一些误解。我想先澄清一下。我觉得你可能只需要两句话,比如你不会——你对你的 API 并没有那么严格限制。显然,从意识形态上讲,你非常认真地对待你作为平台的角色。

Yeah, that's great. While we're on the topic, let's also briefly talk about your privacy stuff, terms of use, which got a little bit of misunderstanding. I just want to clarify that up front. I think it probably takes two sentences from you about like you will not — you're not being that restrictive about your API. Like clearly ideologically you take your role as a platform very seriously.

Diogo

是的。是的。我不知道你指的是什么,但就像这个——我看到过一些关于基准测试的事情。显然我们没有阻止人们——哦天,我应该小心我说的话。我意识到了。

Yes. Yes. I don't know what you're referring to, but like this was — I've seen a couple of things about like benchmarking. Like obviously we're not stopping people from — oh man, I should be careful about what I say. I'm realizing.

Host

不,你公开说过,那是在预览期。你在发布时没有移除它,现在你要移除它。

No, you said you said it publicly that that was in the preview period. You didn't take it out for the launch and now you're going to take it out.

Diogo

团队在做一些我甚至不知道的事情。所以很高兴知道团队沟通了那件事。我让他们和律师确认一下。显然我们没有阻止人们做那种事情。我极其支持。所以我极其反对公开基准测试。我极其支持——对于作为代理的私有基准测试,我持中等态度。

The team is doing stuff that I'm not even aware of. So it's great to know the team communicated that. I asked them to check in with the lawyers about that. Like we are obviously not stopping people from doing that type of thing. I'm extremely in favor. So I'm extremely anti-public benchmarks. I'm extremely in favor — I'm medium about private benchmarks that are proxies.

Host

所以你担心饱和,还是担心在公开基准测试上训练?所以就像很容易作弊。

So are you worried about saturation or like training on public benchmarks? So it's like easy to cheat.

Diogo

不仅容易作弊,还有很多——所以我认为我们或者任何与我们竞争的人,你模糊地知道,就像你可以说——

Not only is it easy to cheat, there's a lot of in — so I think that we are or anyone who's like competition with us that you know vaguely there is like you could say like —

Host

我们有 50 个杰夫克隆。是的。

We got 50 Jeff clones. Yeah.

Diogo

嗯,当然。

Well, sure.

竞争与智能价值 Competition and the Value of Intelligence

Diogo

假设两年后存在竞争,或者干脆假设有一个行业,人们在做和我们类似的事情。我们卖的是每美元或每秒的智能。人们痴迷于成本和速度。我觉得那很酷,但真正重要的是智能。成本和速度是坏事——你付了钱,需要拿回东西,而智能才是真正重要的。智能的问题在于它有一种说不清道不明的东西,对吧?就像好模型的“气味”——我们发布后两小时发生的事情,实际上比视频传播得还广,就像“天哪,这真的能用”。除此之外,发布本身也很疯狂,人们能真正感受到我们有多在乎这件事,而这正是我认为长期的关键。我认为公开基准测试与此背道而驰。它们是一种让人们信任智能的方式,因为智能有一种说不清道不明的东西,但公开基准测试极其、极其容易被钻空子。即使他们努力不这样做,他们仍然会。过去,每个实验室都有一个团队收集看起来像 MMLU 的数据来让它看起来更好,这不过是多此一举的基准测试。所以我相信,从长远来看,需要的是感觉和信任,直到你把它放入工作流中,针对那个工作流进行评估和测量,并对它在真正重要的工作流上的表现有自己的判断。我们的工作是不断推进可靠性的“九个九”。这是我们作为公司需要一直做的事情,我们需要尽一切努力让人们知道这是我们非常关心的事情。如果我们想,我们本可以在一年半前发布 Jev,如果我们想犯傻的话。

Let's say there is competition, or let's just assume there's an industry two years from now of people doing similar things to us. The thing we are selling is intelligence per dollar or per second. People obsess about cost and speed. I believe that's cool, but the thing that matters is the intelligence. Cost and speed are bad things—you're paying for something and you need the thing back, and intelligence is what truly matters. The problem with intelligence is that there's a je ne sais quoi to it, right? Like the good model smell—the thing that happened after we launched, like two hours later, that actually went way bigger than the video, which was like, holy—this is actually usable. Beyond that, the launch was crazy and people could really sense how hard we care about that, and that's truly what I think the long term of this is. I think public benchmarks are antithetical to this. They are a way to get people to trust in intelligence because intelligence has a je ne sais quoi, but public benchmarks are extremely, extremely gameable. Even if they try not to, they still will. Back in the old days, every lab had a team to collect data that looks like MMLU to make it look better, which is just benchmarking with extra steps. So I believe that in the long run it needs to be vibes and trust until you put it into a workflow and evaluate it for that workflow and measure it and have your own sense of how it does on the exact workflow that matters. And our job is to keep moving the nines of reliability. This is an ever-present part of what we need to be doing as a company, and we need to do everything to have people know that this is something we care so much about. If we wanted to, we could have released Jev like a year and a half ago if we wanted to be dumb.

苦涩教训与数据 The Bitter Lesson and Data

Host

哦,就像 bitter lesson,对吧?就像架构和——

Oh, like the bitter lesson, right? Like architecture and—

Diogo

是的,既然你在这里提到了,我就说说。

Yeah, I'll bring it up since you talked about it here.

Host

太好了。

Hell yeah.

Diogo

Sutton 说算法大致上胜过算力。数据显然比算力重要得多,而做正确的任务、拥有一个北极星,是最难也最重要的事情。这在 LLM 领域已经发生了两次,也许 2.2 次。有 RLHF,它把任务转向了指令遵循。没人意识到那是可能的。RLVR 对方向做了微小的调整,现在轮到我们 RLCD,我们有了一个新任务,目标是程序在环。是的,数据重要得难以置信——我怎么强调都不为过。

Sutton says that algorithms beat compute very roughly. Data matters way more than compute, obviously, and doing the right task, having a north star, is the hardest and most important thing. This has happened in LLM land twice so far, maybe 2.2 times. There's RLHF, which shifted the task to instruction following. No one realized that was possible. RLVR did a tiny little edit to the direction, and now us, RLCD, we have a new task, and the goal is programs in the loop. And yeah, data matters so, so unbelievably much—I can't emphasize it enough.

Host

是的,你把自己看作数据实验室而不是模型实验室?这是你们用的说法吗?

Yeah, you consider yourself a data lab rather than a model lab? Is that the wording you guys use?

Diogo

我们是——我们会一直非常关心数据。对我来说,模型能力就意味着数据。数据复杂得难以置信,而正是它带来了“九个九”。你根本不知道数据能改变一切到什么程度。数据太重要了。是的。

We are—we will always care so much about data. To me, model capabilities means data. Data is so unbelievably complicated, and that is what gets nines. You have no idea how much data can shift everything. Data is so important. Yeah.

Host

天哪。

Holy crap.

Diogo

所以如果有人在找工作,我们正在招聘无限多的数据人员。真的是无限。

So if people are looking for a job, we are hiring infinite data people. Actually infinite.

Host

一个好的数据人员是什么样的?显然是那种关心通读转录的人,比如你刚才说的——你们所有数据都是合成的。但这只是表面,对吧?不是合成了就完了。合成,但我们有品味高、非常用心的人查看这些,指出问题,回去重新生成。这就是如今一个好的数据人员吗?

What is a good data person like? Clearly somebody who cares about reading through the transcripts of, for example, whatever you said—that all your data is synthetic. But that's only scratching the surface, right? It's not like synthetic, so what? Synthetic, but we have people with a lot of taste and a lot of care looking at these, articulating what's wrong, going back, regenerating. Is that what a good data person is these days?

Diogo

让我试着理清——这超级复杂,我实际上用一场讲座来培训数据人员,我估计那比这期播客最终的长度还要长。所以我试着说个大概。第一:我们不做人们以为的那种合成数据——实际上第零:数据和合成数据取决于你的任务。你数据的形状、你任务的形状决定了数据。RLVR 的数据是环境,对吧?RLHF 的是人类反馈。每个任务都有自己独特的数据,我们当然也有自己独特的数据。所以第一,我们有那个。第二:我们不想在用户数据上训练的原因,即使我们可以——我们现在可能可以要求任何条款,它会,我不知道会不会有影响。我们真的不想要那个。因为无论如何,真实世界的数据有太多偏差。人们问同样的事情存在幂律,你最终会过拟合,会分裂等等。第二,我们瞄准的是多年后一个完全科幻的未来,这些模型将成为通用基础设施层,一层又一层深入堆栈,直到人们无法想象的东西。我喜欢把我们的模型想成:LLM 是 UDP,我们的模型是 TCP。各种东西都可以建在上面。我们需要能够搞定那些未来用例,让软件开发者真正能构建那些未来东西。而做到这一点的方法是,即使我们拥有当下的所有数据,我们也只会过拟合当下,然后就行不通。我们需要的是——几乎感觉他们像艺术家。他们研究这个认知核心。我们的认知核心比任何人的都更平滑。然后他们找到不平滑的地方,然后以外科手术般的方式解决——你永远无法完美做到,对吧?但他们这样做,能在每一个可能的维度上解决它。

Let me try to figure out how to—it's super complicated, and I literally onboard the data people with a talk that I assume is longer than this podcast will end up being. So I will try to say the high level of it. Number one: we don't do the kind of synthetic data that people think—well, actually number zero: data and synthetic data depends on your task. The shape of your data, the shape of your task changes the data. RLVR's data is kind of environments, right? RLHF's is the human feedback. Each task has its own unique kind of data, and we of course have our own unique kind of data. So number one, we have that. Number two: the reason why we don't want to train on our users' data even if we could—we could probably ask for any terms right now and it would, I don't know if it would make a difference. We truly don't want that. Because no matter what, real-world data has so much bias. There's a power law of people asking the same things, where you'll end up overfitting to it and fracturing to it and all of that. And number two, we are aiming for a complete sci-fi future years from now where these models are going to be the general infrastructure layers and layers and layers deep down the stack to things people can't even imagine. I would like to think of our model like UDP as LLMs and TCP as our models. All sorts of stuff can be built on top of that. And we need to be able to nail those futuristic use cases such that software developers can actually build that futuristic stuff. And the way to do that is even if we had all of the data of the present, we would just overfit to the present and then it wouldn't work. What we need is—it almost feels like they're artists. They study this cognitive core. Our cognitive core is way less jagged than anyone else's. And then they find the jaggednesses and then they address them surgically in a way that—and you can never perfectly do this, right? But they do it in such a way that it addresses it in every single possible dimension.

Host

一般情况而不是精确的——

General case rather than the exact—

Diogo

而这每次都需要大量的智能。

And that requires a lot of intelligence every time.

RLCD与秘密来源 RLCD and Secret Sources

Host

好的,我们提到了一点——你有点批评我的思维受 RLVR 影响很深,这很公平。我们来说说 RLCD。显然你有一些秘密来源。据我所知,你从未真正发表过关于它的论文或类似的东西。

Okay, so we mentioned a little bit—you sort of criticized my thinking as very RLVR-influenced, which is very fair. Let's actually mention RLCD. Obviously you have some secret sources. To my knowledge, you've never actually published a paper or anything like that on it.

Diogo

不,还没有。

No, not yet.

让术语通俗易懂 Making jargon accessible

Host

但人们应该从这得到什么?你能不能给大家一些信心,让他们相信你不是为了听起来酷而编造术语?对我来说,校准(calibration)是一个我觉得已经很好理解的概念,因为我们在播客里讲过,但你说 RLCD 时,我不知道你指的是什么,跟人们熟悉的东西相比。

But what should people get from this? Can you give people some confidence that you're not just making up jargon for the sake of sounding cool? One thing for me is calibration—I do think that's well understood because we've covered it on the podcast—but I don't know what you mean when you say RLCD versus what people are familiar with.

Diogo

这是个好问题,实际上我会给一个相关的问题:什么是 RLHF?

It's a great question, and actually I will give a related question: what is RLHF?

Host

对,实际上 RLHF 意味着多个不同的东西。有最初的 RLHF——我想是 Paul Christiano 教机器人后空翻之类的。不是有那个吗?

Right, and actually RLHF means multiple different things. There's the RLHF of the original—I think it was like Paul Christiano teaching a robot to backflip or something like that. Wasn't there something?

Diogo

是那个吗?

Was that it?

Host

那是原始的 PPO 论文,但我不确定。

That was the original PPO paper, but I don't know.

Diogo

如果我没记错,PPO 不一定来自人类反馈。好吧。但我相信那像是 OpenAI 的对齐工作,可以教难以指定的输出,比如后空翻。我不是 100% 确定。然后实际上还有学习总结。这是一群帮助 Instruct 并共同撰写指令遵循论文的团队的工作,那是在语言模型上做 PPO。

And PPO is not necessarily from human feedback if I recall. Okay. But I believe it was like an OpenAI alignment work that could teach hard-to-specify outputs like a backflip. I'm not 100% sure. And then there was actually learning to summarize. This was work by a bunch of the team that helped with Instruct and co-authored the instruction following paper, which was doing PPO on language models.

Host

这是——抱歉,我在试着操作这个东西。这是 2017 年。

This is the—sorry, I'm trying to manipulate this thing. This is 2017.

Diogo

是的。我不是 100% 确定,但看起来挺对的。如果有一个机器人做后空翻之类的,那可能就是它。

Yeah. I'm not 100% sure, but that looks quite right. If it has a robot doing backflips or something like that, that might be it.

Host

是的。好的,酷。我想我猜对了。太棒了。

Yes. Okay, cool. I guess I got it right. Hell yeah.

Diogo

这就对了。是的,就是那个。

There you go. Yeah, that's the one.

Host

所以想法是你能用它做定义不明确的事情吗?所以那像是版本一。版本二是 OpenAI 做的学习总结工作,实际上是在语言模型上做 PPO 来做一些有点定义不明确的事情。这是人们称为 RLHF 的另一件事,我没有共同撰写。哦,Dario 在那里。酷。

So the idea was can you do ill-specified things with it? So that's like version one. Version two was the learning to summarize work that OpenAI did, which is actually PPO on language models to do something somewhat ill-specified. This is another thing that people refer to as RLHF, which I did not co-author. Oh, Dario's there. Cool.

Diogo

太棒了。

Hell yeah.

Host

还有 Rafford。

And Rafford.

Diogo

是的。是的。是的。向 Alec 和 Ryan 致敬。爱他们。但我所指的 RLHF 是——哦天哪,

Yeah. Yeah. Yeah. Shout outs to Alec and Ryan. Love them. But the thing that I refer to as RLHF is the—oh man,

Host

我对此有评论。

I have comments on that.

Diogo

我对那篇论文有评论,但我们跑题太深了。

I have comments on that paper, but like we're so at tangents deep.

Host

是的。所以真正让我触动的是,我称之为 RLHF 的是指令遵循的任务。不是关于 PPO,那部分不重要。而是关于设定一个北极星,这是一个有价值的方向。有点像苦涩教训的北极星,对我们来说 RLCD 是这个新任务。

Yeah. So the thing that really got to me, the thing that I'm calling RLHF is the task of instruction following. It's not about the PPO, that part doesn't matter. It's about setting a north star of this is a valuable direction. It's kind of like the bitter lesson north star, and for us RLCD is this new task.

Diogo

而且它不是——我不认为这是术语。我试图精确地沟通。只是说嘿,这是另一个北极星,就像 DPO 及其所有后代也做 RLHF,尽管没有使用那篇论文中的算法。所以清楚地陈述北极星是支持可编程 AI,这是我真正抓住的一个词,移除循环中的人类,因为 RLHF 是为了这个而调优,以便你可以自动化一切。

And it is not—I don't see it as jargon. I try to communicate with precision. It's just that hey, here's another north star just like DPO and all of its descendants also do RLHF despite not using the algorithm in that paper. And so clearly stating the north star is being pro-programmable AI is one word that I really catch on to, removing the human in the loop, because RLHF is tuning for this so that you can automate everything.

Host

是的。

Yes.

Diogo

我在北极星的论点中遗漏了什么吗?

Did I miss anything else in the thesis of what the north star is?

Host

有——那是正确的。我在沟通中过于细致。一个细微差别是我们需要务实。我们需要意识到语言模型能做得很好,就像 AI 能做什么,对吧?可能有非常酷的程序化类型,但如果技术还没准备好,那就太傲慢了。

There is—that is right. I am overly nuanced in my communication. The one nuance is that we need to be practical. We need to be aware of what language models can do really well, like what AI can do, right? There could be programmatic types that are sick AF, but if the technology is not ready for it, it's arrogant.

Diogo

不,不,不。我强烈相信你。

No, no, no. I strongly believe you.

Host

酷。听起来很傲慢,但我在有公司之前很久就有这种感觉了。

Cool. It sounds arrogant, but I felt this way since long before I even had a company.

Diogo

我可以证明。

I can vouch that.

Host

我已经谈论这个大约 3 年了。

I've been talking about it for like 3 years.

Diogo

是的,我已经谈论这个很久了,我一直在说,因为我认为它会更容易。他们说他们做事不是因为容易,而是因为他们认为它容易。类似这样。我以为整个项目会花一周。我大错特错。所以我非常抱歉对 OpenAI 的每个人,我以为我在解决这个问题。但我认为悲剧的是——好吧,我认为过度承诺和交付不足也是悲剧,AI 在这个轴上非常极端,我认为 RLVR 是主要的——好吧,RLVR 和 RLHF 都是极端的肇事者。但对我来说,那里有太多潜力。AI 显然非常聪明。我喜欢在我的演讲中问人们,AI 怎么能如此聪明?我们怎么能解决数学中的千禧年大奖难题,但仍然不能自动化最基本的工作?非常基本的写作之类的东西,不需要极其聪明的人来做。这不是一份令人满意的工作。这些人可以做其他事情,但我们仍然需要他们做这些超级基本的不令人满意的事情,因为我们还不能自动化它。但我们有这个超强的自动化引擎,只是没有正确的插头之类的东西来插入所有有经济价值的工作。如果整个 TypeSafe 公司消失,也许需要一两年人们才能真正赶上。我实际上不知道需要多久。如果模型质量重要,那么我们将长期处于非常有利的位置,但它已经完成了,对吧?这已经改变了技术历史的路径。我们将作为一个领域探索那个空间。

Yeah, I've been talking about this for so long and I've been saying it because I thought it would have been easier. They say they do not do things because they're easy, they because they thought it was easy. So something like that. I thought this whole project would take a week. And I was unbelievably wrong. So I am so sorry to everyone at OpenAI that I thought I was like man I'm solving this right now. But I think that the tragic thing is when—well I think overpromise underdeliver is tragic too and AI is super extreme on that axis and I think RLVR is like the main—well both RLVR and RLHF are extreme perpetrators of this. But to me it's like there's just so much potential there. AI is clearly so smart. I love this in my talks, when I ask people like how can AI be so unbelievably smart? How can we solve millennium prize problems in math but still not automate even the most basic of works? Really basic wrote stuff that doesn't take extremely smart people to do this. It's not a satisfying job. There's other things these people could be doing, but yet we need them to do this super basic non-unsatisfying stuff because we can't automate it yet. But we have this supercharged engine of automation that just does not have the right plugs and stuff to plug into all of this economically valuable work. And if the whole company of TypeSafe disappears like maybe it'll take a year or two for people to truly catch up. I actually don't know how long it'll take. If model quality matters, then we are going to be in a very good position for a long time, but it's done, right? This has changed the path of technological history. And we will be exploring that space as a field.

Host

是的,我绝对同意。你创造了可能性。所以我想如果我能转述一下,让人们也能理解:你不应该把 TypeSafe 和 Jeff 的成功当作,好吧,那是一种新模型类型,现在我们完成了,我们回到业务。不,实际上还有五种其他模型类型你应该探索,让千花齐放。

Yeah, I definitely agree with that. You've created possibilities. So I think if I can paraphrase so that people can also understand: you should not take the success of TypeSafe and Jeff as just like well, that is a new model type, now we're done, we go back to business. No, actually there are like five other model types that you should be exploring and let a thousand flowers bloom.

Diogo

绝对,其中一些你可能也会早期互联网能量。我认为它回到了科技乌托邦。不再像哦天哪,有时我的编码代理工作,但所有最好的都被内部囤积,对吧?就像创造又回到了菜单上。不过,这将是一个疯狂的世界。系好安全带。我对此非常兴奋。

Absolutely, and some of which you will probably also early internet energy. I think it's back to tech utopia. It's no longer like oh man, sometimes my coding agents work, but all of the best ones are hoarded internally, right? It's like creation is back on the menu. Though, it's going to be a wild ass world. And buckle up. I'm so jazzed about that.

Host

我的意思是,现在你有资金和动力去做你设想的任何事情。我认为在这么长时间说这些事情之后,看到你拥有这些非常令人欣慰,但实际上向世界展示。

I mean, and now you have the funding and the momentum to do whatever you envision there. Which I think is very gratifying to see you have after so long of saying these things, but actually show the world.

Diogo

我知道。一直是个挑逗者真是太有趣了。

I know. It was such an interesting thing to be a tease the whole time.

秘密总体计划 The Secret Master Plan

Host

感觉我的演讲像是个悬念,因为我没有说自动化会如何发生。

Like my talk felt like it was a cliffhanger because I didn't say how the automation would occur.

Diogo

是的。

Yeah.

Host

Sean 审阅了我们的宣言,他说这些部分有点模糊,你知道,第一步是什么,智能是什么。

Sean reviewed our manifesto and he's like it's a little bit vague in these parts and you know like what's step one, what is the intelligence.

Diogo

我问你要模型,你说模型快来了,嗯,我主要反对“可组合”这个词,但 Bill Pron 说得太棒了。谢谢。我们真的团结在这上面了。我想我们不完全像某些公司那样是个邪教,但我们确实对 Butter 在做的事很兴奋,我的品牌就是务实,我们都超级务实。真的很棒。

I asked you for model and you were like yeah model coming and like well I just I mainly objected to the word composable but Bill Pron got is fantastic. Thank you. I we've really rallied around that. I'd like to think we're not entirely a cult like some companies are, but like we are like jazzed about what Butter doing and like we are like my brand is being practical and like we are all like so super duper practical. It's really great.

Host

是的。所以这里,顺便说一下,这就是秘密大师计划,对吧?机器原生可组合 AI 的形态。

Yeah. So here and by the way here here's the uh the step the the secret master plan, right? The shape of machine native composable AI.

Diogo

制作秘密大师计划是你的主意。

It was your idea to make a secret master plan.

Host

这是埃隆的做法。当他创办特斯拉时,他说:“这就是我们要做的。”

It's a it's a Elon thing. when he started Tesla, he was like, "Here's what we'll do."

Diogo

是的。我正式把功劳归给你。

Yeah. I I'm giving official credit to you.

Host

谢谢。谢谢。谢谢。但是,你知道,你应该告诉我你也要做这个模型发布,因为你告诉了我一半的故事,另一半你当时没有 Doom 演示。你没有任何数字给我。我当时想

Thank you. Thank you. Thank you. Um but like uh you know, you should have told me you're you're also going to do this model launch cuz you like you told you told me half of the story and then the other half you didn't have the Doom demo at the time. You didn't have any numbers to give me. I was like

Diogo

问题是我相信基准测试最大化。没错。所以,这是你需要感受的东西,我认为这是建立长期信任的方式,尽管它伤害了我们很多,你知道,就像去年我们融资时,没人相信我们,他们只想要基准测试之类的,我们说我们不会那样做,我们有原则,我们会坚持立场,那会奖励不良行为者,我不在乎,你知道,你想要什么,这就是我们,我们坚持这一点,所以抱歉

well the problem is I don't believe in benchmaxing. Exactly. Right. So like it is a thing that you need to feel and like I think that this is the way to build long-term trust even though it like hurt it hurt us a lot you know like like last year when we did fund raise no one believed us you know like like and they wanted just benchmarks and stuff and we're like we're not going to do that we are principled we're going to stand by our guns that rewards bad actors I don't give a you know like what do you want like this is who we are and we are standing by that so sorry

Host

是的。嗯,在某些方面,我认为选择艰难的道路,但最终你会创造出你想工作的公司。

Yeah. Well, in some ways I think like choosing the hard path, but you end up making the company that you want to work in.

Diogo

对。否则,如果你出卖自己,那你就像在 OpenAI 工作,但和我的人一起,对吧?那就是

Y right. Otherwise, if you sell out, then you're just working in like Open AI but with my people, right? Which is

Host

是的。是的。我的意思是,我对此没有太多遗憾,显然它结果好得难以置信,你知道,我昨晚谈到我离开 OpenAI 的原因时很情绪化。因为发布后我实际上不得不改变我的措辞。我的措辞是,如果 AI 寒冬真的发生,而我没有尽一切可能去避免它,我会认为自己个人有责任,既包括 RHF 方向,我认为这真的扩大了过度承诺与交付不足,也包括没有全力投入这个,因为我认为这是价值将被印刷的地方。所以,真的很酷,因为我觉得我担心的 AI 寒冬被避免了。你知道,AI 将是有用的。它将用于自动化。不到一周,数字已经不可否认地表明它被用于真正的工作,这就像狂野的西部。是的。

Yeah. Yeah. I mean like I'm I I don't have too many regrets on that obviously like it worked out so unbelievably well and you know like I u I was emotional last night when I was talking about like the reasons I left OpenAI. Um and because like it actually had to change my wording after the launch. Um, my uh my my phrasing was if an AI winter did happen and I did not do every possible thing I could to like avert that, I would see myself as personally responsible both for, you know, the the RHF direction, which I think really widened overpromise versus underdeliver and also not going all in on this because I think this is this is where value is going to just be like printed. So, and it was really cool because I feel like the the AI winter I'm worrying about is averted. You know, like AI will be useful. It'll be used for automation. It's been less than a week and like the numbers are already undeniable that it's like being used for real work and like there's it's it's the wild west. Yeah.

Host

是的。你能,如果你能想起来,你看到了什么数字,比如注册量之类的,你能分享什么?我实际上并不是特别了解一切。是团队在告诉我所有这些事情。

Yeah. Can you sh just just if you have top of your head, what numbers are you seeing like what's what's like signups like what whatever you can share? I'm actually not super on top of everything. Like the team is the ones who are telling me all of these things.

Diogo

我相信它每天都在变化,对吧?

I'm sure it's like changing every day, right?

Host

这有点疯狂。

It's it's kind of nuts.

Diogo

如果有一个里程碑,你就像,是的,这是我们希望的一件事。我们达到了。

If if there's a milestone that you're like, well, yep, that's one thing we're hoping for. We reached it.

Host

我会说我们通过的一个里程碑是每日 token 数。

I will say a milestone that we've passed is tokens per day.

Diogo

嗯,这不是短暂的每日 token 数。这就像即使在晚上也在不断运转。所以,你知道,机器在调用它,而不仅仅是人们在尝试。所以那太酷了。每天一万亿 token 是很多。嗯,所以超过那个很棒。注册对我来说并不重要。实际上,如果我完全诚实的话,这是我们犯的一个小错误。推特上的人称我们为营销天才之类的。嗯,那只是我们。我们没有营销人员也在招聘。嗯,我们只是做我们真诚的、傻气的、不敬的自己。我们只是努力地把人们从等待名单中移除。嗯,我们的平台团队非常厉害。我认为我们的正常运行时间比 Anthropic 更多,同时拥有最前所未有的发布。这有点疯狂。所以向他们致敬。嗯,我们没有意识到的事情,第一,等待名单,等待名单注册对开发者平台来说不重要,在我看来。你知道,我猜其中很大一部分甚至不是开发者。所以他们进去,尝试一些查询,很多人不明白,因为他们不是在编程,对吧?他们就像,什么?这不是聊天机器人,我的 ChatGPT 在哪里,对吧?但如果,我没有精确计算过。我的感觉是,如果世界上每个人只写几个查询,那与一个高级用户的 for 循环相比,那只是个舍入误差,就像创造价值。我们没有意识到等待名单的事情是,我们只是把任何人从等待名单中移除。没关系。可怕的部分是速率限制。然后一旦人们开始从中获得价值,嗯,他们只想要大量的速率限制,因为这就是软件,对吧?就像你前期花精力指定你的任务,然后这个任务创造的价值超过投入,然后现在你有了,是的,你在后台运行它,你让它成为其他东西的依赖,你可以做更高层次的东西,嗯,你就在世界上创造了这么多价值,你知道,早期互联网的人可能没有想象到 2000 年代早期互联网的奇迹,那还不是早期互联网,但就像,无意冒犯,可组合性,所有疯狂的事情都发生了。我只是真的想在我们的宣言中强调这一点。我们追求涌现。我们追求成为催化剂。我们想要赋予人们权力,我们将为此做任何我们能做的。无论是像在我们的市政厅 Discord 中我穿着垃圾袋还是不穿。

Um, and this is not like fleeting tokens per day. This is like even at night like it's constantly churning. So, you know, machines are calling it and not just people trying things out. So that is that is so cool. A trillion tokens a day is a lot. Um so surpassing that is awesome. Signups to me don't really matter. And actually this was like a bit of a mistake we made if I'm like totally honest. People on Twitter were calling us like marketing geniuses and all of that. And um that was just us. We don't have a marketer also hiring. Um and we were just being our genuine goofy like irreverent selves. And we were we were just like offboarding people off the wait list so hard. Um our platform team is so unbelievably cracked. I think we have more up nines of uptime than anthropic while having the most unprecedented launch ever. Like that is kind of nuts. So like props to them. Um and uh the thing we didn't realize so number one weight lists weightless signups don't matter for like a developer platform in my opinion. you know, uh I would guess that a large number of them are not even developers. So they go in, they try some queries and a lot of people don't get it because they are not programming, right? Like they're just like, what? This is not a chatbot where where's my chat GPT2, right? But if like I I haven't exactly calculated this. My sense is that if every single human being in the world like just wrote a couple of queries, that would be a rounding error compared to like one power users for loop that is just like creating value. And the thing we didn't realize with a weight list is like we just off offboard anyone off the weight list. It doesn't matter. The scary part is rate limits. And then once people start getting value from that, uh then they just want tons and tons of rate limits because this is what software is, right? Like you spend effort up front to specify your roach task and then this wrote task creates more value than it takes to put in and then now that you have exactly yeah you run it in the background you make it a dependency to like other things you can make like higher level stuff and uh like you just create so much value in the world you know early internet people probably did not imagine like the wonder of early 2000s internet which is still not early internet but like it's it's through no offense composability um that all of the crazy stuff happens. I just really wanted to emphasize that in our manifesto. We are going for emergence. We are going for like being the catalyst. We're wanting to empower people and we are going to do whatever we can for that. Be it like Discords in our town hall with me wearing a garbage bag or not.

Host

嗯,还有播客,你知道,因为我想做长篇,对吧?就像,是的,我们会越过一些表面的事情,然后我们会深入,人们会真正信任和理解你的使命,你知道,那些会共鸣的人最终会加入你,或者,你知道,购买你,不,抱歉,作为客户,作为客户。

Um and uh and podcasts and and you know getting getting like cuz I want the long form, right? It is like yes, we'll get past some of the superficial things and then we'll go deep and people will really trust and understand your mission and like you know the the people that uh will resonate that will end up joining you or or you know uh uh buying you uh no sorry as as a as a customer as a customer.

Diogo

是的。是的,那很有趣。抱歉。

Yeah. Yeah, that was funny. I'm sorry.

Host

抱歉。我不是那个意思。

Sorry. I didn't I didn't mean to say that.

发布视频观看量与NeoLab标签 Launch video views and the NeoLab label

Host

但没人有像这样亮眼的版本——你的发布视频有 3600 万次观看。

But no one has a version as flattering as this — 36 million views on your launch video.

Diogo

酷。现在到 3800 万了。是啊,我就是个舍入误差。Navia Stokes 有 74,Fable 5 有 57。我没统计最初的 ChatGPT,因为它没有视频。

Cool. Up to 38 now. Yeah, I'm a rounding error. Navia Stokes got 74, Fable 5 got 57. I didn't do the stats for the original ChatGPT, which had no video.

Host

所以你排在前列,对吧?如果要在 2026 年创办一家 NeoLab,我觉得你现在是第一,这相当疯狂。

So you're up there, right? If you were to launch a NeoLab in 2026, I think you're number one right now, which is pretty crazy.

Diogo

是啊。其实我更愿意——我确实有那件 T 恤,“你最喜欢的 NeoLab 最喜欢的 NeoLab”。我根本不在乎是不是 NeoLab。我觉得作为 NeoLab——我们有很多恶搞 NeoLab 的周边。其中一件是“有产品的 NeoLab”,其实那并不是 NeoLab。我真的不在乎这个。我在乎的是成为一个可靠的开发者平台。所以感谢这个比较,但希望我们能超越它们,回到开发者的革命性时刻,以及这个人们可以依赖和信任的稳定事物。

Yeah. Well, I actually would rather — I do have the shirt, "your favorite NeoLab's favorite NeoLab." I don't give a damn about being a NeoLab. I think being a NeoLab — we have a lot of swag that's a parody of a NeoLab. One of them I have is "NeoLab with product," which actually is not a NeoLab. I don't care about that really. What I care about is being a reliable dev platform. So I appreciate the comparison, but hopefully we transcend past them and we go back into a revolutionary moment for developers and this stable thing that people can rely on and trust.

超越正常运行时间的可靠性 Reliability beyond uptime

Host

是的。为此,我觉得你们让我印象深刻的一点是,你们确实谈论可靠性。我以为主要是关于校准,我们谈过 RLCD,但其实也关乎正常运行时间和可扩展性等等,对吧?它们都——

Yes. To that end, I think one thing that really impressed me about you guys is that you do talk about reliability. I thought it was mostly about calibration, which we talk about RLCD, but actually it's also about just uptime and scalability and all those things, right? They all sort of —

Diogo

九个九。就像——

Nines. It's like —

Host

哪个时间在我的——

Which time is in my —

Diogo

但那是其中的一部分。可靠性还在于这个东西有多智能,它有多一致地做你想让它做的事。我认为大型推理模型在我看来非常聪明,但它们仍然缺乏可靠性。我认为有很多用例,它们看起来应该足够聪明来自动化工作。有经济激励去自动化那项工作,但它们仍然不够可靠,连实习生都不如,因为它们针对不同的事情做了优化。所以我认为有能够信任输出的可靠性,也有我们尚未达到的可靠性维度,这让我非常兴奋。我想先自动化简单的工作,再自动化困难的工作。我觉得这是常识。但对我来说,我们会足够——我不知道是否存在“足够可靠”这回事,但我想做到如此之好,以至于人们甚至不需要试用模型就知道它能行。这就像编程中的心流状态,对吧?我只是写查询,因为我需要这里的智能,对于非平凡的分支,我可以直接写在类型安全的系统里,一个查询,然后得到结果,它就能准确分支。那该多好啊。那就是梦想,而那将是一场漫长、漫长的跋涉。

But that's part of it. But there's reliability in how intelligent the thing is, how consistently does it do the thing that you want. I think the big reasoning models are very smart in my opinion, but they still lack reliability. I think there's many use cases where they look like they should be smart enough to automate their work. There is economic incentive to automate that work, yet still they're not reliable enough as an intern because they're optimized for different things. So I think there's the reliability of being able to trust the outputs, and also there are dimensions of reliability that we are not yet at that I'm so excited by. I want to automate the easy work before the hard work. I think that's just a common sense thing to do. But to me we will be sufficient — I don't know if there's such a thing as sufficiently reliable, but I want to get so good that people don't even need to try the model to know that it'll work. That's like what flow state is in programming, right? I'm just writing queries because I need intelligence in here, and for non-trivial branching I can just write it in a type-safe system, one query, and then get the results out and it just branches accurately. That would be so good. That is the dream, and that is going to be a long, long slog.

确定性vs鲁棒性 Determinism vs robustness

Host

是的。我们稍后会讨论你的 API 设计,给大家一些例子,也许还有未走之路之类的。一开始我确实好奇的一个关于可靠性的问题是,我注意到没有种子,所以基本上相同的输入——我总是得到相同的输出吗?如果不是,为什么?

Yeah. We're going to go into your API design in a little bit just to give people examples and maybe path not taken, that kind of stuff. One thing up the front that I do wonder about in terms of reliability is I noticed that there's no seed, and so basically same input — do I always get the same output? If not, why not?

Diogo

哦,好问题。这其实是我们常遇到的一个问题。可靠性其实是个笼统的说法——每当 AI 无法自动化某件事,都是因为某种形式的可靠性问题。可能是类型安全,可能是确定性,也可能只是它参差不齐。所以可靠性是个笼统的说法。我认为它也是北极星的笼统说法。确定性就是相同输入,相同输出。我确实觉得这对单元测试有点意思,但我认为那是错误的北极星。我认为鲁棒性才是人们——我不想告诉人们他们真正想要什么,因为那会显得我有点傲慢。我认为那是更重要的属性。你希望给定相似的输入,得到相似的输出。而语言模型有多不可靠,真是有点疯狂。我们测试的一种方式是,你在提示里放入 UUID,放入小的——我想它们叫 nonce——你希望从所有这些中得到相似的输出,因为它们在语义上确实是同一个问题。而那就是你真正想要鲁棒性的地方——人们就是在 AI 做决策时被坑的。所以我认为那是一个超级重要的属性。我们也可以有确定性——就我能为程序员在脑中建模的范围而言,那是可以提供的。它对某些用例可能有价值。所以请在评论中教育我,或者你——但总的来说,这很容易:确定性是你可以为了更好的成本而权衡掉的东西。我们一直想要处于每美元智能的前沿。为了达到那里,我们正在做绝对恶心的事情。这是——我不该说这个,但这里没人阻止我。

Oh great question. So this is actually a common question we have. Reliability is actually a catch-all — whenever AI can't automate something, it's due to some form of reliability. It could be type safety, it could be determinism, it just could be that it's jagged. So reliability is a catch-all. I just think that it's also a catch-all for what the north star is. Determinism is same inputs, same outputs. I do believe that this is slightly interesting for unit tests, but I believe that to be the wrong north star. I believe robustness is what people — I don't want to tell people what they really want because that would be a little arrogant of me. I believe that is the more important property. You want given similar inputs, get similar outputs. And it's kind of wild how unreliable LMs are. A way that we test this is you put UUIDs in, you put little — I think they're called nonces — in the prompt, and what you want is similar outputs from all of those because it's truly semantically the same question. And that is the part where you really want that robustness — that's where people get burnt with AI making decisions. So I think that is a super duper important property. We could also have determinism — that is a thing that can be available as far as I can mentally model for programmers. It could be valuable for some use cases. So please educate me in comments or you — but in general, it's easy: determinism is something you can trade off for better cost. We are constantly wanting to be on the intelligence per dollar frontier. We are doing absolutely disgusting things to be there. This is — I shouldn't say this, but no one's here to stop me.

公司文化与题外话 Company culture and tangents

Host

你知道,如果你自己批准自己的 PR——

You know, if you sign off on your own PR —

Diogo

这家公司不是这样运作的。我认为这周,我的幕僚长 Kay 是科技界最有权势的人。

That is not how it works at this company. I believe for this week, my chief of staff, Kay, is the most powerful person in tech.

Host

向 K 致敬,组织了这次活动。

And shout out to K for organizing this.

Diogo

天哪——她如此能干和强大。她太不可思议了。我是说,她很烂。别挖她。所以我试着稍微过滤一下,但人们告诉我别叫它模型中的弗兰肯斯坦怪物,因为那有负面含义。我认为弗兰肯斯坦的怪物在这整个故事里是好人——我是说,它是无辜的,对吧?我没读过。好吧,我承认。好吧,那——什么表情我不——

Holy — she is so competent and powerful. She's incredible. I mean, she sucks. Don't poach her. But so I tried to be a bit more filtered, but people are telling me don't call it a Frankenstein's monster of models, because that has negative implications. I think Frankenstein's monster was the good guy in this whole — I mean, it was innocent, right? I didn't read it. Okay, I'll confess. Okay, that — what one facial expression I don't —

Host

如果你想看的话,有部不错的 Jacob Vorti 电影——

Decent Jacob Vorti movie if you want to see —

Diogo

我——不管怎样,那个改编——

I — the adaptation anyway —

Host

你不知道我现在时间有多紧。我的优先事项是睡觉。

You have no idea how little time I have right now. My priorities are sleep.

Diogo

开发者,开发者,开发者。

Developers, developers, developers.

Host

开发者,是的,开发者,开发者,开发者。

Developers, yes, developers, developers, developers.

Diogo

但是的,为了处于每美元智能的帕累托曲线上,我们正在做绝对恶心的事情,而且我们会继续这样做。我们会做疯狂的事情。我认为人们真的需要跳出框框思考。令人惊讶的部分原因是——人们按框框思考,而我们继续这样做。截至目前,我们显然在这方面是最好的,我们想继续在整件事上保持最好。

But yes, we do absolutely disgusting things to be on the Pareto curve of intelligence per dollar, and we are going to keep doing that. We're gonna be doing crazy-ass stuff. And I think people really need to think outside of the box. Part of the reason why surprising is — people thought inside the box and we continue to do that. As of right now we are obviously the best at this and we want to continue being the best at that whole thing.

回到种子与确定性 Back to seeds and determinism

Host

是的。等等,我们是从哪里跑题的?

Yeah. So wait, where did we tangent from?

Diogo

所以我问了你关于种子和确定性的问题,然后你基本上定义了可靠性以及你如何看待它。

So I asked you about will you have seeds in determinism and then you basically define reliability and like how you see it.

Host

是的。

Yes.

确定性模型与每美元智能 Deterministic Models and Intelligence per Dollar

Diogo

我有一个关于鲁棒性的例子,很快。我可以给你看。

I have a robustness example that's real quick. I can show you.

Host

我喜欢这个。我就说一点。如果人们能说服我们那是一件有价值的事,而且我们没有巨大的 GPU 短缺,我们是可以做一个确定性模型的。我们可以很乐意地做所有这些模型。我们活着就是为了取悦大家。

I love that. I'll just say one thing. We can make a deterministic model if people can convince us that that is a valuable thing to do and we don't have a gigantic GPU shortage. We can happily make all of these models. We live to please.

Diogo

你会推翻一切,只不过你做得很好看。

You'll throw over everything except you do it in a nice way.

Host

所以确定性是有可能的。它只是让你每美元买到的智能更少。

So determinism could be on the cards. It just gets you less intelligence per dollar.

Diogo

对。嗯,光是看到 OpenAI 的发展轨迹,你现在就得信我,你会被同行压力逼着去做这件事。人们会想要它。就算你告诉他们不需要,他们还是会想要。所以这就是 TL;DR。

Yeah. Well, just having seen the trajectory of OpenAI, you will just trust me now that you will be peer pressured into doing it. People will want it. Even if you tell them they don't need it, they'll still want it. So that's the TL;DR.

Host

好。好。我很乐意,也许有一天我们会看到那会怎么发生。有人跟我说我——他们说我们品牌的一部分就是不可动摇,他们说那只是“固执”的好听说法。

Okay. Okay. I would love to, maybe one day we will see how that happens. I've been told I'm— they say that part of our brand is being unshakable, and they say that's just the nice way of saying stubborn.

Diogo

固执。对。没错。而我是个非常固执的人。我觉得我们本来做不到那件事。

Stubborn. Yeah. Exactly. And I'm a very stubborn person. I don't think we could have done that.

Host

对。不,但是就像我——好吧,但我以前跟你争论过。

Yeah. No, but so like I— okay, but I have argued with you before.

Diogo

对。

Yeah.

Host

而你在开发者这件事上每次都是对的。好吧。我放弃。你赢了。你赢了。我被说服了。

And you've been right about developers every time. Okay. I give up. You win. You win. I'm sold.

Diogo

不,不。我只是说,我觉得你可以坚持立场,同时如果我给你正确的证据,你也能抛掉自己的先验,然后说“对,这对我来说确实说得通”。所以在这件事上就相信你自己的直觉吧。

No, no. I'm just saying like I think that you can hold your ground while also, if I give you the right evidence, you can sort of throw away your priors and be like, "Yep, that actually makes sense to me." And so just trust your own gut on this.

Host

不过我怀疑我们会在很长很长一段时间里受 GPU 制约。任何每美元智能更少的东西,都意味着同样的智能要消耗更多 GPU。我们的目标不是去签约公司——那是有价值,但我们的目标是让人们去实验、去做奇怪的事,我们需要把它交到尽可能多的人手里,为此掀起一场加州淘金热。

I suspect though that we will be GPU constrained for a very, very long time. And anything that has less intelligence per dollar means it consumes more GPUs for the same intelligence. Our goal is not to onboard companies—like it's valuable, but our goal is to have people experiment and do weird things, and we need to get it to as many hands as possible and start the California gold rush for that.

Diogo

我觉得现在就是。对。

I think there is right now. Yeah.

模型质量与版本承诺 Model Quality and Versioning Promises

Diogo

只是提个醒。我就说出来,因为现在有人正在想这件事,那就是当你说“我们不会承诺确定性模型。为了每美元智能我们会不惜一切代价,而我们正面临 GPU 制约”时,人们会想你可能要量化你的模型。就像你发布时的模型,你可能会量化降级来降低质量,以便腾出内存或带宽之类的。所以你可能应该做出某种承诺——你不必现在做——比如“我们会在发布时维持模型质量”。你在 OpenAI 时发布所有这些 API,甚至 Claude 也是,当它们首次发布模型时,模型字符串并不总是保持同一个模型。你的模型有版本管理,这很好,但你应该公开承诺某种东西,比如一旦某个东西发布了,我们就不改它。

Just a word of caution. I'll just say it because somebody is thinking about it right now, which is when you say things like, "We will not commit to deterministic models. We will do whatever it takes for intelligence per dollar and we are facing GPU constraint," people are thinking you may quantize your models. Like whatever you had at launch, you may quantize down to reduce the quality in order to free up memory or bandwidth or whatever. And so you should probably have some kind of promise, which you don't have to make now, about like, "We will uphold model quality at launch." When you were at OpenAI when you launched all these APIs, and even Claude as well, when they first launched the models, the model strings did not stay the same model at all times. You have versioning in your models, that's great, but you should publicly commit to some kind of like, once a thing is launched, we don't change it.

Host

我们不会在部署模型之后改它们。那太疯狂了。我们在乎开发者。如果你做那种事,那才说得通。再说一次,这就是第一方产品和 API 并存的问题。在第一方产品里你想怎么做都行。随他们去。只要能带来那种体验,都没问题。有了 API,你显然不能那么做。但我要说,我们计划推进的速度比很多人习惯的模型提供商快得多。所以我们会以比人们想象快得多的速度发布新模型,而且我们不承诺对模型提供长期支持,因为我们认为还有很多改进空间。所以存在一种可能,我们会暂时对现在的 Jev 1.13.0 做 LTS。我们可能会这么做,因为用的人太多了,而且我知道开发者讨厌破坏依赖。另一种选择是分裂我们的机队,那对所有人来说都是一种很糟糕的氛围。

We will not change our models when we deploy them. That is insane. We care about developers. It makes sense if you're doing something like that. Again, this is the problem with a first-party product and an API. You can do whatever you want in a first-party product. More power to them. Whatever gets that experience, that is fine. With an API, you obviously can't do that. But I will say that we plan to move a lot faster than many people are used to model providers doing things. So we will be launching new models a lot faster than people think, and we are not promising long-term support for the models because we think that there's lots of improvements to have. So there is a world that we might temporarily LTS what is right now Jev 1.13.0. We might do that because so many people are using it and I know developers hate breaking dependencies. The alternative is fracturing our fleet, and that is a very bad vibe for everyone.

Diogo

就是有 100 个不同版本的模型。

Have like 100 different versions of the model.

Host

没错。如果我们迭代得非常快,也会有很多这样的版本。所以我们确实想要最终不只是有一个 LTS 支持的东西——长期支持。我们想要一种非常酷的做法。我们在这方面有一些研究在酝酿,我觉得那会是有史以来最亲开发者的东西。但它还不是我们当前的模型,我也不承诺我们能保持完全相同的模型。它们每次肯定会变得更聪明。我的感觉是,即使是我们那些已经聪明的模型迭代,模型版本之间的变化往往比把字符串模型调用两次还小。但当我们从“参差不齐”变成“哇”的时候,那才是大变化所在。

Exactly. And if we're iterating very fast, there would be a lot of those versions as well. So we do want to have not just an LTS-supported thing eventually—long-term support. We want a really sick way of doing that. We have research stuff cooking in that direction and I think it's going to be the most pro-developer thing ever. But it is not yet our current models and I'm not promising that we will be able to keep the exact same models. They will get smarter every time for sure. And my sense is that even our model iterations where it already is smart, between model versions the changes tend to be even smaller than calling the string models twice. But when we go from like jagged to like wow, that is where the big deltas are.

移植到其他芯片与每秒智能 Porting to Other Silicon and Intelligence per Second

Diogo

LTS 模型的一个美妙之处是,其实你还可以把它们移植到其他芯片上。我不知道你有没有想过这个。

One thing that's beautiful about LTS models is that actually you can also port them to other silicon. I don't know if you've thought about this.

Host

无可奉告。

No comment.

Diogo

好吧。

Okay.

Host

所以我在乎每美元智能。

So I care about intelligence per dollar.

Diogo

对。

Yes.

Host

对。

Right.

Diogo

但还有速度。

But speed.

Host

什么?

What?

Diogo

速度也算?

Speed as well?

Host

我们走着瞧。

We'll see.

Diogo

对。

Yeah.

Host

我们走着瞧。

We'll see.

Diogo

我是说,这是推理技术树的一整个部分,在过去一年里爆发式增长,对吧?你可以迁移到 Cerebras 和 Etched 之类的上面,获得 10 万倍加速。

I mean, this is a whole part of the inference tech tree that is exploding in the past year, right? That you can move to a Cerebras and Etched or whatever and get 100,000 times speed up.

Host

对。我觉得每秒智能是一个不同的指标,我们甚至谈过像每美元每秒智能这样的指标。我对 Jevan 悖论出现、或者至少 Jev 系列模型的猜测。我在内部喜欢追着人问的一件事是,我不在乎它聪明多少。它必须处于前沿之前。这就是 Jev 这个品牌的意义。它在每美元智能和每秒智能上都是最好的。我们走着瞧。我觉得这是一件很有意思的事。我知道有很多行业极度依赖实时性的东西,对他们来说每秒智能意味着大把的钱,但我们走着瞧。我很想两者都做,让市场以任何方式纠正我。我很想被人告知。

Yeah. I think that intelligence per second is a different metric and we've even talked about things like intelligence per dollar times second and metrics like this. My guess on Jevan's paradox occurring, or at least the Jev series of models. And the thing I like to hunt people down about internally is, I don't care how much smarter it is. It needs to be in the pre-frontier. So that is what the brand of Jev is. It is the best thing at intelligence per dollar for intelligence per second. We'll see. I think that it's an intriguing thing. I know that there's many industries that are extremely dependent on real-time stuff and they will like—intelligence per second means tons of dollars for them, but we'll see. I would love to do both and have the market correct me either which way. I would love to be informed by people.

Diogo

对,完全同意。

Yeah, totally.

实时延迟与规模 Real-time latency and scale

Host

而且这不仅仅是实时性的问题,对吧?还关乎规模,因为在规模上,每一微秒都会被乘以数十亿、数万亿倍。

And it's not just about real time, right? It's also about scale, because at scale every microsecond is just multiplied by billions and trillions.

Diogo

这取决于它运行的后台环境。如果是像数据库 map-reduce 查询这样的大型后台任务,延迟可能不如从中获取智能的成本那么重要。但如果它更像是面向用户的实时应用,你就有 100 毫秒到 1 毫秒之间的预算,那完全是魔法般的。实际上,即使你低于 100 毫秒,如果你把时间减半,那就意味着你可以获得双倍的智能,或者进行顺序智能调用,从而获得非凡的体验。所以这绝对是正在发生的事情。这超级酷。我喜欢每秒智能的用例,但我不认为那会是 Jev 的利基市场。

It depends on the background it's running in. If it's a big background like a database map-reduce query, the latency might not matter so much as the cost to get intelligence from it. But if it's something more real-time like user-facing, you have budgets between 100 milliseconds and 1 millisecond that are totally magical. And actually, even if you were below 100 milliseconds, if you half that time, that means you can get double the intelligence or sequential intelligence calls to have a phenomenal experience. So that is definitely happening right now. It is super duper cool. I love the intelligence-per-second use cases, but I don't think that will be Jev's niche.

更快更便宜的权衡 Faster and cheaper trade-offs

Host

好的。是的。有道理。好的。当思考更快更便宜的承诺时,通常其他模型提供的权衡是更快但更贵。

Okay. Yeah. Fair enough. Okay. When thinking about the promise of faster and cheaper, typically the other trade-offs that other models are offering is faster but more expensive.

Diogo

对。

Right.

Host

所以我认为 Jev 之所以引起如此大共鸣的原因之一,是你做到了更快但更便宜的那个象限,这个象限非常非常空缺,同时保持智能大致恒定。

And so one of the reasons I was thinking about why Jev is resonating so much is that you've done the faster but cheaper side of the quadrant, which is very very unoccupied while holding intelligence somewhat constant.

Diogo

是的。嗯,这是一个非常关键的陈述:同时保持智能恒定。那才是难点,对吧?

Yes. Well, that's a very load-bearing statement: while holding intelligence constant. That's the hard part, right?

Host

不幸的是,基本上你拒绝任何公开基准测试,或者你不喜欢任何关于它的公开基准测试。

Which unfortunately, so basically you refuse to have any public benchmarks or you don't like any public benchmarks about it.

Diogo

而我需要一些内部感知。

And I will need some internal sense.

Host

再说一遍。

Say it again.

Diogo

你需要一些对此的内部感知。

You need some internal sense of this.

Host

哦,当然。当然。

Oh, of course. Of course.

Diogo

我们肯定有自己的内部评估。但需要很多纪律才能不钻空子,而且必须把不钻空子作为最高优先级。我们当然会这么做,对吧?否则我们怎么能保证我们的模型在每美元智能方面领先?我们不会在那里盲目飞行?如果我们用不同的成本或其他东西做完全奇怪的事情,我们怎么比较它们?我们绘制它们并试图找出对用户最好的。所以我们肯定会测量它们。我不反对测量。但当你有任何替代激励时,这极其危险。这是我用铁腕统治的一件事——也许我的同事会认为我用铁腕统治很多事情,但对我来说,不自欺欺人地了解我们的模型有多聪明是最重要的事情之一。我们需要追求真理。

We have our own internal evals for sure. But it takes a lot of discipline not to game those, and it needs to be a top-level priority to not game them. Of course we do that, right? How else can we make the guarantee that our models are in the prede intelligence per dollar? That we're not flying blind in there? If we're doing completely weird things with different costs or whatever else, how do we compare them? We plot them and try to figure out what is the best for the users. So we for sure measure them. I'm not anti-measuring. But it's extremely dangerous when you have any alternative incentive. And this is the one thing that I kind of rule with an iron fist—well, maybe my co-workers might think I rule many things with an iron fist, but to me, not fooling ourselves about how smart our model is is one of the most important things. We need to be truth-seeking.

Host

是的。是的。同意。同意。

Yeah. Yeah. Agreed. Agreed.

API原语与命名 API primitives and naming

Host

嗯,好的,我想过一下 API 选择的一些细节,主要是因为这是唯一会问你这些问题的播客。所以你有三个原语。嗯,choice、score、no。首先,no 是从哪里来的?这是文献中的术语还是什么?

Um okay I wanted to go over some details on the API choices mostly because this is the only podcast that will ask you these kind of questions. So you have three primitives. Um choice score no. First of all no where is that from? Is this like a term in the literature or what?

Diogo

呃,现在它是了。我们为此争论了很多。我们为此争论了很多。它,你知道,它像布尔值,对吧?像真或假。它是——

Uh now now it is. We debated this a lot. We debated this a lot. It is, you know, it is boolish, right? Like true false. It is—

Host

但它是连续的。

But it's continuous.

Diogo

是的。没错。哦,首先这个名字的起源是伯努利。

Yes. Exactly. Oh so first the origin of the name is Bernoulli.

Host

啊。

Ah.

Diogo

是的,所以这就是为什么它甚至拼写得那么奇怪。它像是伯努利名字的一个子集,来自伯努利概率,对吧?实际上就是那样。所以这就是它的起源。我们为此争论了很多。我们喜欢 peool,我们喜欢 pool。我们想叫它 pool party,但没人让我。我们有一堆其他争论,然后我们认为 new 是最好的。我们的理由,实际上 Jev 也是一样,是我们认为我们是一群不敬的疯子,程序员不在乎,你知道,如果 Jev 只是一个字符串,我们没想到它会流行起来,甚至双关语之类的,对吧?实际上,内部对这个名字有很多仇恨。他们都道歉了,除了一人。

Yes, so that's why it's even spelled that weird way. That is like a subset of the name Bernoulli from like a Bernoulli probability, right? Which is actually what that is. So that is the origin of it. We were debating this a lot. We liked peool, we liked pool. We were wanting to call it like a pool party but then no one let me. We had a bunch of other arguments about that and new we figured was like the best thing. Our rationale, and this is actually the same thing with Jev too, is that we think that we are like an irreverent insane bunch and programmers don't care, you know, like if Jev is just going to be a string, we didn't expect it to catch on or even have puns or anything like that, right? Actually, there was a lot of hate on the name internally. They've all apologized except for one person.

Host

仍然坚持。

Still holding strong.

Diogo

是的。呃,我们共同的朋友。嗯,是的。是的。是的。是的。

Yes. Uh, our mutual friend. Um, yes. Yes. Yes. Yes.

Host

我为此尊重她。

I respect her for that.

Diogo

是的。是的。她想让 Jev 叫 Meow。

Yeah. Yeah. She wanted Jev to be called Meow.

Host

她当然会。

She would, of course.

Diogo

是的。当然。

Yes. Of course.

Host

好的。你赢了。你赢了。

Okay. You win there. You win there.

Diogo

但就像,是的。Renewal 是我们必须为这个东西做一个新概念,因为如果它是一个布尔值,会让人困惑。所以实际上这三个都是新概念。这些不是编程中存在的类型,这是故意的,因为它们非常接近类型,但不完全是。分数不是整数。所以如果你有像 instructor 或 pydantic 之类的,把整数或浮点数映射到分数,你会有点被搞糊涂,你知道,我们真的在清晰度方面犯错,而不是让人们容易理解发生了什么。

But like, yeah. Renewal is we had to make a new concept for this thing because if it was a bool it would be confusing to people. So actually all three of these are actually new concepts. These are not types that exist in programming and that was intentional because they map very closely to types but they're not quite that. A score is not an int. So if you had like instructor or pydantic or whatever map ints or floats into scores, you'd get a little bit cooked, you know, and we were really erring on the side of clarity over the side of like making people like easily understand what's going on.

Host

我的意思是,你不担心吗?你不希望东西直接集成到人们已经在使用的东西中吗?

I mean, don't you worry about that? Don't you want things to integrate directly into things that people are already using?

Diogo

是的。是的,我们确实希望。实际上我认为你知道——

Yes. Yes, we do. And actually I think that you know—

Host

你有与其他 SDK 之类的集成,但你有,抱歉,你有自己的 SDK。

You have integrations with like other SDKs and stuff but you have sorry you have your own SDKs.

Diogo

但通常例如作为开发者关系人员,我会非常痴迷于,是的,这里是如何使用 Jev 与 instructor,这里是你知道的那种东西。

But typically for example as a developer relations person I would be very obsessed with like yes here is how you use Jev with instructor here's how you know that kind of stuff.

Host

我们可能在某个地方有,我太落后了。

We might have that somewhere I'm so behind on everything.

Diogo

社区里会有人为你做的。不是说你成功了人们就会说哦那很酷。

Someone will do it for you in the community. Not that you're successful people will be like oh that's cool.

Host

但就像你知道——

But like you know—

Diogo

我也不认为那是二元的。实际上我把成功看作一个分数,总有更多可以攀登的,比如我们能多为社区服务,只是说清楚,我这一节很有压力,因为我没有复习文档,它们不断变化,但对我来说分数确实存在。所以分数类似于 LM 评判,对吧?所以如果你想叫它判断,我想你可以,但那是人们已经使用这种类型的方式,对吧?也许 new 可以像概率,但对我们来说一切都是概率,而 choice 实际上最接近函数调用,但函数调用是一种极其恶心的东西,如果你想 OpenAI juice sauce tea,我们稍后再回到那个,就像 choice 只是暴露 switch match 语句的正确方式。是的。在内部。

I don't see that as binary either. I actually see success as a score and there's always more to climb in like how much we can like be there for our community just to be clear and I'm this section is stressful because I didn't review the docs and they're constantly changing but to me scores do exist. So scores are similar to like LM judging, right? So like if you want to call it like a judgment I guess you could but that is like the way people already use this type of thing right like maybe a new could be like a probability but everything for us is a probability and a choice is actually closest to a function call but a function call is like an extremely disgusting thing that if you want OpenAI juice sauce tea that we should go back into that later like a choice is just like the right way of exposing like a switch match statement. Yeah. Within.

Host

所以它干净地映射到枚举。

So it like maps cleanly to an enum.

Diogo

是的。

Yep.

Host

如果你愿意,你可以选择将其水合为函数。

And you can choose to hydrate it into a function if you want.

Diogo

是的。而枚举选择是其中重要的部分。

Yes. And like the enum choice is the important part of that.

编程原语与结构化JSON Programming Primitives and Structured JSON

Diogo

实际上我认为这些都可以映射到编程原语上,比如 choice 映射到枚举上的 switch 语句。规则映射到 if 语句,分数映射到排序或大于/小于的阈值判断。这一直是我们的愿景:会有更多类型,它们都会映射到编程原语。

And actually I think these map all to programming primitives, where choice maps into a switch statement on an enum. Rules map to if statements, and scores map to sorting or thresholding at a greater than or less than. And this has always been the vision: there will be more types, and they will map into programming primitives.

Host

是的。嗯,还有其他细节你想过一遍吗?这 literally 是给那些决定真正投入 Jev 的人看的。你是专家,对吧?我只是想为他们提供更多关于 API 选择的背景。你知道他们应该如何使用这些功能,比如 legends confidence,在你的测试中有多关键,你知道,就是任何你想提供给人们在这个层面上的专业建议。

Yeah. Um, any other nuance you want to go through? Literally this is for the Jev people who are deciding to really invest in Jev. You are the expert, right? I'm just wanting to provide more background for them on API choices. You know how they should use some of these things like legends confidence, how critical in your testing, you know, just any sort of pro tips that you want to offer people when you're down at this level.

Diogo

谢谢。我喜欢这个。不不,这就是我们在这里的原因。

Thank you. I love this. No no, this is why we're here.

Host

太棒了。我没想到这个,实际上可能有好几个月没人问过我这个了,当我还在 onboarding 我们的 devrel 时。

Hell yeah. I didn't expect this and actually no one has asked me this in probably like months when I was onboarding our devrel.

Diogo

好的。嗯,太酷了。嗯,所以我们的模型是为未来深入计算机程序内部而设计的。我们毫不讽刺地相信,这将比人们今天所考虑的任何东西都要庞大得多。我们的模型可能还没准备好,但我们一直在为那个未来努力。它在这些浅层任务上永远不会足够好。嗯,抱歉,它永远不会——我们不会只是继续攀登浅层任务。我们想要深入程序的内部,因为那才是让软件强大的方式。我们模型内部的所有类型——这实际上是一个输出。嗯,但所有输入部分,比如状态、指令、标准,所有这些都可以是结构化的 JSON 对象。这样程序就可以把它们插入正确的位置,你不需要把它们放进模板里。

Okay. Um so sick. Um so our model is designed for being like deep in the insides of computer programs in the future. We like unironically believe that this will be much more massive than anything people are even considering today. And our model might not be ready for that, but we are like continuously working for that future. It will never be good enough at these shallow tasks. Um, sorry, it'll never be like we're not just going to keep on climbing the shallow tasks. We want to be deep in the guts of programs because that's how you make software powerful. All the types inside of our um this is actually an output. Um, but all the um all the parts of like the the input like the state, the instructions, the criteria, all of them can be structured JSON objects. That way like programs can like insert them in the right spot and you don't need to like put things into templates.

Host

没错。

Exactly.

Diogo

所以我觉得人们没有足够深入地阅读这部分,他们认为都是字符串,那也行。嗯,但这些都意味着——我会说,如果你在使用模板,比如把它变成系统消息之类的,你就是在用旧的方式思考。你知道,我们应该让计算机更容易理解,因为那种结构确实存在,对吧?就像在编程语言中,把所有数字都放进去,然后把它变成字符串,这会很奇怪。通常你只在有人的情况下打印时才这么做,对吧?但在计算机内部,你应该传递嵌套的、处处具有语义的结构。我们真的会优化我们的模型。模型对此已经相当优化了。但问题是,每个不同的嵌套层级都更难推理。我们真的在朝这个方向努力。我认为人们应该继续朝这个方向努力,因为它让代码更易读、更美观,而且与实现细节无关。就像这是我的状态,你知道,这是我的函数状态。把它想象成一个 AI 函数。我的状态的哪些子集,也就是你可用的所有变量,应该传进来?系统消息就像恶心的全局变量,你把所有东西都放进去,一次性放所有指令。然后你希望每条指令都能被准确执行,而不是并行地提问。

So if ever I think people don't read into this part enough and they think it's all strings and that's fine. Um, but these are all meant like like I would say that if you're using like a template like turning it into like a system message or something, you are thinking in like the old way, you know, we should be making things as easy for computers to understand because that structure is truly there, right? Like it would be weird in like like a programming language to have like all of your numbers in and then you pass it into like you turn it into a string. Normally you do that for printing when when you have a human in the loop, right? But for like within the computer, you want to be passing like nested structure that is semantic all around. And we are really going to be optimizing our model. The model's pretty optimized for this. But the thing is every different nested level of structure is harder to reason about. And we want we are really cooking hard in that direction. I think people should keep cooking that direction because it makes the code like so much more legible and beautiful and like agnostic to like the the implementation details. It's like here is my state, you know, like here's my function state. Like think of think of it as like an AI function. Which subsets of my state, which is like all the variables you have available, should I pass in here? System messages are like disgusting global variables where you just put everything in there and you put all the instructions at once. And then you know like you you hope that every single instruction gets nailed instead of asking the questions in parallel.

Host

好的。

Okay.

Diogo

还有,我建议——我真心这么说,不是从赚钱的角度。嗯,我真心建议问很多很多问题,把它们分解,让它们更小,真正地分解,不管模型今天能不能做到。我相信这周发生的事情最大的救赎将是人们的代码库,AI 代码库会变得好得多。你知道,如果你把问题分解成简单的决策,每一个都极其可评估,就像之前 AI 是一个大系统消息,然后也许你有另一个大 AI 来检查它是否真的做到了。这太疯狂了,你知道,有点疯狂。就像那是斯德哥尔摩综合征,对吧?但那有点疯狂。就像如果你想说不读这个子目录,或者不把任何 API 密钥传给 deepsek 或其他什么,那应该基本上可以通过程序保证。你知道,你永远无法保证任何机器学习模型,但通过分解,你实际上可以测量它,对吧?你可以验证它是否真的被调用了

And uh also I would recommend I and I truly say this not from like a like it makes me money perspective. Um, I truly recommend asking lots and lots of questions, break them down, make them smaller and like really decompose like no matter if the models can do it today or not. I believe that the biggest like um saving grace of like what's happening this week will be people's code bases, AI code bases are going to be so much better. You know like if you decompose problems into simple decisions every single one of these things is extremely evaluable like like a AI beforehand is big system message and then maybe you have like another big AI to see like if it actually does this. That's nuts, you know, it's kind of crazy. Like it's that was like Stockholm syndrome, right? But like that's kind of crazy. Like if you want to say like, hey, don't read this subdirectory or don't pass any API keys to deepsek or whatever else like that should be programmatically basically guaranteed. And you know, you'll never have guarantees of any machine learning model, but like by breaking it down, you can actually you can actually measure it, right? you can verify that it was actually called

Diogo

而且我们的模型,我们的模型,接口本身是如此可验证,这应该让人松一口气,你知道,它只会带来更好的工程

and like our model our model like the the interface itself is so verifiable this should be like a sigh of relief you know like it's it's a it's just going to lead to way better engineering

Host

是的,我想我明白了。嗯,所以你知道,过去人们不这样做的原因之一是他们只是调用一个小型 LLM,对吧,它仍然太慢,仍然太贵,相对于把所有东西分块。我自己也做过完全一样的事,对吧,我基准测试过,这里有一个管道,把所有东西填进系统提示,然后得到一个大的输出, versus 把它分解成 100 个不同的东西。它更慢,更贵,效果更差。

yep I think I get that um and and so you know one of the reasons people didn't used to do this in the past was because they would just call a a small LLM right and it's still too slow it's still too expensive versus chunking everything there I've done exactly this myself right like I I benchmark here's a pipeline that fills everything in system prompt and it just gets one big output versus break it down into 100 different things. It was slower, more expensive, not as good.

Diogo

是的。是的。是的。那会发生。是的。而且超级不方便。很笨重。为什么不直接放在一起?你最终会在问题之间重复一些东西。所以,可能效率不高之类的。但结果是非常难以依赖的东西。而软件

Yep. Yep. Yep. And that happens. Yeah. It's and it's like super inconvenient. It's unwieldy. Why not just put it all together? You kind of end up repeating some stuff between questions. So, it's like maybe like, you know, inefficient or something like that. But then it results in something that is very hard to rely on. And software

Diogo

不需要在后台运行。如果我们的东西不能在后台运行,我会心碎。

doesn't need to run in the background. It would break my heart if our stuff couldn't run in the background.

Host

有没有一种你们发现的分解方法,是有效的, versus 你们认为有效但实际上无效的?

Is there a way to break things down that you guys have found that works versus uh what you thought worked and doesn't work.

Diogo

有趣。

Interesting.

Host

因为人们会探索这个,你知道,既然你说了。他们会把这当作参考,然后说:“好吧,这就是我应该使用 Jeff 的方式。”

Because like people are just going to be exploring this, you know, now that you've said it. Like they would use this as a reference and be like, "Okay, like that's how I'm supposed to use Jeff."

Diogo

是的。

Yep.

Host

嗯,那么问题是如何分解?

Um then the question is how do you break things down?

Diogo

有趣。我喜欢把东西分解成最小的语义单元。

Interesting. I I like to break things down into its like its smallest semantic unit.

结构化提示与字面指令 Structured Prompting and Literal Instructions

Diogo

比如最低层级的东西是什么?我尽量从不——我可能是查询模型最多的人,我尽量在我的查询里,第一,这更像是我提示的方式,我把它做得非常非常有结构、非常明确,在问题里我总是——我喜欢用反引号,但它对所有都适用——非常清楚地说明我指的是什么,因为我们希望模型非常字面化。因为当你编程时,你希望指令被非常好地遵循。这就是编程的艺术,而 AI 所做的是扩展可以被遵循的指令种类。所以我是这么做的粉丝。我有时有点懒,会用更混合的方式,但我认为对于真正大型的生产项目,你只想不断添加更多问题。你想让添加更多问题变得非常容易,对所有这些分解非常非常精确,然后让代码具有你想要的确切行为。

Like what is the lowest level thing? I try to never have—I've probably queried the model the most among anyone, and I try to, number one, in my queries—this is a lot more like the way I prompt things—I make it really, really structured and explicit, and in the questions I always—I like the backticks, but it works for all of them—be really clear what I'm referring to, because we want the model to be really literal. Because when you program, you want things that instruction-follow really, really well. That is what the art of programming is, and what AI does is expanding the kinds of instructions that can be followed. So I'm a fan of doing that. I sometimes am a little lazy and I have more hybrid things, but I think that for really big production things, you just want to keep on adding more questions. You want to make it really easy to add more questions, be really, really precise about all of that breakdown, and then have the code have the exact behavior you want.

拒绝作为分解示例 Refusals as an Example of Decomposition

Diogo

如果我可以给一个小例子,就是拒绝,对吧?我不会谈论我们为什么不拒绝。我可能已经做过了。它就像一团模糊。但对于拒绝,我不认为你应该问“我应该在这里拒绝吗?”那真的——我认为答案会相当好,因为那是一个系统一兼容的任务。但我认为你最好问许多不同的独立问题,关于你可以拒绝的不同情况,因为不必只是根据你猜测,你可以实际上指定你想要什么,而且很漂亮。我认为这真的很美。如果你发现一种情况,比如,哦,它没有拒绝是因为这个原因——我没有指定任务的这一部分——那太棒了。这就是软件工程的意义。你通过添加那个问题、添加阈值、也许记住它作为测试用例来修复 bug,现在它就永远解决了。就像你的软件在提示中不会因为上下文腐烂而忘记它。它就在那里,你可以永远测量它。如果模型在某些事情上不完美,你可以根据真实例子为所有这些因素选择你想要的阈值。这就像没有 ML 的 ML。你可以对任何事情这样做。可能有些事情模型还不够好,对吧?比如我看到人们用模型做交易,像自动交易,我有点害怕。

If I could give a tiny little example of this, it's like refusals, right? Like I'm not going to talk about why we don't refuse. I might have done that already. It's like all blur. But for refusals, I don't think you should ask 'should I refuse here?' That's a really—I think the answer will be pretty good because that's a system one compatible task. But I think you're way better off asking many different independent questions about the different situations you can refuse about, because instead of having to just guess based on you, you can actually specify what you want, and beautifully. And I think this is truly really beautiful. If you find a situation where it's like, oh, it didn't refuse because of this reason—I didn't specify this part of the task—that is awesome. That's what software engineering is about. Like you fix the bug by adding that question in, adding the threshold, maybe remembering that as a test case, and now it is just solved forever. Like your software can't forget about that in the prompt because of context rot. It is just there and you can just keep measuring that forever. And if the models are not perfect at some of these things, you can choose what threshold you want for all of these factors based on real examples. It's like ML without the ML. And you can just do it for anything. And there might be some things the model's not good enough yet, right? Like I'm a little bit afraid when I see people doing trading with the models, like automated trading.

自动交易与安全权衡 Automated Trading and Safety Trade-offs

Diogo

它看起来很酷。我只是认为人们应该把它留给专业人士。而且那只是一个非常困难的高层次任务,也许模型还不够好,无法弄清楚。即使它们够好,那么它们会——突然就不会了,因为有效市场。但就像那些事情之一,你可以把它分解成东西并评估它们,你可能会说它在这方面不够聪明。也许我们还不为这个版本部署它,或者我们做出权衡,或者我们偏向安全,或者嘿,模型在检测这种奇怪组合方面不够好,比如讽刺与 VIP 客户的组合,这时我们升级到人类,这就是置信度估计的意义。

It looks cool. I just think that people should leave it to the professionals. And like that's just a very hard high-level task that maybe the models aren't good enough yet to figure out. Like even if they were, then they would—it suddenly wouldn't be because of efficient market. But like that's one of those things where you can break it down into things and just evaluate them and you might be like it's not smart enough at this. Maybe we don't deploy it yet for this version or we make a trade-off or we err on the side of safety or like hey the models are not good enough at, you know, like detecting like this weird combination of like sarcasm with a VIP customer that this is when we escalate to a human and that's what confidence estimates are about too.

校准与微调 Calibration and Fine-Tuning

Host

好的。非常好的回答。我想快速提一件事,我不指望你有太长的回答,就是——嗯,不,不,不,只是具体来说,你仍然依赖阈值作为用户可以拉的杠杆,但如果校准错了呢?对吧,就像你说你的校准是完美的,但——

Okay. Great, great answer. I think one thing I'll mention very quickly, which I don't expect that you have a too long of an answer for, is—well, no, no, no, it's just specifically like you are still relying on thresholding as the lever that the user can pull, but what if just the calibration is wrong? Right, like you're saying your calibration is perfect, but—

Diogo

我没那么说。我没那么说。

I didn't say that. I didn't say that.

Host

所以完美校准意味着较低的值是较低的概率,较高的值是较高的概率,但它可能是错的。它可能在局部上错位,所以那时我会想要微调它或什么的,对吧,而你们不提供。但你可以再次看到这是一个简短的回答,就是你现在没有它。哦,我们想提供微调吗?是这个问题吗?

So perfect calibration means like lower value is lower probability, higher value is higher probability, but it could be wrong. It could be locally misaligned, and so then I would want to fine-tune it or something, right, which you don't offer. But you could again see this is a short answer, which is you don't have it right now. Oh, do we want to offer fine-tuning? Is it the question?

Diogo

那可能是其中一种版本,或者你可以有一个不同的旋钮,对吧?因为就像现在你所说的只是,如果出了问题,技能问题,你应该再次更改提示或进一步分解,或者你更改置信度。那是我的两个选项,对吧?如果你的模型只是弄错了,那感觉不太令人满意。

That could be one version of it, or you could have a different knob, right? Because like right now you're all you're saying is like if something's wrong, skill issue, you should just change the prompt again or break it down even further or you change the confidence. Those are my two options, right? And that doesn't feel super satisfying if your model is just getting it wrong.

Diogo

是的。而且它会弄错很多事情,明确地说。我们有一个报告问题按钮。在 Discord 上向我们抱怨。我们想让它变得更好。每一个模型版本都会明显更好。如果它们没有取得大的改进,我们会很快停止发布它们。所以第一,那完全合理。我认为承认 AI 在某些事情上不完美是务实的,对吧?我确实认为我们会找到它们足够好的用例,而足够好取决于用例,对吧?就像人类可以做很多工作,尽管他们不擅长那项工作,因为他们的 EV 相当高,而且大概有了正确的阈值和一切,可能即使犯了错误,也有大量工作可以完成。关于微调的问题,我可以想象——我可以想象它在计划中。我确实有顾虑,因为在人们需要与人们想要的类别中。就像我认为通用模型往往真的——再次有通用性的精灵,让它擅长一百万其他任务而不是这个狭窄任务,可能会让它在这个任务的边缘案例中更好,我会有点害怕,你知道。是的。我可以想象它,这是我的答案。我在这些事情上无止境地务实。我想要一切——就像我的世界愿景是,我们想构建的东西太多了,但我也像不想发布一个像巨大脚枪的东西,像其他一些 AI 公司会做的那样。

Yep. And it will get many things wrong, to be clear. We have a report issues button. Complain to us in Discord. We want to make it a lot better. Every single model version will be noticeably better. We will stop shipping them quickly if they weren't getting big improvements. So number one, that is totally reasonable. I think that that's simply pragmatic to admit that AI is imperfect at some stuff, right? I do think we'll find use cases that they are good enough at, and good enough kind of depends on the use case, right? Like human beings can do a lot of work despite being bad at that work because their EV is quite high, and presumably with the right thresholding and everything there probably is large amounts of work that could be done even if mistakes are being made. On the question of fine-tuning, I could imagine—I could imagine it in the cards. I do have concerns because like in the what people need versus what people want category. Like I think general models tend to be really—like again there's the genie qua of generality that making it good at like a million other tasks than this one narrow task might make it better at edge cases in that task, which I would be a little bit afraid of, you know. Yeah. I could imagine it is my answer. I'm endlessly practical on these things. I want everything—like my vision of the world is there's so much we want to be building, but also like I would not want to ship something that is like a giant foot gun, like some other AI companies would.

微调回滚与开源 Fine-Tuning Rollbacks and Open Source

Host

嗯,所以 OpenAI 和 Claude,我想甚至 Gemini 都推出了微调,然后又收回了,这是一个有趣的观察,几乎微调现在都在开源模型的领域。

Well, so both OpenAI and Claude and I think even Gemini have rolled out fine-tuning and then took it back, which is an interesting observation that pretty much fine-tuning is now in the domain of open source models.

Diogo

是的。是的。我知道那件事,而且它有点糟糕。所以也许他们收回更好。

Yes. Yes. I do know about that and like it was kind of crap. So like that's probably better that they took it.

Host

它可能只是一个脚枪,告诉人们微调它可能是错误的方式是很好的。另一个有趣的答案可能是,嗯,我们的模型如此不同,就像你知道的,就像量化不适用于我们一样。

It could just be a foot gun and telling people that fine-tuning it is probably the wrong way to go is great. Another interesting answer could be that like well our model is so different, like you know in the same way that quantization doesn't apply to us.

微调与模型级联 Fine-tuning and model cascades

Host

输出 token 对我们不适用。微调对我们也不适用。

Output tokens doesn't apply to us. Fine-tuning also doesn't apply to us.

Diogo

其实我对这种可能性非常开放。这不是承诺,而是一种愿望,说清楚一点。我喜欢非常诚实。我认为,随着每美元智能越来越便宜,我们可能会得到非常小的近似物,希望它们是智能的代理。有没有一个世界,人们不再写 reax 了,因为使用 AI 的每美元智能比 reax 的复杂性更便宜?那会很酷。我会喜欢那样。其中一些狭窄用例可能需要微调才能真正越过阈值。我们拭目以待。我的希望是校准能做到这一点。校准加上模型级联:如果它非常恒定,那么也许就是对的。如果它在中间,那么你就用下一个更大的模型,然后从那里链接下去。我不太确定那会怎样发展,但是的,我可以想象。我还能想象的一件事是,想象你有一系列——我们拥有整个前沿——企业可能想做的事情,或者我认为黑客会愿意处理模型前沿的一部分。也许企业想要更动态的东西。你可以想象拥有不同大小的模型,并根据它在你的技术栈不同部分有多聪明来动态选择哪个模型。你甚至可以想象,因为我们的东西很简单,可以自动微调。

Well actually I'm super open to that possibility. This is not a promise. This is a desire, just to be clear. I like to be really honest. I think that as intelligence per dollar gets cheaper and cheaper and cheaper, we could get really small approximate things that hopefully are proxies for intelligence. Is there a world where people don't write reaxes anymore because the intelligence per dollar that uses AI is cheaper than the complexity of a reax? That would be kind of sick. I would love that. It might require fine-tuning for some of those narrow use cases to really get past the threshold. We will see. My hope is calibration gets that. Calibration plus a cascade of models: if it's super constant, then maybe it's right. And if it's in the middle, then you do the next bigger model and you chain off from there. I don't really know how that's going to go, but yeah, I could imagine it. And something that I could imagine too is imagine you have a series of—we own the entire frontier—something that a business might want to do, or I think a hacker would be okay with dealing with a part of the frontier of models. Maybe a business wants something more dynamic. You could imagine having different sizes of models and dynamically picking which model based on how smart it is on different parts of your stack. And you could even imagine, because of how simple our thing is, some automatic fine-tuning on that.

Diogo

是的。一点都不是承诺。我只是在科幻里烹饪。

Yeah. Not a promise in the slightest. I'm just cooking on sci-fi.

Host

但你会考虑不同大小的 Jeff 模型来提供那个。

But you would consider different sizes of Jeff models so to offer that.

Diogo

绝对会。是的。是的。是的。就像我怎么知道人们需要多少智能?

Absolutely. Yeah. Yeah. Yeah. Like how would I know how much intelligence people need?

Host

对吧?是的。我也不知道。

Right? Yeah. I don't know either.

Diogo

需求是无限的。

Demand is unlimited.

Host

嗯,是的。人们现在告诉我们不要发布东西,因为我们不需要发布东西,因为再次,它已经足够好了。

Well, yeah. People are telling us not to ship things right now because we don't need to ship things because again it's good enough.

Diogo

是的。但那有点逊。我真的很喜欢这句话——我希望人们以此要求我,因为很难收回——我不知道文化就是你所做的而市场不奖励它是否完全正确。我真的很喜欢这一点,因为我认为我们在代表某种东西。也许在未来,我们所代表的东西如此明显,以至于我们相当于无聊的 Visa 之类的,我们只是一个没人真正想到的实用工具,我会穿非粉色的西装或其他什么。但我真的想为此 rally 世界,你知道,就像我想继续做酷的东西,不是因为我们需要,而是因为我想让人们意识到这只是开始,你知道,就像那甚至不是开场白。那有点像低调的研究预览或你想叫它什么都行。我们可以用机器原生智能做更多事情——它会变得疯狂。所以不是唯一的可能不是唯一的大小,可能不是你们发布的唯一模型。

Yeah. But that's kind of lame. And I really like the saying—this is something that I hope people hold me to because it'll be hard to walk back from—that I don't know if exactly the thing that culture is what you do and the market doesn't reward it. And I really like that because I think that we are standing for something. Maybe in the future what we're standing for is so obvious that we're the equivalent of boring like Visa or something like that and we're just a utility that no one really thinks about and I'll be wearing non-pink suits or whatever else. But I really want to be rallying the world to this, you know, like I want to keep doing cool stuff, not because we need to, but because I want people to realize that this is just the beginning, you know, like that wasn't even meant to be the opening salvo. That was kind of like a low-key research preview or whatever you want to call it. And there's a lot more we can do with machine native intelligence—it's going to go wild. So not the only potentially not the only size, potentially not the only model that you guys launch.

Host

绝对不是那些中的任何一个。我想满足我们能满足的任何需求,对吧?但有一个巨大的警告:我不想成为那种把东西扔到墙上的开放产品团队。我希望它在一个统一的愿景下。如果你回到宣言,在我看来,一切都需要在这三件事之一之下。

Absolutely not for any of those. I want to meet whatever needs we can, right? But with a giant caveat: I don't want to be like open product teams that throw stuff at the walls. I want it to be under a unified vision. If you go back to the manifesto, everything needs to be under one of these three things in my opinion.

Diogo

嗯,是的,我还没准备好做这个。

Um, yeah, I'm not prepared to do this.

Host

哦,对不起。对不起问了。我可以只谈谈它。

Oh, I'm sorry. I'm sorry for asking. I can just talk about it.

Diogo

就像我们的东西有三个步骤。听起来像挑逗。我希望一切都在三件事之一之下,以继续推动边界等等。这些不是清单。这些是我们认为构建新技术革命基础的轴。我希望我们做的所有赌注都在其中某个地方。我们会在模型方面做一些非常奇怪的事情。所以,因为机器原生,对吧?就像人类不需要完全理解它。它只需要有价值。

Like we have three steps in our stuff. It sounds like a tease. I want everything to go under one of these three things to keep pushing the boundaries and everything. These are not checklists. These are axes that we think build the foundation of a new technological revolution. And I want all the bets we make to be somewhere in there. And we will be doing some weird, weird stuff model-wise. So, because machine native, right? Like humans don't need to totally get it. It needs to just be valuable.

Host

就给大家一个挑逗或提示——奇怪看起来像什么?什么是奇怪?

Just give people a tease or a hint—what does weird look like? What is weird?

Diogo

我会给大家一个提示。是的。

I'll give people a hint. Yeah.

Host

有些人试图称它们为决策模型。

Some people are trying to call them decision models.

Diogo

好的。

Okay.

Host

我们的原语是决策。我不会那样做,因为我认为还有其他机器原生的类型不是决策。

Our primitives are decisions. I wouldn't do that because I think there are other types that are machine native that are not decisions.

Diogo

好的,我们就到此为止。我认为这是一个相当有趣的提示。

Okay, we'll leave it at that. I think it's a pretty fun hint.

Host

是的。是的。有人说,我以前做过这个。我一年前做了一个决策模型。Jeff 不新。Jeff 不酷。但我认为有分类上的——这是你正在确立的可能性。有性能——对于基准和数字,你仍然在击败,据你所知,你仍然在击败每一个你的克隆。

Yeah. Yeah. There's people saying like, I've done this before. I made a decision model a year ago. Jeff is not new. Jeff is not cool. But I think there's the categorical—here's what you're establishing is possible. There's the performance—for the benchmarks and the numbers that you're getting, you are still beating, as far as you can tell, you're still beating every single clone of you out there.

Diogo

我不在乎基准,说清楚一点。所以即使我们赢了或输了,我想否认——

I don't care about the benchmarks, just to be clear. So even if we were winning or losing, I want to deny—

Host

你建立了类别,对吧?

You establish the category, right?

Diogo

但我也认为决策模型和系统一之间的细微差别,我认为实际上是你试图——

But also I think this nuance between decision models and system one, I think is actually the thing that you're trying to—

Host

是的。我只想让软件工程师超级强大,对吧?用 AI。对我来说悲剧的是,你知道,在那个 AI 冬天的方向。我认为 AI 如此强大却如此未被充分利用,这太悲伤了。这是让我情绪激动的事情。但伙计,我认为那是——我不想只是像纯粹的技术乐观主义者,像所有技术都是好的。我认为现在发生的事情像是一场闹剧。我只想为人们打开那些可能性。是的,我就到此为止。这些天我哭得太多了,不想在记录上哭。

Yes. And I just want to make software engineers superpowered, right? With AI. The tragic thing to me is, you know, in that AI winter direction. I think it is so sad that AI was so powerful yet so underutilized. It's the thing that gets me emotional. But man, I think that is—I don't want to just be like pure techno-optimist, like all technology is good. I think what was happening now was like a travesty. And I just want to open up those possibilities for people. Yeah, I'll just end it there. I've cried too much these last few days to want to do it on the record.

Host

是的。是的。不,我很感激你分享了一点这些,我认为人们可以看到你对此非常真实和热情。

Yeah. Yeah. No, I appreciate you sharing a little bit of that and I think people can see that you're very authentic and passionate about this.

从零到一的困难部分 The Hard Part of Going Zero to One

Diogo

你不一定能从名字里看出这一点,比如 type safe AI。但我认为,一旦人们足够沉浸于“这是你希望世界走向的真正不同方向”,而且你实际上已经完成了从 0 到 1 的艰难部分,那么现在让我们大家一起朝新方向前进。

You don't necessarily get that from the name, like type safe AI. But I think once people immerse themselves enough in here's the genuinely different direction you want the world to go, and actually you have done the hard part about going zero to one on the thing, then now let's all go together in the new direction.

Host

是的,是的。

Yeah, yeah.

Diogo

我不——我确信我不会认为也许我会认为艰难的部分已经完成了。也许我认为还会有很多艰难的部分。就像如果各种东西都自动化了,我们终于看到 GDP 增长,每天都像开派对一样,那么也许艰难的部分就完成了。但我不这么认为。我真的真的认为人们过于关注速度和成本,而对可靠性关注不够。可靠性才是让它令人愉悦的东西。可靠性才是让你能够信任它的东西。

I don't — I am sure that I won't think maybe I will think that the hard part was done. Perhaps I think that there's going to be many more hard parts. Like if all sorts of stuff gets automated and we finally see GDP growth and it's like a JF party every day, then maybe the hard part is done. But I don't think so. And I really, really think that people focus too much on speed and cost and not enough on reliability. Like reliability is what makes it delightful. Reliability is what allows you to trust it.

Host

你有一句话——

You have this line —

Diogo

全要素生产率增长在 5 年内达到 3%。太棒了。冲啊。

TFP growth reading 3% in 5 years. Hell yeah. Let's go.

Host

我从未见过哪个实验室关心全要素生产率增长。

I've never seen a lab care about TFP growth.

Diogo

但这正是经济革命的意义所在,对吧?这实际上与 OpenAI 章程曾经代表的东西极其一致。章程还是一样的,但他们试图把定义挪来挪去,比如变成 1000 亿美元利润之类的。并不是说我讨厌 OpenAI。

But that is what an economic revolution is, right? It's actually extremely consistent with what the OpenAI charter used to stand for. The charter is the same, but they've kind of tried to move definitions around to like 100 billion in profit or something like that. Not that I hate OpenAI.

Host

AGI 是什么并不是一个定义明确的术语,对吧?

It wasn't a well-defined term what AGI is, right?

Diogo

他们试图定义它,对吧?比如完成世界上大部分有经济价值的工作。他们应该回答这个问题:它怎么能解决数学千禧年大奖难题,却对世界上有经济价值的工作贡献为零,只是舍入误差?你知道,我认为现在所有模型大致都并列在零。有可能我们已经开始了,但我猜还不到 1%。

They tried to do it, right? Like doing majority of the world's economically valuable work. And they should have to answer the question: how can it do millennium prize problems in math and zero of the world's economically valuable work, rounding error? You know, I think that all models are roughly tied right now at zero. There's some chance that we have started already, but I would guess that it's not yet 1%.

Diogo

我认为这会体现在——当它真的发生时,会体现在经济统计数据中。那会很棒。它不会导致大规模失业,但会带来一大堆很棒的转变,世界会好得多。而且我真的厌倦了 AI 总是成为事情的前景角色。我认为世界应该只是更令人愉悦,而 AI 应该只是帮助实现这一点。

And I think that will show up in — when it does happen, it will show up in the economic statistics. It's going to be awesome. It will not cause mass unemployment, but it will cause a whole bunch of awesome shifts and the world will be a lot better. And also I'm really tired of AI always being the foreground character of things. Like I think that the world should just be more delightful and AI should just help with that.

Host

就像消失在背景中。

And just like disappear into the background.

Diogo

没错。你知道,我在演讲中说过:2019 年的软件——比如软件、SaaS 之类的——超级有价值,对吧?现在是 2026 年了。尽管 AI 如此厉害,软件怎么基本上还是一模一样,除了有时旁边有个聊天框?那种东西有点用,但不能让你做出公司有利益相关的决策,因为不能信任它做决策。这太疯狂了。你知道,这里有巨大的经济激励。我认为这将会像一场反向的 SaaS 末日。我认为 SaaS 会因此被超级充电。他们最了解哪些东西值得自动化,那将会是一个疯狂的时代。

Exactly. You know, I say this in my talks: how can it be that 2019 software — like software, SaaS, whatever — super duper valuable, right? It's 2026 now. How is the software basically exactly the same despite AI being so freaking awesome, other than sometimes having a chat box on the side? That kind of works but doesn't allow you to make decisions that the companies have stakes in, because they can't be trusted to make decisions. That is nuts. You know, there's so much economic incentive for this. And I think it's going to be like an inverse SaaS apocalypse. I think SaaS is going to be supercharged by this. They are the ones who are most in the know of what things are valuable to automate, and it's going to be a crazy time.

Host

是的,我也这么认为。你解锁了一件美好的事情。

Yeah, I think so too. It's a beautiful thing that you've unlocked.

System 1 vs System 2问题 System 1 vs System 2 Problems

Host

你在这里提到了一件事,我不知道是不是直接在这里,就是什么是系统 1 问题,什么不是,什么是系统 2 问题。

You mentioned one thing here which I don't know if it's directly here, which is what is a system one problem and what is not, what is a system two problem.

Diogo

这是个难题。这是个难题,我的朋友。

That's a hard one. That's a hard one, my friend.

Host

因为现在人们只是试图把所有东西都 Jev 化,对吧?这可能会失败,但有些东西会很好。

Because people now are just trying to Jev everything, right? Which probably is going to fail, but some things are going to be good.

Diogo

把所有东西都 Jev 化是个很有趣的说法。所以我告诉你真相:真相是这是一个经验问题,就像缩放定律是经验性的东西。为什么机器人技术现在真的不奏效,尽管花了那么多钱?我不认为这一定是关于花更多钱。经验结果可能就是不成立。所以从经验上看,我相信这些预训练的超级智能浓缩体从根本上说是系统 1 思考者。我认为系统 1 是最接近描述 LLM 擅长什么的说法。RLVR 对系统 2 思维做出了不可思议的事情。我敬畏。这太酷了。我不认为这会导致 AI 末日,一点也不。当然不是 0%,因为我认为 0% 是校准错误的,但他们所做的真的很酷,他们真的把它推到了极限。好吧,也许他们不这么认为——不是极限的极限,但这对模型来说是一件奇怪的事情,而且它们在这方面非常脆弱。想想人们在 ChatGPT 时代是怎么谈论 AI 的:哇,它真的很通用,它能做很多通用的事情,但它在数学问题上很差,比如 GSM8K 小学数学。然后现在看看人们怎么谈论 RLVR:它太脆弱了,太参差不齐了。为什么它能做这种非常奇怪的事情?实际上,数学不仅仅是尖刺的,它是分形的。这是因为 RLVR 是——如果我们谈论每个东西的北极星是什么,RLHF 是取悦人类,这就是人类反馈。RLVR 是优化基准。所有归入 RLVR 类别的东西按定义字面上就是基准,因为基准是程序化可验证的简单输出,可以表现良好。而 RLCD 是让它对程序化使用可靠。

Jev everything is a pretty funny way of saying it. So I'll tell you the truth: the truth is that this is an empirical problem, just like scaling laws are an empirical thing. Why doesn't robotics really work right now despite all the money being spent on it? I don't think it's about spending more money necessarily. The empirical results just might not be there. So empirically I believe that these pre-trained super condensations of intelligence are fundamentally system one thinkers. I think that system one is the closest thing to describe what LLMs are strong at. RLVR has done incredible things for system 2 thinking. I am in awe. It is super freaking cool. I don't think that it's going to result in AI doom in the slightest. Not 0% of course, because I think 0% is miscalibrated, but it's really cool what they've done and they've really pushed it to the limits. Well, maybe they don't think so — not the limits limits, but it is a weird thing for models to do and they are very fragile at this. Think about how people used to talk about AI back in the ChatGPT days: wow, it's really general, it can do a lot of general things, but it's bad at math problems like GSM8K grade school math. And then now look at how people talk about RLVR: it's so fragile, it's so jagged. Why can it do this really weird thing? And actually, math is not just spiky, it's fractal. And this is because RLVR is — if we talk about what is the north star for each thing, RLHF is please humans, that is what the human feedback is. RLVR is optimized benchmarks. Everything that goes into the RLVR category literally is a benchmark by definition, because a benchmark is programmatically verifiable simple outputs that can do well. And RLCD is make it reliable for programmatic use.

Host

是的,那是——也许我提供一些想法,然后如果我说错了你可以纠正我。我一直在想的一件事也是——所以当你第一天给我访问权限时,我把 Jeff 扔到了一堆东西上,还有多跳推理,对吧?所以单跳非常棒,像最先进的。单跳你绝对不应该用 Jeff 以外的任何东西。

Yeah, that's — maybe I'll offer some thoughts and then you can sort of correct me if I'm wrong. One thing that I've been thinking about is also — so I threw Jeff at a bunch of things when you gave me access on day one, and multihop reasoning, right? So single hop fantastic, like state-of-the-art. You should never use anything other than Jeff for single hop.

Diogo

是的。

Yep.

Host

多跳会开始下降,而且随着跳数增加,它大致是单调递增的。

Multihop is going to start to fall down and it's like kind of monotonically increasing as you increase the hops.

Diogo

是的。是的。是的。

Yep. Yep. Yep.

Host

对。

Right.

Diogo

所以,哦是的,回到那个经验问题,这取决于我们能从模型中提取出什么。对。所以我们想要一切,我们想要挖掘尽可能多的智能,就这样。模型——我把我们看作是在解锁、平滑和雕琢智能,同时添加新能力并填补其中的空白。随着时间的推移,我们会填补越来越多的这些空白。但现实是,我们从事的是挖掘属性的业务。

So, oh yes, back to that empirical question, it depends on what we can pull out of the models. Right. So we want everything, like we want to unearth as much intelligence as possible, period. The models — I see us as unlocking and smoothing and sculpting the intelligence while adding new capabilities and filling in gaps in it. And we will be filling in more and more and more and more of these gaps over time. But the reality is that we are in the business of unearthing properties.

System One与推理 System One and Reasoning

Diogo

这些特性实际上取决于这些压缩核心能提供什么,以及把它们像科学怪人一样拼凑在一起,从而拥有所有东西的特性。但现实是,我们的业务是尽可能多地发掘能力,而 System One 恰好是对有效方法的描述。在那个范式中有效的一切都会是 System One 式的。我们不做所谓的潜在推理——在字符串中推理——是有原因的。我认为模型真正擅长的是模型内部的推理。它并不完全完整;它做得并不好。

Those properties are actually a function of what is available from these condensed cores and franken-signing them all together to have all of the properties of everything. But the reality is we are in the business of unearthing as many capabilities as possible, and System One just happens to be the description of what works. Everything that works in that paradigm will be System One-ish. There is a reason why we don't do what's called latent reasoning—reasoning in strings. I think what models do really well is reasoning within the models. It's not totally complete; it doesn't do great at all.

Host

潜在推理是在字符串中推理?我以为潜在推理是在模型权重内部推理。

Latent reasoning is reasoning in strings? I thought latent reasoning is reasoning inside the model weights.

Diogo

我认为人们过去称之为连续推理。我不完全确定。它被称为潜在推理,是因为过去推理轨迹是保密的。所以它们有点像答案的潜在变量。

I think that people used to call that continuous reasoning. I'm not entirely sure. It was called latent reasoning because it used to be that the reasoning traces were secret. So they're kind of like a latent variable for the answer.

Host

是的。

Yeah.

Diogo

所以现在保密的东西已经转移了。

So what's secret has now shifted.

Host

嗯,对开放的 Anthropic 来说它仍然是保密的,对吧?

Well, it's still secret for open anthropic, right?

Diogo

所以没有推理,Jeb。

So no reasoning Jeb.

Host

是的。

Yes.

Diogo

就你所能做的而言,对吧?因为那违背了 System One 的全部承诺。我的承诺是做任何对机器原生必要的事情。我可以想象有一些形式的推理不那么慢、低效和脆弱,这些是完全有可能的,只是说清楚。所以作为一个务实的人,我不对方法做承诺。我对我的 ROI 北极星做承诺,我会为此而战。这次发布没有发生,我们仍然渴望在世界上的位置。

As far as you will ever do it, right? Because that violates the whole promise of System One. My promise is to do whatever necessary for machine native stuff. I could imagine there are some forms of reasoning that are less slow, inefficient and fragile that are totally on the cards, just to be clear. So as a pragmatic person, I'm not making promises on methods. I'm making promises on what my ROI northstar is, and I'm going to fight for that. This launch didn't happen and we are still hungry for our place in the world.

Host

那太好了。是的。

That's great. Yeah.

Diogo

是的。

Yeah.

Host

我认为另一件事是视觉——另一个你没有的大能力,但也许它永远不属于 System One。

I think the other thing is vision—another big capability that you don't have, but maybe it doesn't ever belong in System One.

Diogo

我认为我有很好的视觉。

I think I have a pretty good vision.

Host

什么?抱歉。

What? Sorry.

Diogo

我认为我有好的视觉。

I think I have a good vision.

Host

不不不,我开玩笑的。我开玩笑的。

No no no, I'm kidding. I'm kidding.

Diogo

哦我的天。

Oh my god.

Host

是的。是的。是的。

Yeah. Yeah. Yeah.

Diogo

因为人们显然首先想要的是视觉,因为 Doom 演示,但还有除了文本之外的一切都是视觉。

Because people obviously the first thing they want is vision because of the Doom demo, but also just everything other than text is vision.

Host

在我看来一切都有可能。实际上这是一个辩论——谁有这个人,你的观众可能是这场辩论的最佳人选。有一个问题是我们应该多大程度地给人们他们认为他们想要的,这就是我们在隐身模式下做了 2 年的事情。我们只是知道这显然会有价值,相对于给他们他们说的想要的,对吧?这有很多维度,对吧?上下文长度就是一个例子,对吧?每一个模型包括我们的——我实际上认为就我所知,我们的在长上下文中不退化方面是最好的。但其他提供商只是像人们想要什么,我们就给他们那个愚蠢的东西。我们需要为此找到平衡,因为如果你把前者——给人们他们想要的——走得太远,你最终会得到 Anthropic 保姆国家式的思维,这非常反开发者,而亲开发者的路线会是给他们他们想要的,但开发者说我们不想把负担放在他们身上去弄清楚智能的基因。所以我们试图弄清楚这个导航:如何快速发布东西,以仍然保持我们的信任品牌,同时像对待成年人一样对待我们的用户,他们可以做出明智的决定,不需要在这些东西上保姆式管理。

Everything is in the cards in my mind. And actually this is a debate—who we have this man, your audience is probably the great one to have in this debate. There's a question about how much do we try to give people what they think they want, which is what we did in stealth for 2 years. We just knew that this is obviously going to be valuable, versus give them what they say they want, right? And there's a lot of dimensions of this, right? And context length is an example of this, right? Every single model including ours—I actually think as far as I can tell ours is by far the best at not degrading in long context. But the other providers are just like whatever people wanted, let's just give them the stupid thing. And we need to figure out a balance for this because if you take the former side too far—give people what they want—you end up with anthropic nanny state style thinking which is very anti-developer, while the pro-developer route would be like give them what they want but developers are like we don't want to put the burden on them to figure out the genes of intelligence. So we are trying to figure out this navigation of how quickly to release things to still have our brand of trust and also teach our users like adults that can make informed decisions that don't need nanny stating on top of this stuff.

Host

是的,我认为这很公平。

Yeah, I think that's fair.

Diogo

老实说,我们不知道答案。我们得弄清楚。这可能是我接下来几天最大的辩论之一。因为我们有很多东西。再次,我们没想到它会火起来。所以,我们想,我们需要一些后续发布。

And we don't know the answer to be honest. We'll have to figure it out. It's going to be that's probably going to be one of my biggest debates over the next couple of days. Because we have a lot of stuff. Again, we didn't expect it to pop off. So, we were like, we'll need some follow-up launches.

Host

是的。我不知道。我不知道你是否没想到它会火起来。我看到了你投入的工作。我从未见过你像过去两个月那样如此专注,对吧?

Yeah. I don't know. I don't know if you didn't expect it to pop off. I saw the work that you put in. I have never seen you lock in so hard as the last two months basically, right?

Diogo

嗯,那也是因为我的幕僚长让我专注。

Well, that's also because my chief of staff made me lock in.

Host

是的。就像我从未——

Yeah. Like it's like I have never—

Diogo

我以为我以前工作很努力。

I thought I worked hard before.

Host

是的。而且不,但就像你出现在我们的写作研讨会上,我当时想,“你在这里做什么?” 而且它很有用。很棒。

Yeah. And no, but like you were showing up at our writing workshops and I was like, "What are you doing here?" And it was useful. It was great.

Diogo

你显然对你的发布非常有意,工作也显示出来了,恭喜。我希望继续专注是我的感觉。我想——我认为我们已经通过了许多我们想要的技术世界的伟大过滤器,但还会有更多。而且天哪,我多么兴奋去打好这场仗。

You clearly were very intentional about your launch and the work showed and congrats. I hope to keep locking in is my sense. I want to—like I think that we've passed many great filters for the tech world that we're wanting, but there's still going to be a bunch more. And holy smokes, am I excited to fight the good fight.

Host

是的,很令人兴奋。在我们扩展到 Typesafe 之外的话题之前,我只想提供任何其他你认为被低估或误解的关于你发布的东西。

Yeah, it's exciting. Before we broaden out to topics outside of Typesafe, I just wanted to offer any other things that you think are underrated or misunderstood about what you have launched.

Diogo

被低估或误解?

Underrated or misunderstood?

Host

是的,你有 panouts,抱歉,模式在这里。也许想进入那个。模型锯齿状,任何东西,你知道,

Yeah, you have panouts, sorry, patterns here. Maybe want to go into that. Model jaggedness, anything, you know,

Diogo

给我一个它的思考。哦天哪,我会对这些都大发牢骚。我真的不应该。我真的不应该。

Give me one noodling of it. Oh man, I would rant about all of these. I really shouldn't. I really shouldn't.

Host

人们可以去你的 Discord,如果他们

And people can come go to your discord if they

Diogo

人们把很多爱投入到 cookbooks 中,这就是我要说的。Cookbooks 有一些火爆的东西。我们曾考虑把一堆这些东西放在主发布博客文章中,但它变得有点长和笨重,而且非常高级用户向。但我们真的真的——我会坦率地说——在发布之前,我们说的每件事听起来都像这个奇怪的外星工具。为什么会有人需要这个?这是一件非常奇怪的事情,我们非常担心教人们关于这个新前沿。它显然成功了,但我们投入了很多工作,因为我们认为教育会是我们巨大的瓶颈。它可能有效,而且不再是问题,因为人们做的事情远远超出了我们能展示给你的如何使用你的模型。

People put a lot of love into the cookbooks is what I will say. The cookbooks have some fire stuff. We had considered putting a bunch of these things in the main launch blog post, but it got kind of long and unwieldy and very power usery. But we really really—I'll be frank—before the launch, every what we're saying sounds like this weird alien tool. Why would anyone need this? It was a very weird thing and we were very worried about teaching people about this new frontier. It obviously succeeded but we put a lot of work because we thought that education would be a gigantic bottleneck for us. It probably works and it's no longer a problem because people are doing things well beyond what we could ever show you how to use your model.

Host

确切地说。但就像他们——是的。而且他们的用例比我们的更酷。就像有一堆东西,我想,天哪,如果那是我们的演示,天哪——那比我们展示的要酷多了。就像计算机使用的东西。天哪,它太酷了。

Exactly. But like they—yeah. And their use cases are kind of cooler than ours. Like there's a bunch of stuff where I'm like, man, if that was our demo, holy—that was way cooler than what we were showing. Like the computer use stuff. Holy smokes is it cool.

发布验证与早期反馈 Launch Validation and Early Reception

Diogo

我们在这上面倾注了很多爱。据我所知,这不是 AI 生成的垃圾。我们投入了很多爱,每一个都是受解决真实客户问题的启发。我们花了功夫帮他们做出很酷的东西。

We put a lot of love into this. This is not AI-generated trash as far as I know. We put a lot of love in here, and each of these is inspired by solving real customer problems that existed. We went through the work of helping them do cool ass stuff.

Host

是的。发布前你们做了多少验证?那个过程是什么样的?那个过程是什么样的?

Yeah. How much validation did you do before launch? Like what was that process like? What was that process like?

Diogo

显然你们做了一些,但显然你们没有像今天这样接触那么多人。

Like clearly you did some but obviously you're not getting in touch with as many people as you are today.

Host

是的,当然。

Yes, of course.

Diogo

其实我觉得反响相当差。团队里非技术的人非常担心。有很多恐惧。就像没人真正理解这个,他们不想要。我们卖的是维生素,不是止痛药。我们是不是该有全职员工来围绕解决那个问题写软件?发布前我们几乎没有收入。技术人员显然是真正的信徒。我们知道这很牛。它的计算特性在这么多维度上都爆表,所以我们觉得,是的,显然它会很大。我肯定超级害怕,这就是为什么我超级投入。但最常见的情况是,我会说超过一半的人玩了之后就是没搞懂。而那些搞懂的人会说,天哪这真的很酷,但怎么通过采购之类的?这相当是一场战斗。我们只知道,好吧,我们的目标市场将是开发者。人们会找到用例,这样所有人都会 FOMO 进来。我不想在人们因为事实改变而改变想法时 rubbing in。我确实想质疑产品市场契合这个概念,因为有一个产品,有一个市场。我们说,嘿你想用这个吗?人们说,我不太确定它是否解决了我们的问题。它爆发了,每个人都说,我们需要尽可能多的速率限制。我们能直接给你 GPU 吗,因为我们现在受限。所以当然营销是其中一部分,但我甚至不认为这是关于营销。我认为这是关于热情的开发者,他们的灵魂基本上以相同的频率共振,那个频率也让其他人兴奋起来。我也希望我们作为一家公司会永远永远感激那些开发者,而不仅仅是那些从开发者开始然后走向企业的公司。

I actually think that the reception was pretty bad. For the non-technical people in the team, they were really worried. There was a lot of fear. It's like no one really gets this and they don't want it. We're selling a vitamin and not a painkiller. Should we have FTEs to write the software around solving that problem? We had almost no revenue before launch. The technical people were obviously true believers. We knew that this was sick. Its computational properties are off the charts on so many axes that we're like, yeah, obviously it's going to be huge. I was definitely super afraid, which is why I locked in super hard. But the most common thing was, I would say more than half the people we had play with it just did not get it. And the people who did were like, man this is really cool, but how do we get this through procurement and stuff like that? It was quite a battle. We just knew, okay our target market is going to be developers. People will find the use cases and that way everyone is gonna FOMO in. I don't want to rub in people changing their minds with the facts changing. I do want to call into question the concept of product market fit, because there was a product, there was a market. We were like, hey do you want to use this? And people are like, I don't really know if it solves our problems. It explodes and everyone's like, we need as much rate limits as we can. Can we literally give you GPUs because we are constrained right now. So of course marketing is an element of it, but I don't even think it's about marketing. I think it's about passion developers whose souls basically resonated at the same frequency, and that frequency got everyone else excited too. I'm hoping as well that we as a company will be eternally grateful to those developers, not just the companies that start off with developers and go to enterprises.

Host

没错。我甚至在想,我们怎么能推出对……天哪,我不知道我该不该说这个,但我会说

Exactly. And like I'm like even thinking about like how can we launch things that are better for oh man I don't know if I should say this but I will

Diogo

对开发者比对企业更好的东西。

better for developers than enterprises.

Host

没错。我们怎么做?我们怎么赋能他们?

Exactly. How do we do that? Like how do we empower them?

Diogo

我有厨师。我有厨师,但这是一件很奇怪的事,我不知道还能怎么表达我的感谢和忠诚。这就是为什么我昨天染了头发。就像我想和他们说话,因为在我们公司最重要的时刻不继续和他们说话,对我来说感觉很不干净。

And I have cooks. I have cooks but it's a very weird thing to do and I don't know how else I can show my thanks and loyalty to that, you know. And that's why I did the dying my hair yesterday. It's like I wanted to talk to them cuz it felt dirty to me during our company's most important times not to keep talking to them.

Host

好。嗯,我的意思是这就是为什么你在这里的原因之一

Good. Well, I mean that's why one of the reason you're here

Diogo

请让我坚持这一点。我努力做到有原则。引用我的话。指出我。如果我变了,你们就拿起干草叉。

hold me to that please. I try to be principled. Quote me on this. Call me out. You have the pitchforks out if I change.

计算机使用演示 Computer Use Demo

Host

呃,我正要简要展示计算机使用的东西。这是这是

Uh I was just going to briefly show the computer use stuff. Is this is this

Diogo

我从未见过,我没有,我见过像航空公司浏览器使用的东西。嗯,在这个新笔记里,让我们把标题设为你好。

I've never seen I haven't I've seen I saw like a airline browser used thing. And um inside this new note, let's make the title say hello.

Host

哇。

Wow.

Diogo

太好了。太好了。好的。嗯,我们继续。你能打开 Arc 浏览器吗?到了之后,你能谷歌搜索 Norbert 吗?

Great. Great. Okay. Um let's move on. And can you open up the Arc browser? And once you're there, can you Google search Norbert?

Host

嗯,现在,你能打开 X.com 吗?

Um now, can you open up X.com?

Diogo

这是这种用例吗?

Is this this kind of use case?

Host

哦,语音用例。这实际上是我见过的第一个。这是

Oh, the voice use cases. This is actually the first one I've seen. This is

Diogo

打开照片。

Open up the photo.

Host

哇。

Wow.

Diogo

哦,等等,等等,等等,等等。哦,你能你能回去一秒吗?你能回去一秒吗?嗯,

Oh, wait, wait, wait, wait. Oh, can you can you go back a second? Can you go back a second? Um,

Host

谣言称 Anthropic 工程师崇拜 Claude 为神。哇。

rumors claim anthropic engineers worship Claude as God. Wow.

Diogo

哇。该死,这挺有趣的。嗯,

Wow. Dang, that's pretty funny. Um,

Host

而你在这里构建产品。哇。这太牛了。

and here you are building prod. Wow. This is sick.

Diogo

呃,是的。所以,所以显然你可以用语音操作整个电脑,以 Jev 作为决策模型。

Uh, yeah. So, so clearly you can operate the whole computer with voice with Jev as a decision model.

Host

所以,就像我反对刷榜一样,我也反对演示。我想确保它可靠地工作。我喜欢人们在玩它。这超级牛。毫无疑问。我想看到这个。我想看到它被使用。我想让我们的团队玩它。我想找到弱点并解决它。我会,天哪,那看起来真的很酷。那看起来真的很酷。我想要那个。我想要那个。就像当我手腕酸痛时,我就 whisper flow 一切。那会很牛。

So, just like I'm anti-benchmaxing, I'm also anti-demos. I want to make sure that it works reliably. I love people are playing with it. This is super sick. Have no doubt. I want to see this. I want to see it be used. I want our team to play with it. I want to find the weaknesses and I want to solve that. And I would Man, that looked really cool. That looked really cool. I want that. I want that. Like when my when my wrists are sore, I just whisper flow everything. That would be sick.

用例与大公司 Use Cases and Big Companies

Host

嗯,如你所知,只是为了完善用例方面,因为我确实得让你走了。嗯,呃,谁,呃,谁是有哪些更大的公司联系过你,并让你对他们想做的事情感到惊讶。

Well, as uh you know, just just to round out the use cases side because I I do have to let you go. Um, uh, who's, uh, who are the who are the bigger companies that have reached out and have surprised you with what they want to do.

Diogo

只是

Just

Host

我对此太脱节了。人们给我看了公司的截图,从我所见,是所有的公司。

I'm so out of touch for that. People have shown me screenshots of companies and from what I've seen, it's all of them.

Diogo

嗯,主要是像,你知道,对于那些在更大公司工作的人,他们不做这种工作,呃,我只想给人们一些例子,比如,你应该去查一下。查一下,查一下。

Well, mostly like, you know, for for those people who work at larger companies and they're not doing this kind of work, uh, I just want to give people examples of like, you should go look that up. Look that up, look that up.

Host

哦,所以呃,就像我认为演示超级超级牛。显然编码智能体是巨大的用例,它们也超级牛。

Oh, so uh, like I think demos are super duper sick. Obviously the coding agents are like gigantic use cases like they are like also super sick.

Diogo

Call 现在全是关于 Jeff 的。

Call is all about Jeff right now.

Host

哦,当然。哦,我能呃,我能稍微岔开一下关于编码智能体的话题吗,如果可以的话?

Oh hell yeah. Oh can I uh can I give a little bit of a a tangent about coding agents if that's okay?

Diogo

是的,请。

Yes please.

Host

哦,让我一下。给我一下。好的,实际上我会回到编码智能体。让我描述一下大的用例家族。就像我们在发布前很久就从第一性原理规划了这些。它们是我们所谓的暗数据。就像人们囤积大数据,但他们不会把 LMS 扔给它,因为那太贵了。所以大公司喜欢这个。他们有一堆数据,希望自己能分析,这就像数据科学家的春梦。所以这是一个巨大的。就像我认为这个加上呃编码智能体是大的赚钱者,因为那是所有量所在的地方,对吧?嗯,有实时的东西,你知道,就像需要智能在循环中的人。

Oh let me give me a second. Give me a second. Okay actually I'll come back to coding agents. Let me describe like the big families of use cases. Like we've mapped this out from first principles like long before release. They are what we call dark data. Like people hoarded big data but they would not throw LMS at it because it just was too expensive. So large companies adore this. They have like piles of data that they wish they could analyze and this is like a data scientist's wet dream. So this is like this is a giant one. Like I think this plus um coding agents are the big money makers because that's what the all where all the volume is, right? Um there's the real time stuff, you know, like people who need like intelligence in the loop.

延迟与AI助手 Latency and AI Assistants

Diogo

我猜那些公司的每一位 CEO,如果不是 CTO 的话,都知道每减少 10 毫秒,他们的产品会好多少。

I would guess that every CEO, if not CTO, at those companies knows how much better their product gets with every 10 milliseconds shaved.

Host

尤其是电商。

Especially e-commerce.

Diogo

是的。哦,或者助手类的东西。有很多 AI 助手类的东西,据我所知,他们真的很喜欢。再说一次,我现在不在客户一线,所以只知道团队告诉我的,但我对此非常兴奋。我真的很期待游戏方面的应用。我特别想玩那种超酷的自走棋,你指挥你的队伍,或者半自动战斗。我觉得那会很酷,但别做得太好,趁我还有工作。

Yeah. Oh, or assistant-like things. There are many AI assistant-like things, and as far as I can tell, they really love it. Again, I'm not in the front lines of customers right now, so I just know what my team tells me, but I'm so excited for this. I'm really excited for this for games. I really want to play sick-ass auto battlers where you're commanding your team, or semi-auto battlers. I think that'd be so cool, but don't make it too good while I still have a job.

Diogo

还有我们所说的“验证一切”,比如验证所有 LLM 调用,有点像可观测性。实际上关于文档,我认为人们应该做的是并行问题非常便宜。所以如果你有大的状态,你想问很多问题,

And there's what we call verify everything, like verifying all LLM calls, kind of like observability. I think actually on the note of docs, what people should be doing is the parallel questions are very cheap. So if you have big states you want to ask many questions on,

Host

就在这里。

Right here.

Diogo

给每条消息加上 ID,然后针对每个 ID 提问。所以当你有一个长状态时,这样你可以为那个状态付费一次,然后针对其中的每条消息问很多很多问题。我认为这是一个节省成本的好方法,而且……

Put IDs on every message and then ask a question about each ID. So when you have a long state, that way you can pay for that state once and ask lots and lots of questions about each message within it. I think that is a great way that saves money and is...

Host

顺便说一句,我一直觉得系统一和系统二的框架很有趣,因为它基本上说明了,你每做一次推理调用,就应该做 1 次、10 次或 100 次 Jev 调用。

By the way, I always think it's interesting framing system one and system two because it basically makes the case that you should always make one or 10 or 100 Jev calls for every one reasoning call that you make.

Diogo

嗯,也许吧。我的意思是,我希望人们花得更少。也许你做一半的推理调用,每次 10 次 Jev 调用,或者类似这样,或者任何能解决原本不可能存在的问题的方法。等等,第四个用例是我所说的智能软件,就像本质上可组合的软件,做以前不可能发生的奇怪有趣的事情。比如编程语言作为 Jev 的东西。我不知道你见过没有。那太酷了,伙计。如果我们知道如何发放 credits,因为我们在基础设施方面还非常早期,我想给所有这些项目 credits。

Well, maybe. I mean, I would like people to spend less. Maybe you do half the reasoning calls and 10 Jev calls each, or something like that, or whatever solves the problem that couldn't have existed otherwise. Wait, and number four use case was what I described as smart software, like software that's intrinsically composable and does weird fun stuff that could never happen before. Like the programming language as Jev thing. I don't know if you've seen that. That is so cool, man. If we knew how to give out credits because we're really early in our infra days, I would want to give all these projects credits.

Diogo

我认为这些就是我们如何规划主要用例的方式。计算机使用也朝着实时方向发展了,那真的非常酷。如果它可靠,我会超级兴奋。我怀疑我们可以让模型在这些用例上做得更好,因为那有点出乎意料。所以那真的很酷。关于编码智能体的事情,这是现在正在发生的一件非常令人惊讶的事情。好的。

And I think those are how we've mapped out the main use cases. Computer use has also come in kind of the real-time direction as well, and that's really really cool. If it is reliable, I am super jazzed about that. I suspect we can make the model a lot better at these use cases because that came out of left field a little bit. So that's really cool. On the coding agent thing, and this is a really surprising thing that is happening right now. Okay.

Diogo

Claude Code 和 Codeex,我相信是赢家,比如第一和第二。我不完全确定。我没有密切关注,但大致如此。但它们是围绕单一模型世界构建的,这对他们来说很合理,对吧?因为一直是一个模型游戏,就像你在购买同一个模型但不同智能。但所有开源编码智能体现在都很兴奋,因为他们得到了他们的 Jevon,而且事情是……我确信他们正在尝试很多奇怪的东西。

Claude Code and Codeex are, I believe, the winners, like number one and two. I'm not entirely sure. I don't follow closely, but it's roughly that. But they're built around a single model world, and that makes a lot of sense for them, right? Because it has been a one-model game where it's like the same model but different intelligence that you're shopping. But all the open coding agents are jazzed right now because they're getting their Jevon, and the thing is there's... I'm sure they're trying a lot of weird stuff.

Host

嗯。

Mhm.

Diogo

但所有编码智能体大致处于同等水平,对吧?因为用 Y 循环做不了太多。但一旦有人找到一个杀手级用例,只能通过那个编码智能体实现,所有人都会涌向它,因为他们垄断了那个东西,但所有开源编码智能体都能复制那个。

But all the coding agents are kind of roughly at approximate par, right? Because there's not so much you can do with a Y loop. But the moment one person finds one killer use case that you can only do with that coding agent, everyone will flock to it because they have a monopoly on that thing, but all the open coding agents will be able to copy that.

Host

对。

Right.

Diogo

我不知道 Claude Code 和 Codeex 会怎么做,因为它们是围绕那个单一模型世界构建的,我认为那将是一件非常有趣的事情。我个人很想能够与它们集成。我想与所有人集成。他们最终可能会成为竞争对手。我不知道。但作为 Sonfire Infrastructure,我的工作不是对此有偏见,对吧?我只想服务世界。但我不知道他们是否会那样做。我认为这会让编码智能体游戏变得超级奇怪。我对此非常兴奋。我确信我现在正让我的团队审阅我写的一份关于设计模式的内部文档,我怀疑这对编码智能体会很有用。所以希望我走回家后就能分享。但我认为世界上有太多值得探索的成熟领域,天哪,如果我没有这个,我现在会很想实验编码智能体。

And I don't know what the Claude Codes and Codeexes will do because they are built around that one model world, and I think that's going to be a really interesting thing. I would love to be able to integrate with them personally. I want to integrate with everyone. They might make competitors eventually. I don't know. But it is not my job as Sonfire Infrastructure to be opinionated on that, right? I want to just serve the world. But I don't know if they would do that. And I think it'll make the coding agent game super weird. I'm so excited for that. And I'm sure I'm getting my team to review right now an internal document I made on design patterns I suspect will be useful for coding agents. So hopefully I can share it right after I walk home. But I think there's just such ripe area for exploration out in the world, and man, if I did not have this, I would love to experiment with coding agents right now.

Host

是的,我的意思是,我相信编码智能体公司也会很乐意与你合作来弄清楚这一点。是的,我确实认为 Claude Code 和 Codeex 与你们合作仍有用例,这很容易探索。好的,我的意思是,你知道,你一直非常乐于纵容所有这些事情。我只想把你从类型安全中带出来,一般性地谈谈,你已经非常清楚地表明了你在眼睛状态上的立场。给你更多空间在对齐安全方面的事情上。

Yeah, I mean, and I'm sure the coding agent companies would love to work with you as well to figure that out. Yeah, I do think that there's still use cases for Claude Code and Codeex with you guys, which it's easy to explore there. Okay, I mean, you know, you've been very obliging in indulging in all these things. I just want to take you out of types safe just generally about and you've made very clear your position on the state of the eye. Give you more room on the alignment safety side of things.

Diogo

哦,我完全没有谈到安全对齐吗?我想我没有。我想也许我没有。

Oh, did I not talk about safety alignment at all? I think I didn't. I think maybe I didn't.

Host

你谈过。我只是觉得,你知道,有很多,你有很多研究者的讨论。我们在欧洲每次都有这个。人们在谈论什么?所以例如,我最近在一个这样的研究者聚会上,人们真的担心节奏,对吧?就像整个话题关于我们应该放慢,因为公众显然还没准备好。我相信你有强烈的感受。我觉得这是那种危险的话题。我很乐意谈论它。我为危险而活。

You did. I just like, you know, I think that there's a lot of, you have a lot of researcher discussions. We have this every in Europe. What are people talking about? So for example, I recently was at one of these researcher gatherings and people are genuinely worried about the pacing, right? Like this whole topic about like we should slow down because the public is clearly not ready. And I'm sure you have strong feelings. I feel like this is the kind of thing that is a dangerous topic to talk about. I'm happy to talk about it. I live for danger.

Diogo

我们公司的品牌是混乱。不是 Jev。是不敬和混乱。

Our company brand is chaos. It's not Jev. It is irreverence and chaos.

Host

而且你知道。是的。

And you know. Yeah.

RLVR与前沿叙事 RLVR and the Frontier Narrative

Host

你当时在 OpenAI,经历了最早、最引人注目的事件之一,也就是那次小插曲,对吧?现在多米诺骨牌已经倒下,以至于每一家前沿实验室都联署了一份文件,说他们想要基于……

And you were at OpenAI during one of the very first very visible incidents, which is the blip, right? And the dominoes have gone down now to the point where every frontier lab has co-signed a document saying that they want to base...

Diogo

有意思。这是一个非常复杂、微妙的事情。我其实确实想更正式地写一篇回应。我有一个简短的版本,那就是:随着你越来越多地做 RLVR……RLVR 其实并不是关于可验证奖励的。这在推理革命之前就已经失败了。这就是任务的奇怪之处。当 RLHF 开始流行时,有三种不同的东西现在被称为后训练,不同的努力,而指令遵循是迄今为止更庞大的孩子。人们不喜欢它,不想考虑它,它很烦人。我跟预训练团队谈过。我说:‘伙计们,这才是魔法。’他们说:‘我们跑了那么多模型扫描。你想让我们等人类评估来决定用哪些模型?’每个人都给代码生成团队大量资源,他们确实取得了一些成功,但他们非常努力地想在单元测试上做强化学习,但显然没有成功,对吧?你需要推理才能做到。所以,明确一点,RLVR 并不纯粹是关于奖励的。它关乎一切的形式,其中一部分是推理被包含在这里,这个潜在变量,你在做事情,当你做事情时,你只是让模型为所欲为,以便让它们尽可能强大,以回答最难的问题。

Interesting. It's a very complicated, nuanced thing. I actually do want to write a response to this more formally. I do have a little bit of a short version of my response, which is that as you RLVR more... RLVR is not actually about verifiable rewards. That has been failing since before the reasoning revolution. And that's the weird part about tasks. Back when RLHF was becoming a thing, there were three different things that are now called post-training, different efforts, and instruction following was by far the vaster child. People didn't like it, they didn't want to take it into account, it was annoying. I talked to the pre-training team. I'm like, 'Guys, this is the magic.' And they're like, 'We'd run so many model sweeps. You want us to wait for human evals to figure out which models to use?' And everyone is giving tons of resources to the codegen team, which did have some successes, but they were trying really hard to do RL on unit tests, and it didn't work, obviously, right? You needed reasoning for that. So just to be clear, RLVR is not purely about the reward. It's about the shape of everything too, and part of it is that reasoning is included in here, this latent variable that you're doing things, and when you're doing things you're just letting the models do whatever they want in order to make them be as powerful as you can to answer the hardest problems.

Host

而整个‘领跑前沿’的讨论,我认为是一个非常狭隘的焦点,因为它假设每个人都需要做更多的 RLVR,对吧?

And this whole pace the frontier discussion I think is like a very narrow focus because it assumes that everyone needs to do more RLVR, right?

Diogo

我显然不认为我们需要在我们的模型上做更多的 RLVR,你懂的?我认为对我们这种形态来说,零是最佳量,对吧?拜托,你知道的?所以这真的,我认为有点像是障眼法,他们说我们实际上想继续做看起来危险的事情,因为它确实做危险的事情,你知道?就像人们说,‘哦,也许沙箱是个问题或者其他什么。’是的,我的意思是显然是的,他们本可以轻松解决,对吧?但他们选择不这样做,因为你让模型在这个‘做任何事’类别中做的事情越多,它就越强大,对吧?所以我认为那里有一些责任推卸,对于他们有意或无意地试图做的,只是一个假设。你知道,我们必须做 RLVR,不仅仅是必须做,我们必须做越来越多,给模型权力去做强大的事情,你知道,在中间做任何他们想做的事情,因为这教会它们在外部变得强大,我们不想限制那些事情,因为那会让它们在这些事情上稍微不那么强大。所以如果你假设所有这些,他们就会说,‘哦是的,我们正走向一个危险的世界,伙计们。’就像每个人都会这样做,这是让 AI 生病的唯一方式。所以基本上,这些都是内部一致的,但实际上从一个有替代方案的前提开始。

Which I obviously don't think I need to do more RLVR on our models, you know? I think zero is the optimal amount for our shape, right? Like come on, you know? So it's really, I think, a bit of a sleight of hand where they are saying that we actually want to keep doing the thing that looks dangerous because it does dangerous things, you know? Like people say, 'Oh maybe the sandboxing was a problem or whatever else.' Yeah, I mean obviously it is, and they could have easily solved that, right? But they chose not to because the more things you let the models do in this do-anything category, the more powerful it is, right? So I think there's some disillusion of responsibility there on things that by design or non-design they're trying to make is just an assumption. You know, we must do RLVR and not just we must do it. We must do more and more and more, with giving the models the power to do powerful, you know, do anything they want in the middle, because that teaches them to be powerful outside of it, and we don't want to limit those things because it'll make it slightly less powerful on those things. So if you assume all of that, they're like, 'Oh yeah, we're heading into a dangerous world, guys.' Like everyone is going to be doing this and this is the only way to make AI sick. So basically it's like these are all internally consistent, but actually starts from a premise that has alternatives.

Host

所以基本上,这些都是内部一致的,但实际上从一个有替代方案的前提开始。

So basically it's like these are all internally consistent, but actually starts from a premise that has alternatives.

Diogo

如果你想想,当然,我认为在‘苦涩的教训’方向上,很少有人做出了正确的任务,你知道,就像 AI 的新方向,或者新的北极星,那是罕见的。再次,就像我认为对于 LLM 本身来说,大概是 2.2 次左右,比如 RLHF,然后是 RLCD,RLVR 在我看来大概是 0.2,我觉得这已经很慷慨了,或者 0.5,或者可能是一整个,我不太在意。但我确实认为人们在这种事情上思维非常封闭,而这里唯一有错的人是研究人员,因为肯定不是大众,你知道?他们只是假设 OpenAI 和 Anthropic 只是在尽力而为,而他们并不是那些知道真正可选性的专家。

If you think about, of course, I think there's like on the bitter lesson direction, I think that there's very few people who've made right tasks, you know, like new directions of AI that is, or new north stars, that is rare. Again, like I think 2.2 times or something for LLMs itself, like RLHF and then RLCD, RLVR is like a 0.2 in my opinion, and I think that's generous, or 0.5, or like it could be one whole one, I don't really care. But I do think that people are thinking very closed-mindedly about this type of thing, and the only people who are at fault here are the researchers, because it's definitely not the populace, you know? Like they just assume that OpenAI and Anthropic are just doing the best they can, and they are not the experts who are aware of the true optionality available.

Host

是的。这很公平。而且你也在尽你的一份力唤醒他们。

Yeah. And that's fair. And you're also doing your part in waking them up.

Diogo

没错。嗯,我在尽力,但我的目标不是说服实验室有其他方向可走。我的目标是激发软件工程师的希望,开始真正自动化他们一直想要自动化的事情。我写过一篇文章,我的团队不让我写,不让我发表,关于我想要的 AI 未来,有很多小事情,比如记得‘做我意思的事’。想象一下,如果一切都能做我意思的事,因为那个演示就是‘做我意思的事’。就像你说,‘是的,不要做我说的,做我意思的。’我们还不能做我意思的事,因为计算机太基础、太字面了,但那个计算机使用的演示就是这样。我认为世界上会出现一些人们不理解的平滑程度。而智能无处不在的承诺,它是……我不想过度承诺。我不认为现在就会发生,但我们会尽我们所能让它发生。

Exactly. Well, I'm doing my best, but my goal is not to convince labs that there's other directions to go down. My goal is to spark hope in software engineers to start actually automating things they've always wanted automated. I had this article that I wrote that my team didn't let me write, didn't let me publish, about the future I want of AI, and there's a lot of little things like remember 'do what I mean.' Imagine if everything could do what I mean, because that demo was 'do what I mean.' Like you say, 'Yeah, don't do what I say, do what I mean.' And we couldn't do what I mean yet because computers are so basic and literal, but that computer use one was just that. And I think that there's levels of smoothness that will happen in the world that people just don't understand. And the promise of smarts all around, it's... I don't want to overpromise. I don't think it's going to happen right now, but we are going to do whatever we can to make that happen.

Host

是的。关于后训练的一般形态还有其他事情吗?你知道,你显然非常深入地参与了中期训练。那是你有评论的事情吗?我想我们从未谈过它。

Yeah. Any other things on the sort of general shape of post-training? You know, you've obviously been very intimately involved in mid-training. Is that something that you do have comments on? I don't think we've ever talked about it.

Diogo

中期训练。我的意思是,这都是一个谱系。

Mid-training. I mean, it's all a spectrum.

Host

是的。

Yeah.

Diogo

对吧。就像我……

Right. Like am I...

Host

这是一个课程,但更花哨。

This is a curriculum but like fancier.

Diogo

是的。我的意思是,这就像一种节省成本的东西,你知道,而不是必须重新预训练。就像有一些有趣的东西。我实际上认为智能在每一个层面上都有两面性,总是超级迷人。就像我是一个形状旋转者,所以我不喜欢发现它,但我喜欢别人发现并教我。但是,你知道,看数据,我们的数据团队非常擅长而我不擅长的事情,我觉得非常非常迷人。我喜欢思考能力是如何被放入模型的,比如短期内,你知道,有微调的快速对齐,长期来看,在反复看到之后,这些东西会越来越深地烘焙到模型中,直到变得稳健,你知道,这就是北极星,而系统一的东西最终会变得稳健。

Yeah. I mean, it's like a cost-saving thing, you know, instead of having to pre-train again. Like there's intriguing stuff. I actually think that intelligence has a Janus-like quality at every single level and it's always super duper fascinating. Like I'm a shape rotator so I don't like finding that but I love it when people find it and teach me about it. But and you know looking at the data, this thing that our data team is so good at that I'm not, I find it really really fascinating. I love actually thinking about like how capabilities are put into the model like over the short term, you know, like there's the really rapid alignment of fine-tuning and over the long term after seeing it over and over and over again like this stuff gets baked deeper and deeper and deeper and deeper into the model until it gets robust, you know, and that is like the north star to surface and like the system one stuff is the stuff that ends up getting robust.

中期训练与预训练 Mid-training and pre-training

Diogo

我觉得中期训练是件很迷人的事。我喜欢各种形式的训练,喜欢各种让新型智能浮现出来的方式。我不会全都自己去做,因为太贵了。我私下说过,也——我该说这个吗?我的哲学是:任何我该私下跟投资人说的话,我也该公开跟人们说,因为这就是我的风格。所以我之前说过:如果你给我十亿美元,我也不会做预训练。我仍然相信这是对的。这是件非常昂贵的事。如果你是工程师,你可以切切割割、做各种事情,你知道,就像科学怪人拼凑,不是最优雅漂亮的东西,但它能解决问题,宝贝。

So I find mid-training to be a fascinating thing. I'm a fan of all forms of training. I'm a fan of all forms of surfacing new types of intelligence. I wouldn't do it all myself because it's expensive. And I have said privately and also — should I say this? My philosophy is anything I should say in private with an investor, I should say in public with people because that is my thing. So the thing I've said before is if you gave me a billion dollars I wouldn't pre-train. I still believe that to be true. It is a very expensive thing. If you're an engineer you can slice and dice and do all sorts of stuff, you know, like Frankensteining is not the most elegant beautiful thing but it solves problems, baby.

全能模型vs专用模型 Omni-models vs specialized models

Host

所以除了重新训练,其他都行?

So anything except retraining?

Diogo

是的,太棒了。我觉得有一个方向确实有趣,就是把你关于这些模型的所有评论综合起来:我们是要一个包含所有这些能力的超级模型,还是把它们进一步拆开?一种说法是,OpenAI 曾朝着全能模型的方向发展。4o 就是其中之一。然后有一小段时间,总有一个主分支:这是聊天调优的模型,这是编码调优的模型。那些是完全不同的东西,极其不同的概念。我来稍微拆解一下。多模态有点不同,因为有时其他模态有帮助,有时有害。比如人们似乎在远离语音,语音和音频不同,因为它似乎不能很好地泛化到其他东西上。这可能会被解决。我喜欢所有这些,但这些都是实证性的真问题。缩放定律不是关于砸钱就能变好。缩放定律实际上是关于一个东西能有多好。有些情况下无论你怎么扩展可能都不够好。所以据我所知,计算机使用目前还没解决。我希望我们能参与解决它,但可能无论我们收集多少数据都解决不了。我们可能需要更好的方法或其他东西。所以你需要非常务实。我喜欢全能模型吗?我喜欢所有形式的智能,但我要直接进入你谈到的与预训练不同的一个东西,那就是后训练,因为我讨厌把智能割裂。那对我来说是坏事。这种聊天优先的推理模式,是因为它迫使智能被割裂。当你为聊天优化时,这往往是纯粹的基于人类反馈的强化学习(RLHF),而 RLHF 中很内在的东西就是做人们自然抱怨的那些事,对吧?

Yeah, amazing. I think one direction that I do think is interesting, just synthesizing all your commentary about these model things, is: do we have a super model that has all these capabilities involved, or do we break them out further? So one way to put this is OpenAI was trending in the direction of the omni model. 4o was one of those. Then for a brief period of time there's always like a main branch: this is the chat-tuned model and this is the coding-tuned model. Those are completely different things, extremely different concepts. I'll break that down a little bit. So multimodality is a little bit different because sometimes the other modalities help, sometimes they hurt. Like, people seem to be moving away from speech, which is different than audio, because it seems to not generalize well to the other stuff. This might get solved. I'm a fan of all of this, but these are empirical real questions. Scaling laws are not about just throw money at it and it gets good. Scaling laws are pragmatically how good is a thing. There are worlds where no matter what you scale it may not be good enough. So computer use is not currently solved, is my understanding. I'm hoping that we can play a part in solving that, but there might be no amount of data we collect that will solve that. We might need better methods or something else like that. So you need to be really practical in all of this. Am I a fan of omni-models? I'm a fan of all forms of intelligence, but I will go straight into one thing you talked about which is different from pre-training, which is post-training, because I hate fracturing intelligence. That is the bad thing to me. And this whole chat-first reasoning mode is because it forces the intelligence to be fractured. When you're optimizing for chat, this tends to be pure RLHF, and it's quite intrinsic in RLHF to do the stuff people like naturally complain about, right?

RLHF与智能分裂 RLHF and fracturing intelligence

Host

哦,你说得太对了。

Oh, you're absolutely right.

Diogo

你知道,谄媚,谄媚,不管这个词怎么念——过度自信、幻觉,甚至那种在 LMArena 上表现出色的风格。粗体、斜体、表情符号,你知道,它不会简单地回答问题。它给出一大段文字,然后问你一个后续问题,这样感觉更像人在跟你说话。所有这些都因为字符串非常奇怪,你知道吗?它们是很奇怪的东西,你需要校准错误。你需要模式崩溃。你需要过度自信才能不失控,因为奖励模型会在发生时狠狠惩罚你,因为那很明显。然后这完全扭曲了概率空间,并且与推理模型的概率空间相互作用,对吧?因为这些模型就像简单的线性东西,倾向于作弊。所以我认为这与暴露智能非常不同,这是我的猜测。而智能的很多艺术在于研究这种微妙之处,我认为至少我在 OpenAI 时人们并没有真正研究这个,因为他们只是聊天聊天聊天,就像现在人们对 Jeff 那样。

And you know, sycophancy, sycophancy, whatever word, how to pronounce that — overconfidence, hallucination, like even the kind of style that excels on LMArena. Bold, italicized emojis, you know, like it doesn't answer the question simply. It gives a long write-up and then it asks you a follow-up question so it feels more like a human talking to you. All of these things come because strings are super weird, you know? They are like weird-ass things and you need to be miscalibrated. You need to mode drop. You need to be hyperconfident in order to not go off the rails because the reward model will punish you so hard when it happens because it's obvious. And then this warps the probability space entirely and it interacts with that of the reasoning models, right? Because the models are like these simple linear things that tend to cheat a bit. So I think that that's very different than exposing intelligence, is my guess. And a lot of the art to intelligence is studying this subtlety that I think at least when I was in OpenAI people were not really studying that because they were just like chat chat chat chat, just like people are with Jeff right now.

System 1 vs System 2 System 1 vs System 2

Host

是的。你给我一个终极目标,你知道,你给我一个目标,我就会去优化它,对吧?但如果你尝试——你知道有句话说你可以有两个目标,你可以同时优化两者,但那实际上就是割裂的行为,对吧?所以是的。

Yeah. You give me an ultimate, you know, you give me an objective, I will just go optimize for that, right? But if you try and — you know like the saying is like you could have two objectives and you could just optimize for both but then that is literally the act of fracturing, right? So yeah.

Host

所以我的意思是,在某些方面你——你也在把智能分解为系统一、系统二,但你只是不同意其他人的割裂方式。

So I mean in some ways you have — you are also factoring intelligence into system one, system two, but you just don't agree with the other people's fracturing.

Diogo

这有点不同。不,如果我可以补充——如果我可以为系统二任务辩护。第一,我们不会扔掉系统二任务。对吧?你可以试着让 Jeff 处理它,实际上有一个智能的答案,那是未知的。系统二任务中有更好和更差的行为,应该非常低置信度、大量不确定性,也许一些启发式方法可以在这里或那里推动一下,但我们也关心它们,只是说清楚。我只是认为那不是智能的本源。所以我们不是要那样割裂任何东西。所有割裂都会让模型变笨。你知道,如果人们让模型说它是 OpenAI 或 Qwen 或 Claude 或其他什么——我不知道现在它说什么。我不会把“你是来自 TypeSafe 的 Jeb”放进模型里,那会割裂它,对吧?我不想要那样。它代表互联网的想法,对吧?要正确。这就是我想要的,因为这样你才能得到平滑可预测的智能。我的意思是,身份是个东西,我想对于第一方产品是的,但对于 API 我不这么认为。人们不想要——如果他们在用 ChatGPT 做聊天机器人,他们不想说它是 ChatGPT,他们想说它像 Chipotle 或其他什么,对吧?

Which is a little different. No, if I could add — if I could defend the system two tasks. Number one, like we don't toss out the system two tasks. Right? Like you can try to make Jeff work on it and there actually is an intelligent answer for that which is unknown. Like there is better and worse behavior in the system two tasks which should be like really low confidence, lots of uncertainty, maybe some heuristics can move the needle here and there, but we care about them too just to be clear. I just think that that is not what the intelligence is native to. So we're not trying to fracture anything like that. And all fracturing makes the model dumb. You know, like if people get the model to say like it is OpenAI or Qwen or Claude or whatever else — I don't really know what it says these days. I am not going to put into the models that you are Jeb from TypeSafe, that fractures it, right? Like I don't want that. Like it represents what the internet thinks, right? Be correct. That is what I want because that's how you get the smooth predictable intelligence. I mean identity is a thing I guess that for a first-party product yes, but for an API I don't think so. Like people don't want — if they're making a chatbot with ChatGPT they don't want to say it's ChatGPT, they want to say it's like Chipotle or whatever, right?

Jeff技能与编码代理 The Jeff skill and coding agents

Host

嗯,你知道,所以你也必须弥补这一点的方式是你有那个技能,对吧?Jeff 技能,那是为了让编码智能体与 Jeff 一起工作。

Well, you know, so the way that you also have to make up for it is you have the skill, right? The Jeff skill, which is for coding agents to work with Jeff.

反思旅程 Reflecting on the journey

Host

好的,几个收尾问题,因为我想让你走了。一个是回顾你的两年旅程。这大约两年,

Okay, a couple of closing questions because I do want to get you out. One is just reflecting on your two-year journey. This is roughly two years,

Diogo

两年多。

Two point something.

Host

与公司一起。我认为这更像是一个四年的旅程,但是的,实际上我在想,记得你在感恩节前后有一次英雄式的冲刺。你取消了一切,因为你说,伙计们,大家都在度假。我要拿走所有空闲的 GPU 去做这件事。

With the company. I think that this is more like a four-year journey, but yeah, actually I was thinking, remembering that you had this hero run around Thanksgiving. You were canceling everything because you were like, guys, everyone's on holiday. I'm going to take all the opening GPUs and go do this thing.

Diogo

是的,那是一段美好的时光。

Yeah, that was a good time.

政变与ChatGPT前时刻 The Coup and the Pre-ChatGPT Moment

Host

那大概就是 pre-type safe 那个时刻吧?那会儿是不是正好在闹那场“政变”?我也不是很清楚。

And that was like the pre-type safe moment, right? Was that when the coup was happening? I don't really know.

Diogo

对,确实是。

Yes, actually.

Host

对,对,应该是这样。我记得。天哪。我现在没时间爆那场“政变”的料,不过那件事倒不算太烦人。

Yeah. Yeah. That sounds right. Yeah, I remember. Oh my god. I don't think I have the time to spill the tea about the coup right now, but that wasn't really annoying.

Host

是“政变”烦人,还是那次运行烦人?

The coup was annoying or the run was annoying?

Diogo

是“政变”烦人。

The coup was annoying.

Host

好吧。对,对。

Okay. Yeah. Yeah.

Diogo

安全团队接管了公司。

Safety took over the company.

Host

对。总之,

Yeah. Anyway,

Diogo

也许下次聊天我再爆那场“政变”的料。嗯,其实这个问题在 ChatGPT 发布之前就一直在我脑子里了。我当时就想:“天哪,ChatGPT 团队在搞大事。他们做的是对的事。他们在做 AI 研究者不擅长、但成功的产品人擅长的事,就是花大量心思在体验上。”你知道,这很罕见。OpenAI 里这样的人很少。而那些人在这件事上做得非常非常好。

maybe next time we chat I'll dump tea about the coup. Um, yeah, actually this problem was one that was in my mind since before ChatGPT even launched. I was like, "Holy, the ChatGPT team is cooking. They are doing the right task. They are doing the thing that AI researchers are bad at but successful product people are good at, which is giving a lot of thought about the experience." You know, it's very rare. Like there's very few people like that at OpenAI. Um, and those guys were cooking on it really really well.

Host

说清楚一点,这是从 GPT-3 到 3.5 的整个历程,其中还包括 AI Dungeon,你之前把它说成是——

And to be clear, this is the whole journey from GPT-3 to 3.5, which included AI Dungeon, which you've talked about as like—

Diogo

嗯,那是一个我们从未预测到的用例。

Well, that's an example of a use case that we never predicted.

Host

对,正是如此。

Yes, exactly.

部署InstructGPT与文案陷阱 Deploying InstructGPT and the Copywriting Trap

Diogo

嗯,哦对,那也是——我当时为了部署 InstructGPT 拼得非常非常凶。其实它早期版本甚至是用一个我们没发表的算法训练的,那算法是我自己做的,因为清洗 PO 数据太慢了。我当时就想,这东西太好了,我们得把它交到用户手里。结果它几乎立刻就占了当时 LLM 市场 50% 的份额。但我们当时以为——我费了很大劲,确保我们发布视频里的每一句话都是真的。嗯,你知道,我真的在想,这是 AGI 吗?因为它在“指令进、指令出”上是超人水平。你当然知道——它不是,但我觉得每个人都应该对“为什么那不是 AGI”有个答案,因为它看起来非常聪明。而我对这个问题的答案,最后发现它只被用来做文案。你知道,Jasper AI、Copy AI,就是写那种现在被叫做网页上的“slop”的东西。嗯,我们当时担心自己把互联网变得更糟了,对吧?于是我回到白板前,想:缺了什么?我们显然很聪明。它缺了某种东西,比如创造价值。缺的是什么?其实我当时做了更多哲学思考,比如,到底发生了什么?答案是,哦,机器。你知道,我问自己的问题是:让我们从一场基于 AI 的经济革命倒推。当那发生时,如果 AI 是一个 API,我们会把 AI 叫做什么?会是人类,还是代码?我判断会是很多个 9 的代码。但所有优化都跑到了人类那一部分。然后我突然想通了。我想,天哪,这就是北极星。我想——我写了一份文档。我跟 Sam 聊这个。Sam 说,这太好了,你应该去做这个。我们说,好好好 Sam,我有工作。你知道,嗯,我当时在做——

Well, oh yeah, that is also—I had fought very very hard to deploy InstructGPT. Um, like actually the early versions of it were even trained with an algorithm we didn't publish that I made myself because it was too slow to clean the PO data and I was like, this is so good. We need to get it in the hands of users. And basically immediately it took 50% of the market share of LLMs at the time. But we thought—I made—I went through great effort to make sure everything in our launch video is true. Um, you know, we—I truly was thinking, is this AGI? Because it's superhuman at instruction in, instruction out. You obviously—it's not, but like everyone I think should have an answer to why that was not AGI, because it looks very smart. And my answer to that ended up only being used for copywriting. You know, Jasper AI, Copy AI, like writing like, you know, what is now called slop on web pages. Um, and we were worried we made the internet a worse place, right? And I went back to the drawing board and I was like, what's missing? We are smart clearly. Something is missing from it like creating value. What is it? Like I actually was doing more philosophy at the time of like, you know, like what is going on? And the answer was, oh, machines. You know, the question I asked myself is like, let's work backwards from an AI-based economic revolution. When that happens, what will we be calling the AI if AI is an API? Will it be humans or will it be code? And I figured it was many nines of code. But like all the optimization was going into the humans part. And then it clicked for me. I'm like, holy, this is the north star. I think—like I wrote a document. I was talking to Sam about this. Sam was like, this is so good. You should go work on it. And we're like, yeah yeah Sam, I have a job. You know, like, um, you know, I was working on—

Host

Sam 都让你去做了。那就去做啊。

Sam just told you to do it. Go do it.

Anthropic、Claude与推理竞赛 Anthropic, Claude, and the Reasoning Race

Diogo

但我当时的猜测是,这太明显了。明显到 Anthropic 肯定已经在做这个了,你知道,我们已经被做掉了。而且其实 OpenAI 更擅长追赶,而不是真正创新。所以 ChatGPT 就是 Claude 的复制品,对吧?嗯,他们内部有个东西,只是没发布。

But my guess at the time is like, this is super obvious. Like it's so unbelievably obvious Anthropic must be working on this already, you know, and like we're already cooked. And like actually OpenAI does better at like catching up than it does at like actually innovating. So like ChatGPT was a copy of Claude, right? Um, like they had an internal thing. They just didn't ship it.

Host

对。Claude 和 Slack。

Yeah. Claude and Slack.

Diogo

但是,呃,但你知道,推理这块我会说他们算是第一个吧。

But, uh, but you know, reasoning I would say first-ish.

Host

对。但那产品到底有多好,可以商榷。

Yeah. But debatable how good of a product that is.

Diogo

对。嗯,不过研究很棒。非常棒的研究。我只是不确定人们是否有那个产品需求。嗯,而且你知道,Claude 也做了编码智能体那套东西。

Yeah. Um, great research though. Super great research. I'm just not sure if people had that product need. Um, and you know, Claude did the coding agent stuff too.

离开OpenAI并创办实验室 Leaving OpenAI and Starting the Lab

Diogo

所以 Sam 那么说,你知道,我就回去干我的活干了一阵子。最终,你知道,指令跟随团队直接说我们赢了。我们解决了指令跟随。我们不用再做东西了。我就想我接下来做什么。我想,你知道,也许我就开始玩玩这个。嗯,我,你知道,做更多哲学、设计和思考。我开始训练模型时以为最后会花一周。结果花了好多年。某个时候我想,天哪,你知道,这里有生命迹象了。这显然没成功,对吧,否则我们就部署了。但我想从研究角度探索,如果全力押注这个会是什么样,你知道,我真的想看看,如果你朝这个方向疯狂地、彻底地全力押注,会是什么样。而且因为我说过的话,你知道,如果 AI 寒冬来了,我会怎么想?我会认为自己个人有责任。我当时跟其他公司聊过,我说:“嘿,我想在这个方向上开一个实验室。”你知道,当时有人感兴趣,我就问他们,多快——哪个更快,这个还是创业公司?他们说,创业公司。我就想,妈的,那就干吧,看来我们要搞点疯狂的了。

So Sam says that and you know I just go back to my job for a while. Eventually like, you know, the instruction following team just says we won. We've solved instruction following. We don't need to do stuff anymore. I'm like trying to think about what I do next. I'm like, you know, maybe I'll just start playing around with this. Um, I, you know, do more philosophy and design and thinking. I thought it would end up taking a week when I started training models. It ended up taking many years. At some point I was like, holy, you know, there's signs of life here. This obviously didn't work, right, otherwise we would have deployed it. But like I want to explore what it would be like research-wise to go all in on this, you know, like I want to really see like what it would be like if you went like absolutely insanely all in in this direction. And because of what I said, you know, like if an AI winter happened, how would I feel? I would consider myself personally responsible. I talked to other companies at the time and I was like, "Hey, I want to start a lab on this direction." And you know, like there was interest and I just talked to them like, how fast—what would be faster, this or startup? And they're like, startup. And I'm like, it man, we ball, I guess we're doing something crazy.

Host

然后你打电话给 Eric 和 Sasha,还有——

And you call Eric and Sasha and—

Diogo

对。嗯,我先打给 Eric。嗯,对 Sasha,其实我并没有想招募她。我尽量做个好人,我只是说,嘿,我是不是疯了?这里是不是缺了什么?你知道,是不是——我是不是太陷在 OpenAI 的泡泡里,以至于没意识到这个问题一定有解?然后 Sasha 说,我加入。我说,Sasha,你在创业公司工作。她说,我现在就把它关了。我说,你要不要考虑一下?她说,哦对,说得好。让我想想。然后她加入了。然后,你知道,两周内我们拿到了融资,然后——有人搬进了我的公寓。那太糟了,因为我有洁癖。嗯,我们就一直搞,最终搞出了那个显示生命迹象的研究。你知道,那是一段疯狂的日子。

Yeah. Well, I call Eric first. Um, with Sasha I actually didn't try to recruit her. I try to be good and I was just like, hey, am I crazy? Is something missing here? You know, isn't there like—am I too much in the OpenAI bubble that I didn't realize there must be a solution to this? And then Sasha was like, I'm in. And I'm like, Sasha, you're working at a startup. And she's like, I'm folding it right now. And I'm like, do you want to think about that? She's like, oh yeah, good point. Let me think about it. And then she joined. And then, you know, within 2 weeks, we had funding, within—we had like people move into my apartment. It was the worst because I'm a neat freak. Um, and we just kept on cooking and eventually we got the research that showed the signs of life. You know, it was a crazy time.

给沮丧前沿实验室研究员的建议 Advice to Frustrated Frontier Lab Researchers

Host

所以问题是,那都是长上下文,而现在的问题是,像你这样的人现在在前沿实验室里,因为拿不到资金或资源,或者随便什么,嗯,关注度,而感到沮丧。嗯,你对他们有什么建议?你知道,他们应该做你做过的事吗?他们应该这么做吗?

So the question is, that was all long context, and now the question is, someone like you is in the frontier lab right now who is frustrated not getting the funding or the resources, whatever, uh, the attention. Uh, what's your advice to them? You know, do—should they do what you did? Should they do it?

Diogo

这是个很有意思的问题。嗯,天哪,我怎么才能不烧桥地回答这个?嗯,我的感觉是,大多数——除非存在某种我其实不太理解的经济逻辑。嗯,我觉得大多数 Neolab 都是垃圾。嗯,我不想把自己和那些东西看作同行。嗯,我不太理解那边到底在发生什么。是因为——第一,我其实不太看重研究者。

That's a fascinating question. Um, man, how do I do this without burning bridges? Um, my sense is that most—unless there's some level of economics I don't really understand. Um, I think most Neolabs are crap. Um, I don't want to see myself with that as peers. Um, like I don't really understand what's going on there. Like is it because—like number one, I don't really value researchers.

重视研究员与北极星任务 Valuing Researchers and Northstar Tasks

Diogo

我重视那些关注我最痛苦教训的人,对吧?不仅仅是我们需要研究人员,而是我们需要他们多思考正确的任务,这才是重要的,对吧?所以当人们看重纯粹的研究血统时,其实是本末倒置的,因为这通常不会创造价值。所以第一,我相信北极星任务,做酷的、真正有用的事情。第二,因为我不看重研究人员,我不推荐走那条路,显然对某些人来说是有利可图的,或者在这种环境下可能有利可图。所以从纯粹务实角度看,我不认为创建 Neolabs 能创造价值。它似乎是在破坏价值,因为他们从头重做工作,真正推动前沿的概率很低。就我和大多数 LEO 实验室交流的情况看,他们并没有真正的方向。他们往往想要钱来玩他们的实验。如果他们有一个方向,我超级支持,说清楚。所以我对某人的建议真的取决于你为什么做这件事。如果你是一个想玩玩研究的研究者,可能实验室是最好的地方。说实话,可能还有其他地方。我不太关注那些政治,但我个人建议不要那样。我认为如果人们被驱动去解决真正的问题,对世界更好。那些问题可能是探索性的,那没关系。但理想情况下,要有你支持的原则。但如果你认为你想做正确的任务,绝对,请做,请打破这种单模态思维。

I value people who look at my bitterest lesson, right? Not just that we need researchers, but we need them to give a lot of thought to the right task, and that's the important thing, right? So it's actually kind of backwards when people value pure research pedigree, because that generally doesn't create value. So, number one, I believe in northstar tasks and doing cool, really useful stuff. Number two, because I don't value researchers, I don't recommend going the well, it clearly is profitable for someone or it might be in this environment. So from a purely pragmatic perspective, I don't see creating Neolabs as something that creates value. It seems to destroy value because they are redoing work from scratch with low probability of actually moving the frontier. And as far as I've talked to most LEO labs, they don't really have a direction. They tend to want money to play around with their experiments. If they have a direction, I'm super in favor of it, to be clear. So my advice for someone is it really depends on why you're doing it. If you are a researcher who wants to play around with research, probably the labs are the best place to do that. TBH, there might be other places. I don't really keep track of that politics, but I would just recommend not being that way personally. I think it's better for the world with people being driven to solve real problems. And those problems may be exploratory. That's fine. But ideally have principles that you stand behind. But if you think that you want to do the right task, absolutely, please do, please break this unimodal mind.

锯齿状前沿与政治定位 The Jagged Frontier and Political Positioning

Host

没错。就像再次,这种对前沿的节奏来自一种对 AI 的看法,看起来像 AI 超级天才,极其参差不齐,而且是可以解决的。

Exactly. Like again, this pacing the frontier is coming from this one view of AI that looks like AI super genius that is incredibly jagged and that is solvable.

Diogo

它是可以解决的,很奇怪,不符合现实,很悲剧,对吧?我认为所有这些真正挖掘技术都是好的。

It's solvable and it's weird and it's not matching reality and it's tragic, right? I think all of these really unearthing technology is just good.

Host

是的。不管怎样,再次,我试图准确代表我交谈过的 Anthropic、OpenAI 的人,顺便还有 SpaceX,的观点,那就是这更多是政治问题,而不是纯粹的技术问题。

Yeah. For what it's worth, again, I'm trying to accurately represent the position of the Anthropic, OpenAI folks I was talking to, SpaceX as well, by the way, is that this is a political thing much more so than a pure X-T thing.

Diogo

是的。所以是的,政治定位,那超出了我的,那远远超出。

Yep. So yeah, political positioning and that's beyond my, that's well beyond.

Host

一旦他们告诉我,我就想,我明白了。这是关于 2028 年选举。

Once they told me that, I was like, I get it. This is about the 2028 election.

Diogo

哦,不。哦,我希望我没听到那个。那真是糟糕的感觉。

Oh, no. Oh, I wish I didn't hear that. That's such a bad vibe.

Host

不,不,不。这不是整个公司。这只是那个房间的讨论。

No, no, no. This is not the whole company. This is just that room's discussion.

Diogo

不,不。那有道理。那让我对人类失去一点信心。但也许我只是一个天真的技术专家。

No, no. That makes sense. That makes me lose faith in humanity a bit. But maybe I'm just a naive technologist.

Host

谁掌管政府,真的开始变得重要,因为政府将帮助监管这些出现的事物。作为一个实验室,你可能应该考虑清楚。

It's really starting to matter who's in charge of the governments that will help to regulate these things as they emerge. And as a lab, you should probably think that through.

Diogo

不,不,我完全同意,说清楚。我认为对此有主见很重要。我个人害怕试图误导人们,因为我认为那会让人吃大亏。我认为人们试图过度自信,我显然不会真的谈论政治。我认为在 CO 发生的事情是人们过于倾向于诉诸权威和过度自信,试图让人们以某些方式行事,显然我们的回应极其不理想。那产生了下游的连锁反应,现在我认为对世界极其糟糕。也许我天真。我认为即使为了更大的善或他们认为的更大的善而误导人们,也只是,我不喜欢。我宁愿不。

No, no, I totally agree with that, to be clear. I think being opinionated on that matters a lot. I personally am afraid of trying to mislead people because I think that bites people in the ass a lot. I think that people trying to be overconfident, I obviously I'm not actually going to talk about politics. I think what happened in CO is people leaned too much in appeals to authority and being overconfident to try to get people to behave in certain ways, and obviously our response was extremely suboptimal. And that had ripples of downstream ramifications that are now I think extremely bad for the world. Maybe I'm naive. I think that misleading people even for the greater good or what they think is the greater good is just, I'm not a fan. I'd rather not.

Host

不管怎样,我不认为这是误导。它就像这就是为什么现在。

For what it's worth, I don't think it's misleading. It is just like this is why now.

Diogo

为什么,是的,就像你知道的。我认为为什么那就是为什么现在,关于风险与目标,那有点误导。其中潜伏着某种程度的狡猾,值得指出,我认为值得承认。好吧,显然他们想要,如果他们想操纵,那么他们不应该承认。那似乎是一个糟糕的策略。但对我来说,那只是对世界感到悲伤。

Why, yeah, like you know. I think how come that is why now that is a little bit misleading about the risks versus the objective. There's some level of sneakiness latent in it that is worth calling out and I think owning up to. Well, obviously they want to, if they want to manipulate then they shouldn't own up to that. That seems like a bad strategy. But that to me is just sad for the world.

Host

是的。

Yeah.

Diogo

希望我永远不会,是的,希望我们永远不会卷入类似的事情。随着我们变大,这可能不可避免。但我想尽可能保持纯粹技术专家的本色。

Hopefully I'm never, yeah, hopefully we are never involved in anything like that. It might be inevitable as we get big. But I want to stay pure technologist to my roots as much as I can.

Host

我是说,杰夫当总统,为什么不呢?我可以,你知道,我相信杰夫的决定胜过我自己。好的。所以,少发帖。更多关于你正在粉碎我对美国和世界的希望。

I mean, Jeff for president, why not? I can, you know, I trust Jeff's decisions over my own. Okay. So, less posting. More about you're just crushing my hopes about America and the world right now.

Diogo

哦,我的天。

Oh my lord.

Host

是的。我是说,我想我看了太多关于阴谋夺取总统职位的电视。你已经选择了你的北极星,你选择了可靠性,然后是可编程和可组合的 AI,还有便宜。

Yeah. I mean, I think I watched too much TV about conspiracies to take over the presidency. You have chosen your northstar, you have chosen reliability and then programmable and composable AI and cheap.

给别人的任务:智能游戏 Tasks for Others: Intelligent Games

Host

第二个或第三个是什么,你想扔给别人作为骨头,你不会去做的?

What is a second or third one that you want to throw as a bone to someone else that you're not going to work on?

Diogo

哦,就像基本上给人们任务。

Oo, like just basically give people tasks.

Host

给人们任务。

Give people tasks.

Diogo

是的。就像你的任务,我想要的有太多。

Yeah. Like your task, there's so many I want.

Host

你已经选择了你的任务,对吧?你明白我的意思吗?

You have picked your tasks, right? You know what I mean?

Diogo

什么?等等,那真是个好问题。天哪。哦,天,我太兴奋了。因为接下来的 50 年你会忙于做你的事情。

What? Wait, that's such a good question. Holy crap. Oh man, I'm so excited by that. Because you're going to be the next like 50 years you're going to be busy doing your thing.

Host

当然。好的。所以,让我给一个有趣的,一个无趣的,也许一个既有价值又有趣的。

Hell yeah. Okay. So, let me give like a fun one and a not fun and like maybe a valuable one that's also fun.

Diogo

我有趣的一个是,我认为如果游戏有智能,它们会超级酷。就像当我看到人们玩 Alli 的 Doom 演示,你可以让 NPC 控制东西,那真的只是一个概念验证。我认为可以做出一些非常酷的东西。它看起来非常非常酷。就像我是 Stardew Valley 的超级粉丝,它非常静态,但仍然引人入胜。我觉得可以发生很多酷的故事。你不需要在游戏循环中调用 Jeb。那可能太贵了。但即使是简单的 NPC 状态机,我认为你可以创造一个如此引人入胜的世界。哦,天。而且,有点难过我不能做这类事情。

My fun one is I think games could be so freaking cool if they were intelligent. Like when I see people play around with like Alli's Doom demo where you can get NPCs to control stuff, that was just really a proof of concept. I think some really cool stuff could be made. It looks really really cool. Like I'm a big Stardew Valley fan, and it's really static and it's still compelling. I feel like there's a lot of cool story that could happen. You don't need to call like Jeb in the game loop. It's probably too expensive for that. But even like simple state machines for NPCs, I think you could make such a compelling world. Oh man. And man, a little sad that I can't work on these types of things.

摆脱KV缓存的编码代理 Coding agents free from the KV cache

Diogo

我的人生路径现在算是定下来了。但我真正、真正想去探索的,是让编程智能体摆脱 KV 缓存的暴政。它可能不如真正的编程智能体那么好,但我觉得有太多奇怪的东西值得思考了。这就是我写那篇《KV Cache Rules Everything Around Me》的原因。信不信由你,我觉得互联网上没人用过这个说法——“cash rules everything around me”,C-A-C-E。我搜的时候一个都没有。我写这个,是想告诉大家编程智能体是怎么工作的、KV 缓存是怎么工作的,等等。它解释了很多东西——为什么路由这么难,为什么子智能体好像不管用,为什么压缩是个这么难的问题。我打算发布一份文档。我的团队可能会否决我,因为信不信由你,我说了不算。但我想发布一份文档,就说:这是我的想法,请大家拿去玩,请把编程智能体一旦摆脱 KV 缓存暴政之后能做的所有事情都摸索出来。

My life path is a little bit set right now. But the thing that I would really, really like to explore is coding agents free from the tyranny of the KV cache. It might not be as good as true coding agents are, but I think there's just so many weird things to think about. That's why I wrote the article "KV Cache Rules Everything Around Me." Believe it or not, I don't think anyone has used the phrase on the internet — "cash rules everything around me," C-A-C-E. When I Googled it, nothing. So I wrote this because I wanted to tell people about how coding agents work, how the KV cache works, and everything. It explains a lot of stuff — why routing is really hard, why sub-agents don't seem to work, why compaction is such a hard problem. And I'm going to try to release a document. My team might veto me, because believe it or not, I'm not in charge. But I want to release a document like, here are my thoughts. Please play with it and please figure out all the ways that we can do things with coding agents once you're freed from that KV cache tyranny.

Host

到底是哪个——它把你锁死了?

Which is it — it locks you in?

Diogo

不是锁死——是把你锁进一个模型里,对吧?而且为了高效,你得不停地往里追加。于是你就不再遵循最好的软件实践了,比如状态管理、抽象、分解。为什么不能给一个更简单的任务?为什么不能给子智能体一个更简单的任务?因为你传递的状态。为了传递这些状态,你需要一种比用来读它的智能便宜得多的智能。为什么不能聪明一点处理?我觉得这里面有大量非常酷、非常有趣的研究可以做,关于不同的编程模式。就像大家在玩递归语言模型那样。我觉得这里面有太多酷东西了。当你想到,哦,我想把一切的状态都显式标注出来。或者想象你有一个子任务——从分解的角度看,编程智能体一次处理一个子任务,我觉得这么说没问题。为什么你需要把所有这些状态都传回父任务?为什么不能聪明地处理它?而且,如果你有一个带标签的子任务层级,为什么不能在需要的时候,在那棵子任务树里搜索相关的上下文?然后还有一件你能做的事——天哪,我忘了写这个。我这里有些想法真的非常酷。希望能发表。我很乐意聊,但那会是一份很长的文档。如果上下文变得便宜了,为什么不能做一些很酷的模式,比如非常廉价地查看你的历史上下文?每次你都从零开始,还得解决一个叫持续学习的问题,这不是很奇怪吗?那其实是个内存管理问题,因为你没有一种聪明的方式来查找记忆,对吧?但如果你能一直这么做呢?或者,当你有并行的子智能体时,它们能读彼此的状态,因为这一切都在你的计算机内存里,你可以聪明地处理谁在读、谁在写,你的编程智能体集群之类的东西对事物加锁,能智能地协调——而不是用那种最基础的锁,什么“你在干嘛、我在干嘛、谁先写,等等等等”。我觉得那里的未来简直疯狂。

Well, not locks you in — it locks you into one model, right? And in order to do it efficiently, you need to keep on appending to it. So now you're not doing best software practices like state management, abstraction, decomposition. Why can't you give an easier task? Why can't you give a sub-agent an easier task? Because of the state that you're passing around. You would need intelligence that is way cheaper than the intelligence you're using to read this, in order to pass this state around. Why can't you be smart about it? And I think there's tons of really cool, fun research to be had there on different programming patterns. Kind of like how people are playing around with recursive language models. I feel like there's just lots of cool stuff in here. When you think about, oh, I want to explicitly label the state of everything. Or imagine you have a subtask — coding agents, I think it's fair to say, work on subtasks at a time from a decomposition perspective. Why do you need to pass all of that state back into the parent task? Why couldn't you do smart things about it? And also, if you had a hierarchy of labeled subtasks, why can't you do a search through that subtask tree for the relevant context when you need it? And then another thing that you can do — oh man, I forgot to write something about this. I have some cooks in here that are really, really cool. Hope to publish it. I'm down to jam about it, but it's going to be a long document. And if that becomes the case where context becomes cheap, why can't you do cool patterns like looking at your historical context very cheaply? Isn't it kind of weird that you start from scratch every time and you need to solve a problem called continuous learning? That's actually a memory management problem, because you don't have a smart way of looking up the memory, right? But what if you could do that all the time? Or what if, when you have parallel sub-agents, they can read each other's states, because you have all of that in your computer memory, and you can be smart about what's reading and writing at the same time, and your coding agent swarm or whatever has locks around things and can coordinate intelligently — not with basic-ass locks like "what are you doing, what am I doing, who should write first, blah blah blah." I feel like the future there is nuts.

Host

让 Jeff 来解决锁的问题?

Jeff to solve locks?

Diogo

我是说,对于多个智能体协同工作来说,那可能非常酷,或者如果你把状态想成智能体形态的话。

I mean, it could be so cool for multiple agents working together, or if you think about state like agent forms.

Host

对。

Yeah.

Diogo

而且你知道,有些东西,比如,是只读进程。有些人喜欢拿到智能体在做什么的摘要。为什么它们不能轻松地共享状态?因为一个只读智能体需要读取部分上下文,弄清楚什么相关,比如实际正在写的是什么,因为探索不是特别重要。或者这是子任务树。我觉得如果有真正聪明的人投入大量时间去重新思考编程智能体的体验,能做出太多不同的有趣东西,那会超级超级棒。

And you know, some things, for example, are read-only processes. Some people like getting summaries of what the agents are doing. Why can't they share state easily? Because a read-only agent needs to read parts of the context and figure out what's relevant to say, like what's actually being written, because exploration is not super important. Or here's the tree of subtasks. I feel like there's so many different fun things that could be done if a really smart person dedicated a whole lot of time to rethink the coding agent experience, and that would be super duper sick.

Host

对。

Yeah.

Diogo

天哪,那会是我的梦想。

Man, that would be my dream.

Host

如果你还没看过 Prime Agent,我会推荐你去看看。它和 RLM 的工作是配套的。我们刚跟你前面那位嘉宾 Alex 聊过,他是 Allen 的朋友。而且是的,这东西有人在做了,但还没特别流行。

I would point you towards Prime Agent if you haven't looked at it. It works together with the RLM work. We just talked to Alex, who is a buddy of Allen's, in the chair before you. And yeah, it is being worked on, but it's not super popular yet.

Diogo

嗯,对,但我的希望是——我想让每个人都去玩玩那些奇怪的东西。我不能保证它会成功,但从技术角度看,它看起来真的非常非常有意思。

Well, yeah, but the hope — I would want everyone to just play around with weird things. I have no guarantees that it'll work, but it seems really, really interesting from a technical perspective.

Host

所以是的,那看起来很酷。等我们搞清楚怎么发放积分之后,我很想给这样的人发积分。你肯定会处在能资助研究的位置上。总之,恭喜你取得的所有成功。从我第一次见到你到现在,你走了好长一段路,整个团队也是。

So yeah, that seems cool. Like once we figure out how to give credits out, I would love to give credits out to people like this. You will be in a position to fund research for sure. Anyway, congrats on all your success. You've come such a long way since I first met you, and the whole team.

Diogo

你也是同一个人。

The same person as well.

Host

对。对。我觉得你身上有一种我从未见过的活力,因为你找到了自己的使命,你知道的。

Yeah. Yeah. I think you are energized in a way that I have never seen you before, because you found your mission, you know.

Diogo

没错。那绝对是真的。

That's true. That's definitely true.

Host

而且你把你的使命讲清楚了,因为很多年来你都在抱怨这些问题,但那时你还没有解决方案,对吧?你有了大致的轮廓,然后你得投入工作。

And you articulating your mission, because for many years you complained about the problems but you didn't have a solution yet, right? You had the rough shape, and then you had to put in the work.

Diogo

我要说,这部分是因为我把自己描述为零分创业者。我不喜欢创业公司。我这辈子从没想过当 CEO。我无法想象有人会做两次。那看起来很可怕。说实话,做一次就已经够糟了。我们最初融资的时候,有个投资人问我:“你敬仰哪些 CEO?”我当时就想:“呃,我为什么要敬仰那些人?”无意冒犯任何人,你知道的?我想做一个真诚、善良的人,我遇到过很多非常好的人,但那些有名的,似乎柜子里都藏着不少骷髅。

I will say that is partially because I describe myself as zeroth-percent entrepreneurial. I don't like startups. I never wanted to be a CEO in my life. I can't imagine anyone doing this twice. It seems horrible. Honestly, doing it once is pretty bad. When we first were fundraising, an investor asked me, "Which CEOs do you look up to?" And I was like, "Ew, why would I look up to those people?" No offense to anyone, you know? I'm trying to be genuine and good, and I've met a lot of really good people, but the famous ones have a lot of skeletons in their closet, it seems.

在OpenAI感到无力 Feeling Disempowered at OpenAI

Diogo

我觉得我在 OpenAI 的时候真的感到无力。我觉得现在说实话更容易一些,因为至少我有一些证据表明这个方向是可行的。我只是觉得在那个疯狂的房子里,每个人都只想着 ChatGPT。我们把 ChatGPT 放在哪里?我们怎么让 ChatGPT 对开发者友好?我就想,你在说什么?函数调用接口太疯狂了。你为什么要部署这个?这太反开发者了。

And I think I just really did feel disempowered when I was at OpenAI. I felt like it's a little bit easier to be truthful now because I have at least some proof that the direction has legs. I just felt like in the insane house where everyone is just like ChatGPT. Where do we put ChatGPT and everything? How do we make ChatGPT good for developers and stuff? And I'm like, what are you talking about? The function calling interface is insane. Why would you deploy this? This is so anti-developer.

Host

有点像在 hack 之上堆 hack,再堆 hack。

Sort of a hacky way on top of hacks on top of hacks.

Diogo

嗯,不只是那样。我经常说的是——这个现在可能也讲不完——但我总是说,如果每个函数没有合理的偏置,我想被移出任何涉及函数调用的项目。所以对我来说,这是一个非常非常简单的请求。

Well, not just that. The thing I often said was—this is also probably too long for right now—but I always used to say I want to be removed from any project involving function calling if you did not get a legit bias for each function. So a very very simple ask on my part.

Host

就像是一个置信度,但不是校准过的,或者是一个概率,对吧?

Which is something like a confidence but not calibrated, or a probability for it, right?

Diogo

比如我们需要给用户控制的能力,比如说动作是拒绝还是允许。Disney 需要设置与 AI Dungeon 不同的拒绝阈值。现在用函数调用控制这个的唯一方式就是恳求。这太疯狂了。对开发者来说,这是一个疯狂的接口。人们已经忍受这个好几年了。他们现在还在用技能这么做。现有的编码智能体高度过拟合于它们现有的框架,因为它们是参差不齐的。它们不太擅长使用外部工具和 MCP,当然是因为过拟合。为什么大公司不能允许轻微的推动来更多地调用这个?这真的很有用。而解决方案就是在系统消息里恳求。这太疯狂了。

Like we need to give users the ability to control, let's say the actions are refuse or allow. Disney needs to set a different refusal threshold than AI Dungeon. The only way to control that with function calling right now is to say pretty please. That's nuts. That's a nuts interface for developers. And people have been dealing with this for years now. They still have that with skills. The existing coding agents are highly overfit to their existing harness because they're jagged. They don't tend to use external tools and MCPs super well because of overfitting, of course. And why can't big companies allow for slight nudges to call this more? It's really useful. And the solution is begging in a system message. That's nuts.

Host

但是好吧,我想我明白你的意思了。天啊,聊这些真的太令人兴奋了。能请你来播客真的很酷。

But okay, I think I get you. And man, it is so exciting to talk about all this stuff. It's really cool to get you on the podcast.

Diogo

是的,你会去做了不起的事情,伙计。我很期待你的下一次大发布。

Yeah, you're going to go do amazing things, man. I'm excited for your next big launches.

Host

哦,当然。你就等着吧。

Oh, hell yeah. Just you wait.

Diogo

你就等着吧,比你想的还要快。

Just you wait, sooner than you think.

招聘与公司文化 Hiring and Company Culture

Host

基础设施的人,我猜。市场人员?

Infra people, I assume. Marketer?

Diogo

看情况。如果你问我,我觉得我是一个相当不错的创始市场人员。但如果你问我团队里的任何人,他们会说:‘闭嘴,Diogo。你需要做 CEO 的事情。’所以,是的,创始市场人员。

Depends. If you ask me, I feel like I'm a pretty good founding marketer. But if you ask anyone on my team, they say, 'Shut up, Diogo. You need to do CEO stuff.' So yes, founding marketer.

Host

这不仅仅是关于香料。我觉得你非常注重香料,这是你独特的才能,但有时候你只需要说——

It's not just about spice. It's like I think you're very spice oriented, which is your unique talent, but sometimes you just need to say—

Diogo

我知道,我知道营销的事情。是的。没有什么比有一大堆事情要做更能教会你委派了。招聘数据人员,或者我们称之为模型能力,但他们是数据人员。数据在行业里有点像贬义词,我想确保——

I know, I know marketing things. Yes. Nothing teaches you delegation like having a tidal wave of stuff to do. Hiring data people, or we call them model capabilities, but they are data people. Data is kind of a slur in the industry, and I want to make sure—

Host

所以我们这里非常支持数据。

So we're very pro data here.

Diogo

是的。但我希望他们成为实际从事模型工作的人中地位最高的。这听起来有点奇怪。我希望每个人都有平等的地位,但我想平衡这一点,我想让大家知道这是非常有价值的。

Yeah. But I want them to be the highest status of the people actually working on the model. That actually sounds a little weird. I want everyone to have equal status, but I want to even that out and I want to know that that's really valuable.

Host

至少我比别人更平等。

At least I'm more equal than others.

Diogo

嗯,我不喜欢奇怪的等级制度。我认为我在公司最自豪的事情之一就是他们不太尊重我,或者他们不表现出来。他们只是调侃我,和我开玩笑,有时对我不好等等。我认为这是一个文化的好迹象。我们正在招聘平台人员,比如到处构建 Jev 的人。我们对位置更加敏感,因为光速更是一个瓶颈。我为欧洲用户感到难过,我们只有三倍的速度,而不是一百倍,因为我们现在在那里没有服务器。这太疯狂了,对吧?

Well, I don't like weird hierarchies. And I think one of the things I'm most proud about in the company is that they don't respect me that much or they don't show that. They just troll me and joke with me and they treat me poorly sometimes and all of that. And I think that that's a good sign of a culture. We're hiring platform people, like people to build out Jev everywhere. We are so much more sensitive to location because speed of light is more of a bottleneck. I'm so sad for the European users that we're only like three times as fast instead of like a hundred times as fast because we don't have servers there right now. And that's insane, right?

Host

没关系。欧洲的生活也慢一点。没关系。

It's okay. Life in Europe goes a bit slower as well. It's okay.

Diogo

哇。我不敢相信是你说的,不是我。或者到处,你知道,如果每秒智能是一个重要的指标,我们会希望它无处不在。我们关心如果他们是建立在我们之上的开发者,我非常关心你。我们正在招聘人员来继续构建更多,不仅仅像目标不是仅仅成为一家像 Jev 这样的公司。目标是发布更多形式的智能。所以我们也在招聘人员来构建这些东西。我们不想只是简单模型的单一技能,但我认为会有一个智能的 AWS,你知道,

Wow. I can't believe you said it, not me. Or everywhere, you know, like if intelligence per second is a metric that matters, we'll want this all over the place. We care about if they're a developer building on top of us, I care a lot about you. And we are hiring for people to keep building more, not just like the goal is not to just be like Jev as a company. The goal is to ship more shapes of intelligence beyond that. So we are hiring people to build those things too. We want to not just be the one-trick pony of the simple model, but I think that there's going to be an AWS of intelligence, you know,

Host

顺便说一下,那将会是你,对吧?是的。

Which is going to be you by the way, right? Yes.

Diogo

我的意思是,这是我想走的方向。说它会是我,那太傲慢了。我会尽我所能确保那发生。我认为那会非常非常酷,你知道?我们现在正在玩系统一智能。想象一下那些层,你知道,这就像它的 TCP。

I mean, that's a direction I want to go down. It would be arrogant to say it will be me. I'm going to do anything I can to make sure that happens. I think that that's going to be so so cool, you know? We are playing with system one intelligence right now. Imagine the layers, you know, like this is like the TCP of it.

Host

是的。还有七层,谁知道还有什么。顺便说一下,我也提出了时间层。我不知道。我们需要把时间层作为七层中的第八层来讨论。但不管怎样,

Yeah. Seven more layers to go and who knows what else. I've also pitched temporal by the way. I don't know. We need to talk about temporal as layer eight out of the seven layers. But anyway,

Diogo

我们可以永远聊下去。你得回去工作或睡觉了。谢谢你来。

We can talk forever. You got to get back to work or sleep. Thank you for coming.

Host

哦天。是的。酷。非常欢迎。很荣幸,伙计。

Oh boy. Yeah. Cool. You're most welcome. It was a pleasure, man.

Diogo

太兴奋了。太兴奋了。我的第一次。

So excited. So excited. My first time.

互动版:逐字朗读 + 针对本期提问 →