Why Jev Is Changing How We Build With AI
打开互动全文版(中英对照 + 朗读 + 问答)→Typesafe 联合创始人兼 CEO Diogo Almeida 解释,Jev 的真正突破在于为机器原生决策而非人类可读字符串优化 AI。
Diogo Almeida, co-founder and CEO of Typesafe, explains why Jev's real breakthrough is optimizing AI for machine-native decisions rather than human-readable strings.
三周前,Typesafe 带着 Jev 走出隐身模式,它迅速成为 AI 领域最受关注的新模型之一。发布时,开发者分成了两派。一些人耸耸肩说它只是个分类器。另一些人则看到了它快速、便宜且易于构建的特点。几天之内,开源模仿者纷纷出现,甚至 OpenAI 也宣布了自己的回应。但据其创造者所说,大部分讨论都错过了重点。Jev 的真正进步不在于速度或成本。而在于对 AI 应该优化什么的不同理念。Diego Almeida 是 Typesafe 的联合创始人兼 CEO。在此之前,他在 OpenAI 工作了四年半,是 InstructGPT 团队的一员,并参与了基于人类反馈的强化学习(RLHF)工作,这些工作帮助语言模型变成了我们今天使用的助手。现在,通过 Jev,他正在追求他所谓的机器原生智能——不是为生成供人消费的字符串而优化的模型,而是为做出软件可以执行的可靠决策而优化的模型。以下是 Diego 讲述他试图弥合的差距。
Three weeks ago, Typesafe came out of stealth with Jev, and it quickly became one of the most talked about new models in AI. When it was released, developers split into two camps. Some shrugged and said it was just a classifier. Others saw something fast and cheap and easy to build with. Within days, open-source lookalikes were showing up, and even OpenAI has since announced its own response. But according to the person who built it, most of this discussion misses the point. Jev's real advance isn't speed or cost. It's a different idea of what AI should be optimized for. Diego Almeida is co-founder and CEO of Typesafe. Before that, he spent four and a half years at OpenAI where he was part of the team behind InstructGPT and the RLHF work that helped turn language models into the assistants we use today. With Jev, he's now pursuing what he calls machine native intelligence, models optimized not for generating strings that people consume, but for making reliable decisions that software can act on. Here's Diego on the gap he's trying to close.
我对 Jev 的电梯演讲就是:自动化都去哪儿了?我认为 AI 聪明得令人难以置信,但在你真正期望它有用的那些事情上,却无用得令人难以置信。据我所知,大多数人并没有真正回答为什么会有这么一整类事情,AI 不仅做得好,而且是超人类地好——比如 ChatGPT、深度研究、Claude Code——但在另一类事情上,它不仅差,而且差到完全无法使用,尽管有巨大的财务激励。
The elevator pitch I use for Jev is just: where is all the automation? I think AI is just so unbelievably smart, yet so unbelievably useless at the kinds of things you'd really expect it to be useful for. And most people, as far as I can tell, don't really have an answer on why there's this entire bucket of stuff where AI is not just good, it's like superhumanly good — like ChatGPT, deep research, Claude Code — but on the other bucket of stuff, it's not just bad, it's like so bad that it can't even be used at all, despite having enormous financial incentive.
我是 Sam Sharington,这里是 Twimmel AI 播客。十多年来,我一直在通过像这样的对话探索塑造 AI 未来的想法和创新,帮助你理解什么是真实的、什么是下一个趋势、什么才是重要的。让我们开始吧。
I'm Sam Sharington and this is the Twimmel AI podcast. For over a decade, I've been exploring the ideas and innovation shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters. Let's jump in.
为什么人们不能用智能体做基本的数据录入和保险核保?实际上,如果你去——哦,我不想针对谁,但比如你去 Anthropic 的网站,试着用他们的智能体,他们的客服智能体,做任何事情,它很擅长把常见问题扔给你,但如果你想做更简单的事情,比如更换信用卡,它就没用了。我认为这是每个与 AI 相关的人至少都应该回答为什么会出现这种情况的问题。我很好奇你是否有一个答案,解释为什么在超级聪明的东西和超级无用的东西之间有这么大的鸿沟。
Why can't people use agents for basic data entry and insurance underwriting? And actually, if you go to — oh, I don't want to be taking stabs at people, but like if you go even to Anthropic's website and you try to use their agent, their customer support agent, to do anything, it's really great at throwing the FAQ at you, but it is not useful if you want to do something even easier like changing a credit card. And that I think is a thing that everyone who is AI adjacent at least should have an answer to why that's the case. Kind of curious if you have an answer why there's such a big divide between super smart stuff and super useless stuff.
我的意思是,我认为部分原因在于我们在现代 AI 中只有这一个工具,对吧?它的核心是一种预测下一个词的 Transformer,周围有一堆机制,让它越来越擅长生成内容、生成词语、生成文本。我们试图把它塞进每一个能塞的地方,但这并不总是每个问题的最佳方案。我猜这可能和你的观点相差不远。
I mean, I think part of it comes down to we've got this one tool in modern AI, right? It's this kind of token predicting transformer at the core with a bunch of machinery around it that makes it better and better at producing content, producing words, producing text. And we're trying to kind of jam that into every place we can, but that's not always the best thing for every problem. And I suspect that that's probably not too far from your viewpoint on this.
有一点吧。但这里面有很多细微差别,对吧?比如,这个吐词的 Transformer 怎么能解决数学中的千禧年大奖难题,却不能自动化会计?这里面有很多细微差别,它真的提出了一些我认为该领域的人还没有准备好回答的问题,而且他们非常有动机——至少 AI 实验室非常有动机——不去如实回答。
A little bit, though. There's so much nuance into it, right? Like for example, how can this token spewing transformer like solve millennium prize problems in math but cannot automate accounting? Like there's a lot of nuance into that where it really begs some questions that I don't think people in the field are really ready to answer and are very incentivized — at least the AI labs are very incentivized — to not answer that truthfully.
这是暗示你准备好回答那个问题了吗?
Is the implication there that you are ready to answer that question?
是的,如果你愿意,我可以。我的意思是,首先,一切都只是观点,对吧?所以我不会说我有证据。我会说我有很强的信念和很多证据。但证明这一点的方法就是真正自动化那份工作。我希望 Jev 是朝着这个方向的第一步。所以我对答案及其细微差别的看法是:你得到你所优化的东西,这对所有机器学习事物都成立,对世界上所有事物也差不多成立。
Yeah, I could if you want. I mean, first of all, everything is just an opinion, right? So I wouldn't say I have proof of that. I would say I have a lot of conviction and I have a lot of evidence. But the way to prove this would be to actually just automate that job. And I'm hoping that Jev is the first step in that direction. So my take on the answer and the nuance of it is that you get what you optimize for, which is just true of all things ML, kind of true of all things in the world.
所有事物。
All things.
所有事物。是的。实际上我认为机器学习的人最不理解这一点,而经常与人互动的人最理解。所以你得到你所优化的东西,基本上所有字符串 LLM 都是为字符串优化的,而字符串是供人类或其他 LLM 消费的。但如果你想要像会计这样的东西,它是供计算机消费的,对吧?你有特定的字段、表格和离散决策,优化字符串和优化这些离散决策非常不同,我相信这是这个差距的最大解释。
All things. Yeah. Like actually I think ML people get it the least and people who interact with humans a lot get it the most. So you get what you optimize for and basically all the string LLMs have been optimized for strings and strings are meant to be consumed by humans or other LLMs. But if you want something like accounting that's meant to be consumed by a computer, right? Like you have specific fields and forms and discrete decisions and it's very different to optimize for strings versus these discrete decisions and that I believe is the largest explainer of this gap.
当你把纳维-斯托克斯方程和会计放在一起时,那是一个刻意构造的例子,直接契合你围绕 Jev 建立的定位以及它真正擅长的东西。但我不认为 AI 擅长某些事情、不擅长另一些事情是这种一维的文本与决策的对立。它是一条锯齿状的边缘,需要很多不同的东西或某种根本不同的东西来平滑那条锯齿边缘,并推动前沿向前。我想知道你对那会是什么样子有什么想法。
When you juxtapose Navier-Stokes and accounting, that is a contrived example that plays directly into the positioning that you've built up around Jev and what that's really good at. But I don't think that AI being good at some things and bad at other things is this one-dimensional text versus decisions. Like it's a jagged edge and there will need to be a lot of different things or something fundamentally different to smooth out that jagged edge and push that frontier forward. And I'm wondering if you have any thoughts on what that looks like.
我能快速测试一下关于锯齿边缘的假设吗?
Can I actually test the hypothesis about the jagged edge for real quick?
当然可以。
Absolutely.
因为我觉得那是个有趣的话题。好的。顺便说一句,我们做的很多工作实际上是专注于让模型成为一个健壮的认知核心,可以插入任何软件,并且我们做的很多工作实际上是专注于如何平滑通过各种奇怪过程引入的锯齿状。但为了测试锯齿边缘存在的假设,你有没有见过 ChatGPT 无缘无故对你粗鲁或侮辱你?
Because I feel like that's a fun one to talk about. Okay. As an aside, a lot of what we do is really focus on making the model this robust cognitive core that can be plugged into any software and a lot of what we do is really focus on how can we smooth out the jaggedness that is put in via all sorts of weird processes. But to test the hypothesis about the jagged edge existing, have you ever seen ChatGPT respond to you rudely or insulting you for no reason?
如果我没让它那样做,就没有。
Not if I didn't tell it to do that.
没错。所以模型在某些事情上实际上是可靠的,对吧?它们不只是——它们在某些事情上超级可靠。我想说,锯齿状并不是 LLM 或 AI 固有的。
Exactly. So the models are actually reliable at some stuff, right? They're not just — they're like super duper reliable at some stuff. I would say that jaggedness is not intrinsic to LLM or AI.
你是说锯齿状是相对于我们将它们应用于特定问题而言的。
You're saying jaggedness is relative to our application of them towards specific problems.
是的。而且相对于我们的优化过程,对吧?就像如果你在某个特定方向上超级努力地优化,你会获得某些类型的可靠性,但会失去其他类型的可靠性,因为你把所有的优化压力都集中在那些真正重要的事情上。比如,如果你为决策优化,在客服聊天中,在不该退款时给某人退款,从准确性的角度来看,和用脸砸键盘完全一样。
Yes. And to our optimization process, right? Like if you optimize super hard in a particular direction, you get certain types of reliability that you lose others because you focused all of your optimization pressure on those things that really matter. Like if you were optimizing for decisions, for example, in a customer service chat saying to give someone a refund when you shouldn't is from an accuracy perspective exactly the same as mashing your face on the keyboard.
你知道,同样是 0% 的准确率,但你永远不会看到它们把你的脸按在键盘上乱敲,因为优化中存在不对称性。这就是 RLHF 所做的,对吧?你使用一个奖励模型,越容易惩罚某件事,就越容易确保它在推理时不会发生。你就再也不会遇到那种情况了,尽管预训练模型实际上可能会输出一堆乱码。所以后训练可以让某些类型的事情变得可靠,而我认为人们没有意识到 AI 从业者在将他们的期望清单设置到模型中时拥有多大的自主权。
You know, it's the same 0% accuracy, but you never get them mashing your face on the keyboard because of the asymmetry in the optimization. This is what RLHF does, right? You use a reward model, and the easier it is to punish something, the easier it will be to make sure that doesn't happen at inference time. And you just never get those types of things anymore, even though the pre-trained models could actually get word salad right. So post-training can make certain types of things reliable, and I don't think people realize how much of an agency that AI practitioners have in setting their menu of desiderata into their model.
听你这么说非常有意思,你几分钟前也说过,每个人都知道你会得到你所优化的东西,但 ML 人员——ML 本质上就是关于优化的,我们在语言模型上的所有收益都是 RLHF,我们一直在谈论成本函数。你认为为什么 ML 人员没有理解这一点?
It's super interesting to hear you say that, and you said that a few minutes ago as well, that everybody gets that you get what you optimize for, but ML people—ML is all about optimization inherently, and all of our gains with language models are RLHF, and we talk about cost functions all the time. How is it that you think that ML people don't get that?
我喜欢这个问题。你看到过我关于“最苦涩的教训”的博客文章吗?
I love this. Have you seen my blog post on the bitterest lesson?
没有,告诉我。我可以——我可以直接解释一下。
No, tell me. I can—I can just explain it.
有 Sutton 的“苦涩的教训”。我为了对自己有用而转述一下。所以这不完全准确,但大致是——我在简化。简化真正好的文字是很难的,对吧?所以我为了有利于我的叙述而简化了一点。
So there's Sutton's bitter lesson. I'm paraphrasing to be useful for me. So this is not exact, but roughly it's—I'm simplifying. It's hard to simplify actually good writing, right? So I'm simplifying a bit to be beneficial to my narrative.
其他所有人都在简化“苦涩的教训”,所以你也不妨这么做。
Everybody else simplifies the bitter lesson, so you might as well too.
是的,我也不妨这么做。但我只想强调,这里面其实有细微差别。但它是这样的。大致来说,算力和 Scaling(规模扩张)比算法更重要。这之所以苦涩,是因为 ML 人员热爱算法。它总是最性感的工作。它有趣。它酷。这是他们所学的东西。而且它苦涩,因为,你知道,谁想做那些系统性的东西,而且搞到 GPU 很难。尤其是在学术界的时候,你没有那么多算力。这催生了一个整个时代的 Scaling(规模扩张)信徒 ML 从业者,原因复杂。你知道,我实际上认为成为 Scaling(规模扩张)信徒有利有弊,但我相信数据比算力重要得多。所以你实际上需要数据来投入算力,而比数据更重要的是你需要正确的任务。所以这是我们做的所有 ML 中最重要的事情,而正确的任务就像是北极星,你会得到你所优化的东西。你知道,像预训练损失——预训练损失事后看来可能很明显。但预训练损失会导致如此棒的 AI 成果,这是一个相当大的信念飞跃,而且绝非保证。所以我认为实际上只有少数人能真正创造一个新任务,这就是“你会得到你所优化的东西”的一部分。如果我可以对 Sutton 的观点稍作细微补充,在游戏的情况下,正确的任务毫无疑问,因为它受规则约束,而且游戏没有数据,因为你在做同策略强化学习,所以你通过像做 rollout 这样的方式从算力中获得数据。所以我猜——我会说,在 RL 游戏的情况下,Sutton 的苦涩教训就相当于我所描述的东西。
Yeah, I might as well too. But I just want to emphasize that there's actually nuance to it. But here it is. So it's roughly speaking that compute and scaling matters more than algorithms. And this is bitter because ML people love algorithms. It's always the sexy thing to work on. It's fun. It's cool. This is what they've studied. And it's bitter because, you know, like who wants to do that systemsy stuff and like getting GPUs is hard. And especially back in the academia days, you didn't have that much compute. And this birthed like an entire era of scaling-pilled ML practitioners for mixed reasons. You know, like I actually think that there's pros and cons to being scaling-pilled, but I believe that data matters a lot more than compute. So you actually need data to put the compute on, and more important than data is you need the right task. So this is the most important thing in all of the ML that we do, and the right task is kind of the north star that you get what you optimize for. You know, something like pre-training loss—pre-training loss is like maybe obvious in hindsight. It was quite a—it was quite a big leap of faith that pre-training loss would lead to such awesome AI stuff, and it was by no means guaranteed. So it's actually a rare few that I think can actually make a new task, and that is the you get what you optimize for part of it. And if I could add like slight nuance to Sutton's take, in the case of games there's no question what the right task is because that's constrained by the rules, and there's no data for games because you're doing on-policy RL such that you get data from compute by like doing rollouts. So I would guess—I would say that in the case of RL games, Sutton's bitter lesson is the equivalent one of what I'm describing.
但推而广之,在现实世界的应用中,数据是比算力更受约束的资源。
But by extension, in real-world application, data is the constrained resource more so than compute.
也要弄清楚数据的方向。所以我现在不一定会谈论 Jev 和 RLCD,因为你知道,那不会有趣。但让我们谈谈 RLHF,对吧?当我们做出——RLHF 有很多定义。有 Paul Christiano 做的原始 RLHF,用来教机器人后空翻。不涉及 LLM。还有另一个 RLHF 学习摘要,OpenAI 团队的一些人——我认为 Paul 也对此至关重要——教模型摘要,像是一个 GPT-2 化的模型来比人类更好地摘要文本。我不确定那部分。然后还有 RLHF,就是现在每个人所指的,即创建指令遵循的任务,比如教模型聊天和遵循指令。所以当我提到 RLHF 时,这就是正在做的任务。在 RLHF 存在之前,互联网上根本没有我们寻找的那种形状的人类反馈数据,因为我们必须创建它,对吧?就像为什么有人会有来自语言模型的、有点同策略的两个完成之间的偏好?因为那是如此奇怪的东西。所以实际上需要团队的一些天才才能意识到这是一个值得走下去的方向。一旦你有了那个方向,首先,当你处于方向的早期时,根据定义你不确定它是否正确,对吧?就像它还没有起作用。你需要大量的数据、计算、算法来让它工作。而你只是有一个强烈的信念,认为这是要走的路。但最终你得到了数据和其他所有东西,但数据需要从头开始获取。而这最终成为世界上某种非凡的东西。这往往就是你在 ML 中获得巨大飞跃的方式。你知道,ChatGPT 是一个巨大的飞跃。O1 是一个巨大的飞跃。而 Jev,至少,你知道,我不想太自吹自擂,但至少从兴奋的角度来看,也是一个巨大的飞跃。
Figure out the direction of the data as well. So I won't talk out necessarily about Jev and RLCD right now because, you know, it won't be fun. But let's talk about RLHF, right? When we made—there's many definitions of RLHF. There was the original RLHF that Paul Christiano did to teach robots to backflip. No LLMs involved. There was another RLHF learning to summarize where some people in the OpenAI team—I think Paul also was pivotal for that—taught the models to summarize, like a GPT-2-ized model to summarize text better than humans did. I'm not sure about that part. And then there's RLHF which is what everyone refers to now, which is creating the task of instruction following, like teaching the models to chat and follow instructions. So when I refer to RLHF, there's the—this is the task that is being done. And before RLHF existed, there was literally no human feedback data on the internet of the shape that we were looking for because we had to create it, right? Like why would anyone have like preferences between two completions from language models that were somewhat on-policy? Because that's such a weird thing to have. So it actually took some genius from the team to realize that this is a direction to go down. And once you have that direction, first off, when you're in the early days of the direction, you are not sure it's the right direction by definition, right? Like it's not yet working. You need like a lot of data, computing, algorithms to make it work. And you just have a strong conviction that it's the way to go. But eventually you get the data and all the other stuff, but like the data needs to be gotten from scratch. And that eventually becomes something phenomenal in the world. And that tends to be how you get like the gigantic jumps in ML. You know, ChatGPT was a gigantic jump. O1 was a gigantic jump. And Jev, at least, you know, I don't want to toot my own horn too much, but at least from like an excitement perspective, was also a gigantic jump.
我们也许退一步,因为你提到了 RLHF。我们上次交谈是在 10 年前——差一个月就 10 年了。那是 2016 年 10 月。我们当时在广泛地谈论深度学习,远在任何这些生成式 AI 的东西之前。给我们简要总结一下你从那以后都在做什么。你在 OpenAI 待过一段时间。
Let's maybe take a step back because you mentioned RLHF. We spoke the last time 10 years ago—just a month under 10 years ago. It was October 2016. And we were talking about kind of deep learning broadly way before any of this Gen AI stuff. Give us a brief summary of what you've been up to since then. You were at OpenAI for some period of time.
我在谷歌,然后我离开了。我退休了一段时间,只是想放松一下,开心一点。做了很多奇怪的事情,比如非常竞技性的电子游戏和即兴表演之类的。最终我意识到,实际上 AI 比那有趣多了。让我加入这个有着奇怪目标的奇怪公司。那就是 OpenAI。这是在语言模型真正成为现实之前。所以我有幸实际上参与了很多 OpenAI 的热门项目。我在那里待了大约四年半。做了很多工作。但真的,即使我在做其他事情时,我一直在痴迷于这个问题,它导致了 type safe 和 Jev,比如为什么 AI 如此难以置信地好,却对简单的工作如此无用?这一直是——你知道,如果我完全坦率地说,这仍然是一种痴迷。直到我们看到大规模的、大规模的变革性自动化和软件做绝对的科幻事情,它才会完成。这里面有太多东西我想深入探讨。比如软件在做科幻事情。就像显然你在和你的手机说话,而你在——推特上的好例子。
I was at Google and then I left. I was retired for a while just trying to like chill out and be happy. Did a lot of weird stuff like very competitive video gaming and improv and stuff. And eventually I realized that actually AI is way more fun than that. Let me join this weirdass company that has weirdass goals. And that was OpenAI. This is before language models were really a thing. So I got the honor to actually work on a lot of OpenAI's greatest hits. I was there for about 4 and a half years. And worked on a whole bunch of stuff. But really, even when I was working on other things, like I've been obsessed with this question that has led to type safe and Jev, like why is the AI so unbelievably good yet so useless for the easy work? And that's been—you know, it's still an obsession if I'm like totally frank. And it will not be done until we see massive massive transformative automation and software doing absolute sci-fi stuff. There's so many things in there that I want to pick at. Like software is doing sci-fi stuff. Like clearly you're talking to your phone while you're—great example on Twitter.
有人在特斯拉里跟 Grok 说话,让一个智能体去生成视频。这挺科幻的,对吧?软件正在做科幻的事。另外还有一点可以挑出来说,就是那些所谓容易的事。容易的事未必真的容易,这正是我们至今还没做到的原因。
Guy is talking to Grok in his Tesla, asking an agent to build a video. Like, that's pretty sci-fi, right? Software is doing sci-fi stuff. And also the other thing to pick at in there is like the easy stuff. Like the easy stuff isn't necessarily easy, which is why we haven't done it yet.
天哪,有——
Oh man, there's—
你想先聊哪个都行。
You can take whichever of those you want to take first.
我不喜欢回避问题,所以两个我都想聊。不过其实我先聊第一个吧。我容易跑题,所以想集中讲这个。我觉得 AI 做的很多事情都特别牛。我用编程智能体,我现在都不怎么自己写代码了,但我每天都用编程智能体。我每天都用聊天机器人。天哪,没有聊天机器人,经营一家公司会难得多。但它们——
And I don't believe in running from questions, so I will try to get them both. But actually, I'll just do the first one. I ramble, so I want to focus up on that one. I think a lot of what AI does is super freaking sick. I use coding agents. I don't even code that much anymore, but I use coding agents every day. I use chatbots every day. Holy smokes, would it be a lot harder to run a company without chatbots. But they are—
顺便说一下,特斯拉那个例子里有个点我没点破,就是当时车正在自己开。
And in case it wasn't obvious with the Tesla example, I didn't say the punch line, which was the car was driving itself at the time.
哦,对。
Oh, right.
对,自动驾驶其实——我觉得有点不一样,我们可以聊聊自动驾驶之间的区别——
Yeah, self-driving actually—I think it's a little bit different and we can talk about the distinction between self—
重点只是科幻感——
The point was only sci-fi like—
哦,好。
Oh, cool.
我对科幻感的看法主要是针对 LLM 的,我觉得现在人们说 AI 时通常就是指这个,因为它似乎是对智能的最大压缩,但实用性却最低。而且它也挺——我不会说它被普遍讨厌,但确实有很多人不喜欢 AI,这很可悲,我认为作为一个领域,如果我们不承认为什么会这样,那就是失职。所以 AI 也有很多阴暗面,我可以列举一些,但我觉得这值得一谈。
My argument with sci-fi was especially related to LLMs, which I think tends to be what people mean by AI these days, because it seems to be like the greatest compression of intelligence while having like the least utility. And it also—it's quite—I wouldn't say it's like universally hated, but there's a lot of dislike for AI, which is tragic, and I think that we would be remiss as a field to not own up to why that is the case. So there's like a lot of dark sides to AI too, and I could name a subset of them, but I think that's worth talking about.
第一,我觉得自动驾驶特别特别酷。而且我把自动驾驶看作一个巨大的工程胜利,未必是 AI 的胜利。我拿 Waymo 当黄金标准。无意冒犯特斯拉粉丝,但他们似乎更安全。而且他们的做法是构建大量软件,像工程一样打造可靠的系统,他们理解其中的各个部分,这些部分有自己分解和抽象出来的模块,确保它们好得难以置信,同时把想要的行为策略编程进去,这样就能泛化到分布之外。所以这特别牛。但所有这些在现在的 AI 里都没有发生。
So number one, I think self-driving is super duper cool. And also I see self-driving as a ginormous engineering win, not necessarily an AI win. I use Waymo as like the gold standard here. No offense to Tesla fans, but they seem to be safer. And also the way they do it is by building a lot of software and engineering like reliable systems that they understand the pieces of, that have their own like decomposed abstracted parts, and make sure those are like unbelievably good while programming in the behavior of the policies you want such that this can generalize outside of distribution. So this is super freaking cool. But none of this is happening in AI right now.
而在 AI 里,你得到的往往是更多——先做个 demo,然后——我这里的愤世嫉俗循环是:做个 demo,融个种子轮 A 轮,说你会把它做可靠,结果从来没做成可靠,然后转向人在环的版本,而不是真正自动化任务。我不是——我不想骂他们,说清楚,鉴于当前 AI 的气候,那其实是正确的做法,因为那些模型不可靠。
When in AI what you tend to get is a lot more—let's just make a demo and—you know, my cynical loop here is: make a demo, raise like a seed Series A, say that you're going to make it reliable, never end up making that reliable, pivot into a human-in-the-loop version of this thing instead of actually automating the task. And it's not—I don't want to hate on them, to be clear, that is actually the right move to do given the current AI climate, because those models are not reliable.
所以即使——我尽量做到尽可能无偏见。我显然不可能完美,但我希望对 LLM demo 和 Jev demo 有同样类型的批判视角。而且我,你知道,我觉得很多 demo 都特别酷,尤其是当它们有很棒的想法时。但我也不能真正为它们背书,因为只有当它们能可靠地在后台运行时,才算解决了工作。而 LLM 离那还差得极远。
So even—I try to be as unbiased as I possibly can. I obviously can't be perfect, but I would like to have like the same type of critical lens to LLM demos as Jev demos. And I, you know, like I think that a lot of them are super cool, especially when they have sick ideas. And also, I can't really vouch for them because they solve work when they actually reliably can be run in the background. And LLMs are extremely not there.
我其实不知道有谁真的把 LLM 当作一个独立的依赖来运行,因为你没法用 LLM 获得抽象,对吧?它可能任意崩溃,你需要让抽象泄漏才能看到,哦,你为什么聊生物学?哦,抱歉,我的 Claude 不能聊生物学。它回退到了另一个模型之类的。对于第一方应用来说这极其合理,但在我看来,对于开发者要构建在其上的东西来说,这超级不合理。
I actually don't know of anyone who like actually runs like an LLM as a separate dependency because you can't get abstraction with an LLM, right? It can just break arbitrarily and you need to have the abstraction leak in order to see like, oh, why did you talk about biology? Oh, I'm sorry, my Claude can't talk about biology. It fell back to another model or something like that. And that is extremely sensible for a first-party app, but super nonsensical in my opinion for something developers are meant to build on top of.
所以我认为那种缺乏实用性仍然存在,尽管 demo 看起来像科幻,因为 demo 总是会显著领先于曲线。OpenAI 从 2020 年就开始展示客服 demo——我相信那大概是六年前了——而客服仍然没被解决。你知道,也许是因为工作很难,就像你说的,但它看起来真的比数学里的千禧年大奖难题容易得多,因为我不认识任何能解决数学千禧年大奖难题的人——我感觉我认识所有人——而我不认识所有能解决客服问题的人,但我认识的每个人可能都能比 LLM 更好地做那份工作。甚至免下车窗口也是。人们一直尝试自动化免下车窗口,一直失败。除了也许免下车窗口比未解决的数学还难之外,没有好的解释。
So I think that lack of usefulness I still think is present there despite demos looking like sci-fi, because demos will always be significantly ahead of the curve. OpenAI has been showing like customer service demos since 2020—I believe that is now roughly six years ago—and customer service is still not solved. And you know, maybe because the work is hard as you say, but it really looks a lot easier than millennium prize problems in math, because I don't know anyone who could solve a millennium prize problem in math—I feel like I know everyone—and I don't know everyone who could solve customer service, but everyone I know probably could do that job better than an LLM can. Or even drive-throughs. People keep trying and failing to automate drive-throughs. And like there's no good answer for that other than maybe drive-throughs are harder than unsolved math.
我觉得有意思的是你拿 Waymo 和特斯拉对比,部分是因为在这个关于“苦涩教训”的对话背景下,对吧?因为从某种意义上说,Waymo 的不同之处在于他们没有采取苦涩教训的路线,而特斯拉更偏向苦涩教训——就像我们要,你知道,我是从三万英尺的高度概括他们的做法,但就像我们要收集大量数据,扔进一个大模型,让它端到端地做,而 Waymo 则是我们要经典地工程化这个东西,把它拆成很多不同的模块,用上所有能用的传感器来获取尽可能多的数据。我在播客上和 Drago 聊过他们怎么做,以及它从根本上多么不像——你知道,我们瞄准的是端到端的单一模型来统治整个驾驶体验。
It's interesting to me that you call out Waymo versus Tesla in part because in the context of this conversation about like bitter lesson, right? Because in one sense, what differentiates Waymo is that they're not taking a bitter lesson approach to it, whereas Tesla's is more bitter lesson—like we're going to, you know, I'm characterizing their approach at a 30,000-foot view, but like we're going to collect a lot of data and throw that into a big model and let it do end to end, whereas Waymo is we're going to like classically engineer this thing, break this thing into a lot of different modules, use all the sensors we can to get as much data as we can. I've had conversations with Drago on the podcast about like how they approach it and how it's so fundamentally not like, you know, we're aiming for this end-to-end single model to rule the entire driving experience.
所以我想强调,那只是其中的机器学习部分。但我不认为那是我们应该构建的东西的极限。我其实极其支持软件和软件工程。我其实认为——我希望,而且我会尽我所能,让 Jev 开启软件工程的另一个黄金时代,有点像早期互联网那种构建疯狂事物的创造能量。而对于机器学习,你需要数据,需要聚焦正确的任务,需要算力,需要算法。但在我看来,机器学习总体上应该是更大系统的一部分。机器学习应该被限定在系统里它真正能做好的那部分。一般来说,获得更高可靠性的方法是聚焦。你知道,如果 Waymo 试图在旧金山端到端地做整件事,它会一样成功吗?老实说,我赌不会。也许五年后这会更好,但为了现在真正交付有用、可靠且安全的东西,我认为你需要工程来研究所有这些,研究每一个机器学习组件的属性,当东西出问题时,能够调试出是哪部分导致的并修复它。
So I want to emphasize that that is just the ML part of it. But I don't think that's the limit of what we should be building. And I'm actually extremely pro software and software engineering. And I actually think that—I hope that and I want to do everything I can that Jev will usher in another golden era of software engineering, kind of like with the early internet creative energy of building crazy things. And for ML you need data, you need to focus on the right task, you need compute, you need algorithms. But ML is meant to be one part of a larger system in my opinion in general. And ML should be scoped to the part of the system that it can really do well. And in general, the way to get higher reliability is to zoom in. You know, like would Waymo be as successful in San Francisco if they tried to do the whole thing end to end? Honestly, I would wager not. Like maybe this is going to be better in 5 years, but in order to actually ship something now that is useful and reliable and safe, I think you need engineering to study all of these things, study the properties of every single ML component, and when things break, be able to debug which part caused it and fix that.
而在“一个模型统治一切”的设定下,这非常难做到。那么 Jev 的想法从何而来?是不是像“嘿,我要创立 TypeSafe,然后朝着我已经知道的东西奔去”?它是逐渐演化的吗?你们是转型进入这个方向的吗?到底是怎么发生的?
And that's very hard to do in a one model rules them all type setting. So where did the idea for Jev come from? Like was this like hey we're I'm going to set up type safe and I'm running towards this thing that I already knew about. Did it evolve? Did you pivot into it? Like how did how did it happen?
哦,我给你讲个有趣的故事。有些投资人——我们一直在融资,有些投资人说:“哇,这是我们听过的最棒的推介。”我告诉他们,那——我尽量不爆粗口——胡说八道。因为去年我给了完全一样的推介,你们却没听懂。你们现在反应的是产品市场契合度。所以,我多年来一直在说同样的话,就是我现在称之为“机器原生智能”的东西,即优化自动化,也就是为机器而非字符串和人类优化,这正是我一直在谈论的。我并不知道它会以 Jev 的形式出现。我们不得不自己取得多项突破,才能走到今天这一步。但北极星始终是:如果 AI 真正为软件实用性而优化,它会是什么样子?而 Jev 是我们对此的初步尝试,也是我们对最大实用性的 80/20 的最佳猜测。
Ooh, I'll give you a fun story about this. Some investors uh we've been fundraising and some investors have said, "Wow, that's the greatest pitch we've ever heard." And I told them that is I'm going to avoid cursing. Bull bull. That's bull. Um because I gave the exact same pitch last year and you guys didn't get it. What you guys are reacting to is the product market fit. So um I've been saying the same thing for years. Um which is what I'm now calling machine native intelligence like optimizing for automation which is optimizing for machines instead of strings and humans is exactly the same thing I've been talking about. I did not know it would be the in the shape of Jev. Like we've had to get our own multiple breakthroughs to even get as far as we've had. But the north star has always been what does AI look like if it's really really optimized for being useful for software? And Jeb is our, you know, initial foray into that and our best guess of like the 8020 of maximum usefulness.
当 Jev 刚出来时,似乎对 Jev 及其所代表的东西有两种截然不同的反应。一派是:“哦,哇,我们创造了分类器。我们创造了逻辑回归。为什么大家都这么兴奋?”另一派则是:“嘿,这东西太棒了。我可以用比 LLM 便宜得多、快得多的方式做我一直想做的事。”我的部分问题是:你怎么看?但更具体地说,我希望你能从技术层面、更细致地谈谈,Jev 与分类器有何不同?它与嵌入模型加分类器有何不同?OpenAI 刚推出了他们的决策模型,那是一个为决策而后训练的小模型,Jev 与它有何不同?是什么让你的方法独一无二?
When Jev first came out, like it seemed like the was a very kind of uh two disperate reactions to to Jev and what what it represented. It you one camp of people was like, "Oh, wow. We've created classifiers. We've created logistic regression. Um you know, why is everyone excited about this?" And there's another camp of people that was, you know, hey, this thing is amazing. like I could do things that I've been trying to do with LLMs for much much much cheaper, much much faster. Um, you know, part of my question is like what do you make of that, but you know, more specifically, I'd love for you to talk about, you know, in kind of a technical level and with more nuance like how is Jev different from a classifier? How is it different from an embedded embedding model plus a classifier? uh you know open AAI just came out with their decisioning model which is like a small model you know post- train for decisions like how is it different from that like what makes it and your approach unique
哦,天哪,我不知道从哪儿开始,因为你留下了太多线索。我就试着即兴发挥一下,如果可以的话。
o um man I don't know where to start there because you've left so many breadcrumbs so I'm going to just try to I'm just going to jam if that's okay
把它们都吃掉,我不全吃,我要吃最美味的那些。
eat them all I don't eat them all I'm going to eat the most tasty ones
好,让我从最美味的那个开始。我已经发过几次这样的牢骚了,不知道能不能不爆粗口,我会尽量控制。分类器,分类器本身没有错。分类器是一种接口。创建分类器不是机器学习创新。在 2010 到 2020 年间,没人在发明新形状的分类器。他们只是在为有用的东西做分类器。分类器在机器学习中并不是为了酷。它们实际上就是实用性的形状。是的,这听起来很傲慢,但这就是你获得……
okay so let me start with the tastiest one of them and I've had this rant a couple of times I don't know if I can do this rant without cursing, so I'll try to control myself. So classifiers, not there's nothing wrong with classifiers. Classifiers are an interface. Like creating classifiers is not an ML innovation. And no one in like 2010 to 20 to 2020 was inventing new shapes of classifiers. They were just making classifiers for useful stuff. Um and classifiers are not like meant to be cool in ML. They are literally the shape of usefulness. Yeah. You know, like that sounds like an arrogant thing to say, but that is how you get
是的。
Yeah.
是的。这就是如何将智能注入软件,让软件变得智能。Meta 靠什么驱动?分类器。也许有些回归器,但可能主要是分类器。Google 靠什么驱动?可能主要是分类器。所以,分类器没有错。实际上,如果有人说我们是零样本通用分类器,那将是我得到过的最大的赞美。对于 Jev,我们还有其他东西,但就 Jev 而言,那会是一个不可思议的描述。就像如果有人能说,我可以给你——随便一个开发者——相当于 2019 年一个十几人的 MLA、MLE 团队,来处理你在编码时临时想到的任何任务,而且瞬间完成。不仅如此,质量还会比 2019 年更好,因为 AI 模型已经进步了很多。它会聪明得多得多。这对开发来说听起来很疯狂。所以,我认为分类器那说法是那些脾气暴躁、不太理解的机器学习人士说的。这完全没问题,因为 Jev 派对是开发者的派对。哦,天哪,我应该说开发者。所以,是的,而且实际上我认为还有另一类人,他们不只是说“我们可以用 LM 更便宜地运行同样的东西”。我认为他们是第二波人。第一波真正兴奋的人是那些觉得“AI 本来就该是这个形状”的人。这显然是让东西有用的方式。我说的是那些像呼吸一样热爱结构化输出的人,比如 DSPI。我不确定 DSP 怎么发音。但就像……
Yes. That is how you put intelligence into software to make the software smart. How is meta powered? Classifiers. Maybe some regressors, but probably mostly classifiers. How is Google powered? Probably mostly gasifiers. So class, there's nothing wrong with classifiers. I actually think it like if someone told us that we were a zerosot general classifier for anything, that would be the greatest compliment I've ever been given. For Jev, we have other things too, but like for Jev, that would be an incredible description of it. Like if someone could say that um I can give you, you know, and you know, random developer the equivalent of a MLA an MLE team of a dozen people from 2019 to work on any task that you think up on the fly while coding and it's instantly done. And not only that, it's going to be better than the quality in 2019 because the AI models have gotten a lot better. It's going to be way, way, way smarter. That sounds insane for development. So, I I think that the classifier thing is um like grumpy ML people who kind of don't get it. And that is totally fine because the JF party is a developer party. Oh man, I should have said developer. Um, so yes and not I actually think there's another class of people who is not just like we can run the same stuff with LM so much cheaper. I actually think that they are the they are the second wave of people. I think the first wave of people who got really excited were the kind of people who felt like this is how AI should have been shaped all along. This is like obviously the way to make stuff useful. And I'm talking about like the people who like live and breathe structured outputs like the DSPI. I I'm not sure how it's pronounced DSP. Um but like
我总是说 DSPY,但我认为 DSPI 是最流行的说法。
I always say DSPY but I think DSPI is the most popular
酷。
cool
的说法。
way of saying it.
但这些人一直在推动这样的愿景:这才是真正构建有用系统的方式。我相信他们在智力上非常正确,但他们一直被模型所束缚,因为模型是为字符串优化的,而不是为他们想要的东西——程序化使用——优化的。所以我认为那些人,不一定是 DSP,但有很多开发者直觉上就是这样思考的。我们之所以如此受欢迎,部分原因是他们在 Twitter 上尖叫,就像幻肢被归还了一样。我没想到会这样,因为我就是这样的人。我设计时就在想:我缺失的幻肢在哪里?结果发现有很多其他人也有同感,那真是太酷了。
But like these people have been kind of um you know trying to push this vision of this is how you actually make useful systems. And I believe that they're intellectually quite correct, but they've been hobbled by the models being optimized for strings instead of them being optimized for the thing that they want, which is programmatic use. So I think those people, not necessarily DSP, but like there's a lot of developers who intuitively think in this way. And part of what made us so popular is kind of on Twitter them screaming like a phantom lib was returned. And I I didn't expect this because I am like this, you know? I I was designing for like where's my missing phantom limb? And it turns out there's tons of other people with this and that was so so freaking cool.
那只是一条线索。还有更多。
That was only one crumb. Like there are more.
我甚至不记得其他线索是什么了。你呢?
I don't even remember what the other crumbs are. Do you?
我只是激动起来了。
I just got worked up.
呃,是的,嵌入模型加分类器。小模型。就像……
Uh yeah, embedding model plus classifier. Small models. Like
问题是,如果你放一个小模型或嵌入模型加分类器。很多人忽略了这一点,我的团队一直告诉我不要再告诉别人,因为这更可能让我们有竞争。但现在我认为大家都忽略了这一点。我想教育世界,即使这不符合我的经济利益。我真正想强调的是,这无关接口。接口很酷。甚至无关速度和成本。速度和成本是坏事。理想速度是即时。理想成本是零。没人会启动一个库然后说“我想把成本和延迟放进这行代码里”。对吧。所以我认为人们就是不明白他们是在为智能付费,如果你在嵌入或小模型上使用逻辑回归,你得到的就是嵌入或小模型上逻辑回归的智能。
the thing is if you do you put a small model or an embedding model plus a classifier. A lot of people miss this and my team has been telling me to stop telling this to people because it makes it more likely we have competition and right now I think everyone is missing this. I want to educate the world even if it's not in my financial best interest. And the thing I really want to emphasize is that it is not about the interface. The interface is cool. It is not even about the speed and the cost. Speed and cost are bad things. Ideal speed is instant. Ideal cost is zero. No one like spins up a library and says like I feel like putting cost and latency into this like line of code thing you want. Right. So the I think people just don't get that they're paying for intelligence and if you use a like a logistic regression on top of embeddings or a small model you get the intelligence of logistic regression on top of embeddings or a small model.
对。但大家都忽略了一点:你优化什么,就会得到什么。当然,如果你真的是唯一在做这件事的人,你就能获得巨大的飞跃。而我们就是唯一在做这件事的人,尽管我一直在喊:“嘿,大家,请做这个。这对世界真的很有价值。这才是让 AI 真正有用的东西。”而且,你知道,已经有很多模仿者了。但我仍然不觉得人们真正理解了。
Right. Yeah. But like everyone misses that you get what you optimize for. Of course, you can get the gigantic jump in this thing if you're literally the only person doing that. And we are literally the only person doing that, despite me trying to shout out like, "Hey, everyone, please do this. This is really valuable for the world. This is what's going to make AI really useful." And, you know, there's been lots of copycats. I still don't think people get it.
我正想说,你说没人在做这个,但就在一周之内,我想就有半打甚至更多的开源项目,还有一堆其他的,比如我昨天在 Dev Days 提到 OpenAI 宣布了他们的决策模型,我想那是个更小的模型。
I was just going to say you say no one's doing this but like within the space of a week I think there were half a dozen or more open-source projects, a bunch of other like I mentioned at Dev Days yesterday OpenAI announced their decision models which is I think a smaller model.
那关乎的是界面、成本和速度,而不是智能,对吧?如果人们真的——
Like that's about like the interface and the cost and the speed not about the intelligence right like if people really—
所以没有人采取和你一样的方法来达到这个目标。
So no one's taking the same approach that you're taking to get there.
不,他们甚至不需要采取同样的方法。他们只需要关心智能,你知道,就像如果你拿一个开源模型,做点什么让它有相同的界面和相同的特性,然后你只是在 NLP 任务上做基准测试之类的,你觉得我没想过吗?这 literally 是这个项目的第一天我就想到了。我以为这个项目会花一周。结果花了——我现在才敢完全诚实地承认,因为它太成功了。承认自己的失败感觉安全多了。我们开始这个项目的时候,我大概是世界上最差的并列,最好的话是专业训练水平。我以为会花一周,因为我为那么多不同的东西后训练过那么多模型。模型在很多事情上似乎非常非常通用,但显然在决策方面不是,这是我艰难学到的。所以我觉得人们在玩这些东西很酷。我担心如果人们玩的是糟糕版本的模型,他们可能会被整个子类型烧伤,然后说:“嘿,这种东西是垃圾。”我希望那不是——
No, they don't even need to have the same approach. They just need to care about intelligence, you know, like if you're taking like an open-source model doing something to make it have the same interface and same characteristics and you just benchmark on NLP tasks or something like that, like you don't think I thought of that. This was like literally day one into this project. I thought this project would take a week. It's been—I only feel okay admitting it now totally honestly because it's been so successful. It feels way safer to admit my failings. At the time we started this, I was probably at worst tied, at best at pro training in the world when we started. And I thought it would have taken a week because I've post-trained so many models for so many different things. And models seem to be very very general at a lot of stuff, but apparently not decision-making, which I've learned the hard way. So I think it's very cool that people are playing around with stuff. I worry that if people play with bad versions of the model, they might actually get burned by the whole subgenre and say like, "Hey, like this kind of stuff is crap." I hope that's not—
像人们说的那么大的问题——
As big a deal as people are saying that kind of—
没错,因为他们想要的是智能。如果模型完全智能——我们的模型不是,别人的也不是——那么你 literally 可以解决一切。分类是 AI 完备的,对吧?任何任务都可以归结为分类问题,你就能解决一切。所以真正重要的是智能,决定了你能自动化什么、不能自动化什么,以及你能构建什么样的科幻。这真的很重要。太重要了。我如此努力,你知道,如果我们不在乎这个,我们本可以早得多发布。即使现在,发布新东西也是一场辩论。当我们对可靠性有如此高的标准时,我们如何继续发布东西?因为我们想值得信赖。我们想真正对开发者友好,而不是对 ML 研究者,他们你知道,对哪个基准测试更好之类的事情很傲慢。我们只是想让人们不必担心他们的智能,因为这就是开发者为了构建东西所需要的。所以我基本上每天都为此激动。这太令人沮丧了,你知道街上的人认出我,他们说:“哦,你是 JEPA 那个人。”我说,不,我的名声变成什么了?你知道,我想成为一个典范,让 AI 可靠到像 SQL 一样无聊,你知道,我想让 AI 如此可预测,以至于你可以写查询甚至不用在评估集上运行,因为你知道这是常识智能会做的,而且每次都这样做。那时软件工程师就获得超能力了。是的,我 obviously 爱 JEPA。我喜欢我们快速且便宜。我们会继续推进。你等着瞧。而且,我怎么强调都不为过,可靠性是人们付费的东西。可靠性是人们想要的。如果你想将智能应用于应用,你需要智能。
Exactly, because the thing they want is intelligence. If the models were perfectly intelligent, which our models are not and no one else's are, then you could literally solve everything. Classification is AI complete, right? Like any task could be boiled down into classification questions and you could just solve everything. So truly intelligence matters in what you can and can't automate and what kind of sci-fi you can build. And it's really important. It's so important. I try so hard, you know, like we could have released so much earlier if we didn't care about this. And you know, even now it's a debate for like shipping new stuff. How do we continue to ship stuff when we have such a high bar for reliability? Because we want to be trustworthy. We want to be just really to developers, not ML researchers who are like, you know, snooty about like which benchmark things are better at. We just want people to like not have to worry about their intelligence because that is what developers want in order to just build stuff. So I get worked up about this basically every day. It's such a frustration and you know people on the street recognize me and they're like, "Oh, you're the JEPA guy." And I'm like, no, what does my reputation become? You know, like I want to be a paragon of making AI so reliable that it's boring like SQL, you know, like I want AI to be so predictable that you can like write queries without having to even run them against like eval sets because you know this is what common sense intelligence would do there and it to just do it every single time. And that is when software engineers get super duper powers. And yeah, I obviously love JEPA. I love that we're fast and cheap. We're going to keep pushing that. Just you wait. And also, I cannot emphasize enough, reliability is what people are paying for. Reliability is what people want. And if you want to put intelligence to applications, you need the intelligence.
那么,让我们深入探讨你如何获得这种智能。我想也许让我们从最终状态开始,比如它是什么,如何工作。我们可以从那里也许触及 RLCD,这是你过程的一大部分。但在这个过程中,我很想听听,你认为会花一周,结果花了更长。长了多少?也许还有一些里程碑和转折点是什么。
So, let's dig into how you get that intelligence. I think maybe let's start at the end state, like what it is, how it works. We can from there maybe touch on RLCD, which is a big part of your process. But along the way, I'd love to hear like you thought it would take a week, it took a lot longer. How much longer? And maybe what some of those milestones and turning points were.
哦,天哪。说到这个,如果你想听一个关于我的有趣尴尬事,当时我有了这个想法,灵光一现,我在 OpenAI 写了一份文件,标题是——也许不是标题——标题是“我们去年本可以拥有 AGI”,它本来是很辛辣的,实际上我当时确实相信。我仍然相信其中的情感,但里面的情感是,相当于 RLHF 努力的一年时间就足以制造这种新型的机器原生模型。而 RLHF 肯定可以在不到一年内完成。事后我们学到了,但事后我认为 RLCD 不可能在一年内完成。它只是难得多得多。我认为让 LLM 更适合做像取悦人类的事情,而不是让 AI 做出校准的稳健决策。当你移除 AI 的精神病态元素时,有太多的参差不齐,这非常——我会说这是大量的工作。我看到北极星,以及我们做事的方式,真的是真的真的试图找到什么方法,就像如果你想象一个认知核心,有这些参差不齐的裂缝,最有效的方法是什么来治愈它,以获得平滑稳健的性能,无论认知核心真正能够达到什么智能水平。所以这很多是关于理解 LLM 的能力。很多是关于看到它们的弱点。很多是关于弄清楚哪些弱点是可以解决的,以及我们如何解决它们。我们的模型不能在 strawberry 中数出 r。幸运的是,这不是现实世界工作的样子。但它可以做一大堆其他事情,因为很多工作是以语义判断的形式出现的,幸运的是——
Oh man. On that note, if you want like a fun embarrassing thing about me is at the time I had this idea and like the aha hit me, I wrote a document in OpenAI entitled—maybe not entitled—titled "We Could Have Had AGI Last Year" and it was meant to be very spicy where it's like and actually I did believe it at the time. I still believe the sentiment to it, but like the sentiment inside of it was like one year's worth of time of like the equivalent of an RLHF effort would be sufficient to make this new class of models that are made to be machine native. And RLHF could have been done in less than a year for sure. In hindsight we've learned, but in hindsight I think RLCD could not have been done in a year. It was just much much harder. I think that making like LLMs are much more suited to making things like human pleasing than AI is to be making calibrated robust decisions. And there's so much jaggedness when you remove the psychopathic elements of AI that it's very—it's been a lot of work is what I would say. And I see the north star and as how we do things is really really really trying to find what are the ways like if you imagine the cognitive core that has like these jagged cracks what is the most efficient way to like heal that to get smooth robust performance at whatever level of intelligence is truly capable within the cognitive core. So it's a lot of like understanding the capabilities of LLMs. It's a lot about like seeing their weaknesses. It's a lot about like figuring out how can—what weaknesses are addressable and how do we address those. Our model cannot count ours in strawberry. Luckily, that is not what real world work looks like. But it can do a whole bunch of other stuff because a lot of work is in the shape of semantic judgments which luckily—
你不能把它分解成一个分类问题。
You can't decompose that into a classification problem.
哦,是的,你可以分解它,但你不能只是原味地问它。而且在某个时候,有点像我只是在开玩笑。
Oh yeah, you can decompose it, but you can't just ask it vanilla. And at some point like it's a little bit I'm just teasing.
是的,某种程度上这算作弊,对吧?比如你把它拆成一个列表,它就能工作,这样 token 就不会被合并。但如果你想要一个合理的分解,对你的工程系统来说有意义,理想情况下应该处于正确的抽象层级,让工程师能理解并且可维护,而不是对任何一个问题过拟合。所以即使我们的模型在完美分解的情况下能解决每一个任务,那可能也不够,因为我们想不断向上移动抽象层级,让人们能在与他们思考方式相匹配的语义层编程,对吧?比如如果你能把“这个人是否想要退款”拆成一百种不同的退款场景,也许那样能行,但那还是很烦人。有些企业还是会这么做,但我不认为那样就算完成了任务。
Yeah, at some point it's cheating, right? Like if you break it into a list then it works, so that the tokens don't get merged. But if you want a sane level of decomposition that makes sense for your engineering system, that's ideally at the right level of abstraction that it's understandable by an engineer and it's maintainable, versus overfit to any one question. So even if our model could solve every single task if decomposed perfectly, that may not be sufficient because we want to keep moving up the abstraction hierarchy so people can program at the semantic layer that matches the way they think about it, right? Like if you could break down "does this person want a refund" into a hundred different refund scenarios, maybe then that'll work, but that's still kind of a pain in the ass. Some businesses will still do it, but I wouldn't consider a job done for that.
我们大致谈过 Jev,可能每个听众都已经见过它、玩过它、对它有所了解,但以防万一,我之前问了你一堆问题。我们来分解一下。第一个是从开发者的角度,你的目标用户,这个东西看起来和感觉起来是什么样的?我大致把它看作一个智能 API。你给它一堆可能的决策和一些输入,它做出其中一个决策,或者告诉你它认为是哪个,并给出一些概率。你会如何调整它?
We've talked broadly about Jev, and it may be that everyone listening has already seen it, played with it, knows a bit about it, but in case not, I asked you a bunch of questions before. Let's break those down. The first is from the perspective of a developer, your target user, what does this thing look and feel like? I think of it broadly as you could call it an intelligence API. You give it a bunch of possible decisions and some input, and it makes one of those decisions or tells you which it thinks it is, and it gives you some probabilities. How would you tune that?
我喜欢这个。很多人用很多不同的方式描述它,只要准确,我都接受。我也喜欢“智能的 SQL”这个描述,或者它像是一个组合,对吧?像一个系统,有接口和模型,我们两者都做了,所以有点令人困惑。但对我来说,我把接口看作构建智能应用的积木,有点像添加第四个逻辑门之类的。所以我真的希望智能能被标准化成一套通用的原语,理想情况下每个人都能让这些东西尽可能智能,而不仅仅是我们。我也喜欢这样的定位:嘿,我们有这个叫语言模型的东西,这里有一个新东西叫决策模型,这有点说明了它的目标,不一定是 Jev,而是某种新兴的类别……
I love it. A lot of people describe it in many different ways, and I take all of them as long as they're accurate. I also like the description of SQL for intelligence, or like it's a combination, right? Like a system, there's the interface and the model, and we kind of made both, so it's a little bit confusing. But to me, I see the interface as the building blocks for making intelligence applications, kind of like adding a fourth logic gate or something like that. So I really want intelligence to be standardized into a common set of primitives, and ideally everyone makes these things as smart as possible, not just us. I also like the positioning as, hey, we have this thing called language models, here's this new thing called decision models, that kind of says a bit about what it aims to do, not necessarily Jev, but like some emerging class of...
是的,告诉我。
Yeah, tell me.
我会给一个小小的警告。我们考虑过“决策模型”这个名字。我不反对这个名字。我认为最好的类别是由社区命名的。但我要声明,我们没有选择它是有原因的,我认为符合这类模型的东西远不止决策。所以我会说,人们与 Jev 的公共接口可以被称为决策模型,但我个人不会这么叫。
I will give a tiny warning about this. We considered decision models as a name. I'm not opposed to the name. I think the best categories are named by the community. But I will give a disclaimer that there's a reason we didn't pick that, and I think there's a lot more that fits this class of models beyond decisions. So I would say that people's public interface to Jev could be called decision models, but I personally wouldn't call it that.
Jev 擅长但又不是决策的例子是什么?
What's an example of something that Jev would be good for that's not a decision?
我不会告诉你,但如果你愿意,我可以给你一个提示。我说过,决策和分类是让智能对软件有用的方式,而我们就是致力于让智能对软件有用。我也说过,可用的不仅仅是决策,这听起来有点矛盾,但我会说有一天它会变得显而易见。我还不能透露,因为这些在路线图中还有点远,但如果你的读者想琢磨一下,这是个有趣的思考题。
I will not tell you, but I will give you a hint if you'd like. I've said that decisions and classification is how you make intelligence useful to software, and we are all about making intelligence useful to software. And I'm also saying that there's more than just decisions available, which sounds like a little bit of a paradox, but I will say it will one day be obvious. And I can't reveal that yet because these are things a little bit more distant in the road map, but it's a fun one to think about if your readers want to noodle on this.
所以这就是它的形态和功能表面。我们再谈谈它的制作。核心是一个模型,我们不叫它决策模型,就叫它 Jev。它是通过一个叫 RLCD 后训练的过程创建的。它是对其他东西的后训练吗?还是从头开始的东西?你能告诉我们关于这个东西的什么?
So that's kind of the shape and functional surface of it. Let's talk a little bit more about the making of it. At the core there's a model that we won't call a decision model, we'll just call it Jev. And it was created through a process called RLCD post-training. Is it a post-train of something else? Is it a ground-up thing? What can you tell us about this thing?
我会以一种过于精确的方式来回答。在 RLHF 存在之前,后训练这个概念并不存在,对吧?我们作为一个团队,没有命名后训练,但我们创造了后训练的概念。当时微调已经存在,人们会说,你只是在字符串上微调模型吗?按照某些定义,我们是在微调模型,但我们做了更多,以至于在某种程度上变得有误导性。所以我实际上会说,事后看来,称它为后训练可能是不准确的,但这个领域还没有术语来描述我们正在做的事情。所以这可能会在未来更多地揭示。
I will answer this in a way that is overly precise. Back before RLHF existed, post-training as a concept did not exist, right? We as a team, we didn't name post-training, but we created the concept of post-training. At the time fine-tuning existed, and people were like, are you just fine-tuning the model on strings? By some definitions we were fine-tuning the model, but we did a whole lot more such that at some point it becomes misleading. So I would actually say that calling it post-training is probably going to be seen as inaccurate in hindsight, but the field doesn't have a terminology yet for what we are doing. So that will probably be revealed more in the future.
如果你现在要发明一个,你会叫它什么?
If you were to invent one right now, what would you call it?
团队里没人喜欢我的想法。所以我在别处说过,所以我可以这么说,但我们回到 bitter 的教训。我们,因为这是机器学习,我们非常关心北极星,我们愿意为它做一切。我尽量坦诚地说明权衡。其中一部分权衡就是在字符串生成方面差得多。甚至像 Jevkin 聊天之类的东西,它很可爱,但显然很笨。这些是我们有意做出的权衡。所以在设计整个流程时,在 LLM 之上进行后训练要容易得多,因为形状非常相似。这里是字符串,那里基本上也是字符串。技术上,像聊天 ML 格式有一些结构,但主要是字符串。所以堆栈中有很多相似之处。我们正在做的是,实际上拆除了很多基本部件,重新调整它们,并使用其他东西的部件。所以我把它描述为模型的弗兰肯斯坦的怪物。我没读过那本书,但有人告诉我它至少是无辜的,如果不是好人的话。澄清一下,因为人们会说,“别把模型叫怪物。”我说,“但这是个好类比。”
No one likes my ideas in the team. So I've said this elsewhere, so I'm okay saying this, but we go back to bitter's lesson. We, because this is ML, we care a lot about the North Star, and we are willing to do everything for it. And I try to be as upfront about the trade-offs as possible. And part of that trade-off is being a lot worse at string generation. And even like the Jevkin chat stuff, it's very cute, but it's obviously dumb. These are trade-offs we intentionally make. So when designing the entire pipeline, post-training was a lot easier on top of LLMs because the shape was very similar. It's strings here, it's basically strings there. Technically there's some structure with like the chat ML format, but it's mostly just strings. So there's a lot of similarity in the stack. What we are doing is we're actually ripping out a lot of fundamental pieces and rejiggering them and using pieces from other stuff. So I have described it as a Frankenstein's monster of models. Which I haven't read the book, but I've been told is at least innocent, if not the good guy. Just to be clear, because people are like, "Don't call the model a monster." And I'm like, "It's a good analogy, though."
你把这个未命名的东西应用在什么基础之上?是像现成的 Quen 模型之类的吗?
What's the foundational thing that you're applying this unnamed thing to? Is it like an off-the-shelf Quen model or something?
很多不同的开源模型,因为它们都有不同的权衡。所以我们实际上到处都在弗兰肯斯坦。但如果我们想要最大程度地精确和有用,重要的是找出北极星。弄清楚你愿意权衡什么,移除所有对那个权衡无用的东西,或者为了你放弃的东西所需的东西,并超级专注于重要的事情。其中很多将遵循 Allah 第一性原理,经过数月的研究。
A lot of different open source models because they all have different trade-offs to them. So we are actually Frankensteining all over. But the important bit if we want to be like maximally precise and useful is figure out the North Star. Figure out what you're willing to trade off and remove all the stuff that is useless for that trade-off or that is needed for the thing you're giving up and hyperfocus on the thing that matters. And a lot of it will follow Allah first principles with months of research.
当然。但总的来说这很宽泛。
Sure. But that's broad in general.
具体来说,有趣的是你并非从零开始。你从 LLM 和 Transformer 出发,在这些已有的构建模块上搭建,明确一下,是 Transformer。
And to be concrete and specific, it is interesting that you didn't start from nothing. You started from LLMs and transformers, and you're building on these building blocks that have established transformers, just to be clear.
我没说 Transformer。我说的是开放权重模型,略有区别,我只是想精确一点。
I did not say transformers. I did say open weight models, slight distinction, which I would just want to be precise.
所以你到底说没说?
So you did or did not?
我没说 Transformer。我说的是开放权重模型,略有区别,我只是想精确一点。
I did not say transformers. I did say open weight models, slight distinction, which I would just want to be precise.
所以我认为,首先,很明显我们没有做预训练。如果我们能用现有的预算比全世界预训练得更好,那全世界就彻底没戏了。我认为自由交易并不是世界上最有效率的事。为可争议的线性收益投入指数级更多的资源,看起来非常次线性。在我看来这是笔糟糕的交易,除非必要,也就是说我们已经把海量内容煮沸浓缩成下游流程可消化的东西。除非你在做根本不同的事,否则为什么要继续做?
So I think that, well, number one, it's very obvious we haven't done pre-training. Like, if we could pre-train better than the rest of the world with the budget we've had, then the rest of the world is totally cooked. I think free trading is not the most efficient thing in the world. And it's exponentially more resources for debatably linear games seems very sublinear. It's kind of a bad deal in my opinion unless it's necessary, meaning we've kind of already boiled the ocean of content down into things that are consumable by some downstream process. Why keep doing it unless you're doing something fundamentally different?
没错。所以我们更专注于能实现的 10 倍或更大的跳跃,我们认为这完全可以在预算内完成。我们显然这么认为,因为我们做到了。但我认为所有 AI 公司都应该被要求达到同样的标准。就我个人而言,
Exactly. So we are trying to focus more on like the 10x or more jumps that can be made and we think that that can be very much done in a budget. Well, we obviously think that cuz we did. But I think that all AI companies should be held to that same bar. Personally,
我之所以想深入探讨这一点,并不一定是想让你说“哦,我们用了 Quen 3.6 之类的”。而是我花了一段时间才理解,这件事的核心是一个 LLM,它是智能的来源,是零样本能力的来源。如果你只从 Twitter 上了解,觉得“哦,这不过是个很酷的新分类器”,你可能会想“那可能是 XG Boost 或其他什么,什么都行,他们做了一些真正的训练”,但接着你会想“哦,那我得对它进行后训练,或者给它一些我领域的信息。我怎么能只给它一堆决策让它决定,智能从哪来?”但如果你意识到底下有个 LLM,你就能开始看到这如何运作。
part of the reason why I'm trying to really drill down on this point is not necessarily for you to say, "Oh, we use, you know, Quen 3.6 or whatever it was." is that it took a while for me to wrap my head around the fact that there was an LLM at the core of this thing and that is the source of intelligence and that is the source of zero shot. And like if you think about it as like, if all you've seen is Twitter and like oh it's like this cool new classifier thing and you're like oh okay well that could be like I don't know XG Boost or like some other, it could be anything and like they've done some like really training thing, but then you are like, "Oh, well then I have to like post train it or I have to like give it some information about my domain or something like that. How am I just going to give it a bunch of decisions and ask it to decide like where is that intelligence coming from to your point about intelligence?" But then you can start to see how that could work if you realize there's an LLM down there somewhere.
我能理解这个类比。我有点想避免 LLM 这个术语,因为严格来说它不是语言模型,它并不对语言建模,尽管人们想强行这么认为。我认为这更广泛,最近这波 AI 一直在创造这些极其棒的浓缩智能片段,但周围有脚手架把它变成 LLM。所以想象一个巨型电池,上面有些东西把它变成语言模型。
I could see that analogy. I kind of, I'm trying not to avoid the LLM terminology because it's technically not a language model, you know, like it doesn't model language despite people wanting to force it into that. I see this as this more broad, this recent wave of AI has been creating these incredibly awesome condensed bits of intelligence, but with a scaffolding around it that turns it into an LLM. So imagine like this mega battery with stuff on top that turns it into a language model.
我只是指它像是互联网的压缩版本。
I only meant it in the sense of like a compressed version of the internet.
没错。没错。所以我会用那个,我觉得
Exactly. Exactly. So I would use that and like I see
你有一个由大量努力创造的互联网压缩版本,你在使用它
you've got a compressed version of the internet that some large effort created and you're using that
正是如此,我认为这是获得我所说的通用智能的唯一途径。这是我目前看到的通往通用智能的唯一方法,或者你知道,是否称之为通用智能有争议,但我认为它是通用的。
exactly that and I think that's the only way to get what I'm calling general intelligence. That's the only way I've seen to get to general intelligence so far or if you know it's debatable if you want to call it general intelligence but I think it's general.
不是要过度简化,你做了很多事情,听起来像弗兰肯斯坦一样拼凑了几个不同的模型。你有一个复杂的过程,我们不会称之为后训练。但有个东西叫 RLCD。那是什么?
Not to oversimplify like you've done a lot of things like you've Frankenstein it sounds like several different models. you you have an involved process that we won't call post training. But there is this thing called RLCD. What is that?
回到 RLHF 的话题,RLHF 不是关于在奖励模型上做 PO,你知道,比如先做 SFT,然后训练奖励模型,再基于策略训练。
So back to the topic of RLHF like RHF is not about PO on reward model on you know like doing first SFT and then reward model training and then on policy training of that.
好的。RLHF 是关于人类决策,而你的东西是关于决策。校准的决策。
Okay. RLHF is about human decisions and your thing is about decisions. Calibrated decisions.
没错。人类偏好。从来不是关于人类决策。是关于赞美人类。没错。人类反馈。
Exactly. human preference. It's never about human decision. It's about praising humans. Exactly. Human feedback.
所以当我谈论 RLCD 时,我指的是专门针对校准决策进行优化的一类算法,因为这对软件有用。我们这里确实有许多算法在尝试。我们实际上讨论过发表一些东西。不过我不会承诺很快,因为团队超级忙。但我认为你可以想象,我认为发表 RLHF 可能对世界很有好处。你知道,它让很多人理解了东西。我们花了几个月才发表。但你知道,我们有一段时间有 instructor GPT,我认为最终在这里发表一些东西也会对世界很有好处。所以我认为这有可能。但任务是任务,算法是算法。明确一下,其他人可能会在同一个任务框架下做出其他算法。
So like when I talk about RLCD, I talk about like the general class of algorithms that optimize specifically for calibrated decisions because that is what's useful for software. So we do have many algorithms in here that we have been playing around with. We actually have talked about publishing something here. Though I will not promise anytime soon because the team is super duper swamped. But I think that you could you could imagine but I think that publishing RHF was probably really good for the world. You know, it allowed a lot of people to understand stuff. It took us a couple of months to publish it. But you know, we had instructor GPT for a while and I think that eventually publishing something here would be really good for the world too. So I think that that could be in the cards. But the task the the t that's the algorithm versus the task. Just to be clear, other people will probably make other algorithms within the same umbrella of the task.
多谈谈任务有用吗?还是你指的是我们迄今为止谈论的那种广义任务?比如在学术背景下,我会问设定,RLCD 的设定是什么?你是怎么定义的?
is it useful to talk more about the task or do you mean the task broadly in the way we've been talking about it thus far? Like what like in an academic context I'd ask about the setting like what is what is the setting for RLCD? Like how have you defined that?
那非常复杂,但我可以给一个简化版。你用过聊天补全 API 中的函数调用吗?函数调用有点像决策,因为它确实是决策,但它在字符串层之上进行。所以所有可用的旋钮都是字符串模型仍然可用的字符串旋钮。字符串旋钮的一个例子是逻辑偏置。逻辑偏置允许你朝采样特定 token 的方向微调。你可以这样做来让某些 token 更可能或更不可能,如果这说得通的话。我的一个主张是,没有函数本身的逻辑偏置,就不应该有函数调用 API。为什么你不向用户提供那个旋钮?因为如果你在做客服,需要决定是否升级到人工,沃尔玛和 Costco 的政策可能截然不同。我对这些政策一无所知,只是随便举例。但现在唯一的替代方案是在语言中编程,比如“请退款”或“请不要退款”之类疯狂的东西,而不是把系统控制权交给实现下游行为的程序员。所以对我来说,RLCD 是关于向构建者暴露正确的接口,让他们能得到想要的属性,而不只是向系统消息祈祷。
That is very complicated but I can give a simplified version of it. Have you ever used like function calling in the chat completions API? So function calling is kind of like decision-making because it is decision making really but it's done on a layer above the string stuff. So all the knobs available are the stringy knobs that are still available from like the string models. And one example of a stringy knob is a logic bias. A logic bias allows you to like nudge in the direction of sampling a particular token. So you can do that to make certain tokens less or more likely if that makes sense. And one of my claims is that no function calling API should exist without a logic bias for the function itself. Like why would you not provide that knob to a user? Because if you are, for example, doing a customer service thing and you need to figure out if you're escalating to a human, your policy is probably extremely different if you're Walmart versus like Costco. I know nothing about these policies, but I'm just baking stuff up. But like the only alternative now then is to like program again into the language like please give refunds or please don't give refunds or you know something like insane instead of giving the controls of the system to the programmer implementing the downstream behavior. So to me RLCD is about exposing the right interface to builders so that they can get the properties they want without just praying to the system message.
像这类基础的东西是基本要求。比如,这最近没发生,但以前 OpenAI 有过过度拒绝的问题。他们不得不回滚模型,因为它拒绝得太多了。我认为那其实是个糟糕的决策。当时的情况是,模型在其权重内做出拒绝或不拒绝的决定。然后系统就只是按照语言模型的想法走,因为它把决策和文本混在一起,然后就这么做了。所以从构造上讲,确实存在一些查询,如果你刷新几次,它有时会拒绝,有时不会。从工程角度看,这简直是疯狂的行为,对吧?为什么你不设一个阈值?为什么你不设一个可调的阈值?为什么你不设一个至少能间接响应人们抱怨的可调阈值?或者如果你在做客服,为什么你决定是否升级的阈值不能依赖于一些非常基本的东西,比如这个客户对你、对我们值多少钱,以及我们的人工客服现在有多忙,因为也许那会改变事情,而今天我们根本没有这些控制。所以我认为 RLCD 真正暴露了最原生工程控制,那就是任务的形状,很多东西都来自那个形状。为了有有意义的旋钮可调、有阈值可切,你需要校准。你需要你的概率是聪明且准确的。你关心概率的长尾,而不仅仅是准确率。任务本身有很多东西要研究,看到权衡,弄清楚正确的行为是什么。一个 ML API 的产品有很多东西大多数人没有意识到,因为他们盲目地追随像 MAU 这样的梯度,呃,月活跃用户。
And like I like these types of basic things are table stakes. Like uh this hasn't happened recently, but this happened back in the older days when OpenAI had issues with over refusals. They had to like roll back the model because they refused too much. And I would argue that that was actually bad decision-making. What was happening is the model within its weights was making a decision refuse or no refuse. And then the system is just going with what the language model thinks because it blends the decision with the text and then just does that. So there does exist queries by construction that if you refresh it several times it will sometimes refuse and sometimes not refuse. That is insane behavior from like from an engineering point of view, right? Why wouldn't you have a threshold? Why wouldn't you have a tunable threshold? Why would you not have a tunable threshold that responded at least somewhat indirectly to how people are complaining, right? Or if you were doing customer service, why can't your threshold on whether to escalate or not be dependent on really basic things like how much is this customer worth to you to us and how busy are our, you know, human workers right now because maybe that changes things and we just don't have any of those controls today. So I see RLCD as really exposing the most native to engineering controls and that is the shape of the task and a lot of stuff comes from that shape. Like in order to have meaningful knobs to tune and thresholds to chop, you need to be calibrated. You need your probabilities to be like smart and accurate. You care about the long tail of probabilities, not just the accuracy. And there's like a lot to the task itself on studying this, seeing trade-offs, figuring out what the right behavior is. There there's there's a lot to the product of an ML API that most people don't realize because they mindlessly follow the, you know, the the gradient of things like MAU's uh monthly active users.
你谈到 RLCD 设置是一种你有一些输入的情况,呃,你知道有用户提供的输入,但也有一些控制系统内部偏差的旋钮。当前的 Jev 产品是否向用户暴露这些,还是这些旋钮是 Jev 工程团队用来创建这个你向世界发布的决策 API 的?
You talked about the RLCD setting being one of you've got some kind of ser some set of inputs uh and and you know there are the inputs that the user provides but there's also these knobs that kind of control internal bias in the system. Does the current Jev product expose those to the user or are those knobs that the Jev engineering team has access to to create this thing that uh you know this decision API that you've published to the world?
据我所知,内部 API 和外部 API 基本上是一样的。所以我们想要,因为我们怎么知道用户应该把这个设成什么,如果我们没有他们的应用,对吧,什么才是正确的,所以置信度是其中的一部分,概率是其中的一部分,大多数人不用概率,但我猜当事情进入硬核工程时,你肯定会每次都使用概率,而不是只选择模型认为最好的,因为即使所有的选择,即使你有完美的概率,对于你作为公司或企业部署这个可以采取的每个行动,甚至一个软件工程师信任它用于自己的家庭系统。每个这些行动都有不同的成本和收益。
As far as I can tell the internal API and the external API are basically the same. So we want we because how would we know what users should set this to if we don't have their application right like what is the right so confidence is part of this probabilities are part of this most people don't use the probabilities but I would guess when things get into hardcore engineering you would definitely use the probabilities every time over just picking what the model thought was best because even if all of the choice even if you had like the perfect probabilities for each of the actions you can take as a company or a business deploying this or even a software engineer like trusting it with your own home system. Each of these actions have different costs and benefits.
嗯,我想回到概率作为输出。我更想的是输入旋钮。但当我,你知道,我粗略地看了一下 Jev API,它像是一个非常简单的 API 形状,但我没有深入文档,看看除了我发送的决策之外,你让我访问的所有可能的旋钮,比如那些旋钮有哪些?
Well, I guess I want to come back to the to the probabilities as outputs. I'm more thinking about like input knobs. But when I've, you know, my cursory look into uh into the Jev API, it's like a very simple API shape, but I've not dug into the docs and looked at all the possible knobs that you give me access to beyond just the decisions that uh I send in like what are some of those knobs?
旋钮是通过输出概率来实现的。我们基本上是信任用户来做这些决定。这相当于想象速度不是问题。这相当于 OpenAI API 为每个 token 输出 logits,然后软件开发者可以选择他们想如何链接采样算法,而不是让 OpenAI 处理一切。所以我们通过给人们原始访问模型来提供这些旋钮。然后其他人可以访问字符串模型。这样他们可以调整他们的行为到他们系统中实际需要的,而不是像 OpenAI 认为的对所有使用 API 或第一方产品的人全球都合适的行为。
The knobs are via outputting the probabilities. We are basically trusting the user to make those decisions. This would be equivalent to imagine speed wasn't an issue. It would be the equivalent to the OpenAI API outputting logits for every token and then the softwares can the software developers can choose how they want to chain the sampling algorithm instead of letting OpenAI just handle everything. So we give those knobs by giving people a raw access to the models. Then other people have access to the string models. And that way they can tune their behavior to what it actually is needed in their system instead of what you know like a open IPM thinks is like a globally okay behavior for everyone using the API or first party product.
这让我想到一个关于你们如何做到的问题。我想这又回到了 RLCD。但历史上,不确定性估计一直非常困难。尤其是当我们谈论这些大型压缩的互联网版本,它们历史上生成文本,以及大型复杂模型一般,所以你们怎么做?
Which brings me to a question I had about how you do that. And I guess I imagine this is circling back on RLCD. Um but like historically uncertainty estimation has been like very difficult. Um and uh you know particularly when we are talking about these you know large compressed versions of the internet that historically generate text and like uh and large complex models generally like so how do you how do you do that?
艰苦的工作,血汗和泪水。而且,你知道,不,这是艰苦的工作。我们隐身了很长时间。你知道,在 AI 领域,隐身两年有点疯狂。
Hard work, blood, sweat, and tears. And also and also like, you know, no it's hard work. Like we were in stealth for a long time. You know, being two years in stealth is like kind of insane in in like AI terms. The amount of
我不是要你泄露家族珠宝。更像是,它是基于什么,比如指导星或原则,帮助我们理解为什么我们应该信任这些概率,而不是他们只是有一个 token 生成器,被训练来生成 0 到 1 之间的数字。
I'm not asking you to divulge a family jewels. It's more like is it is it grounding on something like what's the guiding star or the like the the principle that kind of helps us understand why we should trust these probabilities as opposed to well they've just got some token generator that's been trained to generate numbers between zero and one.
我再给你一根骨头,最后一点,你知道我们在隐身中工作了很长时间。我们错过了很多潮流,尽管我们认为 Jev 对很多潮流会很好。比如,如果我们在 openclaw 的同时推出 Jev 会很酷。会很棒,因为每个人在 claw 上都有很多可靠性问题,而且这些压缩的互联网东西,据我们测量,校准得还不错。它们不是完美校准,但比 RHF 模型校准得多得多。实际上,RHF 和 RLVR 极大地破坏了模型的校准,因为为了很好地输出字符串,你实际上需要极度的过度自信。呃,你需要像模式丢弃损失来使那工作,这完全扭曲了概率,我认为 GPT4 论文开始谈论这个,因为那是模型开始不发布预训练版本,只发布后训练版本的时候。所以我认为 GB4 论文深入了一点,但模型曾经校准得还不错。看看很多 OpenAI 和 Anthropic 的早期发布,你知道测量,用他们所谓的 PI,知道概率它知道事实,这些都很不错,而且肯定不够好,我猜 Jev 会轻松碾压所有这些。问题是,有方法从所有这些中引导校准、可靠性和智能。
I'm going to give you like one more bone in the last thing which is you know we did work in stealth for a long time. We missed so many fads even though we thought Jev would have been great for a lot of those fads. Like holy would have been cool to launch Jev the same time as openclaw. It would have been awesome because everyone has like a lot of reliability issues claw and also like these compressed things of the internet are as far as we can measure decently calibrated. They're not perfectly calibrated but they are way way more calibrated than RHF models are. actually RHF and RLVR destroy the calibration of the models immensely because in order to output strings well you actually need extreme overconfidence. Um you it requires like a mode dropping loss in order to make that work which just completely warps the probabilities and I think that the GPT4 paper started talking about this because that's when models started not shipping their pre-trained versions and only started shipping post-trained versions. So I think the GB4 paper went into it a bit but the models used to be decently calibrated. Look at a lot of opening um anthropics early publishing you know measuring uh with what they called PI know the probability it knows facts like these are all like pretty decent and it's definitely not good enough I would guess Jeb would crush all of those like super easily. The the thing is that there's ways to bootstrap calibration and reliability and intelligence from all of that.
就像 RLVR 引导推理一样,对吧?在理想情况下,RLVR 的推理并非来自人类书写的推理,因为它以不同的方式运作。但如果你有正确的 RL 目标——校准决策,就有办法从种子数据出发,在训练过程中变得越来越校准。我不太确定它叫什么,epoch 阶段之类的。所以有一些粗略的方式,说得通,你可以从并非完全虚无、而是从 AI 今天拥有的美丽原材料中,得到校准的东西,因为 AI 就像——我觉得 AI 仍是一块未经雕琢的钻石。你知道那里有太多事情可做。如果我不是在做 Jev,我可能理想情况下不会创办公司,而是基于 AI 模型做其他很酷的事情。
The same way that RLVR bootstraps reasoning, right? Like in an ideal world, the reasoning of RLVR doesn't come from humans writing the reasoning because it does it differently. But if you have the right RL objective of calibrated decisions, there are ways to get that working from that seed data into becoming more and more and more calibrated over like the training. I don't really know what it's called, epoxing stages, whatever. So there are handwavy ways that makes sense that you can get calibrated stuff out of not quite nothing but out of like the beautiful raw material that AI has today because AI has like a—I feel like AI is still a diamond in the rough. You know there's so much stuff to do there. And if I wasn't doing Jev, I would probably be ideally not starting a company but doing other cool stuff on top of the AI models.
比如你如何实现泛化?大概你是在某个决策语料库上做 grounding,这样你就能教它决策是什么样子,以及如何映射到这些概率之类的。但就像所有训练工作一样,你必须问在分布内的问题,模型的结果才具有代表性。那么,如果你试图实现零样本分类、零样本决策,我怎么知道我的特定问题是在分布内还是分布外?也许更广泛地说,我认为对于工作流、决策以及软件交互这些我们试图自动化的东西,泛化会比仅仅生成文本更难。但我可能大错特错,所以你来告诉我。
Like how do you get to generalization? Like you know presumably you've got you know you're grounding on some corpus of decisions so you can teach this thing like what decisions are like and you know how to map to these probabilities and all this kind of stuff. Um but like that uh you know as with all these training efforts like you know you have to be asking things that are in distribution in order for the results of the model to be representative and like how should we think about you know if you're trying to enable you know zero shot you know classification like zero shot decision-making how do I know that like my particular problem is in distribution versus out. And maybe more broadly like I would think that like generalization would be a harder problem for workflows and decisions and like software interactions and all these things that like we're trying to automate than uh than like just generating text. But like I could be way off on this so you tell me.
这是个有趣的问题,也许作为对话比讲座更有趣。你知道,也许你为什么认为 AI 缺乏泛化?我也有自己的答案,稍后可以说,但你觉得为什么泛化不是 AI 固有的?听起来你在问一个哲学问题,而且
It's an interesting question which maybe would be more fun as a dialogue than a lecture. Do you know like maybe why do you think AI lacks generalization? I have my own answers that I could say later too but like what makes you think that generalization is not native to AI? It sounds like you're asking a philosophical question and
我可以问一个实际版本:ChatGPT 感觉有多通用?我认为这个问题来自一个一般原则:响应的可靠性由训练数据的映射以及你所问内容对模型的熟悉程度决定。用非常宽泛的术语来说,对吧?所以我认为决策比文本更难泛化的原因是,假设我在世界上有足够的文本,如果我让模型生成文本,它开始理解如何把词组合在一起,如果我不要求精确的事实,它就能组合出看似合理的词。而在会计工作流中,我需要做一件非常精确的事情,如果那个工作流与标准做法非常不同,我认为模型会非常自信地犯错。
I can ask a practical version you know how general did chat GPT feel? I think the perspective that you know that the question came from was one of you know as a general principle like the reliability of responses is governed by kind of mapping the training data and like uh you know the degree to which the thing you're asking is familiar to the model. you know, using these terms very broadly and like uh very broadly, right? And so like I think why the take that um you know decisions would be harder to generalize than uh than text is because like presumably I get enough text in the world if I'm asking a model to generate text like it starts to understand like how to put words together and um if I'm not asking it for precise it's facts then you know it can like put words together that are plausible whereas like in the case of you know I've got an accounting workflow and I needed to do a very precise thing like if that workflow is very different from the standard way of doing that workflow like I would think the model would be very confident and wrong.
哦,有趣有趣。所以,ChatGPT 是否通用,这是我想开始对话的一种方式。你觉得它通用吗?我的意思是,我认为它是通用的,这可能走向一个不同的哲学方向,但你知道,我非常强烈地反对用 AGI 这个宽泛术语来描述我们所取得的成就,但过去三个月,也许我越来越能说我们已经实现了这一点,在某种意义上
Oo so fun fun stuff. Um so you know like was chat GPT general is what one one way I would like to start the conversation. Did it feel general to you? I mean I think it I think it is general like and I mean this is maybe going in a a you know a different philosophical direction but like you know I was very very strongly against AGI as like a a broad term for like what we've achieved but like in the past three months maybe like I'm getting closer to being able to say that like we've achieved this like in the sense of
嗯,我可以和一个智能体对话,让它像员工一样在我的业务中做事。AGI 和泛化是不同的概念,但我接受 ChatGPT 是通用的这个想法。是的。我不知道这是否否定了分布内/分布外的问题。我的意思是,有一点,因为分布内/分布外也取决于你在这方面工作了多久、多努力,因为你开始挖掘长尾示例。所以从大局来看,今天的 ChatGPT 即使使用即时模式(关闭推理)也比四年前的 ChatGPT 精细得多。嗯,3 年 10 个月。嗯,当然,但不一定是因为模型更通用,而是因为我们现在构建了这些框架,知道如何出去寻找信息来填充上下文。
um you know I can talk to an agent that you know, you know, conversationally and have it do things in my business like an employee might. Like it it is like AGI and generalization are are different concepts, but like I'm accepting of the idea of chat GPT being general. Yes. I don't know that that like negates the in distribution out of distribution line of question. I mean there there's a bit of it because in an out distribution is also a function of like how long and hard you've worked on that because you start mining the long tail of examples. So in the grand scheme of things ChatGPT of today is a lot even if you do use instant mode so turn off reasoning is a lot more um refined than ChatGPT of four years ago. Well 3 years and 10 months. Well, sure, but not necessarily because the model is more general, but because now we've built these harnesses that know how to go out and find information to fill in context.
我的意思是,如果只谈 ChatGPT,我想确实有泛化收益,但 ChatGPT 已经相当通用了,我认为它已经能做好很多简单的事情。我确实认为有些通用性是事物固有的。我实际上有一篇单独的博客文章,我相信叫《谎言、该死的谎言和基准》,我在其中谈到我们今天看到的 LMS 的许多脆弱性和锯齿状表现来自 RLVR,因为如果你对 RLVR 不宽容,它按定义就是在为基准优化,因为可验证的奖励按定义就是可基准化的。通过在答案前加入潜在推理,你给了模型全世界的力量在这些基准上做得非常好。所以它非常不令人惊讶。它在基准上做得很好。我认为很多锯齿状表现来自这一点。然后你就面临这个挑战:添加越来越多的分布内数据,获得越来越少的回报,使锯齿状前沿更加分形和尖刺,以至于它能非常可靠地做某些事情,但不是超级可靠。我认为看到这如何发展会非常有趣,因为他们在这上面花了很多钱,而我的感觉是它还不是超级通用。我认为从我所见,我们通往通用性的路径会容易得多。但迄今为止最容易实现通用性的是 RLHF。RLHF 极其容易使其通用,因为它非常擅长写人类喜欢的东西,因为那是主观的。所以,我认为这是你得到你所优化的东西,但你需要意识到现实世界。现实世界是这些互联网的压缩版本在呈现智能方面有多好。我认为它们非常擅长写文本取悦人。它们相当擅长系统 1 任务,但要让它们擅长系统 2 任务则相当艰难。因此我们看到的极端锯齿状表现。
I mean, if you talk about just chat GPT, I would imagine that there actually was generalization gains, but chat GPT already was quite general and it would did a lot of simple stuff quite well, I would argue. And uh I do think that there's some amount of generality native to things. I actually have a separate blog post I've called I believe it's called lies, damned lies and benchmarks um where I talk about how a lot of the fragility that and jaggedness we see of LMS today come from RLVR because if you are being non um generous to RLVR, it is kind of by definition optimizing for benchmarks like by definition a verifiable reward is a benchmarkable And by adding latent reasoning before the answers, you give the models all the power in the world to do really well in those benchmarks. So it's very unsurprising. It does very well in benchmarks. And I think a lot of the jaggedness has come from this. And and then you get to this whole like uh like challenge of like adding more and more and more and more and more data in distribution to get more and more diminishing returns to make the jagged frontier even more fractal and spiky such that it can do some things really reliably but not super duper. And I I I think it's going to be really interesting to see how this plays out because they're spending like a lot of money on this and it seems like it's not yet super general is what I would get the sense of. And I think that our path to generality is going to be a lot easier from what I've seen. But by far the easiest for generality was RLHF. RLHF was ridiculously easy to make it general because it was very very good at writing humanping stuff because that is subjective. So, um I think that this is one of those you get what you optimize for, but you need to be aware of the real world. And the real world is how good are like these compressed versions of the internet at surfacing intelligence. And I think that they're incredibly good at writing text to please people. They are quite good at system one tasks, but they are it's it's it's quite a battle to get them good at the system two tasks. Hence the extreme jaggedness that we're seeing right now.
不过,看它往哪个方向发展会非常有意思。我非常期待看到模型不断进化。我不确定这能否回答为什么一个开发模型——嗯,你同意在商业决策这个领域,泛化比在文本领域更难吗?还是你觉得这个前提本身就不对?
It'll be really interesting to see where it goes though. I'm very excited to see the models evolve. I'm not sure that answers why a dev model like — well, do you agree that in this domain of business decision-making, generalization is harder in this domain than in text? Or do you think that premise is incorrect?
我不同意。或者要看你怎么定义文本。如果你问的是,遵循一个简单的商业工作流更难,还是解决一个千禧年数学难题更难?我敢下很大的赌注,数学问题要难得多。别人可能有不同意见、下相反的赌注,但那太违背常识了,我没法说别的。
I would disagree. Or depending on what you mean by text. If you mean, is it harder to follow a simple business workflow or solve a millennium prize problem in math? I would bet quite a lot that the math problem is a lot harder. People might disagree and bet otherwise, but that violates too much common sense for me to say anything else.
好,那我们把这层剥开。问题的一部分是——或者说我的出发点是想做模型与模型的比较,而不是拿 harness 和 Jev 比。这里你可以纠正我。我把 Jev 看作更接近模型,而不是更接近 harness。
Okay. So let's peel this back. Part of the question is — or part of where I'm coming from is, I'm trying to compare model to model, not harness to Jev. And now here's where you can correct me. I'm thinking of Jev as more akin to a model and less akin to a harness.
我明白这个区分。其实我想说的是,想象有一张宏大的图——一张宏大的曲线图。天哪。一张关于泛化程度有多高的宏大曲线图。我想我讲的是不同目标下曲线的斜率。我的主张是,抛开模型今天所处的位置不谈,我的主张是 RLHF 是最容易的,其次是 RLCD,然后是 RLVR。所以我是在对任务的一般空间做判断。就模型而言,Jev 还处在它曲线的极早期,对吧?它字面上就是我们的第一个公开发布版本。应该把它当作早期的 ChatGPT 来看待。我猜在整体鲁棒性方面,RLHF 模型在决策上非常差。所以我猜我们的模型在决策上会不如最大的 RLVR 模型,但会比非推理的 RLHF 模型好很多。这大概就是我认为我们今天所处的位置。随着我们真正扩大投入,我预期我们会大幅超越一切,至少对商业类工作流而言。但再说一次,这是第一个模型,早期研究预览版。所以,如果模型里有蠢的地方,我们很乐意听到并修复它。我们还有很多事要做,我相信把每一个九的可靠性加进模型会是一生的旅程,而每一个九都会解锁各种疯狂的用例。所以请多理解。而且我觉得它已经非常有用、非常聪明了,我也觉得它还能变得好得多,它并没有像其他模型那样处在边际收益递减的极端平台期。
I see the distinction. Actually, what I think I was talking about was, imagine that there's a grand graph — a grand plot. God dang. A grand plot of how much generality there is. I believe I was talking about the slope of the curve for different objectives. And my claim is that, absent where models are today, my claim is that RLHF is by far the easiest, then RLCDs, then RLVRs. So I'm making a claim about the general space of the task. As far as models go, Jev is extremely early down its curve, right? It is literally our first public release. It should be treated like an early ChatGPT. And I would guess that in terms of overall robustness, the RLHF models are very bad at decision-making. So I would guess that our models would be less good at decision-making than the biggest RLVR models, but a lot better than the non-reasoning RLHF models. And that's kind of where I would guess we are at today. And as we really scale up our efforts, I would expect us to greatly overtake everything, at least for businessy workflows. But again, first model, early research preview. And yeah, please, if there's dumbness in the model, we'd love to hear about it and fix it. There's a lot for us to keep on doing, and I believe it's going to be a lifelong journey to add every single nine of reliability into the model, and every single nine will unlock all sorts of crazy use cases. So be understanding. And also I think it's already hella useful and smart, and also I think it can get a lot better too, and it is not in the extreme plateau of diminishing returns like other models are.
我最后那段话又引出了另一个想法,就是模型与 harness 之间的关系。我觉得在 LLM 的情况下,LLM 本身是有用的,但后来我们给 LLM 加上了 harness,它们就变得有用得多。你可以告诉我你同意还是不同意。但就 Jev 而言,Jev 主要是模型、只有一点 harness,很多创新可以通过扩展那个 harness 来实现吗?还是它主要是 harness、只有一点模型?有没有类似的——用同样的模型与 harness 关系来思考,这说得通吗?在 Jev 的情况下,harness 会做什么?
So that last comment of mine prompted another thought, which is this relationship between model and harness. And I think in the case of LLMs, LLMs were useful, but then we added harnesses to LLMs and they became a lot more useful. You can tell me whether you agree or disagree. But with regards to Jev, is Jev mostly model, a little bit of harness, and a lot of innovation can come by expanding that harness? Or is it mostly harness, a little bit of model? Is there a similar — does it even make sense to think about that same relationship between model and harness? And what would the harness do in the case of Jev?
好。我其实要快速把它换成我的术语。那个 harness 就是代码。所以我会质疑说,仅仅在 LLM 上加一个 while 循环和工具调用就让它们变好了,因为我不想抹杀 ML 的人为此付出的所有努力。我当时不在那些实验室里,但我相当确定,2025 年 11 月前后 Claude Code 发生的巨大跃升——Claude Code 在那之前几个月就发布了,当时并没有那么流行,然后后来出现了一次巨大跃升。我猜这就是他们最终把它投入分发时发生的事,因为你优化什么就会得到什么,我认为那才是让 Claude Code 变得非常好的原因。
Perfect. I will actually rename into my terminology real quick. That harness is code. So I would debate that simply adding a while loop on the LLMs and tool calls is what made them good, because I don't actually want to discount all the efforts that ML people put into that. I was not inside the labs at the time, but I'm fairly certain the gigantic jump that happened with Claude Code around November 2025 — Claude Code was released months beforehand and it was not that popular, and then there was a giant jump later. My guess is this is what happened when they finally put it in distribution, because you get what you optimize for, and I think that is what made Claude Code very good.
投入——你说的是把它投入分发,是这个意思吗?
Being put — what you said, put it into distribution, you mean?
编码智能体的分发。字面上,是的。Claude Code,或者像 Claude Code 这样的东西。显然我拿不到他们的数据,但我敢打赌,中间出现了一次巨大的数据跃升,就是他们决定下大赌注把它投入分发的时候。所以,我相信这就是智能体开始能工作的原因。
The coding agent distribution. Literally, yes. Claude Code, or things like Claude Code. I don't have access to their data, obviously, but I would wager that there was a gigantic data jump in between, where they decided to make a big bet to put that in distribution. So, and I believe that is why agents started working.
你是说——你知道,大体上有个说法是 Claude 很棒,但 Claude Code 是那个使用 Claude 模型的 harness,让它对人们有用得多。而我听到你说的是,也许类似这样的情况是巧合,harness 并不是真正的阶跃函数,而是某种模型创新让 harness 变得有用——一种数据创新。
Are you saying that — I think you know, broadly, there's this narrative that Claude was great, but Claude Code was this harness that used the Claude models and made them a lot more useful to people. And what I hear you saying is that maybe something like that was coincidental and the harness wasn't really the step function, but there was some model innovation that allowed the harness to be useful — a data innovation.
对。这符合我对时间线的理解。但再说一次,我是从外部来谈这件事的。Claude Code 发布时,从技术上说,Claude Code 的形态非常强大,就像分类器的形态非常强大一样,但它的强大程度只取决于背后的智能。就我对那个时期的记忆而言,Claude Code 当时像是一个非常小众、奇怪的东西,人们用得不多。人们最终会感受到不可靠带来的痛苦,然后停止使用它。OpenClaw 也发生了类似的事。OpenClaw 显然有一个非常强大的形态。人们现在还在朝那个方向继续创新。问题不在于形态或 harness。问题在于背后的智能。也许合适的 harness 能更好地驾驭智能。哈哈。但在那种情况下,把责任归咎于 OpenClaw 的 harness 就不太厚道了,因为智能本身也会进化,你做这些东西时必须对现实有清醒认识。所以如果这些——我不知道这东西的通用叫法是什么,比如 Claude 助手之类的,不管大家现在在做什么——如果这些开始变得好得多,我不会感到惊讶,因为它们被投入了分发。这不是在抨击 OpenClaw,对吧?OpenClaw 做了一次非常棒的界面跃升,就像 Claude Code 的界面跃升非常酷一样,即使他们没让模型很擅长它。所以我想说的是,我不认为纯粹是 harness 赋予了它这些能力。所以 harness 需要映射——或者说我们称之为 harness 的东西需要映射模型中的能力,才能有用。我其实不会把它本身叫作 harness。我会把它叫作代码,我也会把包裹我们模型的东西叫作代码。我们的模型比 LLM 低层得多。
Yep. That is what fits my timeline of it. But again, I'm talking about it from the outside. When Claude Code was released, technically the shape of Claude Code is very powerful, just like the shape of classifiers is very powerful, but it's only as powerful as the intelligence behind it. And as far as I can recall of that time period, Claude Code was like a very niche weird thing that people didn't use very much. And people eventually feel the pain of unreliability and stop using it. A similar thing happened with OpenClaw. OpenClaw obviously has a very powerful shape. People are continuing to innovate in that direction right now. And the problem was not the shape or the harness. The problem was the intelligence behind it. Maybe the right harness can harness the intelligence better. Haha. But in that case, it would be remiss to blame the OpenClaw harness, because the intelligence can evolve too, and you need to be aware of reality when making these things. So it wouldn't surprise me if these — I don't know what the general term for this is, like Claude assistance or whatever everyone is doing right now — it wouldn't surprise me if these start getting way better because they are put in distribution. And it is not a stab at OpenClaw, right? OpenClaw did a really awesome interface jump, just like Claude Code's interface jump was very cool, even if they didn't make the model very good at it. So I want to say that I don't think it's purely the harness that gave it its capabilities. So the harness needs to map — or the thing we're calling harness needs to map the capabilities in the model in order to be useful. I would actually not call it harness per se. I would call it code, and I would call the thing wrapping our model code as well. Our model is much lower level than an LLM is.
那这个东西,它是一个类型安全的东西,还是说你指的是用户那端访问模型的代码?
And that thing, is that a type-safe thing, or are you talking about the user's code that has access to the model?
用户的代码。
User's code.
好。
Okay.
是的,其实如果你的系统消息里带着你整套业务算法,却让逻辑跑在别人的服务器上,这很奇怪。对 API 来说这很不寻常。通常你会把只有 API 才能做的工作外包给 API,而不会把自己的商业机密之类的东西全交出去。所以这更像是一个普通的 API。我喜欢把它想成:有一天智能会像数据库一样,你需要的时候调用它就行,当你的代码里需要智能的时候。所以我认为类型安全需要代码。我觉得拿模型玩一玩挺有意思,但盯着概率看没什么用。概率只有在进入某个应用、最终让用户满意或产生价值时才有意义。所以没有代码就用不了类型安全。如果你不是程序员,你对我们的模型感兴趣,这很酷。我有点困惑,我很想问问你为什么感兴趣。而且它是为编程而生的。
Yeah, it's actually weird for the logic to be running on someone else's server if your system message has your whole business algorithm on it. That's very unusual for an API. Normally, you outsource to an API the work that only the API can do, and you'd rather not give it all of your business secrets or anything like that. So this is just more like a normal API. I like to think of it as: one day intelligence will be like databases, where you just call it when you need it, when you need intelligence within your code. So I think type-safe requires code. I find it a little bit fun to play with the models, but it's not useful to look at probabilities. The probabilities are only meaningful if they go into an application that then does something that ideally eventually delights a user or provides value. So type-safe cannot be used without code. If you are a non-coder, it's really cool that you're interested in our model. I'm kind of confused. I would love to ask you why you're interested. And also, it's made for coding.
哦,抱歉。它不是用来写代码的,它是用来增强代码本身的能力的。
Uh oh, sorry. It's not made for writing code. It's made for enhancing the power of code itself.
正是如此。这和 LLM 的定位非常不同,因为 LLM 主要用于写代码,而代码的表达能力在过去十年左右基本没变。
Exactly. Which is very different from where LLMs are, because they are mostly used for writing the code, which has the same expressive power as it's had for the last decade or so.
那我总结一下,我听到的是:Jev 本质上是一个模型,你现在并不认为在你和用户之间需要一个由你控制的所谓“外壳”——就是那个 while 循环之类的东西,或者去抓取外部上下文——而是每个用户会把 Jev 模型嵌进自己的代码里,本质上由他们自己提供适合自己的外壳。
So just to pull that back, what I'm hearing is that Jev is at its core a model, and you don't right now see a role for a quote-unquote harness between you and the user that you control — that's like the while-loop thingy, or that's fetching external context — but rather every user is going to embed that Jev model into their code and they're essentially providing the harness that makes sense for them.
对。其实我完全拒绝“外壳”这个概念。我认为对 Jev 这类模型来说,更高层的抽象是可以存在的,但外壳本身只是一种很“无马马车”式的东西,试图把智能变成看起来像人的东西。如果它管用,那当然很酷。毫无疑问——如果它管用,我会做一大堆。但据我所知,它现在还不管用。所以我不用。而且对于我认为需要完成的大量自动化来说,它根本不是必需的。我们不需要到处都用同一个外壳。我们需要的是让用户能把他们的逻辑放进去,这正是软件工程师做的事。他们极其字面地规定自动化的逻辑来创造价值,而我们想扩展他们能放进去的逻辑词汇。我更愿意把它想成 N8N 节点,或者任何所见即所得式的节点编辑器。我们只是其中一个节点,也许你会大量使用它。正是如此。但其余的逻�辑都是你的,数据也全是你的,对吧?这才是让东西产生价值的方式。
Yep. And actually I just reject the notion of a harness entirely. I think higher-level abstractions could exist for Jev-like models, but a harness itself is just a very horseless-carriage type thing of trying to turn the intelligence into something that looks like a human. Which is cool if it works. Have no doubt — if it worked, I would have tons of them. But as far as I can tell, it doesn't yet. So I don't. And that's just not necessary for the massive amount of automation that I think needs to get done. We don't need the same harness everywhere. What we need is for the users to be able to put their logic in there, which is what software engineers do. They hyper-literally specify the logic of automation in order to create value, and we want to expand the vocabulary of logic that they can put in. I'd like to think of it more like N8N nodes, or any kind of WYSIWYG-style node editors. We are just one of the nodes that maybe you put a lot of the place. Exactly. But the rest of the logic is all you, and the data is all you too, right? That is how you make stuff do valuable things.
好。你几次提到 Open Claw。对 Jev 来说,一个有趣的潜在机会是在所有这些智能体外壳、框架,或者随便我们怎么称呼它们的东西里面,因为这些东西要做大量编排决策,而目前这些决策是由基于文本的模型驱动的,也许有更好的做法。你怎么看 Jev 加智能体?
All right. So you've brought up Open Claw a few times. One interesting potential opportunity for Jev is inside all of these agentic harnesses or frameworks or whatever we want to call them, because those things make a lot of orchestration decisions that right now are being driven by text-based models, and maybe there's a better way to do that. What's your take on Jev plus agents?
哦,太多了——你在这里面埋的线索多得惊人。那我分几个层面来回答。第一,我们当然想让它跑通。但第二,我们绝对不是这方面的专家,而看起来这方面的专家们正在摸索,试图找出杀手级用例之类的。所以我们真的很想帮那些人做出成绩、创造大量价值,最好让智能体更可靠或更高效,或者最好两者兼得。那就理想了。快也很酷,但感觉我更偏好可靠性。而很多企业会更偏好成本。所以综合来看,Jev 就是要成为这样一种通用技术:懂常识,能在各处帮上忙。
Ooh, there's so many — that is a surprisingly high number of breadcrumbs you've added in there. So I'll answer it on a couple of levels. Number one, we obviously want to make it work. But number two, we are definitely not experts at this, and it seems like the experts at this are playing around and trying to figure out the killer use cases and everything. So we want to really help those people crush it and create lots of value, and ideally make agents a lot more either reliable or efficient, or ideally both. That would be ideal. Fast would be cool too, but it feels like I would prefer reliability. And a lot of businesses would prefer cost. So given all of that, Jev is meant to be this general technology that understands common sense and can be helpful all over the place.
我确实分享过一篇东西——我其实有一篇相关的博客文章,叫《KV Cache Rules Everything Around Me》,就是 C-A-C-E。不知怎么的,网上居然没人之前用 CAC 说过“cash rules everything around me”。我简直不敢相信我是第一个这么说的。所以第一,我有一篇博客文章讲为什么智能体极度受制于它们的 KV 缓存,以及为什么智能体的很多设计明显来自一个约束:知道 KV 缓存极其昂贵又极其敏感——你不想污染它。所以像为什么人们不更多使用子智能体、为什么路由往往不奏效——这些我都在博客文章里讲了。
I did actually share a write-up — I actually have a blog post related to this called “KV Cache Rules Everything Around Me,” like C-A-C-E. Somehow no one on the internet has said “cash rules everything around me” beforehand with a CAC. And I can't believe I was the first one to say this. So number one, I have a blog post about why agents are extremely beholden to their KV caches, and why a lot of the design of an agent very obviously comes from the constraint of knowing that the KV cache is extremely expensive and very sensitive — like you don't want to pollute it. So things like why don't people use more sub-agents, why does routing tend to not work — all of this I go into in my blog post.
我还发布了一份 Google 文档,我有点后悔,因为我老是收到分享请求,有人想编辑我的 Google 文档之类的。但我当时有一个内部设计,关于 Jev 如何用于编程智能体,因为曾经我们以为我们会更低调一些,以为我们会自己做产品,而现在这个秘密已经完全泄露了,因为我们决定不那样做。但没错,我觉得如果你假设不需要只追加的 KV 缓存,能发生很多很酷的事。KV 缓存对省钱很棒,对吧?但也非常受限。它阻止了很多很酷的编程模式,比如扇出、过滤、重排序之类的。我很喜欢这个面试题:如果 KV 缓存不存在,你会怎么设计一个编程智能体?它真的会让你重新思考智能体可以怎么运作。而且我认为,假如你能随时获得极其便宜又可靠的智能,这是一个非常有用的思考方向,用来弄清楚如何设计智能体。我觉得能做的很酷的事情太多了。
I've also released a Google doc, which I kind of regret because I keep getting share requests from people who want to edit my Google doc or something. But I had this internal design of how Jev could be used for coding agents, because once upon a time we thought we would be a bit more stealthy and we thought we would be building products ourselves, and that cat is now totally out of the bag because we decided not to do that. But yeah, I think there are so many cool things that can happen if you just assume that you don't need an append-only KV cache. A KV cache is great for saving money, right? But also very constraining. It prevents a lot of cool programming patterns like fanning out and filtering and reordering and stuff like that. And I love the interview question: how would you design a coding agent if KV caching didn't exist? It really makes you rethink how agents could work. And I think it's a really useful intellectual direction for figuring out how to design agents if, for example, you had extremely cheap but reliable intelligence on the fly. And I think there are so many cool things that can be done.
KV 缓存有点像到处都有全局变量,但如果你能聪明地管理状态呢?为什么你需要对过去做过的每一个决策都有完整视图,还要用于所有未来的决策?因为大概当你处于某个时刻时,有些额外信息你可能想插入、取出决策、把它从栈上弹出,真正精心策划这个 KV 缓存,甚至做一些超酷的事情,比如在上下文里做对数级查找。没有什么阻止你在上下文里做分层搜索。这样你就不需要一遍又一遍地重复同样的东西。
The KV cache is a little bit similar to having global variables everywhere, but what if you could be smart about your state? Why do you need the full view of every past decision you've ever made for all the future decisions too? Because presumably when you're in a moment, there's extra information you might want that you could insert, get the decisions out, pop it off the stack, and really curate this KV cache, or even do super sick things like log-of-N lookups within your context. Nothing stops you from doing hierarchical searches within your context. And that way you don't need to have the same thing over and over and over again.
另一个巨大的用例其实是工具调用。工具调用第一会导致上下文腐化,因为它们前期需要大量上下文。第二,你并不总是需要它们。第三,因为模型已经对原生 harness 过拟合了,它们调用原生工具的频率远高于你插入的工具。所以模型最终并不怎么调用它们。如果你能给工具调用提供 logit 偏置,从而细粒度地控制什么时候插入、什么时候不插入,而不污染你的上下文,那会怎样?所以有太多酷炫的事情可以做。我真希望我能去做。啊,那会太酷了。那本来会太酷了。
Another gigantic use case is actually tool calls. Tool calls number one result in context rot because they require tons of context up front. Number two, you don't always need them. And number three, because the models have been overfit to their native harnesses, they tend to call the native tools a lot more than the ones you plug in. So the models don't end up calling them that much. What if you could provide logit biases to your tool calls to fine-grainly control when things are inserted or not without poisoning your context? So there's so many cool things that could be done. I wish I could work on it. Ah, it would be so cool. It would have been so cool.
所以在一个 Jev 式的全新世界里,智能体架构可能会非常不同,而你把它留作给听众的练习。
So the agent architecture is potentially very different in a kind of Jev Greenfield world, and you're leaving it as an exercise to the listener.
嗯,我觉得那里有一个有趣的智识方向值得探索。务实地说,我们需要关心什么在现实世界中真正有效。而且 Jev 甚至可能有局限性——对,现在还是 Jev 1.13。也许你需要 Jev 1.16 才能真正让这一切运转起来。谁知道呢?那将是一个完全开放、非常有趣的世界。我的猜测是,未来会是混合体,慢慢加入越来越多的逻辑,因为往 Claude Code 和 Codex 这类东西里加入额外的硬逻辑非常困难,因为它们的核心循环只是一个 while 循环,而且一般来说加逻辑很昂贵。所以需要一段时间才能真正调好,但一旦你有了能工作的软件,它工作起来真的很美好。所以我的猜测是,人们会构建越来越多、越来越多的软件,挂接到这些 harness 上,直到它们转变为更多软件、更少魔法。
Well, I think there's an interesting intellectual direction to look there. Pragmatically, we need to care about what actually works in the real world. And there might even be limitations of Jev such that — right, this is currently Jev 1.13. Maybe you need Jev 1.16 to actually get all of this working. Who knows? It's going to be a wide open, really fun world. My guess is the future is going to be hybrids that slowly add more and more logic, because it's very hard to add additional hardness logic into things like Claude Code and Codex, because their core loop is just a while loop, and in general adding logic is expensive. So it takes a while to really tune, but once you have software that works, it's really lovely when it works. So my guess is people will be building up more and more and more software that latches into these harnesses until they switch over into being a lot more software and a lot less magic whoop.
很早的时候我提到过我想回到一个特定的点,基本上是关于当你考虑围绕决策模型来架构时,或者不管我们想怎么称呼 Jev 所属的这个更广泛的类别,那会需要更传统的工作流导向,而不是我们目前到达的地方——有了智能体和智能体式软件,就是,好吧,让我们生成一些上下文,交给 LLM,让它吐回一些答案,调用一些下一个工具,这些本质上就是下一步。我想探讨的是——在某种程度上感觉像是,如果你先把我们要迈向的那个世界其实并不那么可靠地工作这个事实放到一边,
Very early on I mentioned I wanted to come back to a particular point, fundamentally around when you think about architecting around decision models, or whatever we want to call the broader category that Jev sits in, that would require more traditional kind of workflow orientation versus where we've kind of arrived, that with agents and agentic software is like, okay, well, let's generate some context and give it to an LLM and let it spit back some answers and call some next tools, which are essentially the next steps. And I guess I want to poke at — in some ways it feels like, if you set aside the fact that this world that we move to doesn't really work all that well reliably,
这还挺重要的。这还挺——这还挺重要的。是的。
Which is kind of important. It's kind of — it's kind of important. Yes.
但让我们先把它放到一边。能够把你所有的决策都交给这一个东西,让它在一个循环里运行,让它调用下一步动作,然后回到这个循环,这有一种敏捷性。而把所有东西分解成这些更加静态或僵化的工作流,会随着时间产生大量债务,产生大量复杂性。Jev 是否要求我们回到那里?
But like let's set that aside. There's a certain agility to being able to hand off all your decisions to this one thing and let it run in a loop and let it call next actions and return to this loop. Whereas decomposing everything into these much more static or rigid workflows creates a lot of debt over time, creates a lot of complexity. And does Jev require us to move back there?
我确实会思考这个。我算是把这个问题设好了。
I do think about that. I've kind of set this up.
好吧。是的。你设得非常不公平,因为你把可靠性放到一边了。
Okay. Yeah. You set it up very unfairly because you put aside reliability.
我承认,我承认。但我对此持开放态度,而且实际上我觉得这是个有趣的假设性问题,因为我从来没有不思考可靠性。但我确实认为,对许多实际用途来说,这其实是一个特性。显然存在长尾,会有很多不值得自动化的用途。这就是作为软件工程师做工作的悲剧性问题。我是为这个写一个要花我四小时的 bash 脚本,还是五分钟就做完?而我是那种多半会写 bash 脚本的人。但这里有个两难,对吧?所以我猜对很多工作来说,直接交给不可靠的语言模型然后希望它能工作,大概是值得的,对吧?然后如果不行,你检查一下它的工作就行。但如果你真的想规模化某个东西,比如如果这是一个非常有价值的工作流,你会想让大量人来做,或者你想在后台大量运行它,或者——而这是我认为在 LLM 领域没人真正理解的部分——如果你想把它封闭在盒子里、在后台运行,让它成为一个你可以永远调用的依赖,那这种情况下,前期真正工程化那个系统就是有意义的,而且随着时间维护它也是有意义的,因为那会永远提供价值,而不是每次都要重新发现逻辑。所以我认为,人们就是不把语言模型看作可组合、可分层、可以作为越来越大的认知抽象的基础的东西,因为语言模型在这方面差得难以置信。而在轨道上是一个特性。你知道,你不会想去一个没有轨道的游乐园。或者至少没有轨道、安全带和所有这些,你就做不了最有趣的东西,对吧?所以我认为,对 AI 软件最有价值的部分来说,会是非常类似的情况。
I recognize that, I recognize that. But I'm open to that, and actually I think this is a fun hypothetical question, because I never not think about reliability. But I actually think that it's actually a feature to many actual uses. Like obviously there's a long tail, and there's going to be a lot of uses that are not worth automating. This is the tragic problem of just doing work as a software engineer. Do I write a bash script for this that will take me four hours, or do I just do it in five minutes? And I'm the kind of guy who would probably write the bash script more often than not. But there's a dilemma here, right? So I would guess for a lot of work, it's probably worth it to just give it to the unreliable LM and hope that it works, right? And then you can just check its work if it's not. But if you're actually trying to scale something, like if it's a workflow that is so valuable that you'd want tons of people to do this, or you'd want to run it a lot in the background, or — and this is the part that I think no one really gets in LLM land — if you want it to be closed away in the box and run in the background and have it be a dependency that you just call forever, that's the kind of thing where it makes sense to really engineer that system up front, and it actually makes sense to maintain it over time, because that provides value forever, versus having to rediscover the logic every single time. So I think that people just don't think of LMs as composable and layerable and something that can be the foundation of greater and greater cognitive abstractions, because LMs are so unbelievably bad at this. And it is a feature to be on the rails. You know, like you don't want to go on an amusement park that doesn't have its rails. Or at least you can't do the most fun stuff without the rails and the seat belts and all of that, right? So I think it's going to be a very similar thing for the most valuable parts of AI software.
你是说,用这个有点不可靠的东西来自动化,有点像一片绿洲,你没法完全到达那里。而我们迫使你走向 Jev,就像 Jev 会迫使你走向更传统的代码和工作流之类的,这未必——是一个特性。我想这就是你的定位——
You're saying that automating with this unreliable-ish thing is kind of an oasis and you're not going to fully get there. And the fact that we force you to Jev, like Jev would force you to more traditional code and workflows and all that, isn't necessarily — is a feature. I think that's the positioning that you're —
对自动化来说是一个特性。对即时性的东西来说不是特性。所以如果你在做一次性的临时事情,那当然不值得为它写一个庞大的自动化。
It's a feature for automation. It is not a feature for just-in-time stuff. So if you're doing stuff on the fly that's a one-off, of course it's not worth it to write a big old automation for that.
我们聊了很多关于接下来会发生什么的大方向,但如果要你直接回答“接下来是什么”,你首先想到的是什么?
We've talked a lot about kind of broadly what's next, but if you had to directly answer what's next, what comes to mind?
不吐槽别人就很难说。所以我——
It's hard for me to say without shit-talking other people. So I—
那我们就开始吐槽吧。
We'll start shit-talking then.
天哪。我们的顾问告诉我,我在网上太友善了。我想我还是做真实的自己吧。我对 AI 错误信息当然有怨恨,但我其实不想对任何特定的发布或产品心怀怨恨。所以我真的不想往那个方向走。我们有很多东西想要发布。天哪,我太兴奋了,想告诉人们,又不想告诉人们,想向人们展示为什么我们称它们为系统一模型。那会很酷。我们还有很多储备,比如,嘿,用多个模型发布将会非常复杂。我们先保持简单。所以我也很期待。我很期待能有某种早期访问计划。我想让人们能够试用我们早期的模型,那些不一定有我们可靠性认证的模型,但我们认为它们比市面上任何东西都好。天哪,我只是对技术上的东西感到兴奋。我还有很多公司建设要做,我真的很不喜欢。所以,如果有人想为别人工作,我很想雇佣无数的人,希望这能让我的生活轻松一点。
Oh man. Our advisers have told me that I've been much too nice online. I think that I'd rather just genuinely be me. And I have spite obviously against AI misinformation, but I actually don't want to be spiteful at any particular release or product. So I don't actually want to go down that direction really. We have a lot of stuff that we want to be shipping. And oh man, I'm so excited to tell people and not to tell people, to show people why we call them system one models. That's going to be cool. We have a lot of stuff in the tank as well, like, hey, it's going to be really complicated to launch with multiple models. Let's keep it simple for now. So I'm so excited for that too. I'm excited to have an early access program of some sort. I want people to be able to play with the earlier models we have that don't necessarily have our stamp of reliability but we think are better than anything else out there. Man, I'm just excited for stuff technologically. I also have a lot of company building to do, which I really don't like. So if anyone wants to be working for someone, I would love to hire a jillion people and hopefully that'll make my life a little bit easier.
嗯,瑜伽,能叙旧真是太好了。早就该这样了。这些年来你一直在做一些有趣的事情,但是——
Well, the yoga, it's been wonderful catching up. Way overdue. You've been doing some interesting things over the years, but—
我们应该多聚聚。
We should do this more.
我们绝对应该多聚聚。对 Typesafe 来说是个好时机。恭喜最近的成功和与 Jeff 的合作。
We should definitely do this more. Great moment in time for Typesafe. Congrats on the recent success and traction with Jeff.
谢谢。我相信这只是众多成功中的第一个。所以你就等着吧。你可能比你想的更早会再请我来这里。
Thank you. And I do believe that this is just the first of many. So just you wait. You might have me on here sooner than you think.
好的。我们保持联系。谢谢。
All righty. We'll keep in touch. Thank you.
当然。祝你愉快。
Heck yeah. Have a good one.