How Jev Turns AI Into Software That Gets Things Done
打开互动全文版(中英对照 + 朗读 + 问答)→TypeSafe 创始人 Diogo 讲述 Jev 如何把 AI 变成真正能自动完成工作的智能软件,而不只是更快地写代码。
Diogo, founder of TypeSafe, explains how Jev turns AI into smart software that actually automates work, instead of just writing code faster.
自动化到底在哪里?AI 聪明得难以置信,可它在其他所有事情上却这么没用。
Where is all the automation? AI is so unbelievably smart and yet it's so useless at all other stuff.
不管你用多少 AI 编程智能体,软件本身其实并没有变得更好。也许你写得更快了,但可以说它反而更糟了。
It doesn't matter how many AI coding agents you use, the software actually isn't getting better. Maybe you're writing it faster. It's arguably getting worse.
OpenAI 从 2020 年起就一直在尝试自动化客服。我想要的其实是智能软件。我想扩展软件本身能做的事,让那些本该可自动化的事情真正变得可自动化。
OpenAI has been trying to automate customer service since 2020. What I want instead is smart software. I want to expand what software itself can do such that things that should be automatable can then be automatable.
你们说得最让我喜欢的一句话是:我们造的是产品,不是神。太好了。
My favorite thing that you guys say is we build prod not god. So good.
因为如果是别的大实验室的负责人,哪怕他们也有乐趣,他们也会什么都做。
Because if we had any other kind of big lab leader, even if they had joy, they would cover everything.
而你的观点就完全不同。你说的是:不,我们要创造一个好得多的世界。
And then your view is so different. You're like no we're going to create a way better world.
出于一些微妙的原因,我不认为我们正走在 RSI 那条 SaaS 末日故事的路径上。
For nuanced reasons, I don't think we are on the path of RSI in the SAS apocalypse story.
今天我们请到了 TypeSafe 的创始人兼负责人 Diogo,他对 Martine 和我来说都算是个英雄。是的。
Today we have the founder and leader of TypeSafe, Diogo, with us, who is a bit of a hero to both Martine and me. Yes.
他不仅在打造一个非常有趣的产品,还在开创我们认为非常重要的一场运动。所以我们对今天非常期待。欢迎,感谢你的到来。
He is not only building a really interesting product but creating what we think is a very important movement. So we're super excited about today. Welcome. Thank you for coming.
谢谢。
Thank you.
也许你可以给我们简单介绍一下,什么是 Jev?什么是 TypeSafe?它为什么重要?
Maybe you can give us a brief on just, you know, what is Jev? What is TypeSafe? Why is it important?
这个适合诅咒吗,还是不行?
Is this a curse-friendly or no?
好吧。你说的是?好吧,好吧,行。
Okay. Are you talking about? Okay. Okay. Cool.
其实有人让我准备一段电梯陈述,我这个人容易东拉西扯,讲不好,但我发现我最喜欢的 Jev 电梯陈述就是:自动化到底在哪里?这实在太悲剧了,你知道,这么多智能。AI 聪明得难以置信,却又如此不聪明。我不是在贬低聊天机器人或编程智能体,我自己也很喜欢它们,但它在其他所有事情上都这么没用,这很悲剧。
So I was actually asked for an elevator pitch, which I tend to ramble on and I don't do well, but I realized my favorite elevator pitch for Jev is: where is all the automation? This is so unbelievably tragic, you know, so much intelligence. AI is so unbelievably smart and yet so not. I don't hate on chatbots or coding agents. I love them myself, but it's so useless at all other stuff and it's tragic.
悲剧的是,我们有这么多璞玉,却没有打磨到能干活。但 TypeSafe 是在为软件打造 AI。我们想让 AI 不只是服务于人在回路中,而是真正能构建真正的软件。而 Jev 对我们来说,是我们在整个领域里的第一个模型,让它好得多,去实现自动化。
It's tragic that we have so much diamond in the rough but not polished for work. But TypeSafe is making AI for software. We want to make AI powerful not just for humans in the loop, but to actually build real software. And Jev to us is our first model in this whole space to make it way, way better, to make automation.
这很有意思,因为它在软件圈里火了起来。让我们惊呼“这到底怎么回事”的一件事是,我们认识的每个开发者都打电话来说:哦,这太牛了。太棒了,很快,一切都更好了。
And so it's been interesting because it's kind of caught fire in the software world. One of the things that made us go, what the hell's going on here, is every developer we know is calling us and going, oh, this is freaking awesome. It's great. It's fast. Everything's better.
那这又怎么……因为大家都会想,我们有 Claude Code,有 Codex,不是已经有这些了吗?区别在哪?然后这又如何通向真正的自动化?
And then how does that, because everybody thinks, well, we've got Claude Code, you know, we've got Codex, don't we already have that? Like what's the difference? And then how does that lead to real automation?
我真希望我有些潦草的图示,因为我有一个最喜欢的潦草图示。我喜欢 Claude Code 和 Codex。我很喜欢 Gary Tan 对它们的描述:即时软件。这个描述太妙了。它即时生成软件,你可以用自然语言编程软件,但它的表达力和软件一样。
I wish I had some sloppy visuals because I have a favorite sloppy visual for this. So I like Claude Code and Codex. I love the description from Gary Tan on them. It's just-in-time software, you know, incredible way to describe what they're doing. It makes software on the fly and you can program software in natural language, but it has the same expressive power as software.
我想要的其实是智能软件。与其自动化软件工程,我想扩展软件本身能做的事,让那些本该可自动化的事情变得可自动化。用更华丽的话说,我想表达像“意图”这样的东西。我想扩展我们能做的事的词汇表。我可以聊各种我想要的古怪科幻东西,但编程就是把有价值的东西超精确地指定出来,然后无限复制。这太酷了,我就想让它更进一步。
What I want instead is smart software. Instead of automating software engineering, I want to expand what software itself can do, such that things that should be automatable can then be automatable. And in more flowery language, I want to express things like intent. I want to expand the vocabulary of what we can do. I can talk about all sorts of weird sci-fi things I want, but programming is hyper-specifying valuable things and then infinitely replicating them. It's so freaking cool and I want to just make that more.
哦,有意思。所以一种理解方式是:与其做一个多少替代软件工程师的工具,用一个更快、也许还不那么好的软件工程师,你说的是:不不不。我们要超级赋能现有的软件工程师,让他们写出好得多、有趣得多的东西。
Oh, interesting. So one way to think about it is instead of a tool that somewhat replaces a software engineer with a faster, maybe not even as good software engineer, what you're saying is no, no, no. We're going to super-empower the software engineers we have to write way, way better, more interesting things.
对。我其实,我是说,对,对。
Yeah. I actually, I mean, yeah. Yeah.
顺便说一句,我觉得太多人错过了这一点,它太微妙了,而且把它挑明非常重要,那就是:如果你用 Claude Code 或 Codex,这很棒,或者 Cursor,也很棒,它们会写代码,但那代码和人类会写的是一样的东西。也许更好,也许更差,但基本上还是 10 年前那种代码。
So this is, by the way, I just think so many people miss this point and it's such a subtle point and it's so important to actually tease it out, which is if you use something like Claude Code or Codex, which is great, or Cursor, which is great, they write code but that code is the same thing a human being would have written. Maybe it's better, maybe it's worse, but it's basically still code just like code looked 10 years ago.
对。
Yeah.
而 Jev 的关键在于,不管你是 Claude Code 还是人类,你都有了这个新的原语,这个你嵌进代码里的新东西,它真正扩展了软件的能力。所以它不是写代码,而是你把它包含进你的代码里。
And the thing with Jev is whether or not you're Claude Code or a human, you have this new primitive, this new thing that you stick in your code that actually expands the power of software. So instead of writing code, it is something that you include in your code.
这个,你继续。
Which, go ahead.
顺便说,这很有意思,因为它是个非常强大的原语,你要是解释一下会很好,但它也和程序员的思维方式有点不同。比如,它有概率这个概念,所以,是软件内部的一个智能层。
Which, by the way, is interesting because it's this very powerful primitive, which would be great if you explained, but it's also a little bit different than how programmers think. For example, it has this notion of probabilities, and so, so an intelligent layer inside the software.
对。把它想成一个库,你可以用自然语言描述你想要什么,你给它一个状态机,然后它会带着某种置信度选择做什么,这种东西我们以前其实没有这么普遍地拥有过。
Yeah. Think of like a library that you can use natural language to describe what you want and you give it kind of a state machine and then it will choose what to do with some confidence levels, which we kind of haven't really had before, so ubiquitous.
所以也许,哦,这里面有很多门道。我先跳到一个点上,就是我很喜欢你一开始说的,朝着“自动化到底在哪里”的方向。我太爱软件了。我希望我能整天写代码。不建议大家去当 CEO,不过无所谓。
So maybe, ooh, there's a lot of tricks there. I will jump into one thing first, which is I love the first thing you said, you know, in the direction of where is all the automation. I love software so much. I wish I could be writing it all day. Would not recommend being a CEO to people, but whatever.
而且 AI 这么酷,软件却 10 年没变,这太疯狂了。没人能把这个矛盾解释通,我们最多也就是在旁边加个小聊天机器人,有时能执行一些动作,但不是所有动作,因为有些动作现在还不靠谱。我就想说这么个小插曲。
And also it's wild that AI is so cool and software has been unchanged in 10 years. No one can square this together and the most we can do is add a little chatbot on the side sometimes that can take actions but not all actions because some of the actions are not reliable now. So I just want to give that tiny aside.
我很喜欢这个点。我要跳回那个点,就是这是一种稍微不同的思考方式。是的,我认为机器原生并不完全等同于比特,而这其实正是我们试图做的艺术形式。在我们第一天的入职培训上,我画了一张维恩图,标出 AI 擅长什么、代码里什么有价值,我们就在中间。所以我们不输出外推的浮点数,因为 AI 就是做不好这个。但概率这类东西并不算全新。这和“Jev 是不是只是个分类器”的论点类似。
I love the point. I'm gonna jump back to the point about this is a little bit of a different way to think about it. Yes, I think that machine-native doesn't exactly match bits perfectly and that's actually the art form that we are trying to do. In our onboarding on day one I draw the Venn diagram of what AI is good at, what is valuable in code, we are in the middle. So we don't output extrapolated floats, for example, because AI is just bad at that. But things like probabilities are not exactly novel. And it's similar to the is Jev just a classifier argument.
Jev 绝对是一个分类器。你知道,分类器很酷。分类器被设计成有用的。
Jev is absolutely a classifier. You know, like classifiers are sick. Classifiers were designed to be useful.
是的,它们被设计成有用的。实际上,它和那些机器学习概念有相同的接口,因为这些概念来自试图让系统工作的实用主义者。我现在看到的是,Jev 实际上,我猜 Jev 可能比 2019 年有一个 MLE 团队为你做东西更好,你可以直接即时编程。谁知道能构建什么,因为 2019 年没有那么多好的 ML 团队 MLE 团队来构建狭窄的东西,并能够收集数据集并测量等等。而这只是开始。我觉得有很长的路,顺便说一句,到这一点,你认为这里有一个滑块条吗,一端是像我们今天这样的语言输入语言输出,另一端是像现有的命令式程序,然后你可以在两者之间移动,或者你认为这是设计空间中的点,即语言输入,状态机输出,这将固化为程序员的通用事物。
Yeah, they're designed to be useful. And actually, it's the same interface as some of those ML concepts because these came from practical people who are trying to make systems work. And what I'm seeing is happening now is that Jev actually, my guess is that Jev probably is better than having like an MLE team from 2019 making the stuff for you and you can just program it on the fly. Who knows what could be built cuz like there were not that many good ML team MLE teams in 2019 to build like narrow things and to be able to like collect data sets and measure it and all of that. And it is just the beginning. There's I feel like there's way like like by the way to this point do you think there's a slider bar here where like on one end is like language in language out like we have today on the other end is like a like an existing imperative program and then you can kind of move between the two or do you think like this is like the point in the design space which is language in kind of state machine out which is going to like solidify as a general purpose thing for programmers.
哦,这是个棘手的问题。所以我会说,我心中的答案,是的,我心中的答案是它是一个滑块。所以,实际上,当我为我们拥有的属性设计时,我可能因为个人偏好而犯了错误,但就像每美元智能是我现在的北极星,它可能是错的,明确地说,每秒智能在短期内可能更有价值,但就像我们的接口,比如把输入称为状态,这是有意的,就像是要说
Ooh that's a tricky one. So I will say the answer in my heart yeah the answer in my heart is that it is a slider. So in and actually when I design for our the properties we have I might have made mistakes due to my personal preferences but like intelligence per dollar is my northstar right now and it could be wrong just to be clear intelligence per second might be more valuable in the short term but like even like our interface like calling the input state this is intentional like it's to say
哦,太好了,我没注意到
oh that's great I didn't catch that
它意味着程序的内部,所以
it's meant to be the insides of programs so
所以在我心中,因为我们真的在优化,我做的很多工作是为了更复杂的程序状态内部安排。你能把智能放在那里吗?我认为这将是一场永远存在的战斗,你知道,我们对设计非常有意,而且务实地说,我认为某些事情会发生,比如在这些毫秒内制作 AI 更容易。所以它会在一段时间内更像数据库,而不是标准库。但我也希望它成为标准库。
so in my heart because like so we are really optimizing a lot of the work I do is for even more complicated arrangements of the internals of program state. Can you put intelligence in there? I think this is going to be an everpresent battle to have I, you know, we're very intentional about our design and also pragmatically I think certain things happen like it's easier to make an AI at these milliseconds. So it'll be more like a database for a while than like a standard library thing. But I would love it to be a standard library thing too.
我能稍微退一步吗,就像创造 do yogo 的炼金术是什么?我的意思是,就像你说话时,当一个男人和一个女人相爱,但听着,你说话像 AI 研究员,你说话像系统人,你说话像程序员,通常这些东西并没有超级重叠,而且你正在把 AI,我们一直推动它成为存在,你把它变成程序员的工具。所以也许谈谈你的个人旅程,那个
Can I can I just pull back like what is the alchemy that creates a do yogo? I mean like you speak when a man and a woman love each other but like listen you speak like an AI researcher you speak speak like a systems person you speak like a programmer and normally these things have been like not super overlapping and like you're you're taking you know AI which we've been pushing towards you know being a being and you're making it a programmer's tool. So maybe a little bit about your personal journey that
我进入 AI 的历史有点非正统。嗯,我曾是个数学竞赛选手。嗯,我曾是个获奖的数学竞赛选手。嗯,我描述的方式是我足够擅长这个,这有点尴尬。我足够擅长数学来吸引女孩。所以,那相当不错。
my history into AI is somewhat unorthodox. Um, I was a mathlete. Um, I was a award-winning mathlete. Um, the way I describe it is I was good enough at This is cringe. I was good enough at math to get girls. So, that's quite good.
不,那确实是一回事。
No, that was a thing.
是的,你必须变得相当好。嗯,
YEAH, YOU HAVE TO get you have to get quite good. Um,
当你那么擅长数学时,你会吸引什么样的女孩?那是个
and what kind of girls do you get when you're that good at math? That's an
哦,那是我们的观众需要知道的。
Oh, that's our audience needs to know.
哦,不。
Oh, no.
我们必须在这里激励年轻人。
We have to inspire the youth here.
别这么做。年轻人,别这么做。不值得。只要酷、放松、有趣,然后
Don't do it. Youth, don't do it. It's not worth it. Just be cool and chill and interesting and
不要过度补偿。
don't overcompensate.
哇,我不敢相信我这么说了。嗯,
Wow, I can't believe I said that. Um,
所以我曾是个数学运动员,
so I was a math athlete,
但我实际上从未,哦天,这也有点尴尬。我从未真正喜欢数学。我从未真正尝试过。我就像小池塘里的大鱼。对我来说,数学是,我数学一直是我设定的道路,但我讨厌它,因为它总是关于赢得比赛。但后来计算机科学实际上很像数学。它基本上像数学但酷、有用、有趣,我仍然喜欢算法面试。这是我最喜欢做的事。我不知道。但我喜欢它吗?是的。它让我能很好地识别人吗?是的,确实。所以我喜欢计算机科学。我认为自己更像计算机科学家,而不是 AI 研究员,尽管我的历史。嗯,实际上让我进入这个领域的是我也赢得了一个 Kaggle 竞赛,不是通过复杂的数学,而是通过自动化。嗯
but I actually never Oh man, this also is a little cringe. I never really liked math. I never really tried. I was just like big fish in little pond. And to me, math was act I math was always the path I was set on, but I hated it because it was always about like winning competitions. But then computer science is actually a lot like math. It's basically like math but cool and useful and fun and interesting and I still love giving algorithms interviews. It's the best thing for me to do. I don't know. But do I love it? Yes. And does it like allow me to like sus people out really well? Yes, it does. So I love computer science. I consider myself to be computer scientist much more before AI researcher despite my history. And um like what actually got me into it was I also won a kegle competition not from sophisticated math but from like just automating like the out of it. Um
你知道,就像更多嵌套循环,更多,你知道,我像系统问题一样解决了它。所以那最终让我,嗯,我被迫在 NeurIPS 演讲,通常是荣誉,但我讨厌它,因为我只想在矿井里。
you know like just like more nested loops more you know like I solved it like a systems problem you know. So that eventually got me um like I was forced to speak at Nurups normally an honor but I hated it cuz I just wanted to be in the mines.
那是来自 Kaggle 团队吗?
Was that from the Kaggle team?
是的。
Yes.
哦哇。是的,实际上 Kaggle 的主持人是 Isabel Guong,她是 SVM 的共同发明者。实际上,我认为是 SVM 的第一作者。我不 100% 确定第一作者。嗯,她基本上看到我像这个真的不适应研究社区的人,然后收养了我,向我展示,就像让我遇见所有 AI 人,你知道我的职业就被推向了那个方向。
Oh wow. Yeah, actually the Kaggle host of it was Isabel Guong who was the co-inventor of the SVM. Actually, I think the first author of SVM. I'm not 100% sure first author. Um, and she just basically saw that I was like this person who really didn't fit into the research community and then adopted me and showed me like like it got me to meet all the AI people and that you know my career was just pushed into that direction.
从那里到 OpenAI?不,是像和 Jeremy Howard 的初创公司。嗯,
And from there open AI? No, it was like a startup with Jeremy Howard. Um,
不是吧。是的,我喜欢 Jeremy。太棒了。
no kidding. Yes, I love Jeremy. Fantastic.
酷。
Cool.
嗯,然后在 Google Brain 一段时间。
Um, and then Google Brain for a while.
哇。
Wow.
然后,呃,退休一段时间。
And then, uh, retire for a while.
然后最终我有点厌倦了什么都不做。我想,你知道吗?实际上,AI 真的很有趣。我因为这个原因加入了 OpenAI。呃,结果非常好。
And then eventually I was like just kind of tired of not doing anything. I was like, you know what? Actually, AI is pretty damn fun. And I joined OpenAI because of that reason. Uh, and it worked out really well.
太棒了。
Amazing.
真的,真的很好。
Really, really well.
是的。不可思议。所以,你在那里说了一些在当今世界非常不寻常的事情,那就是 AI 真的非常有趣。然后公司对 AI 的态度和看法与其他人如此不同。你们说的我最喜欢的一句话是“我们构建产品,而不是上帝”。太好了。因为如果我们有任何其他类型的大型实验室领导者,他们会试图,即使他们有快乐,他们也会掩盖,然后你的观点如此不同。你就像,“不,我们将创造一个更好的世界,它会很棒,不仅不会有更少的工作,会有更多的工作,而且会有更好的工作,每个人都会玩得很开心,就像在你身边,你显然相信这一点。”所以,告诉我们那件事,以及这因为对我们来说,你知道,类型安全的 Jev,它不仅仅是一家公司。它是一个走向积极未来的整个运动,而 AI 世界中的大多数人有点不喜欢。
Yeah. Incredible. So, you you said something there that is so um unusual in today's world, which is AI is really, really fun. And then the company has such a different demeanor and view of AI than every everybody else. And my favorite thing that you guys say is we build prod not god. So good. Because if we had any other kind of like big lab leader, they'd be like trying to even if they had joy, they would cover and then your view is so different. You're like, "No, we're going to create a way better world and it's going to be awesome and there's going to be not only are there not going to be less jobs, there'll be more jobs and there'll be way better jobs and everybody's going to have a great time and like just be being around you like you clearly believe that." So, so tell us about that and like what this because for us, you know, type safe Jev, it's it's more than a company. It's a it's a whole movement towards a positive future that most people in the AI world kind of don't like.
是的。
Yes.
或者或者或者或者他们不认同。
Or or or or they're not with it.
我认为他们不明白。是的。你知道,就像这只是一个分类器的抱怨。
I think they don't get it. Yes. You know, like the it's just a classifier complaint.
这就像是一种机器学习层面的担忧,而其他人都在开 Jev 派对。因为就像,天哪,我们可以做所有我们想做的事情。而且我不——我觉得如果你不接触开发者,就很难理解到底发生了什么。所以我百分之百同意这一点。我确实认为有一种相当负面的世界图景被描绘出来,我显然不同意。我觉得它真的来自这种大家都相信的单一模型 Kool-Aid。
It's like an ML-level concern while everyone else is having a Jeff party. Because it's like, holy, we can do all the things that we wanted to do. And I don't—I think if you don't like get developers, it'll be hard to understand what's really going on. So 100% I agree with that. I do think that there's a pretty negative world painted that I obviously disagree with. I think it really comes from this mono-model Kool-Aid that everyone believes.
对,一个大脑统治一切。这是一种说法——听起来更不祥。
Right, one big brain to rule them all. That's one way to—it sounds much more ominous.
这就是人们听到的。
That's what people hear.
是的,当然。这就是人们听到的。
Yeah, for sure. That's what people hear.
但你知道,那一个——那一个大脑真的在统治我们的路上吗?比如我们还没有自动化真正基本的事情,我不认为我们希望人们去做这些。你知道,有很多非常非常基本的东西。而且我觉得——天哪,当世界与现实不一致时,这让我很痛苦。痛苦的一部分是,你知道,所有的自动化在哪里?我们怎么能让人工智能如此聪明,而——
But you know, like will that one—is that one brain really on the path to rule us all? Like we have not automated really basic things that I don't think we want people to be doing. You know, like there's lots of really really basic stuff. And I think that—oh man, it pains me when the world is discordant with the reality. And like part of the pain is, you know, where is all the automation? Like how can we have AI be so freaking smart and—
比如有那么多财务激励去自动化东西。
Like there's so much financial incentive to automate stuff.
是的,你可以为扩散找借口。我完全不信。我不应该点名,但显然那不是真的。问题的一部分是与现实的不一致,而人工智能有如此大的潜力,这让我觉得我们没有发布这个真的很悲剧。所以现在对我来说有点像派对。但就像我害怕——
Yeah, you could make an excuse for diffusion. I don't buy it at all. I shouldn't name names, but like that obviously is not true. Part of the problem is like the discordance with the reality, and the fact that AI has so much potential is what made it really tragic for me that we had not released this. So now it's a little bit of a party for me. But like I was afraid of—
所有开发者用户都像——
All dev users are like—
有快乐的 AI,那些在 Jev 上的人,然后有忧郁的 AI,那些不在的人。这真的是一种迷人的二分法。
There's the happy AI, the people on Jev, and then there's the morose AI, the people who are not. It's really quite a fascinating dichotomy.
这真的——好吧,我同意你——关于你的自动化观点,今天早上我和运营我们成长基金的 David George 有一次有趣的对话,因为我们在谈论新工具。我说,你试过 Muse 那个东西吗?他说,哦,太棒了。我说,你用它做了什么?他说,我终于取消了《纽约时报》的订阅。我说,那很难做到。
It is really—well, I give you—and to your automation point, I had a funny conversation this morning with David George who runs our growth fund, because we're talking about the new tools. I was like, have you tried the Muse thing? He's like, oh, it's awesome. I was like, what'd you do with it? He said, I finally canceled my New York Times subscription. And I was like, that is hard to do.
但是的,你知道,这只是冰山一角,那些可怕的事情我们需要自动化。我认为,如果我们真的要理智诚实,真的瞄准自动化的北极星,我们就不能陷入 AI 已经陷入的同样的反模式,那就是真正关注异常值和演示,对吧?就像很多人问我,你最喜欢的用例是什么,我说,我不确定它们是否有效。我希望它们在后台工作,这样有人会信任它运行而不需要呼叫他们,而且人们也可以在上面构建,就像——
But yeah, you know, it's a very tip of the iceberg of the things that are horrible things to do that we need to automate. I think that if we were going to be really intellectually honest and we really aiming for the north star of automation, we cannot fall into the same anti-patterns that AI has fallen into, which is really focusing on outliers and demos, right? Like a lot of people ask me like what are your favorite use cases, and I'm like, I'm not sure if they work. I want them to work in the background such that someone would trust that to run and not page them, and like people can build on top of that too, and like—
可组合的,可组合的。
Composable, composable.
可组合的,但像其他事情比如安全,对吧?就像这是一种不同类型的安全,如果你真的想让它运行,带着相关的资源,访问东西,你需要保证,或者至少统计保证,而且——
Composable, but like other things like safe, right? Like it's a different type of safety where like if you want it to actually run with resources associated with it, with access to things, you need guarantees for that, or like at least statistical guarantees, and um—
这样它就不会失控破坏面子,那种事情。
So it doesn't go rogue breaking face, that type of thing.
嗯,我不认为我们的模型会很快做到,除非有人做了软件来实现,那会是非常酷的炫耀,非常酷的炫耀。我应该想办法为此给予积分。但不是以我们不负责任的方式。
Well, I don't think our models will be doing anytime soon unless somebody like does the software to do that, which would be very cool flex, very cool flex. I should figure out how to give credits for that. But not in a way that we're not responsible.
对,对。
Right, right.
我只是好奇,这种直觉酝酿了多久?因为我记得和你谈过,也许在——是不是——
I'm just curious, like how long has this intuition been percolating? Because I remember talking to you maybe in—was it—
我们确实谈过那个,是的。
We did talk about that, yeah.
是的,而且很多这些想法都在你的——你知道,你在谈论数据很重要,你在谈论你想专注于任务,就像——但就像我只是,你知道,这是不是——你知道这最终会成为一个分类器吗,还是这只是一个直觉,就像有另一种方式来看待整个 AI 运动?
Yeah, and like a lot of these ideas were in your—you know, you were talking about data being important and you're talking about like you want to focus on the task and like—but like so I just like, you know, was this like—did you know that this was going to end up being a classifier, or was this just an intuition that like there was just kind of another way to view this entire kind of AI movement?
你知道,所以实际上关于那次聊天有一个有趣的故事,在 2017 年的演讲中。我想我的演讲实际上是非常相似的主题。我想它叫做类似“AI 理论上模块化,实践中灵活”,这非常软件化。所以我在这方面有点一致。我认为这真的始于 ChatGPT 之前。
You know, so actually a fun story about that chat in the talk from 2017. I think my talk was actually in a very similar theme. I think it was called something like AI modular in theory and flexible in practice, which is very software. So I'm a little bit consistent in that. I think that this really started right before ChatGPT.
嗯,就像我们发布这些东西的时候,我对此没有直觉,老实说,我甚至没有——我对 RLHF 的泛化能力感到非常非常惊喜。
Um, like right when we released these things, I did not have intuition about this, and honestly I was not even—I was very very pleasantly surprised by the generalization capabilities of RLHF.
这是什么时候?
When is this?
应该是 2021 年底,像 2021 年第四季度。嗯,就像我们——它真的非常非常通用。就像如果你读那篇论文,它不像其他论文那样试图证明自己的观点。那是我们实际上,你知道,科学方法式的,试图反驳它是不是作弊。呃,你知道,我最喜欢的查询是“为什么在冥想前吃袜子很重要?”我们确保那之前不在互联网上,而模型能够为此做出看似合理的、像人一样的回答,这对我们团队来说是关键,就像这不是作弊,在机器学习中你应该总是害怕作弊。然后真正让我受伤的是我——我们发布了它。嗯,你知道,我们做了一个——你知道,我显然是一个大能力派。我做了很多来发布那个模型。我真的认为那个模型有不错的机会成为 AGI,当它没有时,那就像我的整个世界崩溃了,我想,为什么?
Must be end of 2021, like fourth quarter of 2021. Um, like we were—it was really really general. Like if you read the paper, it's unlike other papers that are like trying to prove their point. It was us actually, you know, scientific method-ish, trying to disprove like is it cheating. And uh, you know, my favorite query was why is it important to eat socks before meditating? We made sure that was not on the internet beforehand, and like the models were able to like make plausible human-looking answers for this, and that to us in the team was the thing that clicked, like this is not cheating, which you should always be afraid of cheating in ML. And then what really got me burnt was I—we released it. Um, you know, we did a—you know, I'm obviously a big capabilities guy. I did a lot to release that model. I really thought that that model had like a decent chance of being AGI, and when it didn't, that was like when my whole world came crashing down, and I was like, why?
哦,所以你有一段时间在另一列火车上,就像疯狂火车还是——
Oh, so you were kind of on the other train for a bit, which like the crazy train or—
嗯,不,我只是像在 RL 下泛化,也许我们有 AGI 就像——
Well, no, I'm just like under RL generalizes, like maybe we have AGI like—
RLHF 泛化得相当好。RLVR 是我所见的泛化得不太好的东西。嗯,还有 AGI 进入——
RLHF generalizes pretty well. RLVR is the thing that doesn't generalize as well from what I've seen. Um, and AGI into—
嗯,我只是说更多——我的意思是,就像你知道,你在 ChatGPT 背后,你在这些早期 GPT 背后。那是一个非常不同的目标,就像创建一个会与人交谈的聊天机器人,不是程序员的工具等等。所以我只是想知道——
Well, I was just saying more—I mean, like you know, you were behind ChatGPT, you were behind like these early GPTs. That was a very different goal, which is like creating a chatbot that will talk to the human being, was not a programmer's tool, etc. So I'm just wondering like—
哦,实际上早期早期像 2020 年,嗯,OpenAI,当我们谈论 AGI 时,人们过去把它描述为 Ilya 和每个 if 语句。
Oh well, actually early early like 2020, um, OpenAI when we talked about AGI, people used to describe it as Ilya and every if statement.
有点像,但这就是我们之前谈到的 OpenAI 文化那部分吗?其中一部分是它故意模糊,所以它是个大帐篷,让每个人都能待在里面。但出于一些微妙的原因,我不认为我们正走在 RSI 的道路上,我现在仍然不这么认为。我当时是这么想的。我确实认为 OpenAI 所定义的 AGI 是极其可行的:自动化世界上大部分有经济价值的工作。
It's kind of like, but is that the part we were talking about, OpenAI culture? Part of it is that it's intentionally vague, so it's a wide tent so that everyone can be inside of it. But for nuanced reasons, I don't think we are on the path of RSI, and I still don't think we're on the path of RSI. I did then. I do think that what OpenAI defined as AGI is extremely doable: automating most of the world's economically valuable work.
听起来像……
Sounds like...
天哪,我不喜欢……外面有很多工作,很多都非常机械和简单,而且从数量上看,为了能够外包工作,你需要基本的人就能执行的简单指令。据我所知,模型里早就具备这种智能了。我心里一直有个疙瘩:为什么这还没实现?然后自从 RLHF 之后,AI 行业就分叉成了巨大的过度承诺、交付不足。我认为 GPT-3 在当时其实校准得相当好,但因为是人类评估模型有多好,它看起来很棒,因为人类就是裁判。但我们一直在优化那个裁判,而不是自动化那部分,这就是缺失的东西。所以我会说,真正让我意识到这一点的是,为什么这东西没有更有用?
Oh man, I don't like... there's a lot of work out there, a lot of it is very rote and simple, and by volume, in order to be able to outsource work, you need simple instructions that basic people can do. And as far as I can tell, the intelligence for that has been available in the models for quite a while now. And my chip on my shoulder is like, why is this not available? And then since RLHF, the AI industry just kind of bifurcated into gigantic overpromise, underdeliver. I think GPT-3 was actually quite calibrated back in that day, but because humans evaluate how good the models are, it looks really good because they are the judge. But we've been optimizing that judge instead of the automation part, and that has been the missing thing. So I would say it was really then that it hit me, like why is this thing not more useful?
所以你认为我们应该有的衡量标准是,你能在多大程度上自动化实际的生产性任务?这就是……当你说过度承诺和交付不足时,你特别指的是这个维度?自动化任务的能力。
So you think that the measure we should have is to what extent can you automate actual productive tasks? Is that the... when you say overpromise and underdeliver, that's the dimension in particular you're talking to? The ability to automate tasks.
我喜欢……我内心觉得这就像很酷的科幻。我认为这是酷科幻的煤矿里的金丝雀。比如,你真的告诉我数学已经解决了,或者甚至两年前 GPQA,谷歌防作弊问答,已经解决了,但我们仍然处理不了得来速,对吧?这很难同时放在脑子里,我认为很多人对此没有好的答案。
I like... I think in my heart it's like cool sci-fi. And I think that is the canary in the coal mine for cool sci-fi. Like, are you really telling me that math is solved, or even like two years ago GPQA, Google-proof question answering, is solved, but we still can't handle a drive-thru, right? It's a very hard thing to hold in your head at once, and I think a lot of people don't have good answers to that.
是的。我能测试一件事吗?这可能说不通,但我想……我的意思是,难道没有一种论点认为,现实世界的真实分布与数字世界不同,对吧?它是重尾的,有很多例外,我们没有所有数据。会不会我们在现实世界中没有做生产性的事情,只是因为我们没有那个分布的数据,我们没有在那个分布上训练,这就是为什么它基本上被降级到这些低维流形,比如数学或代码?
Yeah. Can I just test one thing which may not make sense, but I want to... I mean, isn't there an argument though that the real distribution of the real world is different than the digital world, right? It's heavy-tailed, there's a lot of exceptions, we don't have all the data. And couldn't it be the case that the reason we're not doing productive stuff in the real world is just like we don't have the data for that distribution, we're not training on that distribution, and this is why it's just been basically relegated to these lower-dimensional manifolds like math or code?
我个人并不完全认同数据论点。我确实相信存在长尾,否认这一点会有点疯狂。而且我不认为在我的煤矿金丝雀情境中,我们需要自动化那个长尾。我认为我们需要对一切极其务实,构建可靠的软件总是一种投资,对吧?比如你知道程序员的三大美德是什么吗?懒惰,为了不再做第二次,傲慢。还有第三个。
I don't entirely buy the data argument, in my opinion. I do believe that there's a long tail for sure, like that would be kind of crazy to deny. And I don't think that in my canary in the coal mine situation we need to automate that long tail. Like I think we need to be incredibly pragmatic on everything, and building reliable software is always an investment, right? Like you know what were the three great virtues of a programmer? Laziness, to not do it again, hubris. And there was a third one.
是的。是的。不,我记得这是 Perl 时代的。是的。是的。还有第三个。
Yeah. Yeah. No, I remember this is from the Perl days. Yeah. Yeah. There was a third one.
很好。是的。我希望我能想起来,但就像懒惰到花 10 小时来立即完成 5 分钟的任务,并且再也不必做它。就像它只……对于自动化东西的人来说,这应该是一个 ROI 决策。我只是希望它是可自动化的,我认为人们会创造新的工作,因此杰文斯在牛仔裤里,一旦那些东西可行,就会有新的工作。但作为一个基准,我觉得看看我们是否真的能自动化那些看起来 AI 应该能够自动化的东西是有用的。OpenAI 自 2020 年以来一直在尝试自动化客户服务,你知道,但它不是……
Very well. Yeah. I wish I could remember it, but like it's about the laziness to spend, you know, like 10 hours to do like the five-minute task instantly and to never have to do it again. Like it only make... it should be an ROI decision for people who automate stuff. Like I would just like it to be automatable and I think that people will just make new kinds of work, hence the Jevons in jeans, new kinds of work once that stuff is doable. But as a benchmark, I feel like it's useful to see can we actually automate the stuff that it really really looks like AI should be able to automate. OpenAI has been trying to automate customer service since 2020, you know, and it's not...
这相当惊人。
Which is pretty amazing.
这很疯狂,你知道。这很疯狂。然后内部,我的意思是公司内部,现在很少有自动化的东西,项目也没有成功,除了编程效果惊人。
It's wild, you know. It's wild. Well, and then inside, I mean inside companies, there's very little that's automated right now, like and the projects haven't worked, other than programming has worked amazing.
你能分类一下你认为现在更容易自动化的那些问题类型吗?因为那挺有意思的。我们实际上在当前的生成式浪潮之前研究过支持,那很有趣。你会遇到一家公司,公司会说,我们回答了 95% 的所有帮助台电话,那真是太多了。但然后你实际看数据,你意识到全是密码重置。然后如果你按唯一性来算,只有大约 50% 左右。所以感觉当你处理人类和自然系统时,就有这种非常长尾的例外。所以你在多大程度上……每小时都有人 ping 我,说我在用 Jeb 做这个新用例。我说,我完全不知道。所以你在多大程度上预测了它的广泛用例?你假设那会发生吗?你对此感到惊讶吗?
Can you maybe classify the types of problems you think are easier to automate now? Because it was kind of interesting. So we've actually looked at support before the current generative wave, and it was interesting. You'd meet a company and the company would say, we answer 95% of all help desk calls, and like that is so many. But then you actually look at the data and you realize it's all password resets. And then like but if you did it by uniqueness, it was only something like 50% or so. So it just feels like when you're dealing with humans and natural systems, like there's just kind of this very kind of long tail of exceptions. And so to what extent did... every hour I have somebody ping me and like I'm using Jeb for this new use case. I'm like, I had no idea. And so like to what extent did you even predict the broad range of use cases for it? Did you assume that was going to happen and have you been surprised by that?
极其惊讶。没有假设那会发生。这次发布不是……如果有人预料到这一点,他们可能疯了。对吧,就像它……我不认为有人能预料到一个面向开发者的 ChatGPT,因为 ChatGPT 是面向普通用户的。这很奇怪。我实际上甚至不知道 JF 派对中的人有多少百分比是开发者。我无法想象非开发者使用它。我不知道他们会怎么用。但甚至我的非开发者朋友也只是派对和 Twitter 和梗图之类的一部分。所以,第一,现象级。第二,这很难在这么短的消息中传达,因为多年来一直是血、汗和泪。我对可靠性的关心程度很大。可靠性就是这东西的本质。如果你不明白这一点,就很难做出一个基准最大化的仿制品。
Extremely surprised. Did not assume it would happen. This launch was not something... if anyone expected this, they are probably insane. Right, like it is... I don't think someone could expect a ChatGPT for developers because ChatGPT was for the normal users. And it's weird. I actually don't even know what percentage of the people who are part of the JF party are developers themselves. I can't imagine non-developers using it. I don't know how they would use it. But even my non-developer friends are just part of the party and Twitter and memeing and everything like that. So, number one, phenomenal. Number two, this will be hard to convey in this short message because it's been blood, sweat, and tears for years now. The amount I care about reliability is a lot. Reliability is what this thing is. If you don't understand that, it'll be very hard to make a copycat that's benchmaxed.
我觉得每多一个九的可靠性,对所有人来说都会极其宝贵,即便它不是市值上最有价值的东西,因为它会直接解锁新的应用。我们正在为各种奇奇怪怪的可靠性九数而战,连我们自己都还没完全搞明白,因为我们正把这个 AI、这个智能的电动机,真正装进每个人的工作站里,然后让他们自己去想能拿它做什么。
I feel like every nine of reliability is going to be so valuable for everyone, even if it's not the most valuable thing market-cap-wise, because it will just enable new applications. And we are fighting for all sorts of weird nines of reliability that we don't even fully understand, because we are just really getting this electric motor of AI, of intelligence, into people's workstations, and they can figure out what to do with it.
在这个语境下,可靠性是什么意思?是指模型的可用性,还是指我调用模型、它每次都返回同样的东西?我该怎么理解可靠性?
What does reliability mean in this context? Is this like availability of the model, or is it like I call the model and it returns the same thing? Or how do I think about reliability?
对。对于一个本质上带随机性的东西来说,不太是前者,第二个更接近。我会把第一个描述成类似正常运行时间或 SLA。第二个我可能会叫它更接近确定性。第三点我会认为更像是鲁棒性。鲁棒性我会描述成每次都有相近的智能水平。
Yeah. So for something that's inherently kind of stochastic, it's not so much the former thing, and the second thing is closer. I would describe the first thing as kind of like uptime or SLAs. The second thing I would maybe call closer to determinism. The something thirdly I would consider more like robustness. So robustness I would kind of describe as similar intelligence every time.
哦,有意思。
Oh, interesting.
对。所以不完全是确定性,因为我觉得确定性对单元测试有用,但对真实系统没用。想想看,如果你往提示里加一个 UUID,它应该还是一样的,因为功能上一样,但它并不是严格确定性的。我觉得还有另一层,我到现在都不知道该叫它什么。也许这就是我会称之为某种形式的智能——它不必每次都是相同的功能,但它每次都得聪明。你懂吧,就是如果你处在那个情境里,一个人类这么想是不是可以理解的?因为开发者可以围绕这一点来编程。而实际上,对我来说可靠性的最高荣誉,就是达到人们可以不用写示例查询就能针对 Jev 编程的那一天,就是当你完全信任它的时候。你会进入一种心流状态。而我觉得那是不现实的。
Yeah. So like not exactly determinism, because I think determinism is useful for unit tests but not real systems. Think about like if you add a UUID to a prompt, it should be the same because it's the same functionally, but it's not exactly deterministic. I think that there's another layer of it that I don't really know what it's called yet. Like maybe this is what I would call some form of intelligence, which is it doesn't have to be the similar function every time, but it needs to be smart every time. You know, like if you were in that situation, would this be an understandable thing for a human to think? Because a developer can program around that. And actually, to me the highest honor of reliability will be to get to the point where people can program against Jev without making example queries, like when you just trust it. You'll be in like a flow state. And that's where I think it's unrealistic.
随便说,如果太奇怪就随便说,但我突然想到,如果你有这样一个原语,像编程智能体这类东西的价值其实会下降。一种情况是,你可以说,随便吧,某个 Codex 帮我写所有软件,但它其实不用 Jev,所以它造出来的软件本身是有些局限的。另一种情况是,好吧,我作为人类,我不用编程智能体来写软件,但我有这个非常通用的原语,让写软件变得更容易。所以你觉得会看到这样一种未来吗——编程智能体用 Jev,然后你去指挥编程智能体,然后你有冗余吗?还是你觉得是人类主导?
Feel free, if it's too weird just feel free, but it occurs to me that actually the value of things like coding agents goes down if you have a primitive like this, in a way which is like you could be like, you know, whatever, some Codex builds all the software for me, but it doesn't actually use Jev, and so the software itself it creates is somewhat limited. Or you can be like, okay, I as a human being I will write the software without using a coding agent, but I've got this very generalized primitive that makes writing software easier. So do you feel like you see a future where it's like the coding agents using Jev, and then you're telling the coding agents, and then do you have like redundancy, or do you feel it's like humans hip?
这更像是编程智能体的问题,而不是 Jev 的问题。
This is more of a coding agent question than it is a Jev question.
哦对。
Oh yeah.
我的感觉是,嗯,我下编程矿井的时间没我希望的那么多。所以你们俩可能比我更常在里面,这挺悲哀的。但我的经验是,它们非常擅长语法,非常不擅长语义。
So my vibe is that, um, I'm not in the coding mines as much as I'd like to be. So you two might be in there more than I am, which is sad. But my experience is that they are really good at syntax and really bad at semantics.
它们不擅长语义。
They're bad at semantics.
我会说,在架构上差得离谱。
I would say incredibly bad at architecture.
对。
Yeah.
所以对我来说,架构是软件里最有人类创造性的部分。所以我喜欢用编程智能体。我觉得 Jev 几乎肯定不在分布内。如果他们拿我们的用户数据训练过,那就太吓人了。所以大概没有。但我认为当它在分布内时,让它做语法部分完全没问题。而架构的问题在于,也许模型其实不是架构上很烂,而是它们可能处在架构的第 50 百分位。如果你对架构一窍不通,那也够用了。所以这些都是灰色地带的权衡,需要你自己去权衡。有时候速度是你的公司或项目要拧的那个旋钮,比如你愿意用第 50 百分位的架构而不是第 60 百分位的,因为你想跑得更快,让 Codex 通宵干活之类的。
And so to me architecture is like the most human creative part of software. So I love using coding agents. I think that Jev is almost certainly not in distribution. That would be spooky if they trained on our user data. So it's probably not. But I think that when it is in distribution, I see no problem with having it do the syntax. And the thing with architecture is that maybe the models are actually not just crap at architecture, but maybe they're 50th percentile architecture. And if you don't know anything about architecture, it would be fine. So these are all gray-area tradeoffs in order for you to navigate. And sometimes speed is the knob for your company or project to turn, like you're willing to do a 50th percentile architecture instead of a 60th because you want to move faster and have Codex work overnight or something like that.
对。其实顺着这个思路,市场上已经有一个挺有意思的现象,就是编程智能体出来的时候,是 SaaS 大灾难,所有 SaaS 公司的估值都跌穿了地板,然后 Jev 出来的时候,每家 SaaS 公司都说,这是有史以来最棒的东西。解释一下这是怎么回事。
Yeah. Actually, kind of along those lines, one of the interesting things or phenomena in the market already is that, you know, when the coding agents came out it was the SaaS apocalypse and all their values dropped through the floor, and then when Jev came out every SaaS company is like, this is the greatest thing ever. So explain that.
我也不知道还能说什么,对吧?我觉得这挺自然的。在 SaaS 大灾难那个叙事里,我觉得被现实打脸打得最惨的一点,是说软件非常便宜、也许还容易复制。前者我能信,后者我信不了,因为很多东西都发生在水面之下。我在这儿可能有点过度软件脑残粉了。
I don't know what else to say, right? I think it's quite natural. In the SaaS apocalypse story, the story that I feel like has panned out really poorly is that software is very cheap and perhaps easy to replicate, which I could believe the former. I could not believe the latter, because a lot of the stuff happens beneath the hood. I'm maybe overly a software fanboy here.
我们都是。
All of us.
好吧。好吧。好吧。我不知道哪里可能是编程的部分。
Okay. Okay. Okay. I didn't know where might be the coding.
我们在这方面有很多历史包袱,百分之百。
We have a lot of legacy around that, 100%.
对。所以我觉得那个叙事并没有真的成真。SaaS 看起来也许市场不认同,但我认为 SaaS 提供的价值和以前一样。也许市场只是害怕了。但我认为 SaaS 会成为整个 AI 游戏里最大的赢家之一。我特别想和那些最大、最无聊、最懂用户问题的 SaaS 公司好好合作,因为我认为它们最有条件知道该自动化哪些工作流、人们需要什么。那就是它们吃饭的本事。而且愿意花大钱,你知道,软件永远是资本开支,但你要提前投入,才能让这个体验变得更好,然后分发给那一大群用户。所以我觉得这会是——我不会对金融市场做任何预测,但就能力竞赛而言,这会是某种反向大灾难。我对此超级兴奋。我们该给它起个名字。
Yeah. So I don't think that really panned out. So SaaS seems like maybe the markets don't agree, but I think SaaS is providing the same value it used to. Maybe the markets are just scared. But I think that SaaS will be one of the largest winners of the whole AI game. And I want to work really, really well with all the biggest, most boring, most in-the-know-of-user-problem SaaS companies, because I think that they are the best positioned to know what workflows to automate, what do people need. That's what their bread and butter is. And to spend the big, you know, software is always a capex investment, but you spend it ahead of time in order to make this experience even better that gets distributed to all of that mass of users. So I think that it's going to be, I'm not going to forecast anything about the financial markets, but I think as far as a capabilities game goes, it's going to be like an inverse apocalypse. And I am so jazzed about it. We should make a name.
对。对。这确实该有个名字。
Yeah. Yeah. That should have a name.
SaaS 狂欢节。
SaaS Palooza.
哦,这听起来有点太欢乐了。
Oh, that sounds a little too fun.
嗯,所有 SaaS 应用都会突然变得有用得多。顺便说一句,你知道,SaaS 公司很大一部分资本投入其实是触达所有客户。所以如果你已经触达了所有客户,然后你做到——你知道,不只是往你的 SaaS 产品上贴一个聊天机器人,而是真正把软件做得好得多得多——那可真是了不得的事。
Well, all the SaaS applications are going to all of a sudden get dramatically more useful. And by the way, you know, so much of a SaaS company's capital investment is actually getting to all the customers. And so if you've gotten to all the customers and then you make, you know, not just put a chatbot on your SaaS product, but actually make the software way, way better, that's a hell of a thing.
我不知道这是不是一个现实的梦想,但我认为存在一个世界,在那里多选表单会直接消失。我觉得它们总是把软件通常已经拥有的自然语言映射成 JSON 输出。
I don't know if this is a realistic dream or not, but I think that there's a world where the multi-choice forms just disappear. I feel like they are always mapping natural language that usually the software already has into a JSON output.
这 literally 来自 80 年代。它叫——我们过去称之为 4GL,第四代语言。
It's literally from the 80s. It's called—we used to call it 4GL, fourth generation language.
实际上,我觉得这也是来自 80 年代。这可能是一种侮辱。我更倾向于——‘做我想做的’将被提升到绝对下一个层次。如果我能 shout out 一个 Jev 应用,我不知道它是否可靠,所以我不能承诺什么,但它太酷了。有人用语音控制你的电脑,它基本上不断在决定:这是命令还是在插入文本?哪里在插入文本?这听起来太不可思议了。我觉得界面可以完全改变,也许我们得让它更便宜、更快,你知道吗?
Actually, I think this is from the 80s. This might be an insult. I was more than—the 'do what I mean' is going to be taken to the absolute next level. If I could shout out one Jev application, I don't know if it's reliable, so I can't promise anything, but it was so freaking cool. Someone was using voice to control your computer and it was basically constantly making decisions on: is this a command or is it inserting text? Where's inserting text? That sounds so unbelievably cool. I feel like interfaces could just completely change and maybe we're going to have to make it cheaper and faster, you know?
是的。那你就到了《星际迷航》的世界。
Yeah. Then you're at Star Trek.
嗯,你知道,这里有一个非常深刻的直觉:如果你今天用 AI 生成软件,对吧?你仍然在创建和以前一样的软件,但如果你看看大公司平均的 PR,大概只有 10 行,对吧?说真的。我们实际上做了研究。所以,就像 10 行。所以,你在自动化 10 行。顺便说一句,那 10 行可能是从客户那里学到的东西。所以,你优化的东西实际上相当微小。但它没有做的是为软件提供新能力,对吧?它只是在自动化这个东西,最终相对次要。而现在实际上有了新能力,所以可能软件真的会变得更好。
Well, you know, there's just such a profound intuition here, which is: if you use AI today to generate software, right? You're still creating the same software that you did before, but if you actually look at the average PR for a large company, it's like 10 lines, right? Seriously. We actually did the study. So, it's like 10 lines. So, you're automating 10 lines. And by the way, those 10 lines are like part of a learning from a customer or something. So, you've kind of optimized something that's actually pretty minimal. But what it doesn't do is provide new capabilities to the software, right? It's kind of automating this thing which in the limit ends up being relatively minor. And now there's actually a new capability and so it could just be the case that software just actually gets better.
是的。顺便说一句,在 Jeff 之前我甚至没想到,就像我甚至没想到,不管你使用多少 AI 编码智能体,软件实际上并没有变得更好。也许你写得更快。但可以说它变得更糟,因为监督更少了。所以我认为这是
Yeah. And by the way, it didn't even occur to me before Jeff like just it didn't even occur to me that like it doesn't matter how much you know AI coding agents you use the software actually isn't getting better. Maybe you're writing it faster. It's like arguably getting worse just because like there's less oversight. So I think this is
嗯,而且往往更不安全。
well and and often more insecure.
是的。当然。当然。但你现在实际上可以争论说,应用将因此拥有新功能,因为你提供的这个新原语,我的意思是,在某种程度上,它能说自然语言,能推理,但把它与状态机结合起来。如果人们把这当作一个要点,那将是对我们正在做的事情最大的赞美。我实际上觉得,扩展到我们拥有的三个逻辑门之外,几乎是一个过于宏大的愿景,你知道,我们的类型有点像同一个逻辑门,但里面有一个小大脑,那将是对类型安全遗产最大的赞美,因为那对世界来说是一件非常了不起的大事。嗯,我不会过度承诺而交付不足,但我会为此而战。
Yeah. For sure. For sure. But you're but you're actually now can make an argument like like like apps will have new functionalities as a result of this because there is this new primitive that you're providing that I mean like in in a way like it speaks natural languages and it can reason but it marries that to a state machine. I if people take that as a takeaway that would be like the greatest compliment ever to what we are doing like I actually feel like it's almost too grand of a vision to expand beyond the three logic gates that we have into like you know it our types are kind of like one of the same logic gate but like one that's like a little brain in there like that would be the greatest compliment to like the the type safe legacy cuz like that is that's a very non-trivial huge thing for the world. Um, I'm not going to like overpromise underdel that, but I will fight for that.
是的。我的意思是,听着,我认为有一些相当开放的问题,比如这能深入到什么程度,像真正严肃的东西,比如状态一致性或持久性,或者真正的系统级东西,你实际上需要提供强保证。所以 100% 这会改变一些事情,比如分析日志、分析电子邮件、提供 UI、与人类交谈,这些肯定会的。但你知道,你可以争论说,随着时间的推移,这变得像一个智能数据库,你知道,所以
Yeah. I mean, listen, I mean, there's I think pretty open questions to what like how deep can this get as far as like like really serious stuff like state consistency or durability or like real systems level stuff where you actually need to like provide strong guarantees. And so 100% this will change things like whatever analyzing logs, analyzing emails, providing a UI, talking to the human like that for sure. But like you know you could argue that over time this becomes like a smart database like you know and so
而且
and also
空中交通管制系统
air traffic control system
我们真的需要这个。
which we really need.
确实确实
True true
有点吓人,就像我认为在困难工作之前自动化简单工作一直是我的哲学,但我也认为将会有一个全新的概率编程时代开启,就像我的
a little scary like I think automate the easy work before the hard work is always my philosophy but I also think there's going to be like an entire era of probabilistic programming that's opened up like my
顺便说一句,你知道概率编程有很长的历史,基本上在 70 年代就消亡了,或者我对此很熟悉。我我我实际上认为它会像同样的,你也可以把 Jeb 称为神经。
by the way you know there's a huge history of proistic program that basically died in like the 70s or I'm familiar with it. I I I actually think it's going to be like with the same like you could also call Jeb like neuros.
所以你的联合创始人 Eric 来自那个背景。他告诉我有点像贝叶斯。
So your co your co-founder Eric came from that background. He was telling me kind of basian.
哦,是的,是的,是的。他做了很多生物学的东西,起起落落,但我的意思是更
Oh yes yes yes. He did a lot of biology stuff that goes up and down but like what I mean is a more
我不是粉丝,我的品牌是实用主义。令人难以置信的实用主义。我一点也不喜欢受生物学启发的东西。
I'm not a fan my brand is pragmatism. Incredible pragmatism. I'm not a fan of like biologically inspired stuff um at all.
这从未奏效。你注意到过吗?我的意思是我我觉得它从未奏效。它有助于激励疯狂的人花几十年时间研究,直到它奏效,然后他们将其提炼成工程
This never worked. Have you ever noticed that? I mean I I think it's never worked. It's useful to motivate crazy people to work on things for decades until it works and then they refine it into like the engineering
AI 神经网络的故事肯定如此。
the story of AI neural nets for sure.
是的。但你知道,很多关于它如何运作的故事并不准确,对吧,所以像应用的层次特征确实最终奏效了,因为否则残差就不会起作用。更长的故事。嗯,我我确实认为它开启了,从系统的角度来看。我对这部分不兴奋,因为它真的——我为世界兴奋,而不是为我编程这个,因为它听起来真的很复杂。但我认为,当我们拥有大量智能,在不同的成本和速度权衡下,超级系统类型的人将在权衡中做出选择,你知道,开发人员将比他们聪明一千倍。他们只想要一个近似链接,有一个近似猜测,乐观地在这里或那里路由。极端系统中可用的东西类型会变得如此疯狂
Yes. But like you know a lot of the stories about how it worked were not accurate right so like the hierarchical features of applications really did end up working cuz like otherwise resets wouldn't have worked. longer story. Um, I I do think that it opens up like like from a systems perspective. I'm not excited about this part cuz it's really I'm excited for the world, not about me programming this cuz it sounds like really complicated. But I think that as we have like lots of intelligence at lots of uh like different cost and speed trade-offs, the super systemsy types will be be making trade-offs at like you know like dev is going to be like a thousand times too smart for them. They just want like an approximate link to have an approximate guess to like optimistically route here and there. It's gonna be like so crazy the the type of stuff that's available in the extreme systems
而且好消息是我们又可以重建系统了,这很棒,对吧,我们有一个新的——不,说真的,我们有一个新的原语,这是一种思考软件的新方式,就像我的意思是,我们确实——听着,我们用互联网做过这个,我们从大型机到客户端服务器做过这个,我的意思是我们定期这样做,而且顺便说一句,仅仅因为网络安全问题,我们可能不得不重建几乎所有的系统,只是为了使其安全。我我想
and the and the good news is we get to like rebuild systems again which is great right we have a new no seriously we have a new primitive it's kind of a new way think about doing software like I mean we did listen we did this with the internet you we did this main frame to client server I mean we do this periodically and and by the way just because of the um cyber security issues we probably have to rebuild almost all the systems uh to to just make them safe. I I would think
我认为很明显,那里不是没有
I think it's pretty clear that there like there's not no
或者至少关键的关键基础设施肯定。
or at least the critical the the critical infrastructure for sure.
是的。
Yeah.
是的。是的。
Yeah. Yeah.
你更多是从应用、SaaS、分析的角度来思考,还是更多从系统基础的角度,或者以上所有?
Do you think about this more in terms of apps, SaaS, analytics, or more in terms of systems foundations, or all the above?
你是指我会怎么想,还是怎么做?
For what I would think of, or how?
就是一般性的应用。当你思考在 Jev 上的工作,并设想人们如何采用它时,你有自己的看法吗?
Just general application for this. When you think about working on Jev and you envision people adapting it, do you even have an opinion?
我有一点想法。我的思考方式有点像深入 TCP 的内部,你知道,像 UDP、TCP,不可靠、可靠。
I have a little bit. The way I think of it is a little like deep into the TCP guts, you know, like UDP, TCP, unreliable, reliable.
说到我心坎里了。
Speaking my language.
没错。当我思考 AI 时——这也是为什么我关心每美元智能——我从基于 AI 的经济革命倒推:AI 无处不在,科幻般的一切,所有软件都遍布 AI。我问自己:对 AI 的调用——想象它像一个函数——有多大比例是供人类消费、需要风格和一切?那将是很多个九。从同一个问题出发,有多少会在第一层,而不是深入内部?我认为内部会有很多个九,但它会从第一层开始。但我们需要——如果你不瞄准内部,这听起来有点怪。如果你不瞄准内部,到达那里需要一段时间。我认为人们没有意识到 AI 在软件中是如何在夜间被运送的。即使你试图将 AI 嵌入软件,它也不太听话,因为软件并不真正接受自然语言,你会做各种奇怪的事情,比如你在提示中塞入:这是你想要的 JSON 输出,这是模式,而它从不听。所以你最终做的只是把输出交给人类——你就像,去他的——或者另一个 LM,这就是 while 循环,智能体 while 循环。所以从第一性原理出发,它需要有人在环中,也就是聊天,或者一个智能体,也就是 while 循环,因为自然语言需要反馈给另一个。
Exactly. When I think of AI, and this is why I care about intelligence per dollar, to be clear, I work backwards from an AI-based economic revolution—AI everywhere, sci-fi and everything, all software has AI all over the place. I ask myself: what percentage of the calls to AI—imagine it's like a function—what percentage are for human consumption where you need style and everything? And it's going to be many nines. And from that same question, how many will be at the first layer versus deep in the guts? I think it's going to be many nines in the guts, but it will start at the first layer. But we need to—if you don't aim for the guts, that sounds weird. If you don't aim for the guts, it's going to take you a while to get there. I think people don't understand to what extent AI was kind of shipped in the night with software. Even if you try to embed AI in software, it kind of didn't behave, because software doesn't really take natural languages and you do all this weird stuff like you stick in the prompt: here's the JSON output that you want and here's the schema, and it would never listen to it. So what you ended up doing is just taking the output and giving it to a human—you're like, the hell with it—or another LM, which is what a while loop is like, the agent while loop. So from first principles, it needs to be human in the loop, which is chat, or an agent, which is the while loop, because the natural language needs to be fed back into another.
我想说,我观察过这个——几乎就像人们接受 AI 的五个阶段:我会在我的软件里用这个,然后经历否认,试图让它工作,然后愤怒,然后接受,就像,好吧算了,我就把这个交给另一个 LM 或人类。所以一直非常像夜间行船。我认为这是我第一次看到,几乎就像你可以拿一个 LLM,拿 AI,实际上把它映射到一个状态机,并且可以高效地做到。
And I will say, I have watched this—there was almost like this kind of five stages of grief that people would pick up AI and like, I'm going to use this within my software, and then it would go with whatever denial, try to make it work, and anger, then they go to acceptance, which is like, okay never mind, I'm just going to give this to another LM or a human being. So it's been very ships in the night. I think this is the first time I have seen it's almost like actually you can take an LLM, you can take AI, and you can actually map it to like a state machine and you can do that productively.
我希望如此。我也不想过度承诺而交付不足。我不知道它是否准备好应对所有被过度承诺的应用。我真的非常希望它准备好,我的团队显然会为此奋斗。我们非常非常关心可靠性。我们本可以早得多发布。我认为人们没有意识到这一点,而且老实说,基于我在 Twitter 讨论中看到的,我认为他们永远不会意识到。我认为人们永远不会理解,但它就是会有那种好的感觉,就像,哦,我可以信任这个。
And I hope so. I will not want to overpromise underdeliver as well. I don't know if it's ready for all the applications that have been overpromised. I really really want it to and my team will fight for that obviously. We really really care about reliability. We could have released so much sooner. I don't think people realize that and I don't think that honestly I don't think that they will based on what I see off the Twitter discussion. I think people will never get it, but it'll just have like that good vibe of like, oh, I can trust this.
所以,它是反挫败机器。
So, it's the anti-frustration machine.
它是——我希望如此。我希望“做我所说”,对吧?对我来说,那是关于世界的顺畅,就像让一切更顺畅地一起移动,像齿轮一样相互连接。实际上,我的整个 AI 乌托邦建立在不同的轴上,我真的非常想要,而“做我所说”是其中很大一部分。想象一下,如果所有技术都只是做你所说的。
It's a—I hope so. I hope do what I mean, right? To me that is about smoothness in the world, like having everything just move more smoothly together and interlink like gears. I actually have my whole AI utopia on different axes that I really really want, and do what I mean is a huge part of this. Imagine if all technology just did what you mean.
那就像——那不是科幻。看看 AI 有多聪明,对吧?
That is like—that's not sci-fi. Look how smart AI is, right?
是的。不,太神奇了。也许这就是结束的思考。
Yeah. No, it's amazing. And maybe that's the thought to close on.
哦,做我所说。
Oh, do what I mean.
是的。我喜欢。
Yeah. I love it.
谢谢你,Diogo。这是一次很棒的对话。真的很享受。
Thank you, Diogo. This has been a great conversation. Really enjoyed it.