Naval 播客:构建 AI 工厂、超音速喷气机和脑机接口

Naval Podcast: Building Factories for AI, Supersonic Jets, and Brain Interfaces

纳瓦尔·拉维坎特 Naval Ravikant · 纳瓦尔播客 · 2026-06-01 · 约 70 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

三位前沿创始人讨论自建工厂、千倍工程师的崛起以及 AI 如何反映用户判断。

Three frontier founders discuss building their own factories, the rise of 1000x engineers, and how AI reflects user judgment.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 16)

全文 · Full transcript(中英对照)

0. 软件工厂概念介绍 Introduction and the concept of software factories

Host

欢迎收听 Naval 播客,您获取新知识的权威来源。今天我们尝试点新东西。我请来了三位前沿创始人。其实是三位帅哥,还有第四位帅哥 Naval。让我介绍一下大家。Gumo the G Roush。他正在将 Versel 打造成面向智能体世界以及之后一切的 AI 云。

Welcome. You're listening to the Naval podcast, your authoritative source for new knowledge. We're trying something new today. I have three frontier founders with us. Three good-looking guys actually, and a fourth good-looking guy, Naval. And let me just introduce everybody. Gumo the G Roush. He's building Versel into an AI cloud for the world of agents and whatever comes after that.

Naval

很高兴来到这里。

Good to be here.

Host

Blake Shawl,他正在自己的工厂里制造超音速飞机和喷气发动机。Blake 的公司是 Boom Supersonic。还有来自科学界的 Max Hodak。他正在构建一种生物混合脑机接口,在硅上生长活神经元,以恢复视觉等感官功能,但最终目标是探索大脑的新区域和新感官。这三位都不是用现成零件组装产品。他们都在建造自己的工厂。我们更关心的是他们在建造过程中学到了什么,而不是他们具体在建造什么。他们正在产生哪些新知识?他们的阿尔法是什么?他们发现了哪些其他创始人可以借鉴的原则?他们现在正在试图弄清楚什么?还有哪些他们尚未谈论、仍在脑中形成的尖端或疯狂想法?Naval,在我开始问 Gummo 之前,你对这些有什么反应吗?

Blake Shawl, he's building supersonic aircraft in his own factory and jet engines as well. Blake's company, Boom Supersonic. And then Max Hodak from science. He's building a biohybrid brain interface that grows living neurons on silicon to restore sensory functions like sight, but then eventually to explore new parts of the brain and new senses. All three of these guys are not composing their products with off-the-shelf parts. They're building their own factories and you know we don't care as much about what they're building exactly as we do about what they're learning about how they're building. What's the new knowledge they're generating? What's their alpha? What principles are they discovering that other founders can learn from? What are they trying to figure out right now? And also what are the cutting edge or crazy ideas that they haven't even talked about yet and they're still forming in their brains. Naval, do you have any reactions to any of that before I jump into Gummo?

Naval

嗯,我们就尽情享受吧。

Yeah, let's just have fun.

Host

对,你们直接开始吧。

Yeah, you guys should just jump in.

Naval

是的。顺便说一句,我记不清原话了,但我一直深受软件工厂这个想法的影响,工程师的工作就是去上班。过去你直接交付输出,公司里的一切就是,A 交付输出 B 的能力有多强。而现在,我评判你作为工程师的方式是,你是否在建造一个能产生 B 到 Z 的倍增输出的工厂,对吧?这是一个相当大的变化,因为基本上我们过去相信,而且这应该有些争议,存在 10 倍工程师,现在显然有 100 倍或 1000 倍工程师,而世界还没有完全适应这一点。我过去在 Twitter 上因为说存在 10 倍工程师而被喷,这违背了很多平等主义哲学,即人人平等。但现实是,当你在思想领域、智力领域和虚拟数字领域运作时,甚至不是 10 倍,而是 100 倍或 1000 倍,而且一直如此。Satoshi Notch,你知道,发明 JavaScript 的人,Brendan Eich,John Carmack,我的意思是这些是 1000 倍程序员。更不用说如果你选择了正确的事情去做而不是错误的事情,那是一个无穷大的差异,可能只是某个人一开始对做什么有更好的判断。现在显然因为 AI 的杠杆作用,争议少了。有争议的是 token 排行榜,对吧?人们仍然有点困惑,因为他们现在想,嗯,我有一堆 100 倍工程师。看看我支付的所有这些 token。我很好奇你们是否也看到了同样的情况,你们如何衡量 ROI?

Yeah. So, I can't remember my exact quote, by the way, but I've been really pilled with this idea of software factories and the job of the engineer being something that you just show up to work. You used to ship the output directly and everything inside the company was, you know, how good is person A at shipping output B. And now what's happening is the way that I'm judging you as an engineer is like are you producing the factory that would produce multiplicative outputs B through Z, right? And that's a pretty significant change because basically like we used to believe and it should be somewhat controversial that there's 10x engineers like now clearly there's 100x or a thousandx engineers and the world hasn't fully adjusted to this. I used to get flamed on Twitter for saying there're 10x engineers that it flies in the face of so much like equality philosophy that everyone's equal. But the reality is when you're operating in idea domains, when you're operating intellectual domains and virtual digital domains, it's not even 10x, it's 100x or thousandx and it always has been. Satoshi Notch, you know, the guy who invented JavaScript, the Brendan I of the world, John Carmack, I mean these are thousandx programmers. Not to even mention if you choose the right thing to work on versus the wrong thing to work on, that's an infinity difference and it could just be not nemer, just one who had a better judgment on what to work on in the first place. And now obviously it's less controversial because of AI leverage. What's controversial is that the token leaderboards, right? Like people are still getting a little confused because now they think, well, I have a bunch of 100x engineers. Look at all these tokens that I'm paying for. I'm curious if you guys have seen the same like how do you measure ROI?

Host

这就像过去用代码行数来衡量,token 消耗和代码行数感觉都不是直接的范式。

It's like the old measuring lines of code, you know, token consumption with lines of code feel like similarly not direct paradigms.

Naval

我的观察是,Claude 或 ChatGPT 或 GPT 基本上和你在一个领域的能力一样强。所以,如果你是一个非常能干的开发者,那么这些东西就非常强大。如果你是一个初级开发者,那么你会发现它更像一个初级开发者。一方面,这些模型能力惊人。另一方面,你偶尔给它们的反馈似乎极其重要。这些小小的更新似乎完全决定了你从它们那里获得的性能类型。

I mean, my observation has been that Claude or ChatGPT or GPT is about as good as you are in a domain. And so, if you're a really capable developer, then these things are really powerful. And if you're a junior developer, then you'll kind of find it to be like more of a junior developer. Like on the one hand, these models are incredibly capable. On the other hand, the feedback that you give them sporadically seems to be incredibly important. And these little updates seem to totally determine the types of performance you get out of them.

Host

我提供了一种新的支持,就是你来找我,说你没有从模型中得到好的输出,然后我告诉你该用什么提示词。所以,重新提示的质量,我认为你暗示的这一点,极其重要。

There's a new kind of support that I give which is you come to me and like you didn't get good output out of the model and I tell you what to prompt the model with. So like the idea of like the quality of the reprompting which I think you're alluding to is extremely important.

Naval

但我的意思是,明确地说,我认为随着时间的推移,这会变得不那么重要,因为模型会变得非常非常聪明,那么你就能投入更少,得到更多。但至少在这个阶段,它似乎确实反映了用户带来的判断。根据我的经验,

But I mean and to be clear I think that this will become less important over time like as the models get much much smarter then you'll be able to put in less and get more out. But at least at this stage it really seems to kind of reflect back the judgment that the user brings in. In my experience,

Host

我有点抗拒学习所有的技巧和窍门,比如,哦,用 Ralph Wigum,用 Open Claw,用 Hermes,用这个提示引擎,用这个脚手架,插入这个,你知道,总是用计划模式。我完全忽略了这些。我只是假设模型会变得更好,比我学会如何使用它更快。它会比我更快地学会如何使用我。所以我对它们完全粗鲁,我对它们感到沮丧,而且我发现随着时间推移,我输入的信息越来越少,做的工作也越来越少,因为我只是假设我可以蛮力解决,我会反复把 Codex Claw 和 Gemini 扔给同一个问题,浪费 token 来节省时间。我认为无论这些模型看起来多么昂贵,它们仍然比人类便宜得多。所以我会说,浪费 token,节省时间。不要把 token 看作输入或输出。只看你的时间和最终输出。即使它们写的是低质量代码,我知道在很多情况下确实如此,不一定是生产质量或可扩展的代码。当我想把它投入生产时,我会投入更多 token。我会说,'好吧,现在检查一下,重写它,而且它们每一代都会变得更好。'所以,是的,我不认为这会在哪里停止。只要我们有可验证的领域并解决问题,它们就会解决这些问题。

I've kind of resisted learning all the ticks and tricks and tips like, you know, there was a, oh, use Ralph Wigum, use Open Claw, use Hermes, use this prompt engine, use this scaffolding, plug in this piece, you know, always use plan mode. I just ignored all of that. I just assumed the model's just going to get better faster than I would figure out how to use it. It would figure out how to use me faster than I would figure out how to use it. And so I've just been completely hamfisted with them and I get frustrated at them and just sort of I found myself typing less and less information and doing less and less work as time goes on with the models because I just assume I can brute force my way through it and I'll throw Codex Claw and Gemini at the same problem over and over and just waste tokens to save time. And I think no matter how expensive these models might seem, they're still way cheaper than a human. So I would say just waste tokens, save time. Don't look at the tokens either as inputs or outputs. Just look at your time and look at the final output. And even if they're writing low quality code, which I know in many cases they are, it's not necessarily production quality or scalable code. When the time comes and I want to ship it to production, I'll just throw more tokens at it. I'll say, 'Okay, now go through look at it, rewrite it, and they're just going to get better every generation.' So yeah, I don't see where this necessarily stops. As long as we have verifiable domains and solve problems, they're going to resolve those problems.

1. 模型作为同行 vs 初级工程师 Model as Peer vs. Junior Engineer

Host

这属于未解决问题的领域,也许你处于创造力的前沿,需要与模型非常协作、仔细和紧密地工作。但我在软件工程方面还没达到那个水平。不过你可能是团队里最极端的软件工程师,对吧?在这群人里,你大概是最硬核的、从软件背景出身的人。你觉得这些模型在能力边界上表现如何?

That's in the unsolved problems domain where maybe you're at the cutting edge of creativity that you need to be working very collaboratively and carefully and closely with the model. But I'm not at that level in software engineering. But you're probably the most extreme software engineer on the team, right? Like out of this set, you're probably the one who most hardcore came up from a software background. How are you finding these models at the edge of their capability?

Naval

嗯,最近发生了一件事,和你说的非常吻合。以前你给模型一个提示,它就像经典的下一个词预测那样,顺着你的想法跑。而现在模型已经能做到这种直观的规划模式,甚至不需要显式规划,它会回来跟你说:‘你看,你问我的事,有三条路可以走。这里有这样一组权衡。’这就是人们在 X 上说的‘哦,现在我们有了博士级别的工程师模型’。很明显,模型在某个时刻毕业了。它们以前是初级工程师,现在成了首席工程师,因为它们会带着一组权衡回来找你。当然,有时它们的预测非常糟糕,这很搞笑。但显然,现在我已经更尊重模型了,把它们当作同行,和它们进行智力上的来回交流。但仍然有很多差距。所以如果你是一个非常熟练的工程师或架构师,我觉得你还能榨出更多价值。所以 Max 提出的问题是:‘如果你是初级,你得到的是初级吗?’显然不是,因为初级工程师会获得他们自己永远写不出的更高级的代码知识。但经验丰富的架构师能获得 10 倍提升,而初级工程师只有 2 倍?这我还在摸索。

Well, there's one thing that's happened recently that what you're saying resonates strongly with. It used to be that you would give a prompt to the model and it kind of does the classic next-token prediction thing and runs away with your idea. And models now have been doing this intuitive planning mode without even having to plan, where it comes back to you and says, 'Look, what you're asking me for, there are these three routes we can take. There's this set of tradeoffs that we're going to go down.' That's a moment where people do the whole thing on X like, 'Oh, now we have a PhD-level engineer model.' It's very clear that the models at some point graduated. They used to be junior engineers; now they're principal engineers because they come back to you with a set of tradeoffs. And obviously sometimes they make really bad predictions, which is hilarious. But clearly, it's now this—I respect the models a lot more as a peer, like I'm going back and forth intellectually with them. But there are a lot of gaps still. So if you're a really proficient engineer or architect, I think you're still extracting more juice. So the question that Max was positing of, 'If you're junior, do you get junior back?' Well, clearly not, because a junior gets more advanced knowledge in code that they would never have been able to write by themselves. But doesn't an experienced architect get 10x whereas a junior engineer gets 2x? That's what I'm kind of trying to figure out still.

Host

对,对。但我觉得,架构决策很重要。所以当你考虑发展时,我现在看到团队里一些初级软件工程师的下一步职业发展是什么。从为某个功能写实现,到选择技术,比如在 Postgres 和其他数据库之间做选择,或者在 ZMQ 和其他消息队列之间做选择。模型可以给出建议,但你会看到它然后说:‘不,不,我想用这个别的。’这就是我说的那种小反馈,在当前阶段对你得到的输出类型真的很重要。

Yeah. Yeah. But I mean, I think there's architectural decisions. So when you think about the development, I'm seeing this now with some of our junior software engineers on the team, of what is the next step in their career progression. It's going from writing implementation for a feature to picking technologies, like choosing between Postgres versus some other database, or picking between ZMQ versus some other message queue. And those—I mean, the models can suggest them, but that's the thing where you'll see it and you'll be like, 'No, no, I want to use this other thing.' That's the type of little feedback that I'm saying really matters in the types of output that you seem to get at this point.

Naval

品味和判断,对吧?品味和判断。话虽如此,你可以问它们该用哪个以及为什么,它们什么都知道。它们会给你很好的权衡。这就是我所说的最近发生的变化,你会说:‘嘿,去把这个超高基数的遥测数据放进 Postgres’,然后它说:‘不,兄弟,我们不把那种数据放进 Postgres。你应该考虑 ClickHouse 或 Athena 之类的。’这种事我遇到过很多次,真的很令人印象深刻。但我仍然在纠结的是,显然人类还在补充模型。有没有可能反过来?比如人类是那个接收指令的人:‘去给我拿这个 API 密钥,因为只有你能做’,或者‘给我弄到这笔投资所需的资金’。你看着就知道,我们显然还没到那一步。

Taste and judgment, right? Taste and judgment. That said, you can ask them which one should I use and why, and they know everything. They'll give you really good trade-offs. That's the change I was saying has happened recently, where you would say, 'Hey, go and put this super high cardinality telemetry data into Postgres,' and it's like, 'No, bro, we don't put that kind of data into Postgres. You should consider ClickHouse or Athena or whatever.' That's happened to me a lot, which is really impressive. But the thing I'm still kind of struggling with is clearly the human is still completing the model. At one point, is it the other way about? Like the human is the one sort of getting the instructions back: 'Go get me this API key because it's something that only you can do,' or 'Get me this amount of capital for my next set of investments.' You just watch that clearly we're still not there yet.

Host

那只是暂时的偏差。很快每个好的 SaaS 公司或托管服务商都会有一个 CLI 和 API 接口,模型可以直接对接。它们甚至不一定需要 API——只要是基于文本的 Unix 系统,智能体就能自己破解出 API。至于钱的部分,你插入加密代币,放比特币,放什么都行,模型就去支付它需要的任何东西。我觉得有人在研究这个。

That's a temporary aberration. Pretty soon every good SaaS company or hosting provider will have a CLI and API interface that the models can meet directly. They don't even necessarily need an API—as long as it's text-based Unix-based, the agent can hack its own API. And then the money part, you insert crypto tokens, put in Bitcoin, put in whatever, and the model goes and just pays for whatever it needs. I think there are people working on this.

Naval

但我在思考的是:纯软件死了吗?纯软件工程过时了吗?这就像说英语。模型现在会说英语了。我们以前必须学代码才能和模型交流。现在模型会说英语了,而且它们像人类一样说模糊、不严谨的英语,并且能理解事物。那么创始人的护城河在哪里?硬件是福音,你知道。如果你要造硬件,同时建一家软件公司很难。正如 Patrick Collison 所说,软件是艺术,很难雇到艺术家。所以现在作为硬件创始人,太好了,你可以很快开发出非常好的软件。如果你在创建模型,也许那才是新的软件工程:训练模型、调整模型、后训练和微调模型。但经典的软件工程——它死了吗?纯软件还值得投资吗?纯软件还能围绕它组织公司、团队并试图获得杠杆吗?

But the thing I am now thinking through is: is pure software dead? Is pure software engineering an obsolete thing? It's like saying speaking English. The models now speak English. We had to learn code to communicate with the models. Now the models speak English, and they speak fuzzy sloppy English like a human, and they understand things. So where's the moat for a founder? Hardware is a boon, you know. Now if you had to build hardware, it was hard to build a software company alongside. As Patrick Collison says, software is art and it's hard to hire artists. So now as a hardware founder, great, you can have really good software developed fairly quickly. If you're creating models, maybe that's the new software engineering: training models and tweaking models and post-training and fine-tuning models. But classic software engineering—is that dead? Is pure software investable? Is pure software something to organize a company, a team around and try to get some leverage?

Host

你们看到 Mitchell Hashimoto 那篇叫‘积木经济’或‘构建块经济’的文章了吗?类似这样的标题。他的论点是,现在对智能体最有用的东西是强大的可重用构建块。因为用 Max 的例子,你不会希望你的 Clanker 每次发邮件都重新发明一个 Q 基础设施系统。它需要引入适合任务的正确大小的构建块。然后说:‘好吧,这次用 ZMQ。’我质疑那种希望智能体从第一性原理重新发明整个宇宙、并且与社会和文明其他部分不兼容的想法。这几乎就像为你重新发明高速公路、法律、政策等。即使有额外优化的潜力,能榨出更多价值,但大规模合作的价值在于说我们都依赖 Postgres 13.2,这仍然非常非常有价值。我认为这些智能体将要使用的基础设施软件和构建块——显然这就是我们在构建的东西——看起来极其有价值。而且我不认为智能体很快就会出现。顺便说一句,你甚至可以用我一直在用的另一个比喻:模型可以重用的智能体已经创建好了,就像一个令牌缓存,因为你不想消耗一万亿个令牌来重现已经存在的东西。

Did you guys see the article by Mitchell Hashimoto called 'The Block Economy' or 'The Building Block Economy'? Something like that. His argument is that the most useful thing for agents to have now is really powerful reusable building blocks. Because to Max's example, you wouldn't expect your clanker to reinvent a Q infrastructure system every time he needs to send an email. It needs to bring in the right building block that's the right size for the task. And say, 'Well, okay, for this one, it's ZMQ.' I challenge the notion that I would want the agent to reinvent the entire universe from first principles in a way that's incompatible with the rest of society and civilization. It's almost like reinventing highways, laws, policies, etc. just for you. Even if there's a potential for extra optimization, extra juice that you can get out of it, there's still a sort of cooperation at large scale value of saying we're both depending on Postgres 13.2, and so that's still really, really valuable. I would say the category of infrastructure software and building blocks that these agents are going to use—obviously this is what we're building—seems extremely valuable. And I don't see the agent anytime soon. And by the way, you could even use another metaphor I've been using: the agent's already been created that the models can reuse is like a token cache, because you don't want to churn through a trillion tokens to reproduce what's already existing.

2. 氛围编程与软件开发 Vibe Coding and Software Development

Host

嗯,所以模型总是有可以分叉的起点,但这将深刻地改变事物。

Uh and so there's always starting points that the model can fork off from, but it's going to change things quite profoundly.

Naval

所以这些就像是模型用的库和依赖项。

So these are like libraries and dependencies, but for models.

Host

是的,特别是针对智能体。

Yes. For agents specifically.

Naval

回到 Naval 的问题,我小时候就学了编程。那件事贯穿了我的青少年和二十多岁,我沉迷其中,连续编码 20 小时,超级有趣,我了解所有关于编程语言的知识。但现在我已经很久没写过一行代码了。部分原因是我的工作变了,但也是从去年 12 月开始,我用 AI 构建了大量软件,现在每天都在用。有些项目我幻想了好几年,现在真的用上了,而且我一行代码都没写。我无法想象再回去手写代码。虽然我本来也不太可能那样做,但总的来说,我看不到手写代码是未来的一部分。

To Naval's question though, I mean I learned a program when I was really little. And I like that was the thing that through all of like being a teenager and in my 20s like I get like sucked into it and just like code for like 20 hours and it was super fun and I knew all this stuff about programming languages. I haven't written a single line of code in quite a while now. And I mean, partly that's because my job is different, but also since December, I've built a huge amount of software that I now use every day. There's all these projects that I've kind of fantasized about for years that now I'm like using um that I've actually built and I didn't write any of that. And I just can't imagine going back to like actually writing code by hand anytime. Like I mean, I'm unlikely to do that anyway, but just like in general, I see that I have a hard time seeing that as part of the future.

Host

是的。有一点很酷,就是你理解各个部分如何组合在一起。我觉得任何理解 API 是什么、数据如何流动、输入输出和性能的人,因为你必须围绕模型来定位,比如我对这个操作有某种期望,这总是比写代码有用得多。就像一个优秀的工程领导,通过 Slack 或一对一交流进行所谓的“氛围编码”,你传达你的意愿、意图和经验,让别人去执行。只是现在我们用智能体来做同样的事。所以我认为这就是你成功的原因。但我不确定每个人都能达到同样的成功水平。

Yeah. There's something really cool is that you understand how the pieces click together. Like I feel like anyone that understands what an API is and how data flows, inputs and outputs, performance because you kind of you have to orient the model around like this is a certain level of expectation that I have out of this operation like that's that's always been infinitely more useful than um than writing code. Like I feel like a really good a proficient engineering leader has been quote unquote like vibe coding through people on Slack or one-on-ones because you're transmitting your will, your intent, your experience and you're letting others run with it. Uh it's just that now we do the same but with agents. Uh and so I think that's why you've been successful with it. But I don't know that everyone sees the same level of success.

Naval

我从 20 年没写代码到现在一直在编码,但通过智能体,我构建了大量软件。结果发现,只要理解软件工程和算法的基本原理,就能走得很远。我停止编码的原因是我没时间搞懂最新的语言、架构和基础设施组件,Vercel 让事情简单了很多,但即使如此,开始阶段只是拼凑组件、组装基础设施,非常烦人。真正改变的是,以前你可以构建很多东西,很多都很直接,但你会遇到一些随机问题,然后花不确定的时间调试某个小问题。现在有了智能体,你不再卡住了,这太棒了。

I mean I went from not having written code in 20 years to I'm coding all the time now. but through agents and I'm building tons of software and it turns out that just understanding the basic principles of software engineering and algorithms actually gets you a long ways because the reason I stopped coding was because I didn't have time to figure out the latest language latest architecture infrastructure pieces to plug into and I know Vercel makes it a lot easier but even then just getting started was a bare like just plugging pieces together assembling infrastructure was just so annoying. The thing that really changed is I mean it used to be that you could build a lot like you like there's a lot that was straightforward but then you would hit some random thing and then you could spend some indefinite period of time debugging some narrow thing and now with the agents what happens is you just don't get stuck anymore which is pretty amazing.

Host

或者它们卡住了。

or they get stuck

Naval

不,我是说它们能相对快速地找到正确的方法。以前我记得朋友学编程时,觉得编程本质上很令人沮丧,那是学习的一部分感觉,但现在不再是这样了。

it's removed well no I mean like relatively quickly they can find like the right way to do things and it used to be that like I remember when their friends learned a program be like nope it's just like intrinsically frustrating like if like that's part of the feel that's how you learn and that just isn't true anymore.

3. 在 Boom Supersonic 的应用 Application at Boom Supersonic

Host

Blake,你在 Boom Supersonic 是怎么应用这些东西的?

Blake, how are you applying all the stuff at uh Boom Supersonic?

Blake

是的。我发现这完全改变了软件和硬件开发者的角色。我们从第一天起就尝试将许多传统工程工作流,特别是硬件工程工作流,转化为软件。如果你不熟悉硬件工程,我尽量解释清楚。很多硬件工程发生在工程师笔记本电脑上的 Excel 表格里,孤立地进行,这些表格非常复杂,有时包含 VBScript 代码,所有这些都是软件,但被当作不是软件来对待——没有版本控制,没有自动化测试。如果你想从空气动力学家那里交接给结构工程师,那是通过邮件发送表格手动完成的,就像 90 年代一样。太糟糕了。所以我们开始构建软件框架,来自动化和可重复硬件工程流程,目的是降低迭代成本。但进展缓慢,因为我们永远雇不起足够的软件工程师。现在我们进入了一个令人震惊的不同模式:软件工程师创建架构,因为他们理解系统、算法和关注点分离;然后硬件工程师可以“氛围编码”他们的部分,因为他们懂硬件工程。结果是小团队的生产力惊人。举个例子,设计涡轮叶片:经典流程中,叶片冷态时开始,但运行时变热膨胀,所以必须设计空气动力学和结构,使其在冷态和热态下都能工作,需要在冷热之间转换,在结构和空气动力学之间转换。这需要一个工程师一天时间,为一个叶片做一部分分析,而喷气发动机有上千个叶片,所以做不了太多。现在,通过软件和硬件人员的结合,我们创建了解决方案:你可以改变叶片几何形状,实时看到结构和空气动力学结果。这使得两个工程师就能设计整个喷气发动机,这完全不同。

Yeah. What I found is it completely changes the role of software and hardware developers. The thing that we did from day one was uh try to take a lot of traditional engineering workflows and I mean hardware engineering workflows and turn them into software. And so if you haven't been around hardware engineering, let me see if I can make this more clear. there uh there's a lot of engineering hardware engineering that happens in Excel spreadsheets on engineers laptops in a silo and it's very complex uh spreadsheets sometimes like VBScript code and all of this is actually software but it's it's treated as if it's not software there's no there's no source control there's no automated testing if you want to hand something off from like an aerodynamicist to a structures engineer that's done manually with like a spreadsheet over email like it's the 1990s. It's terrible. And so we we started building these kind of like software frameworks that can automate and make repeatable hardware engineering flows. The idea we could reduce the cost of iteration. Um but it was it was slowgoing because we could never get enough we could never like afford enough software engineers. And what we've gotten into is this uh mind-blowingly different model where the software engineers actually create the architectures because they understand systems, they understand algorithms, they they understand, you know, division of concerns. Uh and then the hardware engineers can vibe code their pieces because what they know about hardware engineering and the result is just like mindblowingly different productivity for small teams. like give an example like like if you're designing a turbine blade like classically so a turbine blade starts like uh cold but when it runs it's hot so it gets bigger and so you have to design both the aerodynamics and the structural design of the thing to work on it cold shape and this hot shape and so you have to convert between cold and hot and you convert between structures and aerodynamics and this takes like one engineer one day for one blade for one piece of the analysis and there are like a thousand blades in a jet and and so you can't do much and we literally now with a combination of software and hardware people created the solution you can change blade geometry you can see in real time the structures and aerodynamics results and so it allows two engineers to design an entire jet engine which is just wildly different.

Host

你提到的一点是,软件工程师为其他工程师创建工具和架构。对我来说,这是企业软件最大的颠覆:不再有初创公司能卖给你硬件协作工具,因为内部你随时都在编码你需要的东西。甚至电子表格也过时了,对吧?电子表格成功的原因是没人能构建定制软件。所以最接近定制软件的东西就是带有 VBScript 函数的电子表格。我个人几乎完全从 Excel 转向了 Python 模型,在那里我可以得到可信的模拟。

One of the things you mentioned is that you have software engineers creating the tools and architectures for the rest of the engineers. That to me is the biggest u the cataclysm of enterprise software is that there's no like startup that builds hardware collaboration tools that can sell you anything anymore because in internally you're just coding the right things that you need at any given time. Even spreadsheets are kind of cooked, right? Because the reason spreadsheets were successful is that no one could build custom software. So the thing that approximates custom software the most is a spreadsheet with a bunch of EV script functions. I personally have moved almost entirely from uh Excel to Python models where I can actually like get like believable simulations of things.

Blake

是的。我认为 AI 还没有做到但未来一年内(大概 26 年)会实现的事情非常令人兴奋:现在它能生成软件,但很快它将能生成 STEP 文件和 PCB 布局。

Yeah. I mean the thing that that AI hasn't come to yet that I think it it will within the next year like probably within 26 that will be very very exciting is right now it can generate software but soon it'll be it will generate step files and PCB layouts.

4. AI 对硬件和软件的影响 Hardware and software impact of AI

Host

至于机械和电气工程,那将是另一件我们尚未见过的事情。那会非常非常酷。

And when it comes to mechanical and electrical engineering, that will be a whole other thing that we haven't seen yet. That'll be very, very cool.

Naval

是的。在硬件方面,我认为这对所有那些软件写得很差的小工具公司和零件公司来说确实是个福音,因为它们做不出好软件,而现在它们将能够做出足够好的软件,或者甚至可能不是人类前端的软件。它可能完全是智能体式的,供智能体访问,你只需通过语音和控制硬件来操作。这也是我认为中国现在大力投入开源模型的原因之一。它们基本上全力以赴,因为它们拥有硬件优势。它们有非常复杂的供应链和组件链,它们基本上在说:‘嘿,如果我能按需生成软件,那么我就不再对硅谷有这种劣势了。’所以这不是它们做开源的唯一原因。我认为它们也落后了。它们在蒸馏模型。它们在追赶,你知道,它们在协作资源。但我认为中国政府有资助项目的传统,这些项目会帮助整个生态系统发展,尤其是在网络效应业务中。

Yeah. On the hardware side, I think it's really a boon for all these little gadget companies and part companies that write really bad software because they can't make great software and now they're going to be able to make good enough software or it may not even be software that is a human front end. It might just be completely agentic for an agent to access and you just talk to it through voice and control hardware. And this is one of the reasons why I think for example China is big into open-source models right now. They're basically going all in on it because they have hardware superiority. They have these very complex supply chains and component chains and they're basically saying, 'Hey, if I can just generate software on demand, then I don't have this disadvantage anymore against Silicon Valley.' So that's not the only reason why they're doing open source. I think they're also behind. They're distilling models. They're catching up, you know, they're collaborating on resources. But I think the Chinese government has a history of funding efforts that then sort of help their entire ecosystem along, especially in network effect businesses.

Host

所以我认为它们想调动所有资源,在 AI 上迎头赶上,并用它来给它们的硬件带来优势。讽刺的是,它们在做所有这些开源的事情,因为 OpenAI 并不开放。你知道,Grok 发布模型,但我认为它们落后一两个模型。谷歌有一些本地模型,但没有什么真正有竞争力的。Anthropic,据我所知,我甚至不知道它们有任何开源模型。所以,所有的开源重担都来自中国。这帮助了我们所有的硬件创始人,但更帮助了它们的硬件创始人和工厂等等。但所有那些你从亚马逊买来在懒散的周六下午摆弄的小玩意儿和零碎东西附带的糟糕软件,正在迅速变得好很多。我认为每个人都意识到了,没有优秀的前沿编码模型,你就无法自我改进。所以想象一下中国作为一个整体没有能力生产前沿的一切,对吧?不仅仅是生产软件;在硬件管道的任何环节,就像 Blake 说的,你需要生成软件。如果你在生成软件的能力上落后,你就会在生成一切的能力上落后。

And so I think they want to pull all their resources, catch up on AI, and use it to give their hardware stuff an advantage. And ironically, they're doing all the open source stuff because OpenAI is not open. You know, Grok publishes models, but I think they're a model or two behind. Google has some local models, but nothing really that competitive. Anthropic, to my knowledge, I don't even know of any open source models from them. So, all the open source heft is coming from China. It helps all our hardware founders, but it helps their hardware founders and factories and so on that much more. But all the crappy little software that goes with all the little random knickknacks and thingamajigs that you buy off of Amazon for to tinker with a lazy Saturday afternoon, that software is getting a lot better very quickly. I think everyone's had the wakeup call that without great frontier coding models, you don't have self-improvement. And so imagine China as a whole not having the ability to produce frontier everything, right? It's not just producing software; it's in any piece of this hardware pipeline like Blake was saying, you need to generate software. If you fall behind on your ability to generate software, you fall behind on the ability to generate everything.

Naval

我很好奇一件事,因为每个人都喜欢谈论中国模型,你们用中国模型吗?你们认识用中国模型的人吗?这其实是我昨天的一个争论,晚餐时有一个人声称,你会用 DeepSeek 处理 97%的事情,因为它太便宜了,如果你需要更多智能,你就反复运行同一个问题,你只会用 OpenAI、Anthropic 等模型处理最先进的任务。我当时想,我不知道,我认为智能是纯粹的善。你总是想要更多智能,当这些模型犯错时你并不知道,而且它总是比真人更便宜、更实时。所以,你只会用最智能的可用模型,这不一定是个好消息,因为这意味着你最终会在 AI 中造成垄断或寡头局面。但我总是想要最智能的程序员。我总是想要最正确的答案。我总是想要最好的判断。考虑到我将通过资本、代码、人和营销投入的杠杆,我希望每次都做出正确的决定。而且通常在两个模型之间,比如说我有一个模型我知道比另一个稍微聪明一点,它们都给出答案。通常我实际上不知道哪个是正确的,对吧?所以如果我知道一个模型稍微聪明一点,我会选择那个答案,最终我会停止问我认为不那么聪明的模型。但我不知道,你们有没有发现这些所谓不那么智能的模型的用途?

One thing I'm curious about from you guys is, because everyone loves to talk about Chinese models, do you use Chinese models? Do you know anybody that uses Chinese models? This is an argument I had yesterday actually, which is one person at the dinner table was claiming that you'll just use DeepSeek for 97% of things because it's so cheap, and if you need more intelligence you'll just run it over and over again on the same problem, and you'll only use the OpenAI, Anthropic, etc. models for the most advanced tasks. And I was kind of like, I don't know, I think intelligence is an unalloyed good. You always want more intelligence, and when these models make a mistake you don't know it, and it's always cheaper than a real person and real time. So, you'll just use the most intelligent model available, which isn't great news necessarily because it means that you're going to end up creating a monopoly or oligopoly kind of situation in AI. But I always want the most intelligent programmer. I always want the most correct answer. I always want the best judgment. And given the amount of leverage that I'm going to pour into it through capital and code and people and marketing, I want to make the right decision every time. And often when between two models, let's say I have one model that I know is a little smarter than the next one and they both give me answers. Often I actually don't know which is the correct answer, right? So if I know one model's a little smarter, I'm going to go with that answer and eventually I'm going to stop asking the model that I think is less intelligent. But I don't know, have you guys found a use for these so-called less intelligent models?

Host

我们看到有用途,因为我们有 AI 网关数据,基本上每个应用、智能体等都经过它。所以肯定有开源模型的使用,但顶尖的仍然被前沿智能主导。还有一个子类别或一个警告,那就是合理成本和性能下的前沿智能在大规模上很出色。所以人们不会对 Gemini 特别兴奋,但它们推出的这些模型在正确的性能成本组合下非常智能,而且对于编码以外的许多任务,有趣的是,它们是最好的模型。它们就像最好的工业生产模型。你可以把它们用于支持任务或浏览器自动化。比如我总是会把 Gemini 模型放在那里。我会为这类事情寻找中国模型。但任何时候我在推动前沿,你需要最好的编码模型。那基本上现在就是两三个模型,而中国肯定不在其中。

We see uses so that we have the AI gateways data that basically every application, agent, etc. goes through. And so there's definitely usage of open models, but the top is like heavily dominated by the frontier intelligence. And there's a subcategory or there's like a caveat to that, which is that frontier intelligence at reasonable cost and performance like slaps at scale. So like people don't get really excited about Gemini, but they put out these models that are like super smart at the right performance cost combination and for a lot of tasks other than coding actually interestingly enough. They're the best models. They're like the best industrial production models. You can throw them at like support tasks or browser automation. Like I would always put a Gemini model there. And I would look to Chinese models for those kinds of things. But anytime I'm working to push the frontier, you need the best possible coding model. And that's basically now like two or three models and the Chinese are certainly not in it.

Host

嘿,Max,你在大力推动垂直整合和极度紧迫感。你想谈谈这个吗?

Hey Max, you're pushing pretty hard into vertical integration and extreme urgency. Do you want to talk about that?

Naval

是的,我的意思是,很多事情我们买不到,所以你必须自己制造。我们的偏好总是购买,比如如果有供应商以优惠价格提供服务,比如 PCB,我们不制造 PCB,它们基本上是免费的,你可以从亚洲无限量购买。但我们的产品越接近一块共价键合的物质,它们就会越好:功耗更低、体积更小、性能更高、寿命更长。而且就是没有现成的组件。为了进行那种集成,为了能够真正创新,而不仅仅是拼凑现成的东西,那真的非常有限,我想你必须学会自己去做。这表现为垂直整合。所以我们在东海岸拥有一个自有的 MEMS 代工厂,我们买下它是因为没有其他办法做我们想做的封装和组装工作。我认为所有这些将在未来几年受到 AI 的严重影响。现在还没到那一步。

Yeah, I mean for many things we can't buy it so you got to make it somehow. Our preference would always be to buy something, like if there's a vendor that offers a service at a great price, like for example PCBs, we don't make PCBs, those are basically free, you can buy them in unlimited quantity from Asia. But the closer that our products get to being like a single block of covalently bonded matter, the better they'll be: lower power, smaller, higher performance, last longer. And there's just like the components aren't available. And in order to do that type of integration, to be able to actually innovate beyond just piecing together things that you can buy off the shelf, which really is very limiting, I guess you have to learn it to do it yourself. And that shows up as vertical integration. So we own a captive MEMS foundry on the east coast which we bought because there was really no other way to do the type of packaging and assembly stuff that we wanted to do. And I think that all of this is going to be affected heavily by AI over the next few years. It's not quite there yet.

5. 监管与合规中的 AI AI in Regulatory and Compliance

Naval

事实上,讽刺的是,我们在公司内部看到的 AI 最大的影响之一是在监管互动方面,因为如果我们能做像生成文档这样的事情,或者如果我们能问,比如我们想改变,我们想改进这个产品,有成千上万的 ISO 标准可能适用,我们必须遵守哪些,并追溯这些。过去这需要你跟着整个监管质量团队好几个月,他们追溯这些,现在 AI 基本上就知道了。但当我想到像外科手术项目或 MEMS 晶圆厂这样的事情时,我认为最终软件仍然需要手。它会比我们更聪明,但如果它不能制造东西,那么这些就是真正的边界。所以我们已经在我们的晶圆厂以及公司的许多其他部分进行了仪器化,随着这些模型变得更好,这应该会立即体现在我们正在做的细胞工程和我们正在开发的材料科学上。

In fact, ironically, one of the biggest impacts that we've seen of AI inside the companies is in regulatory interactions, because if we can do things like generate documentation, or if we can ask like we want to change, we want to evolve this product, like there's thousands of ISO standards that might apply, which ones do we have to comply with, and trace this through. This used to be like you're following a whole regulatory quality team for several months as they trace this, and now the AI just kind of knows. But when I think about stuff like the surgical program or the MEMS fab, I think ultimately the software still needs hands. Like it's going to be smarter than us, but if it can't make things, then those are real boundaries. And so we've instrumented our foundry as well as many other parts of the company in ways where, as these models get better, that should show up pretty immediately in things like the cell engineering that we're doing and the material science that we're developing.

Host

这让我意识到,我已经有一段时间没有用律师来生成基本法律文件了,对吧?我不再让律师起草保密协议,或者签这个、研究那个,所有基本的法律任务也消失了,因为你知道那个老笑话,法律就像意大利面条式的代码,他们有一套非常复杂的代码,试图用英语表达,却与这里的代码矛盾,必须适应那里的代码,而且没有真正的 API。但对于初级工程师和初级工程来说,我应该说初级工程师基本上晋升为高级工程师,而初级工程被智能体接管了。同样,我认为从某种意义上说,坏处是你可以看待法律,说律师助理刚被解雇,或者你可以说律师助理刚晋升为高级律师,现在他们可以花时间思考法律。思考软件工程如何与律师一起演变实际上很有趣,因为律师,你永远不知道他们到底在这些文件中放了什么。你只是信任他们。比如,嘿,律师,你能看看这份文件吗?你能告诉我它是否合法吗?你能做红线标注吗?随便。归根结底,你与律师关系中所看重的是他们是一个值得信赖的权威。他们上过法学院,并且把自己的声誉押在上面。

It sort of makes me realize that it's been a while since I've generated a basic legal document using a lawyer, right? I stopped asking lawyers for NDAs and you know agreement for this and sign that and research this, and like all the basic legal tasks are gone too, because you know the old joke that law is like spaghetti code, you know they have this very complicated code that they try to put in English and it contradicts this code over here and has to fit into that code over here and there's no real APIs for it. But for just like junior engineers and junior engineering, I should say junior engineers basically got a promotion to senior engineers, and junior engineering got taken over by agents. So the same way, I think in a way the downside is you can look at law and say you know paralegals just got fired, or you could say paralegals just got promoted to senior lawyers and now they can spend their time thinking about the law. It's actually kind of interesting to think about the parallels of how software engineering is evolving with lawyers, because lawyers, you never know what they put into these documents exactly. You just trust them. Like, hey lawyer, can you look at this document? Can you tell me if it's legit? Can you do red lines? Whatever. Like, at the end of the day, what you're valuing in the relationship with a lawyer is that they're a trusted authority. They went to law school and they're putting their reputation on the line.

Naval

我认为这与当今软件工程中最大的问题有相似之处:那些最终变成 PR 的庞大代码堆,然后人们说,就像 Twitter 上所有的梗一样,过去我们读 PR 的每一行代码。嗯,在我的基础设施世界里,我希望工程师能够说“我理解”并不一定意味着你读了 PR 的每一行。你需要能够说“我签字确认理解这个 PR 的后果”,或者我编写了测试工具、模拟、证明、类型检查器等,以便能够说即使没有阅读这个,我有信心签字确认它在生产中是安全的。所以这很有趣,因为有一个世界,我们接受一切都会是意大利面条式的代码,我们不完全理解它,但我们编写了给我们信心的评估器,然后我们依赖像基础设施生产工程师这样的人来说“好的,我同意把它投入生产”。你知道,归根结底,如果你的系统宕机,会有人被传呼。我认为人们低估的另一件事是,创建软件从零到一很容易,但想想一千天后,你的软件是什么样子?它安全吗?经过测试吗?是生产级的吗?性能好吗?你还有动力投入所有这些 token 来维护它在生产中吗?

I think there's a parallel with the biggest problem in software engineering today: these mountains of slob that end up as a PR, and then people say, like there's all these memes on Twitter like way back in the day we used to read every line of code of a PR. Well, in my world infrastructure, I want engineers to be able to say I understand doesn't necessarily mean that you've read every line of the PR. You need to be able to say I am signing off on understanding the consequences of this PR, or I wrote the test harness, the simulations, the proofs, the type checkers, etc., to be able to say even without reading this I have confidence I can sign off on it's going to be safe in production. And so it's kind of interesting because there's a world in which we embrace that everything is going to be spaghetti code and that we don't fully understand it, but we write the evaluators that give us confidence, and then we rely on people like the infrastructure production engineers to say okay I'm fine sending this into prod. You know at the end of the day, someone is going to get paged if your systems go down. And I think another thing that people are underestimating is that creating software is really easy zero to one, but think about a thousand days from now, what does your software look like? Is it secure? Is it tested? Is it production grade? Is it performant? And are you still motivated to invest all of those tokens in maintaining it in prod?

Host

我的意思是人类正在成为验证者,对吧?这就是我们如何用好的验证数据训练这些模型,现在我们需要人类验证者。所以是的,我认为很多人的旧职能,律师、工程师、运营人员,转移到验证整个堆栈,并说“是的,这大致正确,我大致支持它,如果出了问题我会支持你”。

I mean humans are becoming verifiers, right? And that's kind of how we train these models with good verification data, and now we need human verifiers. So yeah, I think a lot of the old function of people, lawyers, engineers, operations people, move to verifying the stack and saying yeah this is roughly correct and I'll roughly stand behind it and I'll support you if it goes wrong.

Naval

我们看到的与监管相关的一件事是,它极大地减少了变革厌恶并改善了迭代。举个例子,假设你要认证一架飞机。你必须做的无数事情之一是证明它能承受雷击,而测试计划的监管文档长达 200 页。你通常的做法是雇用一个,老实说,不是特别聪明的工程师,他愿意在那里像键盘上的猴子一样写 200 页的合规文档。这需要几个月。顺便说一句,如果你现在改变飞机,你会想哭,因为又需要两个月来重写这些合规文档。我们发现我们可以构建一个 RAG,让我们基本上通过提示在几分钟内完成所有这些工作。一阶效应是,哦,你节省了很多时间。二阶效应是,如果你改变飞机的规格,现在只需要几分钟而不是几个月。所以你实际上愿意改变。三阶效应是,你现在基本上可以摆脱那些不太出色的工程师,而拥有少数真正有创造力的人。他们可以快速迭代,因为变更成本降低了,从某种意义上说,整个监管负担,这确实损害了迭代能力,消失了。我认为这是目前 AI 中一个被严重低估的故事。我认为硅谷的共识是监管很糟糕,我们想走得更快,我们想实现这个惊人的未来,我们想要富足,我们想要繁荣,而任何减缓那个未来的东西都应该避免。当然,我认为我们过度监管了。我们让建造东西变得不可能。在很多地方,建造任何类型的东西,无论是物理的还是其他的,所需的工作简直疯狂。但你知道,很多监管本身并不是问题。如果你真的读过很多这些东西,比如没有雾霾的城市很好,能在许多河流中游泳很好。很多这些东西都是进步。

One of the things we see related to the regulatory is it massively reduces change aversion and improves iteration. So to give you an example, let's say you're going to go certify an airplane. One of the zillions of things you have to do is prove that it could withstand a lightning strike, and the regulatory documentation for the test plan for such a thing stretches on for say 200 pages. And what you would classically do is hire a, let's be honest, not super bright engineer who's willing to be there monkey at keyboard writing 200 pages of regulatory compliance documentation. And it takes a couple months. And by the way, if you change the airplane now, you want to cry because there's another like two months of rework of this regulatory compliance documentation. And what we found is we can build a RAG that will enable us to basically prompt our way through all of that work in let's call it minutes. The first order effect is oh you save a lot of time. The second order effect is if you change the specification of the airplane, it now takes minutes not months. So you can actually be willing to change. And the third order effect is you can now basically get rid of the not very great engineers and have a small number of really creative ones. They can iterate rapidly because the cost of change goes down, and in a certain sense the entire regulatory burden which really hurts the ability to iterate drops away. I think that this is a really undersold story in AI right now. I think the consensus in Silicon Valley is that regulation sucks, like we want to go faster, we want to realize this amazing future, we want abundance, we want prosperity, and stuff that slows down that future is just kind of to be avoided. And certainly I think we've overregulated. We've made it impossible to build stuff. It's just totally crazy what goes into getting building any type of thing in a lot of places, either physical or otherwise. But you know, a lot of the regulations themselves are not the problem. Like if you've actually read a lot of these things, like having non-smog choked cities is great. Being able to swim in many rivers is great. Like a lot of these things were progress.

6. 监管摩擦与 AI 代理 Regulatory friction and AI agents

Naval

问题在于,人类要理解和遵守这些规定非常困难,每次和政府部门通信都要等上几个月。如果你能把我们学到的很多东西变得完全无摩擦,那实际上会非常酷。我认为这是目前 AI 领域一个被低估的故事。

The problem is that it is really difficult for humans to deal with understanding and complying with this and that every time you have to exchange a letter with the government, you wait months. And if you could take a lot of the things that we've learned and kind of make them totally frictionless, that would actually be pretty cool. And I think that is an under-told story in AI right now.

Host

是啊。直到监管机构开始向我们吐出大量文本,然后你开始收到大量需要遵守的文件,这就成了智能体之间的战争。

Yeah. Until the regulators start spewing tokens back at us and then you start getting huge amounts of documents from the regulators that you have to comply with, and it's agent on agent wars.

Naval

但这基本上就是我们现在的情况。

But that's basically what we have now.

Host

是啊,但这是一场公平的战斗。

Yeah. But it's a fair fight.

Naval

我认为这比我们现在的情况有所改进。现在糟糕的事情之一是,如果你建造任何实体建筑,你必须获得建筑许可证。这就像你有罪直到被证明无罪。我们遇到的最糟糕的是消防部门,因为他们有从燃烧的建筑中救人的道德权威,但实际上他们只是几个月地折腾你的建筑设计。如果我们能用智能体取代消防队长,快速批评你的建筑计划,即使它的反馈过度,也会比今天的延迟好得多。

I'd argue that's an improvement from where we are now. Like one of the terrible things right now is if you build anything physical, you have to get a building permit. It's like you're guilty until proven innocent. And the worst thing we've run into is the fire department because they have like the moral imprimatur of people pulling people out of burning buildings. And yet what they actually do is just mess with your design for buildings for months. And if we could replace the fire marshal with an agent that would critique your building plan quickly, even if its feedback was overdone, it would be massively better than the delays that exist today.

Host

当 Max 谈到所有这些监管可能是一件好事时,我想到了让智能体成功的因素:人类或其他智能体设置正确的测试护栏。很多人对 SLGOLO 感到兴奋。我不知道你们有没有玩过那个,或者像 Ralph 循环那样,你告诉模型‘去做这个,这是你的退出标准’。好吧,我告诉 Blake:让我们都变成超音速。你的退出标准是你已经遵守了所有这些规定。所以完全存在一个世界,我们说监管很好。它们就像我们的测试套件。只要它通过了这些测试,没有引发矛盾,而且规定实际上合理等等,它们实际上是一个很棒的护栏。否则,我们就会直接把垃圾产品推向市场。

When Max was talking about this potentially being a good thing to have all this regulation, my head went to the things that make agents successful: humans or other agents setting up the right testing guardrails. A lot of people are really excited about SLGOLO. I don't know if you guys have played with that or like Ralph loops where you tell the model 'go do this and this is your exit criteria.' Well, I'm telling Blake: go make us all supersonic. Your exit criteria is that you've complied with all of these regulations. So there's totally a world where we say the regulations are great. They're like our test suite. As long as this is passing these tests for one that's not incurring contradictions and the regulations are actually reasonable, etc. They're actually an awesome guardrail to have. Otherwise, we would be shipping slop directly into the air.

Naval

是啊,但这会变成一场红皇后竞赛,对吧?他们会有智能体,我们也会有智能体。我认为我们可能有更好的智能体,这很好,而不是人类对人类的对抗,但无论如何,他们的周期时间、响应时间可能会更低。就像应用商店被垃圾信息淹没一样。我敢肯定专利局现在也被垃圾信息淹没了。所以这些机构会缓慢地采用 AI。他们会被聪明的企业家用大量文件进行 DDoS 攻击。这些东西的审批时间可能会因为突然的泛滥而延长。这创造了真正改变监管模式的机会。想象一下,如果我们像今天建造东西一样在城市里开车。在你开车去任何地方之前,你必须写一个计划,寄给某个监管机构,你的计划必须说明我们要走某条路线,以这个速度行驶,使用转向灯,在每个停车标志前停车,永远不闯红灯,等等。然后三个月后,你收到批评意见,比如‘我们认为你应该走另一条街’。最终你获得批准。你永远无法开车去任何地方。这太疯狂了。你哪儿也去不了。然而,这正是我们在美国建造实体基础设施的方式。有罪直到被证明无罪。我们实际上应该做的是让更多这些东西基于执法而非预先批准。

Yeah. But this is going to turn into a red queen race, right? They're going to have agents. We're going to have agents. I think we might have better agents, which is good, as opposed to having to do human versus human, but if anything, their cycle time, their response time may get lower. Like the app store is drowning in spam. I'm sure the patent office right now is drowning in spam. And so these agencies, they're going to be slow adopters of AI. They're going to get DDoSed by clever entrepreneurs just overloading them with documents. It's possible that the approval time for this stuff might extend out as this suddenly gets flooded. It creates the opportunity to really shift the regulatory model. Imagine if we drove around a city the way we build things today. Before you could go anywhere, you'd have to write a plan up, ship it to some regulator, and your plan would have to specify we're going to take such and such a route and we're going to drive this speed limit. We're going to use our blinker and we're going to stop at every stop sign and we're never going to run a red light, blah blah blah. And then 3 months later, you get back critique. It's like, well, we think you should drive on this other street. And eventually you get approval. You never go drive somewhere. It's insane. You can never go anywhere. And yet that is absolutely the way we build physical infrastructure in this country. It's guilty until proven innocent. And what we should actually do is make more of these things enforcement-based rather than pre-approval based.

Host

我的意思是,我不知道。我不想太过分。比如,如果我向很多人发货医疗设备,那需要……那里有未知因素。就像我们负责任一样。我们做了临床试验。我们报告了所有数据。

I mean, I don't know. I don't want to be under too much. Like, if I ship a medical device to a lot of people, there needs to be... it's like there's unknowns there. It's like we were responsible. We did clinical trials. We reported all the data.

Naval

但是 Max,这就是为什么现在医疗领域创新如此之少,因为 FDA 的审批过程是一场噩梦。事实上,过去十年硅谷科技领域最大的两个进步,AI 和之前的加密货币,都在数学领域,因为这是最后一个不受监管的领域。当他们开始监管前沿模型和 GPU 时,那也会停止。你知道,Peter Thiel 感叹物理领域没有创新。正是巨大的监管壁垒阻碍了创新,你总能找到像疫苗或医疗这样的可怕例子,对吧?但监管无处不在。触角无处不在,有各种不同且矛盾的监管机构。你看到 SpaceX 先是因为没有足够的……我忘了是什么,移民或难民之类的被起诉,但另一方面政府规定他们不能雇佣这些人,因为他们不是公民。这不像逻辑代码,必须在同一个地方编译。这些都是到处随意编造的规定。你可能遵守了一个州的规定,却违反了另一个州的规定,在这里违反了联邦规定,惹恼了这个人,那个人选择起诉 50 个人中的一个,而那个人是他的朋友。这非常武断,非常反复无常。

But Max, this is why there's so little innovation in medical right now because the FDA approval process is a nightmare. In fact, the two biggest advancements in tech in Silicon Valley in the last decade, AI and before that, crypto, they're both in the math domain because it's the last unregulated domain. And when they started regulating frontier models and started regulating GPUs, that stops as well. You know, Peter Thiel laments about how there's no innovation in the physical domain. What's been held back by just the huge regulatory barriers and you can always find a scare version like vaccines or medical like famous ones, right? But the regulations spread everywhere. The tentacles are everywhere and there's all these different contradictory regulatory bodies. You saw how SpaceX got sued first for not having enough... I forget what it was, migrants or refugees or whatever, but they're not allowed to hire them by government regulation on the other side because they're not citizens. This is not like logical code that has to compile in one place. These are made-up random regulations all over the place. You might comply with one state, you violate another state, you violate federal over here, you annoy this guy over here, that guy chooses to prosecute one out of 50 people who are his friend. It's very arbitrary. It's very capricious.

Host

而且,认为这会让事情更安全的想法,我认为完全是个神话。就拿波音来说吧。他们认证了 737 Max,这架飞机有一个单一的传感器,对飞机的俯仰姿态拥有完全控制权。没有一个实习生会蠢到认为这是个好主意。然而,它却一路通过了认证系统。这些东西实际上并没有让我们更安全,只是让我们更慢。

And moreover, the idea that this makes things safer, I think it's just a complete mythology. Like just watch Boeing as an example. They certified the 737 Max, which had a single sensor that had complete authority over the nose up, nose down attitude of that airplane. No intern is dumb enough to think that's a good idea. And yet, it got all the way through the certification system. This stuff doesn't actually make us safer. It just makes us slower.

Naval

嗯,我的意思是,这里肯定有功能失调。我认为其中一些确实让我们更安全,比如 NRC 让我们更安全,他们的工作是确保核能安全。他们通过自 70 年代以来直到大约一年前都不批准任何核电站来实现这一点。如果我们从不建造任何核电站,那当然是绝对安全的。我想明确表示,在很多方面我支持放松监管。我同意 Blake 的观点,很多事情可以更高效地完成。但我也认为,仅仅说‘哦,这就像 FDA’或者笼统地归咎于机构,有点过于轻率了。问题更深层,以至于当 FDA 批准了 10 种非常重要的药物时,他们得不到任何赞誉。一个病人死了,他们就被拖到国会面前挨骂。

Well, I mean, there's definitely dysfunction here. I think that some of this makes us safer in the sense that the NRC makes us safer, which is that their job was to make sure that nuclear energy was safe. They did this by permitting zero plants until I think like a year ago since the 70s. It will be perfectly safe if we never build any of it. And I want to be really clear that I'm on the side of deregulation on a lot of this. I agree with Blake that a lot of this can be done a lot more efficiently. But I also think it's a little too dismissive just to say it's like 'oh this is like the FDA' or even it's in the agencies in general. And the problem is deeper to the degree that when the FDA approves 10 really important drugs, they don't get any credit for that. One patient dies and they get hauled before Congress and yelled at.

7. 监管中的不对称激励 Asymmetric incentives in regulation

Naval

所以这里的激励是严重负偏的。我认为现实是,这反映了美国人民的信念。在人体研究中的风险感知和新药上市速度之间存在权衡。如果我们行动更快,我们确实会学到更多。这完全是不对称的:如果你批准了一件坏事,你的职业生涯就完了;如果你阻止了一件好事,没人会注意到。这就造成了不对称的放缓。我认为这是监管体系中最需要解决的问题。

And so they have very negatively biased incentives here. And I think the reality is that this is reflective of the beliefs of the American people. There's this trade-off between the perception of risk taken in human subjects research and the rate at which we get new medicines. It is absolutely true that if we move faster, we would learn. It's totally asymmetric. If you approve a bad thing, your career is over. If you block a good thing, nobody notices. So it creates this asymmetric slowdown. I think that is the most important problem to solve in the regulatory state.

Host

但这是一个非常深刻的问题,因为这就是选民的立场。我们对我们正在做的一些事情进行民意调查,以了解美国人民的态度。如果你逼得太紧,有很多方法可以绕过它。你可以去繁荣。有很多方法可以尝试加快速度。但如果你被视为不良行为者,你就会被社会排斥。这才是你需要回答的问题,比仅仅说“哦,我们需要监管改革”要深刻得多。

But this is a very deep problem because this is where the voters are. We poll some of the stuff we're working on to understand where the American people are on it. If you push too hard, there are all kinds of ways to work around it. You go to prosper. There are all kinds of ways to try to go faster. But if you're seen as a bad actor, then you're rejected from society. That is the thing you need an answer for, which is deeper than just saying, 'Oh, we need regulatory reform.'

Naval

你说得很深刻,Max。就是选民,对吧?

You have a deep point there, Max. It's the voters, right?

Host

这就是公民的立场。我们喜欢责怪政客。你在 X 上总能看到人们说“哦,这个政客,那个政客”,但他们是选举出来的,多数票。这就是人民真正所在。那就是一揽子方案。那就是他们选择的组合。你可能不喜欢这种配置,但如果你移除这一个,会有非常相似的东西取而代之,因为选民会把他们再选回来。在文化上,大多数人很难理解我们失去了什么,错过了什么。例如,法国:有一个法国企业家在 X 上哀叹 57%的 GDP 被政府吸走,所以你无法创建公司。但对普通法国公民来说,这是看不见的。他们没有注意到他们错过了什么。他们只知道他们比美国稍微穷一点。《经济学人》刚刚发表了一篇文章,说美国如何超越所有人,增长更快。但他们立刻又说这是因为海洋、自然资源,除了资本主义之外的一切。他们不想说那个肮脏的“资”字。不知为何,所有这些杂志在某个时候都变成了马克思主义者。他们无法想象如果我们再自由一点、再开放一点会是什么样子。

This is where the citizens are. We like to blame politicians. You'll see on X all the time, people say, 'Oh, this politician, that politician,' but they're elected, majority vote. This is where the people literally are. That's the package. That's the bundle they've chosen. You may not like this configuration, but if you were to remove this one, something very similar would take its place because the voters would just vote them right back in. Culturally, it's very hard for most people to understand what we lost, what we missed. For example, France: there's a French entrepreneur on X lamenting that 57% of GDP gets sucked up by the government, so you can't create companies. But to the average French citizen, that's not visible. They don't notice what they're missing. They just know they're slightly poorer than the US. The Economist just did a piece on how the US is outstripping everybody and growing faster. But then they immediately turn around and say it's because of the oceans, natural resources, everything but capitalism. They don't want to say the dirty 'c' word. For some reason, all these magazines became Marxist at some point. They can't envision what could have been if we had just been a little more laissez-faire, a little more open.

Naval

所以我希望看到在 50 个州之间进行真正的实验:不同的法规,不同的税收结构。不是说现在联邦税收结构和联邦法规主导一切,但想象一下,如果你得了癌症,你可以去某个小州,尝试每个人以“买者自负”的方式研发的每一种药物,你可以自己做研究。这就是所谓的实验区。同样适用于无人机,同样适用于——嗯,飞机有点难,因为你要穿越很多区域。

So I would love to see a true experiment among the 50 states: different regulations, different tax structures. Not because right now the federal tax structure and federal regulations dominate everything, but imagine you could go to some small state if you had cancer and you could try every drug that everyone was cooking up in a caveat emptor way, and you got to do your research. But this is known as the experimental zone. Same way for drones, same way for well, aircraft is a little harder because you got to cross a lot of areas.

Host

你认为创新区的概念有什么神奇之处吗?因为我们有一个巨大的邻避问题。但如果你创建了自愿加入的迎臂区,它们就创造了一个实验框架。根据定义,它发生在人们同意的地方,你可以尝试不同的规则或没有规则,或者不同的执行方式,或者无罪推定,然后看看实际发生了什么,创新后果是什么,安全后果是什么。然后成功可以传播。但就 Naval 的观点而言,创新区并不能解决药物发现的问题。不久前通过了《尝试权法案》。我们拥有“单患者 IND”这条路径已经很久了。FDA,如果你的医生打电话说“嘿,我想给我的病人用这个未经批准的药物”,他们批准了超过 99%的请求。他们甚至可以在电话里批准。问题是,要给患者用药,你仍然需要临床级别的药物,而唯一拥有它的实体通常是正在进行临床试验的知识产权所有者。他们投入了数亿美元来制造这个东西。问题是,如果你的病人——可能本来病得很重——出了什么事,FDA 会做出不利推断。这将被视为该药物的一个属性,是全球性的,与你的创新区无关。所以有两个问题:一是你需要让知识产权所有者给你一些他们的药,他们不会这么做;二是你需要防止全球监管机构对他们给你药后临床试验可能发生的情况产生怀疑。

Do you think there's something magical in the notion of innovation zones? Because we have a huge NIMBY problem. But if you create opt-in YIMBY zones, they create an experimentation framework. By definition, it happens where people are consenting, and you can try different rules or no rules, or different ways of enforcing, or innocent until proven guilty, and then see what actually happens and what are the innovation consequences and what are the safety consequences. Then the successes can spread. But to Naval's point, an innovation zone would not solve the problem in drug discovery. There was the Right to Try Act passed a little while ago. We've had this pathway called Single Patient IND for a lot longer than that. The FDA, if your doctor calls and says, 'Hey, I want to give this to my patient, an unapproved drug,' they approve over 99% of those. They can even grant them over the phone. The problem is that in order to dose a patient, you still need clinical grade drug, and the only entity with that is typically the IP owner who's in the middle of running a clinical trial. They're investing hundreds of millions of dollars into making this thing. The problem is that the FDA will draw an adverse inference if something bad happens to your patient, who's probably really sick to begin with. That's going to be seen as a property of the drug which is global, not related to your innovation zone. So there are two problems: one, you need to get the IP owner to give you some of their drug, which they're not going to do; and two, you need to prevent the global regulator from casting doubt on what might happen with their clinical trial if they give you some.

Naval

在医学上你会怎么解决这个问题?

How would you address that in medicine?

Host

哦,这个特别具体,非常内行。我认为必须禁止 FDA 对不同用户使用同一胶囊做出不利推断,例如。有很多具体的方法可以通过相对轻度的监管来真正加速创新,只要防止这种偏执驱动我们的决策。

Oh, well, that in particular is just very inside baseball. I think the FDA has to be prohibited from drawing adverse inferences across different users of a capsule, for example. There are a bunch of specific ways that you could really accelerate innovation with a relatively light regulatory touch by just preventing this paranoia from driving our decisions.

Naval

有没有比 FDA 更好的东西?我们用什么基准来比较这些监管机构?还是说这不是一个有趣的问题,因为每个人都跟着 FDA 走?

Is there anything better than the FDA out there? What are we benchmarking these regulators against? Or is it not an interesting question because everyone follows the FDA?

Host

我对此有两个扩展。第一个是欧洲,它并不比 FDA 好,但他们的系统不同。他们有这些公告机构,基本上是由东道国政府认可的私营企业,负责认证火车、飞机或医疗设备等。公告机构系统在审查层面创造了稍好的激励,因为他们可以雇佣人员,可以发展,公告机构之间存在竞争。他们自己必须遵守东道国政府为认证设定的条件,但这意味着他们可以有比美国多出数千倍的审查员。

I'll give two expansions to that. The first is Europe, which is not really better than the FDA, but they have a different system. They have these notified bodies, which are basically private businesses that are blessed by their host governments to certify things, whether it's trains or planes or medical devices. The notified body system creates slightly better incentives at the review layer because they can hire people, they can grow, there's competition among the notified bodies. They themselves have to be compliant with the conditions placed by their host governments for certification, but it means that there can be many thousands more reviewers than you might have in the US.

8. 中国 BCI 与低成本 China's BCI and lower costs

Naval

我要说的第二点是,目前确实有一个获批的可植入式脑机接口(BCI)在中国,中国药监局(CFDA)有自己的考量。他们确实有一套系统,我认为如果我们不小心,它会让我们面临激烈竞争。而且他们处理的方式非常不同。

The second thing I'll say is there actually is one approved getting paid implantable BCI today which is in China and the CFDA is thinking for itself. And they really do have a system that I think is going to give us a run for our money if we're not careful. And they handle it very differently.

Host

他们怎么处理的?

How do they handle it?

Naval

我的意思是,将药物或设备推向市场的成本要低得多。你可以在人体上试验,也可以在市场上试验。所以我最近花了很多时间思考的问题是:20 年前,我们购买的笔记本电脑和手机少得多,每台都贵得多。现在它们更便宜了,数量也更多了,我们买得更多了,总支出却上升了。这很好。高通、三星、苹果等公司的股价大幅上涨。每个人都很高兴。他们用手机和笔记本电脑产生的额外财富来购买手机和笔记本电脑。这在医疗保健领域不会发生。在医疗保健领域,由于存在这种报销机制,就像企业销售一样,我们用于购买医疗保健的资金池基本上是固定的。它不会像技术增长行业那样,随着产生更好医疗效果的东西增多而增加。这意味着医疗保健支出的增长率大致与税收收入的增长率相同。所以,假设人工智能蓬勃发展,取得了重大进展,两年后我们在人工智能上的支出是现在的 10 倍,这可能是好事。但如果两年后我们在医疗保健上的支出是现在的 10 倍,那将是一场灾难。这与成为技术增长行业根本矛盾。随着时间的推移,有更多的东西可以花钱来延长和改善患者的生活质量,比如我们可以恢复 80 岁失明者的视力。我们也许能将寿命延长到远超以往。我们可以恢复年老体弱患者的能力。但你如何支付这些费用?医疗保健领域存在一个普遍问题,实际上都是同一个问题:将这些产品推向市场太昂贵了。这就是中国正在解决的问题。出路不是单一支付方或修改健康保险。而是降低成本,这样人们可以用信用卡购买,最坏的情况像买车一样分期付款。要做到这一点,我们必须让这些产品更便宜地推向市场。中国正在这样做。这将使他们能够以 1 万或 10 万美元的价格出售这些产品。医疗保健领域没有私人市场。因为没有私人市场,人们有时会用什么类比?想象一下,不去餐馆付钱,而是去所有餐馆,然后在月底把所有收据和账单寄给你的保险公司或政府,他们会报销你。那么,每家好餐馆外面都会排长队。每家差餐馆都会有空位。等待时间会很糟糕。产品不会改进。你基本上是在一个更大的资本主义社会内部运行一个小型共产主义社会。这就是我们在医疗保健领域所做的。

I mean the costs to bring a drug to market or a device to market are just much lower. I mean you can try things in humans and you can try things on market. So the problem that I've spent a lot of time recently thinking about is like 20 years ago we were buying far fewer laptops and phones, each one was much more expensive. Now they're cheaper, there are far more of them, we buy more of them, the total spending has gone up. This is great. Stock prices of things like Qualcomm and Samsung and Apple are way up. Everybody's happy. They're using the excess wealth generated by the phones and laptops to buy the phones and laptops. This doesn't happen in healthcare. In healthcare, because you've got this reimbursement mechanism in the way where there's this kind of enterprise sale happening, the bucket of money that we use to buy healthcare is basically fixed. It is not increasing as there is more stuff that is producing better healthcare outcomes like we see in technological growth industries. And so this means that the rate of spending on healthcare grows at roughly the rate of growth of tax receipts. And so if let's say that AI is booming and there are major advances happening and two years from now we're spending 10 times as much on AI as we are now, this could be great. But if in two years we're spending 10 times as much on healthcare, this would be a catastrophe. And this is fundamentally at odds with being a technological growth industry. And so as time goes on and there's more things to spend money on that extend and improve quality of life for patients, like we can restore vision to people who go blind in their 80s. We might be able to extend life far past where it's been before. We can restore capability to patients that are older and in worse condition. But how do you pay for that? There's this kind of omni problem in healthcare which is all really the same problem: it's just too expensive to bring these things to market. And that's what China is getting at. The way out of this is not single-payer or some revision to health insurance. It's to bring down the costs so that someone can buy this with a credit card, finance maybe like a car worst case. And to do that, we have to make it cheaper to bring these things to market. And China's doing that. That will allow them to sell these things for $10,000 or $100,000. There's no private market in healthcare. And because there's no private market, what was the analogy people make sometimes? Like imagine instead of going to restaurants and paying, you would basically go to all the restaurants and then at the end of the month, you would send all the receipts and all the bills to your insurer or the government and they would reimburse you. Well, there'd be a line outside every good restaurant. Every bad restaurant would be available. The waits would be terrible. The product wouldn't improve. You're basically running a small communist society inside a larger capitalist society. And that's what we're doing in healthcare.

9. 道路类比与医疗计划 Roads analogy and healthcare plan

Host

我们在道路上也是如此,这就是为什么会有交通堵塞。道路上的情况完全相同。这就是为什么高速公路没有可变定价。这就是为什么它总是拥堵。

It's also what we're doing on roads, which is why we have traffic. It's the exact same situation on roads. It's why there's no variable pricing for getting on the highway. It's why it's always clogged.

Naval

如果你想触碰医疗保健的雷区,想想这个医疗保健计划。告诉我它有什么问题。想象一下,你年收入的前 20%是你的医疗免赔额。没关系。如果你身无分文、无家可归,那就是零。如果你很富有,那就是数百万美元。但无论你的年收入是多少,前 20%是你的医疗免赔额。然后剩下的部分由政府或保险系统支付,直到他们今天设定的通常上限。你会很快创造一个私人市场。所以在牙科、整形外科和许多可选医疗程序中,你实际上会得到竞争局面。你会得到改进。如果你看看眼科中的 LASIK,牙科中的贴面和牙套以及所有牙科手术。或者如果你看看整形外科,这些领域似乎确实在进步,因为它们是私人支付者。人们用钱投票。所以我们需要在正常的医疗保健系统中做类似的事情。但人们会失去理智。他们不想超前思考。他们会说,‘不,不,不。那穷人怎么办?’好吧,穷人没有收入。所以他们又说,‘嗯,20%对有些人来说太多了。’好吧,你可以设置一些免赔额。但总的来说,如果没有一个私人市场,人们大部分时间都在为医疗程序付费,你就不会得到你所说的反馈循环。你就不会得到向系统投入更多资金的能力。现在,非常富有的人不能自愿向系统投入资金,但价格不存在。费率卡不存在。系统不是为此设计的。就像如果你去购买医疗服务,想自掏腰包,有时他们会给你报一个比他们向保险公司收取的价格高 10 倍的价格。

If you want to step on the third rail of healthcare for a moment, think about this healthcare plan. Tell me what's wrong with it. Imagine that the first 20% of your annual income was your healthcare deductible. It doesn't matter. If you're broke and homeless, it's zero. If you're rich, it's millions of dollars. But whatever your annual income is, the first 20% is your healthcare deductible. And then the rest is paid by the government or the insurance system up to the usual caps that they have today. You would create a private market pretty quickly. And so in dental and plastic surgery and a lot of optional medical procedures, you would actually get a competitive situation. You get improvement. If you look at optometry with LASIK, you look at dental with veneers and braces and all that dental surgery stuff. Or if you look at plastic surgery, those fields do seem to be advancing because they're private payers. They have people who are voting with their money. So we need to do some equivalent of that in the normal healthcare system. But people lose their minds. They don't want to think one step ahead. They're like, 'No, no, no. What about the broke person?' Well, the broke person has no income. So they're like, 'Well, 20% is too much for some people.' Okay, you can put some deductible in there. But generally, if you don't have some private market where people are paying a lot of the times for medical procedures, you're just not going to get this feedback loop that you're talking about. You're not going to get this ability to spend more money into the system. Right now, very wealthy people can't spend voluntarily into the system, but the prices aren't anywhere. The rate cards aren't anywhere. The system's not designed for it. It's like if you go shopping for medical care and you want to pay out of your pocket, sometimes they'll quote you a price that's 10 times what they charge the insurance company.

10. GitLab 的 Sid 故事 Sid's story from GitLab

Host

你听说过 GitLab 的 Sid 的故事吗?你知道 Sid 吗?

Have you heard Sid's story from GitLab? Do you know Sid?

Naval

所以,他是……

So, he was...

Host

他进行了非常成功的 IPO,然后被诊断出患有罕见癌症,并且活过了预后时间,真正地把命运掌握在自己手中。我想他从一线化疗开始,然后有一种替代疗法可用。他用尽了那种疗法,医生们说,‘我们无能为力了。’从那以后,我想大概有六七家公司从中诞生,现在他的治疗阶梯上有 20 或 30 种药物。他还活着。几年后,他状态很好。我前几天见到他,他基本上创建了自己的个性化药物和治疗方案。

He had a massively successful IPO, then was diagnosed with a rare cancer and has achieved, has lived way past the prognosis, has really taken it into his own hands. I think he went from frontline chemo and then there was one alternative that was available. He exhausted it and the doctors were like, 'We've got nothing for you.' Since then, I think like six or seven companies have come out of it, there's now 20 or 30 drugs in his escalation ladder. He's still alive. Years later, he's doing great. I saw him the other day and he basically created his own personalized medicines and treatment plan.

Naval

是的,我现在已经听过一些这样的轶事了。

Yeah, there's a handful of these anecdotes that I've heard now.

11. 高端医疗与患者自主权 High-end medicine and patient agency

Naval

我非常清楚,在高端领域,如果你有资源,想要使用现代科学的全套工具,那么疯狂的结果是可能的。如果你去问医生‘如果我这样做会发生什么’,他们只会大喊大叫、摔东西。但很明显,高端领域可以实现疯狂的事情。我认为这种针对个体的医疗将成为研究如何构建更具可转化性事物的丰富来源。

It is really clear to me that at the high end, if you have the resources and you want the full toolbox of modern science, outcomes are possible that are crazy. If you go ask your doctor what will happen if I do this, they will just start shouting and throwing things. But it is clear that crazy things are possible at the high end. I think this type of n-of-1 medicine will end up being a rich source of research for understanding how to build more translatable things.

Host

这要求患者在自身最脆弱的时刻拥有极大的主动性,这很讽刺。我朋友因癌症去世,他最不想做的就是研究个体化医疗,因为他每周都在濒死。但这正是 AI 应该大放异彩的地方,提出正确的解决方案,并让人们在那种情况下真正能做的事情民主化。从知识角度来看,很少有人能获得这些,而不仅仅是金钱方面,这太疯狂了。

It requires a ton of agency from the patient at a moment when they're at their weakest, which is ironic. My friend passed away from cancer, and the last thing he wanted to do was research n-of-1 medicine because he was dying by the week. But this is where AI should really shine and come up with the right solutions and democratize what you can actually do when you find yourself in that situation. It's crazy how few people get access to this from a knowledge perspective, not just monetarily.

12. 组织中的自主软件 Autonomous software in organizations

Naval

你们组织中有多少自主运行的或接近自主并能自我改进的软件?

How much autonomous software do you guys have in your organizations that's running on its own or near autonomous and improving on its own?

Host

对我们来说,很多基础设施已经是自主的。我们有一个在发现异常时触发的功能。我建议每个人都创建一个类似的版本。当任何异常发生时,大多数工程组织通过手动设置警报或监控阈值来响应,这很疯狂,但整个行业就是这样运作的。我们已经自动化了很多 SRE 的工作。任何变慢、变快或吞吐量变化的指标都会触发异常警报。一个智能体进行调查,可以决定创建事件,相关人员被拉入,智能体开始修复。我们做所有事情,除了给智能体改变生产的工具,但我们把解决方案放在银盘上递给工程师。另一个效果很好的事情是自主优化流程和自主安全研究。我们开源了一个叫 DeepSeek Coder 的工具,非常棒。我们用它针对整个单体仓库运行了 10000 个并发智能体,在几天内发现了几个季度的安全研究进展,花费了 14000 美元的 token。这相当于几个月的红队和安全研究。我们定期进行自主安全研究,因为网络安全正变成噩梦。漏洞太多,对手太强大。SRE 和优化工作非常明显。你可能在推特上看到有人将代码库从语言 A 翻译到语言 B。很多工作,比如优化或用原生编程语言重写,现在用前沿模型都可以做到。

For us, a lot of the infrastructure is already autonomous. We have a capability that fires off upon finding anomalies. I recommend everyone creates a version of this. Upon anything anomalous happening, most engineering organizations respond by setting up alarms or monitoring thresholds by hand, which is insane, but that's how the industry works. We've automated a lot of the SRE job. Any metric that slows down, speeds up, or changes throughput fires off an anomaly alert. An agent investigates, can decide to create an incident, people get looped in, and the agent begins remediation. We do everything except giving the tools to change production, but we serve solutions on a silver platter to engineers. Another thing working well is autonomous optimization processes and autonomous security research. We open-sourced a tool called DeepSeek Coder. It's incredible. We run it against our entire monorepo using 10,000 concurrent agents in the cloud, and it found several quarters' worth of security research progress in a couple of days and $14,000 worth of tokens. That's months of red teaming and security research. We're running periodic autonomous security research because cybersecurity is becoming a nightmare. There are too many vulnerabilities and powerful adversaries. SRE and optimization work are very obvious. You've seen people translating code bases from language A to language B. A lot of work like optimizing or rewriting in a native programming language is now doable with frontier models.

Naval

对于我自己用 vibe 编码的应用,我为 TestFlight 用户建了一个 bug 报告队列。他们可以在应用内报告 bug,上传日志和截图。他们也用它来提功能请求。我有一个简单的守护进程,编译所有 bug 报告,在后台主动分析和修复它们,然后给我发一个 TestFlight 版本,在发布给测试者之前先试试。对于功能请求,我现在只是编译它们,但我可以想象未来一个应用可以由用户构建。我不是说这是个好主意,可能会一团糟,但至少它可以处理 bug 报告。

For my own vibe-coded app, I built a bug reporting queue for my TestFlight users. They can report bugs from inside the app, which uploads logs and a screenshot. They also use it for feature requests. I have a simple daemon that compiles all bug reports, proactively analyzes and fixes them in the background, and ships me a TestFlight version to try before shipping to testers. For feature requests, I compile them now, but I could see an app in the future being built by the users. I'm not saying that's a good idea; it might be a mess, but at least it can take bug reports.

Host

顺便说一句,我们应该发布那个,就当是个社会实验,看看会发生什么。

We should ship that, by the way, just to see what happens as a social experiment.

Naval

是啊,社会实验。你最后会得到那辆荷马·辛普森的车,有雨伞、手电筒、小丑喇叭,什么功能都有。但肯定对于修 bug,你可以这么做。我们做了一个类似版本的实验:我让整个公司停止所有项目工作一周,说每个人,从接待员到工程师,构建你认为最重要的事情。唯一的要求:使用 AI,并在全公司演示。我预期会有大量愚蠢的项目和少量能推动进展的项目。结果我们得到了大量能推动进展的项目和非常少的愚蠢项目。有两三个是改变公司轨迹的。最让我惊讶的是接待员,那个负责从卡车上卸货并在货物入库时发邮件的收发员,她为此构建了一个自动化工具,我们实际上正在使用。结论是,每个人都有一些能让世界变得更好的想法,但很多时候他们的初步想法很蠢,而且他们无法预见到这一点。但如果他们有能力从想法变成实际的东西,如果不行,他们可以反应和迭代。如果你给他们一周时间,到周末他们就能构建出有意义的东西。

Yeah, the social experiment. You end up with that Homer Simpson car with an umbrella, flashlight, clown horn, every feature. But definitely for bug fixing, you could do that. We did a version of that experiment where I stopped all project work across the entire company for a week and said everybody from receptionist to engineers build whatever you think is most important. The only requirements: use AI and demo it for the whole company. I expected a large number of silly projects and a small number of needle movers. What we got was a large number of needle movers and a very small number of silly projects. Two or three were trajectory-changing. What surprised me most was the receptionist, the shipping and receiving associate whose job was to take packages off a truck and email people when their stuff came in, built an automation for that that we're actually using. The conclusion is that everybody has some idea of what could exist that would make the world better, but many times their first-order ideas are stupid and they can't project that out. But if they have the ability to go from idea to an actual thing, if it's not working, they can react and iterate. If you give them a week, by the end of the week they've built something that makes sense.

Host

但想象一下,如果所有工作都像那样。你如何建立一个不直接做工作的劳动力?他们所做的就是训练替他们工作的智能体。我们也做过这个。你必须提醒大家,举办黑客马拉松:‘嘿,我们来构建智能体吧。’

But imagine if all work was like that. How can you set up a workforce that does not do the work directly? All they do is train the agent that does the work for them. We've done this as well. You have to remind folks and create hackathons: 'Hey, let's build agents.'

13. 自主公司与文化转变 Autonomous Company and Culture Shift

Host

嗯,显然有很多人,文化正在发生变化。有很多新来的人直觉上就知道,他们的工作不是直接做事情,而是训练做事情的智能体。但我很好奇,未来的自主公司会是什么样子?

Uh and obviously there's a lot of people there's a culture change happening like there are a lot of people that are just coming in who intuitively know their job is to not work on the thing is to actually train the agent that works on the thing. But I'm curious about like, you know, what does the autonomous company of the future look like?

Naval

可能会变得更疯狂。也许你只需打开所有摄像头,智能体就会观察一切。它发现发货和收货效率很低,然后它创建了应用,呈现出来。

It could get a lot crazier. Maybe you just turn on all cameras and the agents just watching everything that's happening. It see the shipping and receiving thing is very inefficient and it creates the app presents the app.

Host

你看到了吗?Zach 把这个东西装到了每个人的机器上。他在考虑这件事。我们也看到了这个,我们很可能要在 AI 网关中推出一项功能,允许人们选择保留输入和输出,然后你可以说,对于我所有的输入和输出,你能提取我从工作中学到的技能,然后把它作为技能导出,这样我甚至可以自己下载。但你可以想象,公司里的人会想要分享并整合这些。

Did you see that? Zach installed this thing into everyone's machines. He's thinking about it. It's like we saw this too like we're um we're we're likely going to ship a feature into AI gateway that allows people to opt in into preserving inputs and outputs and then you can say for all of my inputs and all my outputs can you extract the skills of the things that I like learn from my work and then dump it as skills uh so that I can even download them for myself. But you could imagine people in companies wanting to share and pull this together.

Naval

有趣的是,对我来说,这在我的工作中是难以想象的,因为我的工作不是重复性的。我寻找可以自动化的事情。我自己的工作几乎没有什么可以自动化的了。我希望每个人最终都能达到那种状态,对吧?你始终在创造力和兴趣的最大区域工作。如果还有任何可以自动化的事情,你就应该自动化它。把它从你的生活中清除出去。它会解放你,让你变得有创造力,而那就是你创造所有价值的地方。但我认为,在职业心态中很难看到这一点,因为你雇佣人是为了反复做同样的事情,而这种情况正在消失,这真的很可怕,因为人们会问,那我该做什么?嗯,你会做有创意的事情。你会想出新的东西,你不必每天都想出新东西。那是不可能的,对吧?但你会偶尔想出一个新东西,然后它会创造出别的东西,成为你的一个杠杆点。但这对人们来说确实是一个可怕的时期。如果你已经十年如一日地做同一件事,现在突然之间,你要训练一个智能体并把它自动化掉,那很可怕。

It's funny because for me that's so unimaginable for my own work because my own work is not repetitive. I look for things to automate. There's almost nothing left for me to automate for my own work. And I hope that's where everybody ends up, right? You just work in your maximum zone of creativity and interest at all times. And if there is anything left to automate, you should automate it. Get it out of your life. It'll free you up to be creative and that's where you generate all the value. But I think that's very hard to see in the job career mindset because you hire people to do the same thing over and over and that's going away and that's really scary because people like, well, what am I going to do? Well, you're going to do creative things. You're going to come up with new things and you don't have to come up with a new thing every day. That's impossible, right? But you're going to come up with a new thing once in a while that will then create something else, some point of leverage for you. But it is a scary time for people for sure. If you've been doing the same thing over and over for 10 years and now all of a sudden it's like, well, now you're going to train an agent and automate it away, that's scary.

Host

我认为历史上回报大约是 70% 的智力,30% 的能动性,而现在将是 70% 的能动性,30% 的智力,随着模型越来越好,这种情况会进一步转变。

I think historically it was the returns were like 70% intelligence, 30% agency and now it's going to be 70% agency, 30% intelligence and that will shift further as the models get better and better.

Naval

我其实不太确定,Max。我要提出相反的观点。我认为是 99% 的智力和 1% 的能动性,因为那样智能体就会行使能动性,对吧?你实际上会说,'嘿,智能体,我在做聪明的决定,思考大问题。你去执行吧。' 事实上,有时我想在通过 vibe coding 流出的应用上构建功能。我会问智能体,'接下来我应该构建什么功能?' 你知道,去看看日志,看看用户,我该怎么做?

I'm actually not sure about that, Max. I'll take the counterpoint on that. I think it's 99% intelligence and 1% agency because then the agents will exercise the agency, right? You will literally be like, 'Hey agent, I'm making smart decisions and thinking big thoughts. Just go implement stuff.' In fact, sometimes I want to build features on apps that I'm flowing out of vibe coding. I'll ask the agent, 'What features should I build next?' You know, go look at the logs, go look at the users, what should I do?

Host

明确一下,我说的是对人类的回报。最适合未来的人类将是那些更有能动性的人,也就是说,那些能进来就想着'我要打开云端,然后想我该构建什么',而不是看 YouTube。这里有一个有趣的实验,我打赌我们现在都认识很多以前不编程的人现在在编程,包括很多情况下我们自己,对吧?所以生态系统中程序员的数量百分比可能已经增长了 10 倍,没错,现在编程的人可能确实是一年前的 10 倍。

To be clear, I'm talking about the returns to humans. The humans that will be best fit for the future will be the ones that are more agentic which is to say like the ones that can come in and just have the thought of like I'm going to open cloud and be like what should I build versus watch YouTube and here's a fun experiment I'll bet you we all know a lot of people now who are coding who weren't coding before including many cases ourselves right so the number the percentage of coders in the ecosystem has probably gone up by might be 10x right yeah it might literally be 10 times as many people are coding now than were coding a year ago

Naval

太疯狂了,我们的注册人数爆棚,出现了新一类非工程师的人。他们只是使用基础设施。但我想可能是播客主、YouTuber 和在 X 上发帖的人。大多数人仍然不写代码。比如我去找别人,我说,'哦,老兄,vibe coding 太有趣了。它比我有过的一个小游戏群还有趣,我以前玩电子游戏和 FPS 来发泄压力。' 我完全停止了玩游戏。所有时间都花在了 vibe coding 上。它更有娱乐性。你能从中得到真实的东西,而且反馈循环同样紧密甚至更好。我去找其他朋友,我说,'嘿,你应该试试 vibe coding。' 他们只是茫然地看着我,我说,'不,不,你不明白。构建东西容易多了。' 但我认为对他们来说,这始终是一个后台的黑箱过程。他们从未理解过。他们以为你可能一直在和电脑说话。所以他们看不到什么变了。他们没有意识到它容易多了。对他们来说,正如 Max 所说,开始是那么难以想象和困难,他们不会去做。所以我们可能从 0.01% 的人口写代码,现在变成了 1%,可以说是 100 倍的增长,但 99% 的人仍然永远不会写代码。所以我们处在一个奇怪的空间。

It's wild our signup numbers are through the roof and there's this new class of people who are not engineers. They just use the infrastructure. But I think it might be like podcasters and YouTubers and like people posting on X. The majority of people are still not creating code. Like I go to people and I'm like, 'Oh man, vibe coding is so much fun. It's more fun than like I had a little gaming group that I used to play video games and FPS's to blow off steam.' I completely stopped playing. All that time went into vibe coding instead. It's more entertaining. You get something real out of it and but the feedback loop is just as tight or even better. And I went to my other friends and I was like, 'Hey, you should be vibe coding instead.' And they just gave me this blank look and I'm like, 'No, no, you don't understand. Building things is so much easier.' But I think to them it was always like some blackbox process in the background. They never understood it. They assume maybe you were just talking to computer all along. So they don't see what's changed. They don't realize it's a lot easier. To them just that starting to Max's point the starting is so impossible to imagine and hard they don't do it. So we might have taken you know 0.01% of the population writing code to maybe now it's 1% call it a 100x increase but 99% still never going to write code. So we are in this weird space.

Host

太疯狂了。它就像一个电子游戏,一个很棒的电子游戏,但能产出真实的东西。

It's crazy. It's like it's a video game and it's a great video game but real stuff comes out.

Naval

是的。我的未婚妻昨晚整晚没睡,因为她睡不着,她在捣鼓什么东西,当然她没写任何代码。但它就是让人上瘾,编程对我来说已经十多年没有这种感觉了。这太棒了,因为它就像人们的彩票。我认为普通人已经更多地进入了 vibe coding,但通过更多是媒体模型,比如视频模型,对吧?更多的人可能尝试制作视频和图像,而不是写代码和应用。问题是,视频本身也有问题,对吧?也许有一天,我们可以说'给我做一部关于 X 的好电影',然后它就会吐出一部好的纪录片,但现在它们还没有品味或判断力。这是我和 Andre Carpathy 的一个赌注:哪一年你可以直接扔进一本书,然后得到一部电影?我认为更近了,尽管自从我们几年前打赌以来,他的时间线已经大幅缩短。到 2030 年,我们会有几十部《指环王》。会有某个粉丝觉得他拍错了,我要自己拍一版。比如那些著名的故事,或者我另一个衡量 AI 进步的基准是,我是《苍穹浩瀚》系列的超级粉丝。有一个电视剧和九本书,他们拍了前六本书,但没有拍最后三本,而且有重大的分歧,我还没读过书。

Yeah. My fiance was up all night last night because she couldn't go to sleep because she was hacking on something and of course she wasn't writing any of the code. But it's just like it's addictive in a way that programming hasn't been for me for like over a decade. It's amazing because it's like a lottery for people. I think the normies have gotten a little more into the vibe coding but through models that are more media models, video models for example, right? More people probably fooled around making videos and images than they did writing code and apps. The problem is like I don't video has its own issues, right? Maybe someday we like make me a great movie about X and I'll just spit out a good documentary, but right now they don't have the taste or the judgment. This is a bet that I have with Andre Carpathy was like what's the year that you'll be able to just dump in a book and get a movie out? I think closer although I think he has come down substantially in timeline since we made this bet a few years ago. By 2030 we're going to have like dozens of Lord of the Rings. Like there's going to be some fan who's like he did it wrong. I'm going to make my own take. Like the famous stories or like the one of my other benchmarks for progress in AI is I'm a huge fan of a series called The Expanse. There's a TV series and there's nine books and they've made the first six books but they haven't made the last three books and there's meaningful divergences and I just I haven't gotten in I haven't read the books.

14. 人类创造力 vs AI Human creativity vs AI

Host

比如我期待能输入最后三本书,基于电视剧,然后生成最后三季。

Like I'm looking forward to I can dump in the last three books conditioned on the TV series and be like generate the last three seasons.

Naval

这就要来了。

This is coming.

Host

那是个很棒的功能。是的。但某种程度上这很容易,因为已经有这么多参考素材。当你说给我下一部《指环王》时,我真的很兴奋,因为我们还没有在想象力上取得突破。

That's a great feature. Yeah. But that's in a way it's easy because there's already all this reference material. When you said get me the next Lord of the Rings, I was really excited because we haven't really had a breakthrough in imagination.

Naval

哦,我们会在文化中看到像《哈利·波特》和《指环王》这样的作品。我对此非常兴奋,那会是——我同意那会是更激动人心的。

Oh, we're going to see that in culture the likes of Harry Potter and Lord of the Rings. I'm really excited about that and that will be the I agree that that will be the more exciting one.

Host

人类能独特地做什么?这回到了核心问题。人类将能独特地做什么?我认为马克斯,你是一个 AGI 最大化主义者。所以对你来说,什么也没有。智能体会做一切。

What can humans uniquely do? This gets back to the core issue. What are humans going to be able to uniquely do, right? And I think Max, you're an AGI maximalist. So for you it's nothing. Agents will do everything.

Naval

我不是反人类,但我只是觉得,如果你的身份认同取决于你有多聪明和多有创造力,你会过得很糟。

I'm not like antihuman, but I just like I think it's going to be we will have to find like if your identity is how smart and creative you are, you're going to have a bad time.

Host

是的。我想我仍然站在另一边。我认为创造力仍然是环境中让你惊讶的东西。你跳出系统,做一些在系统内甚至无法想象的事情。它超出了训练数据,超出了输入系统的数据分布。我认为总有空间留给这一点。你有没有注意到每个 Claude 网站看起来都一样?人们基本上调出了一个 Claude 网站的样子,一旦你从模型中得到足够多的生成。比如有一种外观:衬线字体,棕色和奶油色,使用等宽字体和特定间距。过了一段时间,你得到这个分布,你会说这不具创意。这是 Claude 产出的垃圾。这不会是人与计算机的对决,而是人与计算机联手对抗纯计算机。

Yeah. I guess I'm still on the other side of that. I think that creativity is still the thing in the environment that surprises you. You step out of the system and do something that wasn't even imaginable within the system. It's outside of the training data. It's out of the distribution of data that was fed into the system. And I think there'll always be room for that. Have you noticed that every Claude website looks the same? And people like basically like dial in what a Claude website looks like once you get enough generations out of the model. Like there's a look it's this serif font. It's brown and cream and they use monospace fonts with a certain amount of spacing. Like after a while you get this this distribution that you say well this this is not creative. This is slop that came out of Claude. It's not going to be human versus computer. It's going to be human with computer versus just computer.

Naval

纯计算机最终会发生,但我们离得很远。

Just computer will eventually happen, but we're pretty far away.

Host

但计算机将能够产生这些疯狂的超级刺激,它将制造娱乐。我的意思是,我们在 TikTok 上看到了一种弱形式。所以当你思考时,我个人对艺术的定义是有意义的分布外行为。所以这在某种程度上是令人惊讶的。感觉就像你在 Z 轴上移动。你惊讶于这件事被实现了,而且有意义。

But the computer is going to be able to produce these crazy super stimuluses that like it's going to be it's going to make the entertainment. And I mean, we kind of see a weak form of this in TikTok. And so when you think about the going like my personal definition of art is meaningful out of distribution behavior. And so this is something that kind of is surprising in some way. Feels like you're kind of moving in the Z-axis. Like you're surprised that the thing was realized, but meaningful.

Naval

有意义意味着,对我来说,它某种程度上改变了你未来在宇宙中的轨迹。你的生活因为思考和反思它而变得不同。嗯,我对艺术的定义完全不同,导致完全不同的结果。抱歉打断。

Meaningful means that like it somehow to me means that it somehow changes your like future trajectory through the universe. Like your life is somehow different for having thought about it and reflected on it. Well, my definition of art is completely different and leads to a completely different outcome. Sorry to interrupt.

Host

不,有趣的是,仅仅通过你的定义,你就得到了不同的前提。那是公理的推论。

No, it's interesting how just by your definition, you get to a different premise. That's the extrapolation of the axiom.

Naval

是的。我的意思是,我喜欢我的定义的一点是它非常宽泛。比如,可以有军事演习,你会说那是艺术。我认为我们会一直看到这个。我们会到处看到第 37 手。不过,我很好奇你对艺术的定义是什么。

Yeah. I mean, one of the things I like about my definition is that it's so broad. Like, there can be like military maneuvers that you can be like, that was art. And I think we're going to see this all the time. We're going to see move 37s all over the place. Although, I'm curious what your definition of art is.

Host

我的意思是,我有多个定义,所以没有一个具体的、打包成一个东西的定义。但我确实认为艺术是你传达情感的东西。你把你感受到的东西传达给另一个人。所以你创造了一个物体或东西,它捕捉了你内心的情感。所以对我来说,计算机几乎从定义上就无法做到这一点。完全相同的艺术品,如果没有背后的意图,就有点无意义。现在你也可以说自然是艺术,比如自然中的美,比如你看到的日落,对吧?不是人类。所以那个我会称之为没有动机的纯粹智能在运作。比如日落中有美,因为那里有智能。有一个复杂系统在运作,你的大脑识别它,那里没有动机。所以没有自我介入。但更人性化的艺术,我认为是某人感受到了什么,他们想让你感受到那个东西,或者他们想再次感受那个东西,或者他们想捕捉他们与那个东西的感觉,所以他们创造了那个东西。归属权非常重要。比如一张美丽的照片,对吧?如果一个人拍了照片,而 AI 生成了完全相同的照片,直到最后一个像素,那么拍照片的人对我来说更有意义。

I mean, I have multiple definitions, but so it's not like a concrete I haven't packaged into one thing. But I do think of art as something where you convey emotion. You convey something you felt to another person. And so you create some object or something that creates that takes an emotion that you felt inside. And so to me, a computer almost by definition is incapable of doing it. The exact same piece of art without intent behind it is sort of meaningless. Now you can also argue nature is art like beauty in nature like you see a sunset, right? Not let's say human. So that one I would call it's pure intelligence working without motive. There's beauty for example in a sunset because there's an intelligence there. There's a complex system at work there and your brain recognize it and there's no motive there. So no ego gets involved. But art in kind of the more human sense I think of as someone felt something and they wanted you to feel that thing or they wanted to feel that thing again or they wanted to capture the feeling they had with that thing and so they created the thing. Attribution to who created it is going to be really important. So like for example a beautiful photo, right? If a person takes the photo versus AI generates the exact same photo down to the last pixel, the person taking the photo will have more meaning for me.

Naval

我刚刚投资了一家初创公司,它用硬件在站点上做可验证性,证明某个人类确实拍了照片,这将有很多非常酷的用例。

I just invested in a startup that does verifiability with hardware at the station that someone some human actually took a photo which is going to have a lot of really cool use cases.

Host

我们会被垃圾淹没。毫无疑问。

It will we will be drowned in slop. No question.

Naval

你还记得一两年前的 ControlNet 吗?有一个特定的场景,像一个中世纪村庄,里面有一个漩涡。你记得吗?

Do you remember the control net stuff from like a year or two ago? There was like there's one particular scene of like it was like a medieval village. It had like a swirl in it. Do you remember?

Host

是的。

Yeah.

Naval

那是 AI 生成的,那是我第一次看到这个并觉得它真的很酷。无论你是否想称它为艺术。

That was AI generated and that was one of the first times I looked at this and thought it was really cool. Like whether you want to call it art.

Host

但那个没有打破你的前提,因为有人类提出了训练和提示,才得到那个很酷的谜题。顺便说一句,AI 未来也完全可能做到。但我把功劳归于想出那个视错觉 ControlNet 想法的人,而不是 AI。

But that one doesn't that one break your premise because some human came up with the training and the prompt to arrive to that really cool riddle. By the way, it's totally possible that an AI can also do that in the future. But I give whoever came up with that idea of the optical illusion control net, I give them more credit than the AI.

Naval

我认为门槛会被大幅提高。需要越来越多才能让你惊讶。它必须越来越令人印象深刻。这已经发生了。

I think the bar is going to be raised massively. Like it's going to take more and more to surprise you. It's going to have to be more and more impressive. Like that's already happened.

Host

是的。比如 OpenAI 为所有人毁了吉卜力工作室。没人想再看吉卜力工作室的作品了,对吧?已经做过了。所以——

Yeah. Like OpenAI destroyed Studio Ghibli for everybody. Nobody wants to see that Studio Ghibli work ever again, right? It's been done. And so

Naval

哦,那个我也有一个反驳点。你看过真正的吉卜力工作室吗?它实际上比 OpenAI 产出的垃圾好得多。现在再看一遍。它令人印象深刻。

Oh, that one also has I have a counter point to that one. Like have you watched real Studio Ghibli? It actually looks so much better than the slop that OpenAI put out. Like watch it again now. It's impressive.

Host

是的。当你在互联网上到处看到大量吉卜力风格的东西时,它现在已经在分布内了。它不再令人惊讶。艺术价值已经存在了。

Yeah. At the point where you've seen tons of Studio Ghibli things everywhere all over the internet. It is now in distribution. It's no longer surprising. The art value has been around.

Naval

没错。没错。不,你的惊讶定义仍然成立。我只是认为人类是能够完全从数据分布之外产生惊讶的。而且我认为他们可以带着意图做到,我确实认为意图对意义很重要。所以针对你的意义点,你说意义和惊讶,对吧?我想说的是,人类可以——他们从系统之外产生惊讶。

That's right. That's right. No, your surprise definition still works. I just think that that humans are the ones who can generate surprise completely out of the data distribution. And I think they can do it with intent and I do think intent matters for meaning. So to your meaning point, right, you said meaning and surprise, right? And I guess what I would say is that humans can steal the ones generate surprise out of the system.

15. 人类创造力 vs AI 的分布外想法 Human creativity vs AI out-of-distribution ideas

Naval

举个例子,假设你训练了一个 AI,让它精通数学,对吧?一个完美的数学 AI,它完全在数学的形式系统内运作。然后库尔特·哥德尔出现了,他提出了完全在系统之外的东西,对吧?哥德尔不完备定理。它完全跳出了系统,利用物理学的属性来打破系统。所以我认为 AI 无法做到那种事。因此,外部总有创造力的空间。惊喜,而意义来自于人类的参与——他们为了某个目的而做,并传达了某些东西。所以也许我可以按自己的方式解读你的定义,但我们会看到结果如何。我对人类更乐观一些。

For example, let's say you took an AI and you trained it to be perfect at mathematics, right? The perfect mathematics AI and it's within the formal system of mathematics. And then Kurt Gödel comes along and he has something completely outside of the system, right? Gödel's incompleteness theorem. It completely stepped out of the system and used attributes of physics to basically break the system. So that kind of thing I don't think an AI could get to. So there's always room for creativity outside. Surprise and then the meaning comes from the fact that a human was involved that they did it for a purpose and they conveyed something. So maybe I can interpret your definition my way, but we'll see how it plays out. I'm a little more optimistic about humans.

Host

那么,如果你训练一个 AI 模型,它是在某种数据分布上训练的,在一些 token 上训练。然后它学习语言的某种分布及其内部结构。LLM 或 Transformer 有可能走出分布,产生一个训练集中不存在的新想法吗?

So if you train an AI model, it's trained on some data distribution. It's trained on some tokens. It then learns some distribution of language and the structure within that. Is it possible for an LLM or transformer to go out of distribution, to have a new idea that was not present in the training set somehow?

Naval

嗯,训练集如此之大,很难想象有哪些想法不在训练集中。但如果存在,它们可能存在于物理、互动、感觉、情感和进化等自然领域,而这些是 AI 不受影响的。所以我确实认为语言之外仍有东西,但语言确实涵盖了很多。语言是一个很好的压缩器,我们拥有大量的语言数据。

Well, the training sets are so large that it is hard to imagine ideas that are not within the training sets somewhere. But if they exist, they probably lie in the natural domain in physics and interaction and feeling and emotions and evolution in things that the AI is not subject to. So I do think that there are still things outside of language, but language does encapsulate a lot. Language is a great compressor and we've got a lot of it.

Host

但我的意思是,你可以通过自我对弈达到这些其他东西。

But I mean you can get to these other things through selfplay.

Naval

自我对弈和传感器,比如摄像头是传感器,就像我们的眼睛是传感器一样。

Selfplay and sensors like cameras are sensors like our eyes are sensors.

Host

是的。我的意思是,我认为问题是如何在没有随机性的情况下走出分布?所以在强化学习的情况下,你可以获得随机性,比如你可以从动作空间的分布中采样一个动作,然后随机性可以带你进入新的领域。但我想反过来问,人类能走出分布吗?任何新想法从何而来?我们是否也依赖随机性来进入这些新领域?

Yeah. I mean I think the question is how do you go out of distribution without randomness? So in the case of RL you can get randomness like you can sample an action from a distribution of an action space and you can get randomness that can take you down these walks into new territory. But I think the real question, to turn this around, is can humans go out of distribution? Where does any new idea come from? Are we also dependent on randomness to get us into these new territories?

Naval

我们并不依赖纯粹的随机性,就像自然选择通过纯粹随机性运作一样,对吧?你只是突变一个基因,然后看看会发生什么。但人类似乎有能力穿越无限空间,你知道,直接消除巨大的区域。所以我们的创造力在更大的图景中是有意义的。这似乎是我们独特的能力之一。也许 AI 开始在边缘做到这一点,就像我们看到它在解决一些数学问题那样。但即使是数学,也是一个非常有限的领域,尽管它很大。我不是说它永远达不到。我没有那种信心。但我认为至少目前,我会说真正跳出框架、让人惊讶,仍然是人类的领域。我认为人类加 AI 是未来的方向。没有 AI 的人类,算了吧。纯粹的 AI,我认为还没到那一步。但人类加 AI,我们已经处在这个时代了。我们会在这个时代停留多久?我打赌会比人们想象的要长。我认为人类将拥有巨大的价值。事实上,更多的价值。我们所有人,这里的每个人,我们的生产力都飙升了。基本经济学通常说,当一个人的生产力更高时,他们更富有,状况更好。你实际上会雇佣更多这样的人,而不是更少。也许你们有些人不再招聘初级员工了,尽管我不确定这是否真的如此。我不认为这是初级与高级的问题。如果有人非常擅长 AI,而且非常聪明和有创造力,我比以往任何时候都更想雇佣他们,因为我从他们身上获得的杠杆作用将是不可思议的。

We're not dependent on pure randomness like natural selection works through pure randomness, right? Where you just mutate a gene and then see what happens. But with humans, we seem to have this ability to cut through infinite space and get, you know, just eliminate huge swaths. And so our creativity makes sense within the larger scheme of things. That seems to be one of our unique capabilities. And maybe AI is starting to do that at the edges as we're seeing with solving some of these math problems. But even math is a very bounded domain, but it's a big one. I'm not saying it'll never get there. I don't have that confidence. But I think at least at the moment, I would say that truly stepping outside, surprising people, that's still a domain of humans. And I think humans plus AI is where it's all moving to. Like human without AI, forget it. Pure AI, I don't think is there yet. But I think human plus AI, we're in that era. How long we stay there, I'm betting is longer than people think. I think humans will have an enormous amount of value. In fact, more value. All of us, everyone here, our productivity has gone through the roof. And basic economics normally says that when someone's productivity is higher, they're wealthier. They're better off. You actually hire more of them, not less of them. Maybe some of you are not hiring junior people anymore, although I don't know if that's necessarily true. I don't think of it as junior versus senior. If someone is really good with AI and they're really smart and creative, I want to hire them more than ever because the leverage I'm going to get out of them is incredible.

Host

这是一个新的要求。我们正在招聘初级和超级资深员工,只要他们非常擅长智能体,非常擅长 AI,并且能快速适应。

That's a new requirement. And we're hiring juniors and super seniors as long as they're really good with agents and really good with AI and quick to adapt.

Naval

而且很多人不再需要被雇佣了。他们可以自己创造东西。

And a lot of them don't need to be hired anymore. They can create their own thing.

Host

我的假设是,我们最终会有更多的小团队。完成给定任务所需的人数大幅下降。你知道,那些只看到一阶效应的人会说‘哦天哪,所有工作都消失了,因为我可以两个人造一个喷气发动机。我不需要一千个人,998 个工作没了。’但实际上这意味着你可以创造许多不同的喷气发动机。

My hypothesis is we end up with a larger number of smaller teams. Like the number of people required to accomplish a given task drops by a lot. And you know people who only see first order effects say 'oh my gosh all the jobs are disappearing because I can do a jet engine with two people. I don't need a thousand, 998 jobs are gone.' But what it actually means is you can create a lot of different jet engines.

Naval

我认为完全正确。

I think that's exactly right.

Host

我认为这将回到 Naval 的观点。我认为人类独特的是创造力,而一直缺失的是,很多人可以有创造力,但他们不知道如何将愿景变成改变现实的真实事物。所以我认为我们将迎来创业的爆发,创始人的爆发,以及大量非常小的团队,因为你不需要很多人来完成某件事。

I think there will be, goes back to Naval's point. I think the thing that's uniquely human is the creativity and what's been missing is that a lot of people can be creative but they don't know how to turn their vision into a real thing that's changing. So I think we'll have an explosion of entrepreneurship, an explosion of founders and a very large number of very small teams because you don't need many people to accomplish something.

Naval

是的。我认为 AI 提供了基础智能和领域知识,并穿透了所有行话,而现在智能体实际上提供了大量的能动性。所以剩下的主要是创造力、品味,是的,你需要足够的能动性来开始,坚持的能动性,但你不一定需要花 20 年学习一件事才能深入并做出贡献。所以随着这个障碍降低,通才们如鱼得水。归根结底,我们都是通才。我们都喜欢思考一切。我们不喜欢被困在一件事上。就像 Max 在这里谈论意识、FDA、脑科学和创造力。我们所有人都一直在试图思考一切。所以,那些在 Twitter 上总是喜欢说‘专家、资历、来源’的人,对吧?那些人正在受到伤害,因为专业知识不再重要。你花了 5 年、10 年获得某个领域的博士学位,你知道,希望你能发展出创造力、直觉、品味和判断力,因为如果它只是帮你记住一大堆东西和行话,学习一些框架性的东西,那么 AI 会直接穿透它。就像一个计算器乘以十亿,或者一个加速的思维自行车。所以我认为这是拥有 AI 的人与没有 AI 的人之间的较量。因此,你现在能为自己做的最好的事情就是真正精通这些工具,熟悉它们,并始终知道它们能力的边界,知道它们能做什么和不能做什么。而这是一个移动的目标。

Yeah. I think AI provided base level intelligence and domain knowledge and cut through all the jargon, and now agents actually provide a lot of agency. So the main things left are creativity, taste, and yes, you need enough agency to get started, agency to stick with it, but you don't necessarily need the agency to spend 20 years learning one thing before you can dive into it and make a contribution. And so that barrier going down, generalists are having a field day. And at the end of the day, we're all generalists. All of us like to think about everything. We don't like to be just trapped in one thing. Like Max is here talking about consciousness and the FDA and brain science and creativity. Like all of us are trying to think about everything all the time. And so, people on Twitter who are always fond of saying 'experts, credentials, sources,' right? Those are the guys getting hurt because the expertise doesn't matter. You spent 5 years, 10 years getting a PhD in XYZ, you know, hopefully you developed your creativity and your instincts and your taste and your judgment because if all it did was help you memorize a whole bunch of things and jargon and learn some scaffolding stuff, well, AI will cut right through that. It's like a calculator times a billion or a bicycle for the mind but accelerated. So I think it's about people with AI versus people without AI. And so the single best thing you can be doing right now for yourself is just getting really good with these tools, getting comfortable with them and always knowing the edges of the boundaries of what they're capable and what they're not capable of. And that is a moving target.

互动版:逐字朗读 + 针对本期提问 →