Building Randomness: Thermodynamic Computing and AI-Driven Chip Design
打开互动全文版(中英对照 + 朗读 + 问答)→Thomas Ahle 探讨热力学计算、AI 生成的 Verilog 以及验证芯片正确性的挑战。
Thomas Ahle discusses thermodynamic computing, AI-generated Verilog, and the challenges of verifying chip correctness.
是啊,那为什么不试着造一块本质上就是随机的芯片呢?这位是 Thomas Ahle。我在苏黎世和他碰了面。他是那种罕见的“银河大脑”型人物,在概率机器学习、形式化验证和芯片设计方面都游刃有余。不过,有个小问题:当“Token 之神”递给你一个看起来能用的东西时,你怎么知道它真的正确?
Yeah, so why not try and build a chip that's just inherently random? Meet Thomas Ahle. I caught up with him in Zurich. And he's one of these rare galaxy brain people who's comfortable in probabilistic machine learning, formal verification, and chip design. However, there's a small problem. When the token god hands you something that looks like it works, how do you know it's actually right?
我的背景是理论计算机科学。我以前研究高维数据算法、局部敏感哈希。后来我转向热计算,开发热计算技术,也是为了加速贝叶斯智能。有时候我觉得这就像芯片设计的“可爱版”。我们从头到尾负责:从你的意图出发,经过设计、优化,再到形式化验证,一直到流片。
So my background is in theoretical computer science. I used to do algorithms for high-dimensional data, locality-sensitive hashing. Then I moved to thermal computing to develop thermal computing, uh also to speed up Bayesian intelligence. And sometimes I think about it as the lovable for chip design. So we take it all the way from your intent uh through the design, through optimizing your design to formalizing and verifying your design, uh all the way to tape out.
我以前并没有完全意识到这一点。如今,芯片不一定从工厂开始,它可以从代码开始。工程师用 Verilog 这种语言设计整个电路,几乎就像写软件一样,直到很晚之后才变成物理硅片。但首先,这段代码必须经过模拟和形式化验证,必须被证明是正确的。因为一旦芯片制造出来,如果有任何 bug,那就麻烦大了。几个月前,Thomas 在博客上写了用一群 AI 智能体协作构建 Verilog 模拟器,43 天生成了超过 50 万行代码。他需要这么做的原因是商业软件贵得离谱,而且对智能体不太友好,每个席位要 1 万美元左右。
Now I didn't fully appreciate this before. These days, a chip doesn't necessarily start in a factory. It can start as code. So engineers design the whole circuit in a language called Verilog, almost written like software, and only much later does any of it become physical silicon. But first, that code has to be simulated and formally verified. It has to be proven correct. Because once a chip is fabricated, if there are any bugs, you're in big trouble. So a few months ago, Thomas blogged about building a Verilog simulator using a swarm of AI agents collaborating with each other, and it generated over half a million lines of code in 43 days. Now the reason he needed to do this is that commercial software costs a ridiculous amount of money and isn't very friendly to using agents. $10,000 per seat or something.
对,大概是一个 CPU 内核的价格。
Yeah, probably for one CPU kernel.
如果 AI 能生成芯片设计、证明或可运行的程序,你怎么知道它真的正确?
If an AI can generate a chip design or a proof or a working program, how do you know it's actually correct?
它通过测试的百分比,或者像那个寓言一样,比如“哦,它通过了 70% 到 80% 的测试”,但问题是,我和做基准测试的人聊过,他们说“但它真的做对了吗?”你懂的。
The percentage of test that it's got right or like even like the fable it was like oh yeah it got like 70 80% test correct and it's like yeah but I talked with the people who made the benchmark and like yeah but it did did it get any of them actually right you know like
在芯片设计中,噪声是敌人。制造商花大价钱消除它,但在热计算中,情况恰恰相反:噪声就是计算。它直接来自概率机器学习,后者本来就运行在随机性和不确定性之上,而热计算试图让芯片本身成为一个随机微分方程。让芯片自身的噪声稳定下来,它就能得出通常需要花费巨资才能计算出的答案。
Now in chip design noise is the enemy. Manufacturers spend a fortune getting rid of it but in thermodynamic computing the opposite is kind of true. The noise is the computation. It comes straight out of probabilistic machine learning which already runs on randomness and uncertainty and thermodynamic computing tries to make the chip itself a stochastic differential equation. Let the chips own noise settle into place and then it can land on answers that would normally cost a fortune to compute.
然后你注入所有这些噪声,它就会开始按照这些随机微分方程来运行。实际上,它的行为有点像按照那个矩阵的逆矩阵来运行。
And then you infuse all of this noise um and it'll start to behave according to um to to these uh stochastic differential equations. It actually behaves sort of according to the inverse of that matrix.
他们已经发布了这项技术的第一个版本,是一款名为 CN101 的芯片,已经流片到硅。现在还在早期阶段,针对的是相当狭窄的概率工作负载,但真正的考验是,一旦他们扩大规模,基准测试会怎么说。我们 MLST 非常感兴趣的一点是,如果我们让 AI 构建一切,我们在整个过程中到底保留了什么理解?
They've already released their first version of this technology. It's a chip called CN101 all the way to silicon. It's still early days and it's aimed at a fairly narrow band of probabilistic workloads but the real test is what can the benchmarks say once they scale this up. And something of huge interest to us at MLST is if we're going to let AI build everything what kind of understanding do we actually keep in the overall process?
不仅仅是它变得更聪明了,人类也在变得更笨。
It's not just that it's getting smarter. It's also that humans are getting dumber.
现在我要说明一下,Normal Computing 非常慷慨地承担了我们这期节目的制作和差旅费用,但我们保留了完全的编辑控制权。
Now I should disclose that normal computing very kindly offered to cover our production and travel costs for this show but we kept full editorial control.
这些是设备制造的工具。比如,RTL 的编译器,RTL 就像你使用的编程语言。所以,如果你在构建芯片,你通常不会画所有不同的门电路等等。你用 Verilog 这样的编程语言来写,它有一些结构,有点像超级并行编程语言,很多结构能很好地映射到硬件上。但是,这个系统没有好的开源编译器。所以整个硬件行业没有软件那种开源氛围,软件里好东西都是免费的,工具市场充满活力。这里更封闭,被大供应商垄断。但人们编译后得到网表,可以开始构建原理图,也可以查看。你会得到这些巨大的东西,最终可以送到晶圆厂。但还需要做模拟,比如测试,也可以做形式化验证。
These are tools for device manufacturing. Um so we're it's things like I mean it's compilers for RTL. Um it's like the where RTL is like the programming languages you use. So, if you're building a chip, um, you typically don't draw all the different, uh, gates and so on. You write it in programming language like Verilog, um, which has sort of constructs that are quite It's like a super, um, parallel programming language. Um, which a lot of constructs that map well to hardware. Um. But then, there actually no good, uh, open source compilers for this system. So, the whole um, the whole hardware industry it it doesn't have the same open source, um, feeling as like software where all the good stuff is free and there's like this big vibrant market of of tools. Here, it's it's much more locked down to like big, um, big providers. But, people then, yeah, they compile this to and they get like a netlist. They can start building their schematics. They can also look at it. You get like these giant things that you can eventually send to, uh, to the fab. Um. But, yeah, so this But, they also need to do like do simulations, for example, for testing. You can do like formal verification.
让我复述一下,确保我理解了。所以 Verilog 有点像描述电路的程序设计语言。在送去制造之前,我们可能想做模拟,因为制造非常昂贵。所以我们可以设计电路并做模拟。
Let me just play that back so I understand it. So, so Verilog, it's a bit like a programming language for describing a circuit. And what we might probably before we go and get that fabricated is we want to do simulations because it's very expensive to get it fabricated. And it's So, um, so yeah, we we can design the circuit and do simulations.
是的,完全正确。而且非常昂贵。可能会花掉你很多钱,比如 90 年代英特尔那个著名的除法 bug,可能让他们损失了 5 亿到 20 亿美元。在 90 年代那是很多钱。后来还有其他例子,差点让一些公司破产,比如微小的 bug 竟然通过了制造。这和软件世界真的很不一样,软件里人们直接在线上修东西,快速行动,打破常规。硬件行业的人要谨慎得多,这也是为什么形式化验证、形式化方法在这里能蓬勃发展的原因。
Yeah, exactly. Um, it's it's super expensive, too. It It could It could cost you what they said like there was this famous, uh, Intel, uh, division bug in the '90s that probably cost them between 500 million and 2 billion dollars. And I mean, that's a lot of money in the '90s. Yeah. And there's I think there are other, um, other examples later on that nearly bankrupted some of these companies like tiny bugs that managed to to make it through, um, to the to the fab. And it's Yeah, it's it's a really different world again from software where we just um you know, people fix stuff in prod. You just move fast and and break things. Um that's definitely like Yeah, people are much more paranoid in the hardware industry, which is also why it's a it's a cool place for formal verification, formal methods to um to thrive.
好的,所以你是说目前有这些商业验证器和模拟器,它们贵得离谱。是不是每个席位 1 万美元左右?
Okay, so so you're saying at the moment there are these commercial, you know, uh verifiers and simulators, and they cost a ridiculous amount of money. So is is it something like $10,000 per seat or something?
对,一个 CPU 内核的价格。
Yeah, per for one CPU kernel.
明白了。
Right.
对,比如你想在数据中心扩展,跑一百万个智能体,我是说,计算机已经很贵了,但也没那么贵。那是什么?100 亿美元左右,对吧?仅仅是为了那些许可证。我认为这也是 AI 在硬件领域还不那么流行的原因之一,因为他们无法训练,模型没有针对这类工作负载进行训练,因为不可行。比如你没有那么多开源代码来开始训练,而且你也没有他们需要学习的工具。
Yeah, like say you want to scale this up in a data center with a million agents running like you're going to like I mean, computers are already expensive, but not that expensive. Uh what what Yeah, what is that? $10 billion or something, right? Just for for those licenses. And I think actually that's one of the reasons also why AI still is not as popular in the um in the hardware space because um they haven't been able to train like the the models are not as trained to this kind of workloads because they just um it's not feasible. Like you don't have all of the open source code out there to start the training, but you also don't have the tools that they need to learn to use.
你并没有掌握全部任务。你不能只是在上面做强化学习。我的意思是,他们显然在做一些这方面的工作。我们也能看到,从一代模型到下一代模型,它们确实在变好,但差距就像 Python 和 JavaScript 一样,天壤之别。我们一直在内部用 AI 开发这些 EDA 工具。我大概保持着最长运行智能体的世界纪录,有大约 20 个 GPT 智能体已经运行了 6 个月左右。它们还在取得进展。
You don't have all of the task. You can't just do all of the reinforcement learning on top of it. I mean, they clearly are doing some of that. We can also see from model generation to model generation that they are getting better, but it's kind of day and night like Python or JavaScript. We've been developing these EDA tools in-house using AI. I probably have the world record for the longest-running agents, having some 20 GPT agents running for around 6 months now. They're still making progress.
所以,据我理解,和 Anthropic 类似,Colina 发过一篇博客文章,他们有一个 C 编译器的功能规格说明,用了大约 4 万个智能体,复现了一个 C 编译器。你也为智能体做了类似的事情……
So, as I understand correctly, similarly to Anthropic, Colina had this blog post out, and they had a functional specification of a C compiler, and they had about 40,000 agents, and they reproduced a C compiler. You did a similar thing for an agent for a...
顺便说一句,是 4.8,对吧?所以……
4.8, by the way, right? So...
嗯,这也很有意思。你知道,这正是我想谈的,对吧?有些领域已经进化得如此完善,以至于我们有了功能规格说明,而且相当连贯,这很诱人。所以我们的想法是,用一大堆智能体,基于这些测试来复现这个软件的功能,并且递归地、智能体式地去做。但我的异议是,我认为关键不在于你最终到达哪里,不在于功能和测试是否通过,而在于你是如何到达那里的,以及它的结构如何。智能体式编码有一种倾向,就是构建出一个看似能用的意大利面条式代码。但如果你想想你谈到的这个项目,我记得你说过有 50 万行代码。而 5 年后,所有这些代码,你们可能大部分都没读过。这听起来像是个问题。
Well, that's an interesting thing as well. You know, because this is what I want to get to, right? It's tantalizing that there are some domains that are so well evolved that we have a functional specification, and it's reasonably coherent. So the idea is we get a shitload of agents and we reproduce the function of this software based on these tests, and we do it recursively, agentically, and so on. Now, my contention with this is that I think it's not about where you end up. It's not about the functions and the tests passing. It's about how you got there and how structured it is. There is this tendency with agentic coding to build a spaghetti monster which seems to work, but if you think about it in this project you're talking about, I think you said there's like 500,000 lines of code. And in 5 years' time, there's all of this code that probably you folks haven't read most of it. That sounds like a problem.
是的,我觉得对于很多这类事情,我们寄希望于某种摆脱代码复杂性的逃逸速度,即模型改进的速度会超过代码变糟的速度。我有过 3 天的 Fable 访问权限,真的很好。它确实清理了一些 GPT 5.5 搞乱的东西。
Yeah, I think for a lot of this stuff, we're relying on the hope of some kind of escape velocity from code complexity, that the models are going to keep improving faster than our code gets messed up. I had those 3 days of Fable access, and it was really good. It definitely cleaned up a few things that GPT 5.5 had messed up.
但真的如此吗?你觉得它会不会有欺骗性?因为当 Fable 出来时,4.8 突然就显得很糟糕。而当 4.8 出来时,它和 4.6 相比看起来很棒。所以这里面有某种欺骗性,但就像一种障眼法。
But did it though? Do you think it could be deceptive? Because when Fable comes out, all of a sudden 4.8 looks terrible. And when 4.8 came out, it looked amazing compared to 4.6. So there's something deceptive about it, but it's like a parlor trick.
是的,没错,确实如此。原则上,这一切都可能是烟雾和镜子。但我认为这就是为什么有硬性测试也很好,对吧?所以我也可以看到,实际上我突然在更客观的测试上取得了进展,而这些进展我有一段时间没看到了。对于其他模型,所以我想这给了我一些信任。但奇怪的是,无法对你代码中的所有内容有同样深入的理解,我确实很怀念那种感觉。
Yeah, no, it's true. In principle, it could all be smoke and mirrors. But I think that's why it's good to have also the hard test, right? So I could also see that it was actually suddenly I was making more progress on the more objective tests than I hadn't seen for a while. With the other models, so I guess that gives me some trust. But it is weird to sort of not have the same in-depth level of understanding everything in your code, and I definitely miss it.
我对这带来的后果很感兴趣,因为你说你今年早些时候发了那篇博客文章。然后我们在电话里聊天时,你说你注意到有些东西是错的。而且有一种倾向是积累理解债务。当这种情况发生时,我认为从进化的角度来看,你就被困住了。因为我认为对事物本质的深入理解是下一个设计决策、下一次进化的基础。那么,你是否会陷入这种鱼缸效应,即你处于无人地带,真的不知道该做什么?
And I'm interested in what the consequences of that are, because you said you put this blog post out earlier in the year. And then when we had a chat on the phone, you said that you noticed there were some things that were wrong. And there's this tendency though to accumulate understanding debt. And when that happens, I think from an evolutionary point of view, you're stuck. Because I think deep grounded understanding of how things are is the basis for the next design decisions, the next evolution. So do you get into this kind of fishbowl thing where now you're in no man's land and you don't really know what to do?
是的,我确实认为尽可能多地理解、并花时间去理解是很重要的。我认为对于编译器来说,有一些非常重要的架构设计决策。然后还有很多只是实现这 100 个不同的函数。它们都只是,尤其是现代编译器,一层又一层。你必须把这些东西从前端层降低到中间解释层,再到后端层。而其中很多你肯定不想自己写。它也会让人瘫痪,我想。我确信这也是 Tropical 选择它的原因之一。我们谈过这个 Program Bench 基准,我记得是 Facebook 发布的,它是个基准测试,任务是他们拿了 150 到 170 个程序。其中一些非常复杂,比如 FFmpeg 就是其中之一。AI 只需要在没有互联网的情况下重新实现它们。是的,基本上当它出来时,所有 LLM 都得了 0%,因为没有一个能通过所有测试。但我看到人们,我很少看到人们发布这个基准的结果时会说“哦,它通过了百分之多少的测试”。或者即使是 Fable,也是“哦,它答对了 70%、80% 的测试”。但问题是,“是啊,但我和制作这个基准的人聊过,他们说‘但它真的做对了吗?’”因为如果程序只通过了 70% 的测试,那它很可能不对。
Yeah, I do think it's important to understand as much as you can and have time for. I think for compilers, there are some very important architectural design decisions. And then there is a lot of just implementing these 100 different functions. Like they all just, especially modern compilers, it's layer by layer by layer. You have to lower these things from the front end level to the middle interpretation to the back end layer. And a lot of that you definitely don't want to have to write. It also paralyzes, well, I think. I'm sure this could be one of the reasons why Tropical also picked it. We talked about this Program Bench thing I think Facebook released, where it's a benchmark where the task, I think they took 150 to 170 programs. Some of them are really complicated like FFmpeg is one of them. The AI just has to re-implement them without internet access. Yeah, basically when it came out, all of the LLMs got 0% because none of them were able to pass all of the tests. But I've seen people, I very rarely see that when people post benchmarks for this thing. They always post it's like, "Oh, the percentage of tests that it got right." Or even like the Fable it was like, "Oh, yeah, it got like 70, 80% tests correct." And it's like, "Yeah, but I talked with the people who made the benchmark and like, 'Yeah, but did it get any of them actually right, you know?'" Because if the program only passes 70% of the tests, it's probably not right.
我知道,在某种程度上这就是我们之前谈到的线索。在过去 60 年里,回到行为主义,在结构、能力和预测之间一直存在这种线索,对吧?所以本质上,Program Bench 是在论证你可以从程序的外部行为中学习到它的面相。所以你可以学习到深层结构、约束等等。而我怀疑这是不可能的。当然,除非 LLM 已经知道源代码,因为它就在训练数据里之类的。但你对此怎么看?
I know, in a way this is the thread that we were talking about before. For the last 60 years, going back to behaviorism, there's been this kind of thread between structure and competence and prediction, right? So essentially this Program Bench is making the argument that you can learn the physiognomy of a program from its external behavior. So you can learn the deep structure, the constraints, and so on. And I suspect that that is not possible. Unless, of course, the LLM already knows about the source code because it's in its training data or whatever. But what do you think about that?
是的,我认为人类其实一直在这样做。我觉得有整个逆向工程领域,人们……是的,实际上我听了一些 FFmpeg 的人聊的东西,他们谈到他们放在里面的所有这些编解码器,很多时候他们根本不知道它们是做什么的。就像它只是一团晦涩的代码,他们可能甚至没有访问代码的权限。
Yeah, I think humans do it all the time actually. I think there's the whole field of reverse engineering where people... Yeah, actually I listened to something with the FFmpeg people where they were talking about all these codecs that they had that they put in there, and a lot of the time they have no idea what they do. Like it's just this obscure blob of code, like they maybe probably don't even have access to the code.
他们只有几个用这个编码的示例视频片段,也许有人分享过电影片段的截图,然后他们就得去尝试读取这个编码后的数据块,然后想,它在做什么?我们怎么能写一个程序来解码这个东西?不知怎么的,人们居然能做到。你知道,我们肯定有某种先验,对吧?比如人们通常会在视频编解码器里放什么?哦,可能有一些傅里叶变换之类的东西,或者他们可能会分块处理,或者我不知道。我没怎么做过视频编码,但是,就像完全黑盒一样,只是不断改进你的程序,可视化输出,直到电影结尾出来,你就知道它大概是什么样子,电影应该是什么样子。
They just have a couple of example video clips that were encoded with this, and maybe some people had shared screenshots with clips from the movie, and then they had to go and try to read this encoded blob and be like, what is it doing? How could we write a program that decodes this thing? And somehow people are able to do it. You know, we must have some sort of prior, right? Like what are the typical things people put in a video codec? Oh, there's probably some Fourier transforms of some stuff, or they probably chunk things up, or I don't know. I haven't done that much video encoding, but yeah, just like a completely black box, just try and keep improving your program and visualize the output until the end of the movie comes out and you have an idea of how this probably looks, how the movie is supposed to look.
是的,这也是一个有趣的事情。所以,我们可以在部分知识的情况下进行爬山搜索,对吧?所以我们可能有一个验证器。正如你所说,很多科学都是一种智能形式,你使用 MATLAB,做拉普拉斯变换或图像绘图,你取一个分布,然后它一步一步成形。所以,你是在向未知迈步,每一步都像是生成与判别的不对称性,对吧?所以,我们能很好地判别,但还不能生成,但我们迈出这些步子,当我们到达那里时,我们把它压缩成一个模型,那就是智能的产物。
Yeah, well, that's an interesting thing as well. So, we can hill climb in a partial knowledge regime, right? So, we might have a verifier. And as you say, like a lot of science is a form of intelligence where you're using MATLAB and you do like a Laplacian or you do an image plot, you take a distribution and it takes shape one step at a time. So, what you're doing is you're taking steps into the unknown, and every single step, it's like there's a generation discrimination asymmetry, right? So, we can discriminate well but we can't generate yet, but we take these steps and then when we get there, we now kind of collapse that into a model, and then that's the artifact of the intelligence.
对,这就像 A* 搜索之类的,你边走边修剪搜索树,比如到目前为止什么看起来最成功。
Right, it's like an A* search or something where you prune your search tree as you go, like what seems most successful so far.
正是如此。但说到你的观点,我们每一步都在使用工具箱。所以,你知道,有些人很擅长谜语或智力测试,这很大程度上是技巧,对吧?因为他们有一个抽象技巧的工具箱,但然后就有这个问题:也许仅此而已。也许有这些自然的模式,这些抽象,我们可以组合起来,处理智能环境中的任何新颖性。所以,因此,也许语言模型可以自主做到这一点。我认为它们可能可以,是的。
Exactly. But to your point, every single step of the way we're using the toolbox. So, you know, there are some folks that are really good at riddles or intelligence tests, and a lot of it is skill, right? Because they have a toolbox of abstract tricks that they use, but then there's this question of well, maybe that's all it is. Maybe there are these natural patterns, these abstractions that we can compose together and handle any kind of novelty in the intelligence setting. So, therefore, maybe language models could do this autonomously. I think they possibly could, yeah.
这是一个好观点,因为我曾经试图在 Twitter 上从 Noam Brown 那里得到答案,因为你看到,你从更好的预训练模型开始,然后他们开始对它进行强化学习,似乎强化学习的效果越来越好,取决于模型一开始有多好,在某种意义上。
This is a good point, because I tried to get an answer from Noam Brown at some point on Twitter about this, because you see how the better pre-trained model you start with, and then they start reinforcement learning on it, and it seems like the reinforcement learning works better and better depending on how good the model was to begin with, in a sense.
是的,非常有趣。我的意思是,我同意你的看法。我认为说 LLM 不智能是错误的,对吧?因为存在某种组合闭包,这意味着从 LLM 的原始元素中,我们可以进行爬山搜索,构建一些计算结构来解决问题,但还不止于此。它可以发生在不同的抽象层次上。所以,如果 LLM 有更高的抽象,那么它们可以遍历更高抽象的组合闭包,解决这个问题。但似乎发生的是,我的意思是,没有持续学习,所以那些不会被添加到库中供以后重用,但也没有抽象。所以,人类会做的是,他们会看这个计算图,然后说:“啊,我看到那只是这里这个东西的类比。”我现在要把它压缩成一个新变量,屏蔽掉所有复杂性。为什么语言模型不这样做呢?
Yeah, super interesting. I mean, I agree with you. I think it's wrong to say that LLMs are not intelligent, right? Because there is some kind of a combinatorial closure, and that means from the primitives in an LLM, we can do hill climbing and we can build some computational structure to solve problems, but that's not quite it. And it can happen at different levels of abstraction. So, if the LLMs have higher abstractions, then they can traverse the combinatorial closure of the higher abstractions and they can solve the problem. What seems to happen though is, I mean, there's no continual learning, so those don't get added to a library and reused later, but there's also no abstraction. So, what humans would do is they would look at this computational graph and they would say, "Ah, I see that that's just an analog of this thing over here." And I'm now going to compress it into a new variable that screens off all of that complexity. Why do language models not do that?
我认为你触及了持续学习这类事情,这显然是我们缺乏的一大块。我认为语言模型在预训练期间确实会这样做。我确实认为,例如,Mistral 在我们的代码库中表现出色,就是因为它不知何故看到了大量代码,它似乎对 bug 和问题等有这种直觉,那一定是类似那样的,你知道,它以前见过这些模式。它有某种抽象……但确实,当你更实时地运行它们时,它们做得不太好。比如它们不会,我认为现在很多人在思考这个问题,我们如何做得更好的持续学习,这样它就能在飞行中很好地构建所有这些抽象和这些东西。我的意思是,我知道有些公司积极反对这样做。
I think you're getting into continual learning type thing, which is obviously a big thing we lack. I think language models do that during pre-training. I do think that's why, for example, Mistral is so good at doing stuff in our code bases, because it has somehow seen a lot of code, and it just seems to have this sort of intuition for bugs and problems and so on, and that must be kind of like that, you know, it's seen the patterns before. It has some kind of abstract... But it's true, like when you actually are running them more live, they don't do it very well. Like they don't, and I think a lot of people these days are thinking about this, how can we do better continual learning so it can just build all of these abstractions and these things really well on the fly. I mean, I know some companies are actively against trying to do that.
为什么?
Why?
比如我认为 Anthropic,对吧?我认为 Dario 说他也认为这是一个大安全问题,因为如果智能体在飞行中学习太多,你很容易失去所有对齐工作,然后它就会离安全检查点越来越远。但我认为有足够多的人在研究这个,所以它可能会发生。但这绝对是一个大的未知数,对吧?我认为这也必须改变我们服务模型的很多方式,例如,因为突然你需要能够,如果我们所说的持续学习是指实际实时更新权重,那将给当前范式带来很多麻烦,因为那样你就突然不能对所有客户使用相同的权重了。我的意思是,我想 Thinking Machines 有这样一个东西,他们有一个共享模型,然后上面有一些 LoRA。
Like I think Anthropic, right? I think Dario said he considers it a big safety issue, too, because you could easily lose all of the work that has gone into alignment if the agent is learning too much on the fly, then it gets further and further away from the sort of safe checkpoint. But I think there are enough other people working on it that it'll probably happen. But it's definitely like a big unknown, right? And I think it'll also have to change a lot of things in how we serve models, for example, because suddenly you need to be able, if we mean by continual learning that we actually update the weights live, it's going to cause a lot of trouble for the current paradigm because then you suddenly can't use the same weights for all customers, too. I mean, I guess Thinking Machines has this thing where they have one shared model and then some LoRAs on top.
嗯。
Mhm.
所以也许类似的东西可以行得通,但仍然需要你不知何故一直这样做。而且我们也在常规计算中研究替代计算形式。
So maybe something like that can work, but it's still like you have to somehow keep doing this all the time. And so we're also working in normal computing on alternative forms of computation.
是的。
Yeah.
比如这些热力学芯片,或者只是某种非常规计算。实际上,有时很难在这些模拟电阻器中保持记忆,除非你在推理的同时持续学习。比如你必须使用这些奇特的基板存储器,如果它们必须是永久的,但如果你只是使用像 DRAM 中的基本电容器,它们需要不断刷新。所以要么你必须花费大量能量,要么你就想让学习永远持续下去。这样它们就会自动刷新。这可能也是最接近大脑工作方式的方式,对吧?就像它们永远不会冻结。
Like these thermodynamic chips or just sort of unconventional computing. It actually becomes sometimes hard to keep the memories in these analog resistors unless you keep learning at the same time as you're doing inference. Like you have to use these fancy substrate memories if they have to be permanent, but if you're just using basic capacitors like you have in DRAM or something, they require constant refreshing. So either you just have to spend a lot of energy on that, or you just want to keep the learning going forever. That way they kind of automatically refresh. And that's probably the most similar to how brains work, right? Like they never just freeze.
我觉得像突触这些东西,它们总是在适应。
I think like the synapses and stuff, they're always adapting.
对,没错。我觉得在所有用来类比智能的词里,适应性是第一位。从某种意义上说,当我们看 Claude 时,你可以想,Claude 是模型本身吗?还是整个生态系统?我觉得是生态系统,对吧?所以它是自适应的,因为……
Yeah, exactly. I think out of all the words we can use to analogize intelligence, adaptivity is the number one. And in a sense, when we look at Claude, you can look at, you know, is Claude the model? Is it the ecosystem? I think it's the ecosystem, right? So it is adaptive because...
是性格或者人设之类的东西。
The personality or the persona or something.
嗯,你知道,人们倾向于拟人化,但不是的,比如有数百万人在用 Claude Code。因为这和我们如何在世界中形成因果抽象有间接关系。我们的方式是,我们身处世界,是智能体,能做决策,能调和不确定性,能说如果我这么做会发生什么,我能和朋友分享,然后有一个美妙的渗透过程,抽象就这样被嵌入。现在,Claude Code 也发生了同样的事,对吧?因为有人在他们的机器上使用它。他们实际在跑测试,在做反事实推理,然后所有这些都被重新训练进下一版 Claude,然后别人会复用那个抽象。所以,我们作为一个系统具有适应性。
Well, you know, there's a tendency to anthropomorphize, but no, like millions of people are using Claude Code. Because this is tangentially related to how we come up with causal abstractions in the world. So, the way we do it is we are in the world and we are agents and we can make decisions and we can reconcile uncertainty and we can say what would have happened if I did this and I can share it with my friends and there's this wonderful percolation process where the abstraction just becomes embedded. Now, that does happen with Claude Code, right? Because there are people using it on their machines. They're actually running tests, they're doing counterfactuals, and then all of this gets retrained in the next version of Claude, and then someone else will reuse that abstraction. So, we have the adaptivity as a system.
共享大脑,所有经验都汇入那里,叫什么“主体”之类的。再回溯上去,然后对。
Brain shared brain where all of the experiences go where it's called the bulk or something. Go back up and then yeah.
对,但有趣的是要讨论,这本质上是一样的吗?比如,如果我们真有一个假设中的实时自适应、发散的 Claude,它会比我们已有的 Claude 好很多吗?这也和你们做的工作有关,因为,你知道,在我看来,智能的过程就是创造这些粗粒度化,也就是技能。这有点像你们用 ASIC 在做的事。你们在构建定制硬件,让某些类型的计算变得非常快。这几乎就像是智能的结果。所以你说:“我要拿一个非常非常复杂的东西,把它削减,用我最好的方式表示,然后把它烧进一个不可适应的硬件基底。”这么说公平吗?
Yeah, so but it's interesting to discuss is that effectively the same? Like if we actually had some hypothetical real-time adaptive divergent Claude, you know, would it be much better than the Claude that we already have? And it's also related to the work that you guys do because, you know, the process of intelligence in my view is the creation of these coarse grainings, you know, skills. And that's kind of like what you guys are doing with ASICs. So, you're building this customized hardware for making certain types of computation go really really quickly. And it's almost like that's the result of intelligence. So, you say, "I'm going to take a very very complicated thing, and I'm going to whittle it down and represent it in the best way I can, and then I'm going to bake it into a non-adaptable hardware substrate." Is that fair?
我觉得这正是我们现在在做的事情之一,就是尝试构建最适合我们现有模型的硬件,并真正实现超高效的推理。你可以说,过去至少问题在于,如果你把硬件锁得太死,就会阻碍自己在软件方面的创新,对吧?比如英伟达的芯片对创新来说一直很不错。它们相当灵活。当然,它们也在很多方面引导了 AI 的发展方向,比如偏向矩阵乘法等等,但人们仍然能够创新。但另一方面,我认为随着这些 AI 辅助 EDA 工具的出现,硬件制造正变得越来越容易,所以这有点缩短了循环。也许现在人们开始考虑制作自己的 CUDA 内核,对吧?不久之后,我们可能就不会再写 CUDA 内核,而是直接为我们想要的每个东西设计定制电路,这并非不合理。所以,如果你想让新算法跑得飞快,就为它设计一个专用电路。AI 会帮你优化并验证正确性,当然你仍然需要晶圆厂,但人们在批量投片方面也越来越擅长。但我认为拥有真正灵活的硬件会非常有趣。比如在芯片上学习,理想情况下,它越有适应性,我们就越不需要把东西烧进去。我是说,我不知道,我只是觉得那更酷。
I think that is one thing we are working on right now, this idea of trying to build hardware that best fits the models we have right now and just really make super-efficient inference. You could say in the past at least the issue has been you stop yourself from innovating on the software side if you lock down your hardware too much, right? Like the Nvidia chips have been pretty good for innovation. They're pretty flexible. Of course they have also guided the way we do AI in a lot of ways like towards matrix multiplications and so on, but still people have been able to innovate. But on the other hand, I think hardware is getting easier and easier to make with these AI for EDA tools, so it's sort of shortening the loop. Maybe now people are thinking about making their CUDA kernels, right? It's not so unreasonable to think that soon we'll just be doing our AI for making instead of CUDA kernels, we'll just make some custom circuits for every single thing we want. So, if you want your new algorithm to run really fast, you just design a specialized circuit for that. And AI helps you optimize it and check that it's correct and of course you still have the fab, but people are getting better at batching things for the fab. But I think having really flexible hardware is going to be very interesting. Like the stuff where it's learning on chip and ideally maybe the more adaptability that it has and the less we need to bake in. I mean, I don't know, it's just more cool, I think.
对,但这不就是递归式自我改进的绝佳例子吗?因为正如你所说,我们在构建 AI,然后 AI 又帮我们构建更好的内核、更好的软件、更好的硬件,这反过来又让 AI 变得更好,然后,你知道,就形成了这样一个循环。
Yeah, but isn't that a wonderful example of this recursive self-improvement because as you say we are building the AI and then the AI is helping us build better kernels, better software, better hardware, which then in turn makes the AI better and then, you know, you get this kind of a loop.
我的意思是,这在某种程度上就是递归式自我改进,对吧?
I mean, this is recursive self-improvement in a way, right?
没错,没错。但我们应该把这个落到实处。所以,有个概念叫自动形式化。你们在网站上基本上说,你们在构建芯片方面做了类似 AlphaProof 的事情。粗略地说,AlphaProof 就是,你知道,他们在 2024 年国际数学奥林匹克竞赛中获得了银牌。当时他们做的是用语言模型生成一堆 Lean 代码。然后显然有点乱,所以部分形式化是手工完成的,他们做了多次迭代等等。然后他们用 Lean 做验证。而你们在为芯片做的事情有点像那样。
Exactly. Exactly. But we should bring this to life. So, there's this concept called auto-formalization. And you guys on your website say basically that you have done something similar to AlphaProof in respect of building chips. And roughly speaking, AlphaProof is when, you know, they won silver at the IMO in 2024. And what they did back then was they used a language model to generate a bunch of Lean code. And then it obviously was a little bit messy, so some of the formalization was by hand and they did multiple renditions and so on. And then they would do verification with Lean. And you're kind of doing something like that for chips.
对,对,所以这很有趣,因为自动形式化,我们定义为把人类的规格说明写出来,转换成形式化规格,比如用 Lean。然后当然,你还需要证明步骤,证明你的代码,无论是什么,确实满足那个规格。我认为有些公司,比如 Axiom,他们非常专注于这部分,对吧?而且在某种意义上,AlphaProof 的主要工作也是从形式化开始的。然后难点在于训练模型提供证明或反证。我认为 AlphaProof 中一个非常巧妙的技巧是,当他们做证明的形式化时,他们是否做对其实并不重要,因为他们只是让模型提供证明或反证。所以如果他们搞错了,命题不再为真,那么它就会证明该命题不成立,或者它仍然可用。当然,当他们实际做 IMO 挑战时,他们希望自动形式化是正确的。所以他们手工完成了。但他们训练时并不需要它,我认为这有助于规模化。我觉得是的,顺便说一句,我们也可以在硬件上做类似的事情。拿一个芯片设计,然后基本上可以提出一些可能成立也可能不成立的属性,然后训练模型尝试证明或反驳该属性是否成立。
Yeah. Yeah, so it's interesting because there is the auto-formalization, which we define as taking human specifications and writing them up and turning them into formal specifications in Lean, for example. And then of course, you also need the proof step where you prove that your code, whatever you have, actually satisfies that specification. And I think some companies like Axiom, for example, they're very focused on this part, right? And also in some sense, AlphaProof, that's also the main thing it did was it started with a formalization. And then the hard part was training the model to provide a proof or disproof. I think it's actually a really nice trick in AlphaProof that when they did the formalization of the proof, it didn't really matter if they got it right or wrong because they just asked the model to provide a proof or disproof. So if they got it wrong and it was no longer true, then it would just prove that it was not true or like it would still be usable. Of course, when they actually did the IMO challenges, they wanted the auto-formalization to be correct. So then they did it by hand. But they didn't need it for the training, which I think helps scale it up. I think yeah, so we can do a similar thing with hardware by the way. It's pretty easy to just take some chip design and then you can basically just come up with some properties that may and may not be true and then you can train the model to try to prove or disprove that this thing holds.
所以,那都是关于创建证明,但自动形式化在某种意义上更难,因为为它创建训练数据更难,对吧?我认为这也是过去两年强化学习以来 AI 的故事。任何你能为它创建好的强化学习环境的东西,你大概都能学会,但其他任何东西现在都遥不可及。而且有些芯片有数千页的规格说明。如果你想把它变成形式化模型,如果某个地方错了一两个词或一两个数字,那就不起作用,或者你证明的东西就不重要或不相关。我认为这在芯片行业一直是个问题。他们试图通过正交团队来解决。一个团队设计芯片,一个团队设计测试,另一个团队设计测试的测试,比如功能覆盖率,他们衡量测试测试了什么,然后逐项核对。希望如果三个团队都以同样的方式阅读和理解,他们就有正确的想法。你可以尝试用 AI 做类似的事情。你可以争论如果同一个模型做这三项工作,是否真的正交。我不认为它完全正交。我确实认为这些模型在长时间运行时存在足够的熵。你肯定能通过多次做事并看它们之间是否一致来发现很多错误。但这绝对是一个信任练习和人类练习,要弄清楚如何让硬件工程师信任,以及如何让他们容易验证我们的形式化模型是否符合他们的想法。你可以想出各种技巧,既可以让你的理解更直观,比如你的理解是什么,是否匹配,也可以提问,看你是否同意这些问题,或者展示例子。你也可以尝试某种来回。可以尝试很多技巧,但这是个有趣的问题。我觉得没有人,比如 AlphaProof 和所有其他人,真正解决了这个问题,因为他们的数学陈述只有一段,不是数千页长。
So then, but that's all about creating the proof, but then the auto-formalization in some sense is harder because it's harder to create the training data for it, right? I think that's also a story about AI in the last two years since reinforcement learning. Anything you can create a good RL environment for, you can probably learn, but anything else is out of reach right now. And so some of these chips have thousands of pages of specifications. If you want to turn that into a formal model, if you get just a couple of words wrong somewhere or a couple of numbers, then it doesn't work, or what you prove is not important or relevant. And I think it's always been an issue in the chip industry. They've tried to solve it by having orthogonal teams. So one team designs the chip, one team designs the tests, and another team designs tests of the tests, like functional coverage, where they measure what the tests test and check everything off. Hopefully, if all three teams have read something the same way and understood it the same way, they have the right idea. You can try to do something similar with AI. You can argue whether it's really orthogonal if it's the same model doing each of the three jobs. I don't think it is quite orthogonal. I do think there is enough entropy in these models when you do long runs. You can definitely find a lot of bugs by doing things many times and seeing if there's agreement between them. But it's definitely a trust exercise and a human exercise to figure out how we get hardware engineers to trust and how we make it easy for them to verify that our formal model fits with what they thought. You can come up with all kinds of tricks, both to make it more visual, like what your understanding is and whether it matches, and ask questions and see if you agree with the questions, or show examples. You can also try some kind of back-and-forth. Lots of tricks you can try, but it's an interesting problem. I feel like no one, like AlphaProof and all these other people, never really solved this because their math statements were just one paragraph, not thousands of pages long.
哦,是的。有趣。因为 AlphaProof,我想他们想用 Lean 4,但 Lean 4 几乎没什么东西,所以他们创建了一个从 Lean 3 到 Lean 4 的转换器,他们需要很多人来修复它。但我想我的观点是,他们需要在大量的 Lean 代码上微调一个语言模型,然后他们用 Lean 作为中间语言。有趣的是,他们的新模型在 IMO 上赢得金牌,他们根本没有做任何验证,不是在形式意义上。但那么是否存在一个频谱?因为你刚才描述的方式,它不是二元的,对吧?你谈到了测试覆盖率和不同的视角,盲人摸象。所以有这些功能测试和这些描述,我们可以做视觉检查等等。那么最终会不会是我们总是在与我们不完全理解的东西搏斗,但尽可能多地使用信号?
Oh yes. Yeah, interesting. Yeah, because with AlphaProof, I think they wanted to use Lean 4, and there was hardly any stuff for Lean 4, so they created a converter from Lean 3 to Lean 4, and they needed lots of human people to fix it. But I guess my point is that they needed to fine-tune a language model on a ridiculous amount of Lean code, and then they were using Lean as an intermediate. Interestingly, with their new model that won the gold on IMO, they weren't doing any verification at all, not in a formal sense. But then is there a spectrum? Because the way you were just describing it, it's not binary, right? You were talking about test coverage and different perspectives, the blind men and the elephant. So there are these functional tests and these descriptions, and we can do visual inspection and so on. So will it end up being a case where we're always wrestling with something we don't completely understand, but we're using as many signals as possible together?
是的,我认为我们现在的发展方向也是试图涵盖更多的规格创建过程,对吧?显然,当我们开始时,人们已经写下了所有这些规格,我们想帮助他们处理这些。但当你这样做时,你也错过了进入整个规格创建过程的全部意图。你不知道里面某个地方是否有一些数字,你不知道他们为什么选择这些数字而不是其他数字。我认为通过内化更多的过程,在某个时候你至少可以像任何人类对这块芯片那样安全,对吧?而且可能也有一些没人关心的歧义。
Yeah, I think where we're going now is also trying to encompass more of the spec creation, right? Obviously when we start out, people already have written down all of these specs, and we want to help them out with those. But when you do that, you also miss out on the whole intent that went into the whole process of creating the spec. You don't know if there's just some numbers somewhere in there, you don't know why they chose those numbers and not some other numbers. I think by internalizing more of the process, at some point you can at least be as safe as any human could have been about this chip, right? And there might also just be some ambiguity that no one cares about.
歧义这件事很有趣。Eric Curiel 有一个精彩的演讲叫《数学不表征》,他谈到了广义相对论和它的四种完全正交的表征。实际上 AlphaProof 也有类似的情况。那么在 Lean 中有多少种方式来表示非负数?显然有很多。谷歌的人,只是为了刺激模型,用不同的方式表示问题。这在某种程度上几乎是训练数据里的事,对吧?如果他们用不同的方式自动形式化同一个问题,他们实际上得到不同的问题。
The ambiguity thing is interesting. There's a wonderful talk by Eric Curiel called 'Math Does Not Represent', and he was talking about general relativity and four completely orthogonal representations of it. It was a similar thing actually with AlphaProof. So how many ways are there in Lean to represent non-negative numbers? Apparently there's quite a few. The Google guys, just to stimulate the model, were representing the problems in different ways. Is this almost an in-the-training-data thing in a way? If they auto-formalize the same problem in different ways, they actually get different problems.
是的,既在智能方面,我们如何引导模型做得更好,也在最终输出的可读性和抽象性方面。我想当人们想到自动形式化时,我们几乎有一种理想主义的观点,认为存在一个唯一真实的表征,它将是可读的,等等,但这似乎相当模糊。
Yeah, both in terms of the intelligence, so how can we bootstrap the model to do better intelligence, but also in terms of the legibility and abstraction of the final output. I guess when people think of auto-formalization, we have this almost idealist view that there is a one true representation and it's just going to be legible and everything, but it seems quite vague.
是的,存在唯一真实的表征吗?
Yeah, is there one true representation?
不,我绝对认为不存在唯一真实的表征。对于较小的芯片,比如浮点或加密芯片之类的,规格实际上相当简单。更多的是当你遇到大型系统级的东西时。我认为软件在很多方面领先于硬件。在形式化方面,他们落后,但在结合 AI 思考架构方面,我认为现在很多人会把他们的架构文档交给 Claude 或其他模型,问:“嘿,你觉得这个怎么样?我们应该在这里移动东西吗?”我认为这整个讨论有助于模型理解你的意图,你关心什么,不关心什么。
No, I definitely think there's not one true representation. With smaller chips, like say you have just a floating point or crypto chip or something, the specification is actually pretty simple for those. It's more when you come up to the big system-level stuff. I think software is in many ways ahead of hardware. In terms of formalization, they're behind, but in terms of thinking about architecture together with AI, I think a lot of people now run their architecture documents by Claude or somebody and ask, 'Hey, what do you think of this? Should we move things around here?' And I think this whole discussion helps the models understand what your intents are and what you care about and what you don't care about.
为了举个例子让它更生动,你发表的那篇关于 DRAM 的文章,Alon Eyal。那篇文章谈到了这些定时 Petri 网。
And just to give us an example to bring it to life, there was that DRAM article you published, Alon Eyal. The article was talking about these timed Petri nets.
我在网上查了一下。显然,这是 20 世纪 60 年代用来描述分布式系统的一个概念。请解释一下。
And I did look this up on the internet. Apparently, it's a thing from the 1960s for describing distributed systems. Explain that.
我们之前讨论的正式化方法,非常偏向 RTL 层面,也就是那种周期级层面,你关心的是证明每个时钟周期都精确发生预期的事情。但当然,还有另一种正式化方法,对加法器、算术电路或加密这类东西非常重要。但人们遇到的很多难题,更多是系统级问题或协议级问题。在软件领域,更流行的是证明这类性质,比如这个系统永远不会死锁或活锁,或者协议级属性,以及所有不同部分之间的时序要求是否合理,是否存在内部不一致。目前,针对这些情况存在不同的形式语言。对于周期级的东西,人们使用 SystemVerilog 断言,也就是 SVA 这类东西。对于更高级的协议级东西,有更经典的如 TLA+,我想是 Lamport 提出的。我们尝试用这两种方式来进行形式化,因为它们对不同的问题各有用途。有可能在某个时候它们可以合并,如果你提供某种东西,就像你在做非常高级的数学,然后你可以把它一路化简到具体操作。但我认为这也很有趣——这是一个非常新的领域,所以我们正在探索不同的方法。而这些时间 Petri 网就是表示这种超级并行系统的一种方式,比如在存储器中,所有这些不同的 bank——你可以执行 activate,比如,如果你激活其中一行,你必须等待数据流到底部才能读取它。然后过一段时间,你必须刷新它,因为这些电容式 DRAM 单元必须不断刷新。但与此同时,你也可以在另一个存储 bank 中工作——芯片上可能有很多这样的 bank。所以存在某些 bank 间依赖和 bank 内依赖,还有 bank 组。人们确实在里面构建了非常疯狂的东西。
What we talked about before with the formalizations is very much at the RTL level, at the very sort of cycle level where you care about proving that the exact thing happens at every single clock cycle. But of course, there is another kind of formal that's very important for things like adders, arithmetic circuits, or crypto. But a lot of the hard problems people have are more like system problems or protocol-level stuff. In software, it's been more popular to prove things like this system can never deadlock or livelock, or protocol-level properties, and that all the timing requirements between different things make sense and there are no internal inconsistencies. Currently, there exist different formal languages for these things. For the cycle-level stuff, people use SystemVerilog assertions, SVA-type things. For the higher-level protocol things, there are more classic things like TLA+, I think by Lamport or something. We try to formalize things in both ways because they can be useful for different things. It's possible at some point they can all merge if you supply it, kind of like you're doing very high-level math and then you can reduce it all the way to the actions if you want to. But I think it's also interesting—it's a very new field, so we're trying to explore different ways. And these time Petri nets are one way to represent these super parallel systems that you have, for example, in memories where all these different banks—you can do activate, like for example, if you activate one of these rows, you have to wait for the data to run down to the bottom before you can read it. Then after a while, you have to refresh it because these capacitor DRAM cells have to be refreshed all the time. But you can also be working in a different memory bank—there might be lots of them on this chip—at the same time. So there are certain interbank dependencies and intrabank dependencies, and also bank groups. People build really crazy stuff in there.
非常酷。非常酷。对我来说,真正令人兴奋的是,有可能制造出能在某些类型任务上快几个数量级的芯片。这就是为什么我想稍微谈谈热力学计算,对吧?所以,你知道,显然,不是强迫晶体管稳定在 0 或 1,而是让噪声进行随机游走,并对其进行偏置,使芯片成为一个随机微分方程,对吧?我的意思是,这听起来很疯狂。这到底是怎么工作的?
Very cool. Very cool. I mean, what's really exciting to me is that it's possible to build chips that can do certain types of things orders of magnitude faster. And that's why I want to talk a little bit about thermodynamic computing, right? So, you know, apparently, instead of forcing transistors to settle at zero or one, you let noise do a random walk and bias it so the chip is a stochastic differential equation, right? I mean, that sounds crazy. Like, how does that work?
是的,我的意思是,这其实是当初让我加入 Normal 的原因之一。在那之前,在加入 Normal 之前,我在 Facebook(当时我们这么叫)的研究组做概率计算。我们当时在做贝叶斯神经网络,就是假设所有权重都有概率分布,然后尝试从你观察到的先验数据中推断后验。但很多这类技术都比较慢,因为你必须要么用不同的随机种子做大量重复实验,要么尝试用解析方法。那样可能有点失去了概率的意义。但有趣的是,你手头有这些芯片,而芯片制造商花了大量时间把系统中的每一点噪声都消除掉,为所有东西设定极其严格的余量,精度极高。这可能是世界上精度要求最高的行业。然后我们拿它们做什么呢?我们只是在各处添加随机性。所以为什么不尝试构建一个本质上就是随机的芯片呢?我认为大脑可能就有很多随机性。但在这里,我们做的第一块芯片就是这样一个电容器阵列,它们之间有一些可编程的电阻。然后你注入所有这些噪声,使其按照这些随机微分方程运行。然后你想,好吧,我们能拿它做什么?这是一个新的计算范式,我觉得探索起来非常有趣。你可以用它做的一件事,实际上结果是,你放到芯片上的权重微分矩阵——随机微分方程实际上表现得像是该矩阵的逆。所以我们可以尝试捕获它并取平均。
Yeah, I mean, this was one of the things that really got me to Normal in the first place. Before that, before Normal, I was at Facebook, as we called them, in the research group that does probabilistic computing. We were doing Bayesian neural networks, where you assume probability distributions on all your weights and you try to infer the posterior from the prior data that you look at. But a lot of these techniques were kind of slow because you had to either try to do lots and lots of repetitions with different random seeds, or you were trying to do it analytically. That kind of maybe takes a little bit of the point out of the probability. But then it's funny because you have these chips, and the chip manufacturers spend so much time getting out every single little piece of noise out of their systems and having these extremely sharp margins for everything, so much precision. It's probably the most precise business in the world. And then what do we do with them? We just add randomness everywhere. So why not try and build a chip that's just inherently random? I think the brain probably has a bunch of randomness. But here, the first chip we made was this array of capacitors, and you have certain resistances between them you can program. Then you infuse all of this noise to behave according to these stochastic differential equations. Then you think, okay, what can we do with that? It's a new computational paradigm that I found very interesting to explore. One of the things you could do with it, it actually turns out that the matrix that you put onto the chip in the weights differential—the stochastic differential equation actually behaves sort of according to the inverse of that matrix. So then we could try and capture that and average it out.
而且我们在电话里聊到这个时,你说了一些非常有趣的话。因为你知道,我们经常在这方面说得头头是道。前几天我和 Michael Jordan 聊过,你知道,就像,是的,我们需要不确定性量化,我们需要自适应计算。
And when we spoke about this on the phone, you said something very interesting. Because you know, we often talk a good game about this. I was talking with Michael Jordan the other day, and you know, like yeah, we need uncertainty quantification, we need adaptive computation.
是的,我认为这算是我们遇到的问题之一,因为我认为贝叶斯机器学习在生成式 AI 出现之前的某个时期确实非常强大,因为那时只有一个输出,所以把分布作为输出是有意义的。但现在你有了这些序列,正如你所说,你不断输出这些 token,我认为没有人真正关心某个特定 token 的不确定性,或者在其上有一个更好的分布。你真正想知道的是,在模型考虑了 10 个不同选项、回溯了所有这些之后,最终给出一个答案时,这个答案有多可信?你要么必须非常深入地研究机制可解释性,试图真正把所有的不确定性一路传递下去,要么你必须尝试使用一些更拟人化的方法,这些方法基于人类如何估计自己的不确定性,并尝试在非常高的层面上应用它。
And yeah, I think this was kind of one of the issues we had, because I think Bayesian machine learning was really strong for a certain amount of time, at a certain point in time before generative AI, because you had the one output and then it made sense to have the distribution as the output. But now you have these sequences, as you're saying, you keep putting these tokens out, and I think no one really cares about what is the uncertainty about one particular token and having a better distribution on that. You really want to know, after the model has thought about 10 different options and backtracked and all this stuff, and it comes out with a final answer, what is—like how much can I trust this answer? You either have to go really deep on mechanistic interpretability to try to really carry all of the uncertainty all the way through that, or you have to try and use some more maybe anthropomorphic methods that are based on how humans would estimate their uncertainty, and try and apply that at a really high level.
因为你有贝叶斯背景真的很酷,因为我对此的反思方式是,你知道,当你有一个想法,你对自己有多自信有一些直觉,这似乎是因为你有一个深层结构,所以你可以内省,你可以合理化,你可以说,好吧,有这个组成部分和那个组成部分,还有这些约束,我对那部分不太确定,那似乎是缺失的环节。
Because it's actually really cool that you've got a Bayesian background because the way I introspect about this is you know like when you have a thought and you have some intuitive notion of how confident you are and it seems to be because you have a deep structure so you can introspect and you can rationalize and you can say okay well there there's this component and this component and there are these constraints I'm not quite sure about that bit and that seems to be the missing link.
是的。是的,我们实际上做了一些实验,嗯,回到 Transformer 刚流行的时候,我们试图做,嗯,那是在我们确切知道公司要做什么之前,它将是用于硬件的 AI,嗯,所以我们想做预测,我们只是把贝叶斯定律构建到模型中,所以对于某个特定问题,它会尝试找到很多证据,然后它会说,如果答案是肯定的,如果答案是不同的东西,那么我看到这个证据的概率是多少,它会做很多次,然后最后,你知道,你可以用贝叶斯定律说,那么原始陈述为真或假的概率是多少。嗯,它实际上效果很好,我们做了这个,嗯,我做了一些内部的预测游戏,它击败了所有人。嗯,我们有两个版本。其中一个版本,嗯,我们有点像,可以说是神经符号的,你只是让模型提出所有这些概率,然后你手动计算。然后我还尝试了另一个版本,我只是告诉模型现在你用贝叶斯定律来做,或者什么的,不知何故它实际上做得更好一点。所以,我不知道它是怎么做到的,或者也许它只是能够,嗯,是的。也许是因为它能够回头思考,实际上其中一些值我可能是在胡扯,我应该忽略它们,或者嗯,但是,嗯,是的,这很有趣。嗯,我知道现在有一些人在训练,比如现在有一些基准测试,人们试图用 LLM 预测 Polymarket 这类东西,他们有基准测试,但我不太确定他们是否使用这些技术,还是都是端到端的强化学习,他们希望它自己学会一个好的方法论。
Yeah. Yeah, we actually did some experiments um back also when transformers first became popular where we tried to do uh it's before we knew exactly what the company was going to do that it was going to be AI for hardware um so we wanted to do like predictions and we just sort of building base uh law into the model so um for for some particular question it would try and find like lots of pieces of evidence and then it would say what is the probability that this like if if the answer is yes and if the answer is like different things like what is the probability that I would see this evidence and it would do that lots of times and then at the end you know you could use base law to say okay then what is the probability that the original statement was true or false. Um it actually worked really well like we did this um I did some like internal like prediction game and it and it it beat everyone. Um I'd we had two multiple versions. One of the versions um we sort of it was more uh neuro-symbolic you could say where you would just count like have the model come up with all of these probabilities and then you would manually calculate it. And then I also tried another version where it you I just told the model now you use Bayes' law and do it or something and and somehow it actually did a little bit better. So, I don't know how it how it did that or or maybe it just was able to uh, yeah. Maybe it's because it was able to look back and think actually a couple of these values I probably bullshitted and uh, I should just ignore that or um, but um, yeah, it's interesting. Um, I I know there are some people training like there are some benchmarks now for uh, where people are trying to predict polymarket and this kind of stuff with um, um, with LLMs and they have benchmarks for that, but I'm not quite sure if they use these sort of techniques or it's all just end-to-end reinforcement learning and and they hope it just picks up a good methodology by itself.
但这是这种活动专业化循环的一个绝佳例子。
But it's a wonderful example of this kind of um, activity specialization loop.
是的。
Yeah.
所以,你知道,嗯,在第一个月,嗯,有一个关于五编码的评论。所以,在第一个月,嗯,每天都有一个误报。你知道,所以我创建了一个技能,然后 Claude 会重新训练它,然后嗯,最终它有点收敛了。这就是五编码的看涨理由,对吧?你知道,你只是每天修复它,修复它,修复它,最终你会进入着陆轨道。而且它效果很好。但在某种程度上,这类似于你们正在做的事情,对吧?因为你们有这种外层循环,它处理复杂的东西,然后你压缩它,优化它,然后你把它烘烤成一种结晶的硬件。
So, you know, um, for the first month um, there was a comment with five coding. So, for the first month um, there was a false positive every single day. And, you know, so I I just created a skill and then Claude would just retrain it and then um, eventually it just kind of converges. And and this is this is the the bull case of five coding, right? That, you know, like you just every single day you fix it, fix it, fix it, fix it and eventually you'll you'll come into the landing track. And and it works really well. But in a way this is similar to what you guys are doing, right? Because you you you have this kind of outer loop which takes something complex and then you kind of um, compress it and you optimize it and then you bake it into a kind of crystallized hardware.
是的,我在想有一条埃隆·马斯克的旧推文。我不知道人们是否在嘲笑他。我的意思是,我不知道。他发了推文,但他说,嗯,为什么 LLM 不直接写二进制或汇编语言呢?
Yeah, I I was thinking about there is this older tweet by Elon Musk. I don't know if people were laughing at him. I mean, I don't know. They he tweeted it out, but it was he was saying like um, why don't uh, LLMs just write the binaries directly or the assembly
我看到了
I saw show
是的。我不明白为什么我们不能做到。这只是我们是否想这样做的问题,或者我们是否认为,因为我认为在世界上任何地方总是存在一些基本的计算问题。我们并不总是,近年来我们不太关注它们,因为我们太兴奋于 AI,我们只对 AI 能做得好的所有事情感兴趣,但我的意思是,显然像密码学或一些基本算法,你永远不希望 LM 去做,即使它们能做非常大的数字乘法。那太效率低下了,或者你不如为那个优化一个电路或一段代码,编译也有类似的问题,还有芯片综合,你试图探索所有这些不同的设计等等。所以你不一定希望 LM 去做,因为它太慢了,相对于一些超级优化的循环。嗯,是的,不要跳得太远,但再次思考国际象棋很有趣,当然你有 AlphaGo 那样的事情,一切都通过神经网络运行,对吧?但今天最先进的是像 Stockfish,当然,他们做了更像混合体,所以他们采用了神经网络,他们制作了一种特定类型的神经网络,当你改变状态时可以非常快速地更新,然后他们将其与超快速搜索结合起来,它实际上优于最好的开源 AlphaGo 类型的国际象棋引擎。
Yeah. I I don't see why we wouldn't be able to do it. I just it's a question of if we would want to do it or like do we think because like I think there's some fundamental computational problems always and in like everywhere in the world. We don't always like recent years we don't focus so much on them because we're so excited by AI and we want to we just interested in all the stuff that AI can do really well, but I mean obviously things like cryptography or like some of these like basic algorithms like you never want to want going to want the LM to just do it even if they can do very large number multiplications. It's just super inefficient or like you might as well optimize like a circuit or a piece of code for that and compilation has some of the same problems where you want to and also and chip synthesis where you're trying to explore all these different designs and so on. So you don't necessarily want to um have the LM do it because it's so slow versus some super optimized loop. Um Yeah, not to jump too much in things, but it's interesting to think about chess again like of course you have like the AlphaGo thing where it was like very everything ran it through a neural network, right? But but today the state of the art is in like Stockfish of course is that they took um they did more like a hybrid so they took the neural networks um and they made these like um like a certain type of neural network that can update really fast when you change the the state and then they they combine it with like just super fast search uh and it actually outperforms the best like open source like AlphaGo type chess engines.
哦,真的吗?使用类似自适应微调?
Oh, really? Using like adaptive fine-tuning?
嗯
Um
还有结构化推理。
And structured inference.
我不知道你是否称之为自适应。更像是他们采用了经典的国际象棋搜索引擎,并用神经网络替换了评估函数,但那种非常浅而宽的神经网络,评估起来非常非常快。嗯,是的,我认为对于综合和编译之类的事情,可能也有类似的情况,你可以尝试全部用 LLM 来做,但在某个时候速度是瓶颈,你可以从拥有更多知识和更多直觉中获益,但在某个时候,也有一个硬性的对话问题,你只想暴力破解一些东西。那时,你希望能够切换到更经典的算法。
I don't know if you call it adaptive. It's more like they took the classic chess search engine and they replaced the evaluation function with a neural net, but like kind of a very shallow wide neural net that they can that that is really really fast to evaluate. Um and yeah, I think that it could be that for something like synthesis and compilation, there's a similar thing where it's like you could try and do all with LLMs, but at some point that's the speed is a bottleneck and like you can throw you you get a benefit from having more knowledge and more intuition and all of this stuff, but at some point there's also just a hard conversation problem where you just want to brute force some stuff. And at that point, you want to be able to switch to like more classical algorithm.
完全正确。现在,之前我们谈到了热计算。所以你们做了一些工作,这些工作对于像马尔可夫链蒙特卡洛和扩散模型之类的东西有着巨大的潜力。但你上次和我谈话时确实说过,在某些情况下,这可能是一种虚假的经济。比如,你可能有一个扩散模型,它可能有一个不同类型的神经网络在末端。你可能会发现,你做扩散所获得的好处可能会被模型的另一部分所瓶颈。
Absolutely. Now, before we were talking about you know, thermal computing. And so you you you folks have have done some work that is incredibly like has huge potential for things like you know, Markov chain Monte Carlo, I think and diffusion models. But you did say to me when we spoke last time that in some cases it can be a false economy. Like for example, you could have a diffusion model which might have like a different type of neural network on the end of it. And and you might find that the benefit that you have doing the diffusion might be bottlenecked by another part of the model.
那么在实践中,我们能在哪里看到巨大的提升?
So in practice, where can we see a huge uplift?
我认为当你构建硬件时总是很有趣,因为总是存在与算法的协同设计问题,对吧?算法非常依赖于 GPU 和我们现有的硬件。所以你可以尝试针对瓶颈,比如我们可以像之前讨论的那样,制造对内存超高效的热 DRAM。那可能会加速推理,但只能加速这么多,比如现在推理时 GPU 利用率大约只有 10%。你可以在那里期望有 10 倍的提升,但你不确定如果突然有了那么多不同的架构,你能构建什么样的新算法。如果你真的全力投入的话。扩散模型也有类似的问题,对吧?现在的架构在随机性方面效率不高。比如 GPU,人们并不真的想到处采样大量的高斯随机变量。所以他们可能也会构建架构,试图通过不过分关注这些来消除瓶颈。然后我们可以构建新的硬件,使其更高效,并获得一些性能提升。但要真正充分利用它,你还必须提出新的模型,全力投入其中。
I think it's always interesting when you build hardware because there's always this co-design problem with the algorithms, right? The algorithms are so based around the GPUs and the hardware we have now. So you can try and target bottlenecks like you can make hardware like we talked about with the thermal DRAM that's super efficient for the memory. That might speed up inference, but it can only speed it up like say we have some like 10% GPU utilization now for inference. You could hope to have a 10x there, but you don't know if you suddenly have access to that much different architecture, what new algorithms could you build? If you really went all in on that. And there's some similar thing with the diffusion stuff, right? Now the architecture is not that efficient in terms of randomness. Like the GPUs, people don't really want to sample tons of Gaussian random variables everywhere. So they maybe also build architectures where they try to remove that bottleneck by not focusing too much on these things. And then we can build new hardware that makes them more efficient and that has some performance gain. But to really make the most of it, you also have to come up with new models that sort of go all in on that.
这些 Notion 和应用很有趣。你想过这个吗?现在所有应用都想成为你的中心 AI。比如 Notion 有这个,他们想与其他应用集成,Linear 也有 AI,他们想与其他应用集成,每个人都试图成为你的中心 AI 助手。然后所有其他应用都只是工具调用。
It's interesting these Notion and apps. Have you thought about this? How all of the apps now they want to be like your central AI. Like Notion has this, they want to integrate with the other apps and you have Linear have an AI and they want to integrate with the other apps and like everyone is trying to capture like being your central AI assistant. And then all the other apps will just be tool calls.
我知道。我得小心点,因为 Notion 现在赞助了一些广告。不过不,我实际上经常用它,因为它有一个很棒的智能体式界面。它有一个类似 CLI 的界面。所以我从 Claude 那里使用它。但他们想让你做的是付费使用内置的智能体,按 API 价格。而我真的看不出这样做的理由,因为我已经有了 Notion MCP,我有 Notion CLI。
I know. I have to be careful because Notion are now sponsoring some ads. But no, I actually use it a lot because it's got an amazing agentic interface. It's got like a CLI interface. So I use it from, you know, from Claude. But what they want you to do is to pay them to use the agents that are built in at API prices. And I don't really see the reason in doing that when I've just got a Notion MCP. I've got a Notion CLI.
我认为 API 价格确实在扼杀那个领域的创新,对吧?我的意思是,我认为 Codex,我认为他们实际上开放了更多,你可以使用你的账户。我昨天刚看到,我想也许 Anthropic 也开放了一些,你可以使用更多像 SDK 这样的东西,用你的账户,但价格差异太大了,如果你必须使用 API 定价,那么没有什么是有竞争力的。
API prices stuff is really killing innovation, I think, in that space, right? I mean, I think Codex, I think they are actually opening up more that you can use your accounts. I saw just yesterday, I think maybe also Anthropic opened up some that you could use more like the SDK you can use with your accounts and but it's just like the price difference so big that if you have to use API pricing like nothing is competitive.
嗯,我的意思是,你知道,我们不必深入探讨这个,但我认为创造力就是尊重约束。而且,你在牛津学过语言学,乔姆斯基总是说语言能力和语言表现之间有区别。所以,你知道,他多年来一直用推土机的类比,你知道,他说我也喜欢推土机。它们非常适合清除积雪。它们不是对科学的贡献。他甚至说过关于深蓝的事,计算机以那种方式赢得深蓝,有点像推土机赢得奥运会举重比赛。
Well, I mean, you know, we don't have to go too deep into this, but I think creativity is all about respecting constraints. And I mean, you studied linguistics at Oxford and Chomsky always says that there's a difference between linguistic competence and linguistic performance. So, you know, he's using this bulldozer analogy for ages, you know, he said that I love bulldozers too. They're great for clearing this snow. They're not a contribution to science. And he even said about Deep Blue that a computer winning at Deep Blue in the way it did is a little bit like a bulldozer winning the weightlifting competition at the Olympics.
我认为在某种程度上是这样。就像深蓝,我的意思是,现在也许现代国际象棋 AI 有点不同了。我实际上在国际象棋引擎上工作了 15 年。但我记得他也说过那句话,我有点惊讶,因为你会想,如果你的科学是生物学,或者你想理解人类如何做语言,那么当然,LLM 如何学习语言可能不太相关。但如果你的科学更抽象,什么是语言,我想你会真的有兴趣看到不同的系统发展语言,你可以比较,看看什么是共同的,什么是相同的,什么是不同的,等等。那会给你一个更广泛的理解,什么是这个东西的概念。我的意思是,当然你必须对 AI 语言有一些尊重,才能把它包括在内,而不是认为它完全是随机鹦鹉之类的东西。也许你会说,我真的不在乎它。我不想把它包括在我的语言模型中。但如果你确实认为它实际上在做语言,那么我认为它是否以与人类相同的方式做并不重要。也许如果它以不同的方式做会更有趣。
I think it was in a way. Like with Deep Blue, I mean, now maybe modern chess AI is a bit more different. I actually worked on chess engines for 15 years. But I remember him saying that too and I was a bit surprised because you'd think that if your science is biology or you want to understand how humans do language, sure then maybe the LLMs is kind of not so relevant how they learn language. But if your science is more the abstract of what is language, I thought you would be really interested in seeing different systems developing language and you can comparing and see what's common, what's the same, what's different and so on. That would give you a wider understanding of what is the concept of this thing. I mean, of course you have to have some kind of respect for AI language to even include it in just think it's like completely you know, just some stochastic parrot thing. Maybe you're like, I don't really care about it. I don't want to include it in my model of language. But if you do think it's actually doing language, then I think it doesn't really matter if it does the same way as humans or not. Maybe it's more interesting if it's in a different way.
我知道。我知道。顺便你提到了思维链。我们能从思维链中解读出多少?好吧,因为你知道,有些人只是称之为思维链缺失,比如斯巴鲁·卡马哈蒂,但实际上它可能是现在做可解释性的操作方式,对吧?为了真正理解他们在想什么。但你可以争辩说,思维链就像新闻秘书,而不是指挥者。所以它几乎是事后虚构。但那不太对,不是吗?因为那听起来有点像思维链和语言模型输出之间没有因果关系。那不是真的。那么我们能从中解读出多少呢?
I know. I know. You mentioned chain of thought by the way. How much can we read into chain of thought? All right, because you know, some people just call chain of thoughtlessness, like Subaru Kamahati, and but actually it is probably the modus operandi now for doing interpretability, right? For actually understanding what they're thinking. But you could argue that chain of thought is like the press secretary, not the orchestrator. So it's almost a post hoc confabulation. But that's not quite right, is it? Because that sounds a little bit like there's no causal link between the chain of thought and what the language model outputs. That's not true. So how much can we read into it?
是的,就像你想给模型一个思考的地方,对吧?当你做强化学习时,我的意思是,强化学习之前的思维链和强化学习之后的思维链,我认为是非常不同的,因为之前你试图提示它,一些技巧有时有效,有时无效,但一旦你做了强化学习,你需要它像图灵机一样,对吧?它需要有无限的内存,并且能够有东西可以操作。而当你只有 Transformer 和单次输出时,就像我的意思是,可能在乔姆斯基层级中,它会是一个完全不同的系统类型,比如我不知道,自动机之类的,对吧?那只能做有限的计算,但现在它可以做任意多的计算,它只需要学会如何去做。比如,你可以想象构建一个非常简单的 LLM 类型系统,如果有了思维链,它就会是一个通用图灵机。所以现在在这一点上,它只是不,因为它可以有这种状态,它可以不断读取和放置,我认为构建这样的结构会相当容易。但问题是它是否能学会它。
Yeah, it's like you want to give the model somewhere to think, right? When you do reinforcement learning, I mean, the chain of thought before the reinforcement learning and after reinforcement learning, I think it's very different because before it was kind of you try and prompt it and some tricks they kind of work and kind of doesn't work, but once you do the reinforcement learning, you need like it's like a Turing machine, right? It needs to have infinite memory and being able to have something to operate on. Whereas when you just had the transformer and just like single-shot output, it's like I mean, probably in the Chomsky hierarchy, it would be a completely different type of system, like I don't know, an automaton or something, right? That's like only finite amount of computation it can do, but now it can do as much computation as it wants, and it just has to learn to figure out how to do it. Like, you could imagine building a really simple LLM type system, and with access to chain of thought, it would be a universal Turing machine. So now at that point, it's just No, because it can have this state, and it can keep reading and putting I think that would be pretty easy to make a structure like that. But then the question is whether it can learn it.
但至少现在它有了表征能力,所以它能做到,而且显然有些东西在改进并起作用。
But at least now it has the representation capacity, so it can do it, and clearly something is improving and working.
这实际上是生态系统中更大的问题,即软件总是被设计成降低复杂性并引入渠道化的东西。电子表格就是一个很好的例子。现在,会计师使用电子表格,每个人都在用,它创造了一个人人都用的界面,并降低了系统的复杂性。这种智能体式 AI 到处制造混乱。这也是人们不发布产品的原因之一,因为它产生了一些短暂而复杂的东西——有点像加强版的 bash 脚本。所以我创造了一个只有我能理解的网络,而且它随着时间变得越来越专门化,所以我无法与他人分享。整个生态系统变得非常混乱。
This is actually the bigger problem with the ecosystem, which is that software is always designed to be a thing that reduces complexity and introduces canalization. So, spreadsheets are a great example of this. Now, accountants use spreadsheets, and everyone's using them, and it creates an interface that everyone uses and reduces complexity in the system. This agentic AI just creates spaghetti everywhere. And this is part of the reason why people aren't shipping, because it creates something ephemeral and complex—it's a little bit like bash scripting on steroids. So I've now created this web that only I understand, and it's becoming more and more specialized over time, so I can't share it with other people. And the entire ecosystem is becoming very messy.
是的,不,我绝对认为很多事现在正因为这个而崩溃。我的意思是,很多开源的东西就像每个人都从头写新代码,而不是试图聚在一起打磨这些共享库。但我的意思是,对很多事情来说,这比走那条路更高效,而且你可以让一切完全符合你的要求。但是,是的,我实际上觉得你谈到的电子表格和这种渠道化非常有趣,对吧?因为这些工具内部有很多领域知识。而这实际上也是我认为那些构建这些工具的人可能担心的,比如提取这些知识有多容易,以及人们构建他们工具的克隆。比如现在每个人都在锁定所有数据,对吧?因为他们意识到数据变得如此有价值。他们想自己保留或自己构建一些东西,而不是让别人使用。我认为很多专业工具也可能出现类似情况,因为工具基本上就是数据,它们经过很长时间的开发,找到所有正确的模式,一切都内置其中。那么你如何避免其他人进入并智能体式地提取和蒸馏它呢?
Yeah, no, I definitely think a lot of things are breaking now because of that. I mean, a lot of the open source stuff is like everyone just writes new code from scratch instead of trying to come together and hone these shared libraries. But I mean, it is just for a lot of things more efficient than going through that, and you can get everything just the way you want it. But yeah, and I actually think it's very interesting what you're talking about with the spreadsheets and this channeling, right? Because there's a lot of domain knowledge inside of these tools. And that's actually also what I think people might be worried about who build these tools, like how easy it is to extract that knowledge and people building clones of their tools. Like if you have everyone these days locking down all the data, right? Because they're realizing data is becoming so valuable. They want to keep it themselves or build something themselves rather than have other people use it. And I think it could be a similar thing with a lot of specialized tools, because the tools basically are data, like they're developed over so much time, finding all the right patterns, everything is built in there. So how do you avoid other people from going in and agentically extracting and distilling it?
我确信这发生在你身上过。网上的陌生人会说:“我刚生成了这篇论文。我刚生成了这段代码。你看看。”现在有大量的 AI 精神病,即你做一些稍微超出你专业领域的事情,Claude 会说服你,你的东西不是平庸的,而是很棒的。然后你与他人分享,其他人立刻看穿。在我看来,这也是一个严重的问题。因为我认为有些专家对事物有非常清晰的想法。他们一直在做软件工程,他们一直在科学领域,当他们使用 Claude 时,如果他们勤奋,那会很棒,对吧?因为他们实际上可以使用好的抽象和表征。但现在那里有一场污染的海啸。
I'm sure this has happened to you. Random people on the internet will say, "I've just generated this paper. I've just generated this code. Have a look at it." And there's massive amounts of AI psychosis out there, which is that you do things that are slightly outside your domain of expertise, and Claude will convince you that your stuff is not mediocre, it's great. And then you share it with other people, and other people immediately see through it. And this is a bit of a serious problem as well, in my opinion. Because I think there are experts out there who have really clear ideas about things. They've been doing software engineering and they've been in science, and when they use Claude, it's brilliant if they're diligent, right? Because they can actually use good abstractions and representations. But there's now a tsunami of pollution out there.
是的。
Yeah.
而且它破坏了社会契约,对吧?比如过去,如果我写了东西请你读,你至少可以假设我花在写作上的时间比你阅读的时间多 10 倍。但现在你对任何事情都持怀疑态度,因为我为什么要花时间读一些你自己可能都没读过的东西呢?是的,我的意思是,但很多这些都是社会问题,对吧?比如我们如何保护开源生态系统免受 AI 生成的 PR 的影响,这些 PR 从来没有人读过,然后这些志愿者维护者不得不阅读所有的垃圾。
And it breaks the social contract, right? Like in the past, if I wrote something and asked you to read it, you could at least have assumed that I would have spent 10 times more time writing it than you reading it. But now you're really skeptical about anything, because why would I want to spend time reading some stuff that you didn't even read yourself, maybe? Yeah, like I mean, but a lot of this is social problems, right? Like how do we protect the open source ecosystem from AI-generated PRs where no one has ever read them, and then these volunteer maintainers have to read through all the slop.
是的。
Yeah.
而且这真的不好玩。你和任何人谈过这个吗?比如人们谈论需要一个类似 GitHub 的社会信用系统,一个业力系统,这样人们可以给你降级,如果你的业力不够高,也许他们就不会读你的 PR。
And it's not really fun. Have you talked with anyone about that? Like people talk about needing a GitHub-like social credit system, a karma system, so people can downrate you, and if you don't have high enough karma, maybe they'll just not read your PRs.
我们确实需要那个。
We do need that.
是的。而且 archive 最近也对人们上传内容设置了门槛。
Yeah. And archive recently put gates on people uploading things there as well.
对。是的,他们有这个禁令。如果你有幻觉……就像一年禁令。
Right. Yeah, and they had this ban. It's like a one-year ban if you have a hallucinated...
是的。
Yeah.
但在某种程度上这很可悲,因为它让新人尤其是年轻人更难进入那些领域,因为他们没有任何业力或可展示的东西。
But it's sad in a way because it makes it harder, especially for new people and young people to break into that stuff, because then they don't have any karma or anything to show.
我知道,但我认为我们遇到这个问题的根本原因是,这项技术是人类历史上创造的最具欺骗性的东西,对吧?这是一个严重的问题,因为它完全是关于认知主观性,即你生成你不理解的东西,它说服你它是正确的,而你看不到那些瑕疵。显然,专家能一眼看出瑕疵。但它创造了依赖性。你知道,当你开始发布你不理解的东西时,你想要保持一致,对吧?你现在已经声明了我了解这件事,所以我会继续做下去。然后奇怪的是,人们对此也非常防御,因为如果你批评他们用 Claude 做的工作,他们会觉得是针对个人的。所以它只是创造了一个永久的循环。
I know, but I think the broad reason we have this problem is that this technology is the most deceptive thing ever created in human history, right? And it's a serious problem because it's all about epistemic subjectivity, which is that you generate things that you don't understand, and it convinces you that it's correct, and you can't see the glitches. And obviously, an expert can look at it and see the glitches straight away. But it creates dependency. You know, when you start posting stuff that you don't understand, you want to be consistent, right? You've now made a statement that I know about this thing, so I'm going to keep doing it. And then weirdly, people are very defensive about it as well, because if you criticize work that they've done with Claude, they take it personally. So it just creates this perpetuating cycle.
而且它创造了一种你可能并不真正拥有的理解感,对吧?就像参加考试或抄袭别人的作品——你可能觉得是你写的,但你没有,你的大脑没有经历那些过程。而且我认为如果你担心 AI 接管,这也是危险的,因为不仅仅是它变得更聪明,而且人类也在变得更笨。就像我们不再……
And it creates a feeling of understanding that you might not actually have, right? And it's like doing an exam or copying somebody else's work—you might feel like you wrote it, but you didn't, your brain didn't go through the motions. And I think it's dangerous also if you're worried about AI taking over, because it's not just that it's getting smarter, it's also that humans are getting dumber. Like we no longer...
是的。
Yeah.
我们在理解事物方面变得懒惰。我们不读论文。你只是把它们放进 AI,然后说:“哦,给我解释一下这篇论文。”但是,是的,这很奇怪,对吧?因为另一方面,它可以是一个巨大的加速器,它可以让你加速这么多,所以很难不使用它。但要弄清楚什么时候应该停止,什么时候应该再次开始使用它。
We get lazy in terms of understanding stuff. We don't read the papers. You just put them in AI and be like, "Oh, explain this paper to me." But yeah, it's strange, right? Because on the other hand, it can be this big speed-up, it can speed you up so much, so it's hard not to use it. But finding out when you should stop and when you should start using it again.
我知道,这让人想起埃隆·马斯克的那条推文,他说特斯拉没有研究员,只有工程师。而对我来说,这其实才是大问题。这是个悖论,因为你可以用语言模型来增长知识,这是事实,对吧?如果目的是增长知识,那么如果你是个好奇的人,你就可以一直挖下去,学到很多东西。那为什么平均来看,它们反而侵蚀了我们的知识呢?我觉得这是埃隆的错,但也不是他的错,你知道,他说工程学是达到目的的手段。所以我们在造这个东西,它需要通过这些测试。从某种意义上说,我不在乎你的知识在这个过程中是否被侵蚀,因为那不是我在衡量的东西。我们也不应该非黑即白,因为显然工程师们会在某些地方犯错,也会在过程中学到很多。但这似乎相当趋同,你知道,当你为了知识本身而追求知识时,你似乎会打下更深的基础,发现新事物。
I know, it's reminiscent of the Elon Musk tweet, you know, he said that there are no researchers at Tesla, only engineers. And for me, this is the big problem actually. It's a paradox because you can use language models to increase your knowledge. That's a fact, right? If the purpose is to increase your knowledge, then if you're a curious person, you can just dig and dig and dig and learn a hell of a lot. So why is it that on average they erode our knowledge? And I think it's Elon's fault, but not his fault, you know, he says that engineering is a means to an end. So we are building this thing and it needs to pass these tests. And in a sense, I don't care if your knowledge erodes during the process because that's not what I'm measuring you on. And we shouldn't be binary about it because clearly engineers trip up on things and learn a lot along the way. But it seems to be quite convergent, you know, when you're pursuing knowledge for its own purpose, you seem to build deeper foundations and discover new things.
是的,我想这就像资本主义里的注意力之类的,你知道,公司不一定非要培养员工。我的意思是,也许如果他们能看到利润动机,但归根结底他们只是想完成工作。而且当然,如果你花很多精力培养员工,然后他们离开了,那总是有点难受。我不知道有没有办法强制它。但我认为这也有点自作自受,因为学习东西本身就很难,可能会让人沮丧,而且它不总能让你进入像提示词那样的心流状态,对吧?我一直在尝试用这个——我记得是卡帕西建议的——当他从 LLM 学习新东西时,他不会复制粘贴,也不会并排窗口,而是总是手写所有代码。
Yeah, and I guess that's like an attention with capitalism or something, you know, companies aren't necessarily trying to develop their employees. I mean, maybe if they can see a profit motive in it, but at the end of the day they just want to get the job done. And of course, if you spend a lot of energy developing your employees and then they leave, it's always a bit tough there. I don't know if there's a way to force it. But I think it's also a little bit self-inflicted because learning stuff is just hard and can be frustrating, and it doesn't always get you in the same kind of flow as just prompting or something, right? And I've been trying to use this—I think it was Karpathy who suggested this—when he was learning some new stuff from an LLM, he didn't copy-paste or have side-by-side windows, but he would always write all of the code by hand.
是的。
Yeah.
我做了一个内部应用,因为我们要招很多 AI 工程师——很难找到既是优秀的硬件工程师、又是优秀的 AI 工程师、还是软件工程师的人,所以当然你必须是这样。所以我们必须扩展人员。我们必须让硬件人员扩展到 AI,让 AI 人员扩展到硬件。所以我做了内部培训工具等等,我在里面放了同样的块,如果你愿意,试图阻止复制粘贴,至少给人们一个警告,说你应该试着自己打字。不知怎么的,你不会这么想——如果我只是看着然后自己打字,为什么会学得更好?但经验表明,不知何故,我们的神经元之类的东西,经历这些时刻、经历这些动作很重要。
I made an internal app to try and—because we hire a lot of AI engineers—it's hard to find people who are both really good hardware engineers, really good AI engineers, and also software engineers, so of course you have to be that. So we have to scale up people. We have to scale up hardware people on AI, and AI people on hardware. So I made internal training tools and so on where I try to have the same blocks in there, if you like, trying to block copy and paste and at least give people a warning and say, you should maybe try and type it yourself. Somehow, you wouldn't think so—it doesn't feel like anything if I'm just looking at it and typing myself, why would I learn it better? But it's just empirical that somehow our neurons and stuff, it's important to go through the moments, to go through the movements.
你必须弄清楚哪些项目需要深入理解,哪些只是你在做的,哪里不那么重要。我认为我们也会受到 AI 的诱惑,开始并行启动更多项目,对吧?因为,哦,我可以在做那个的同时再开一个,我再开一个,再开一个,然后突然你就想,好吧,显然我无法真正深入理解所有这些,因为我现在开了这么多项目。所以你必须阻止自己,想,也许不是启动第五个项目,而是回去试着理解前几个项目里到底发生了什么。我想也许会有一些演变——我们还在学习使用这些工具的所有社交方面,什么是有效的,什么是无效的。希望我们甚至可以在工具中构建一些东西。你说的范围界定也是这个意思吗,试图让它们变得更好?
You have to figure out which projects you need to understand in depth and which ones you're just making, where it doesn't matter as much. I think we also get tempted with the AI to just start way more projects in parallel, right? Because, oh, I can just start another one while I'm working on that, I'll start another one and another one, and then suddenly you're like, well, there's obviously no way I could actually understand all of these in depth because now I started so many projects. So then you have to stop yourself and be like, maybe instead of starting the fifth project, I'll go back and try and understand what's actually going on in the first couple of projects. I think maybe there'll be some evolution—we're still learning all the social stuff around using these tools and what is effective and what's not effective. Hopefully we can even build some stuff into the tools. Is that what you're saying also with the scoping, to try and make them better?
而且这很难——我的意思是,我认为现在另一个真正困难的社会问题是,你如何做团队合作?这也是我作为管理者会思考的事情。如果我团队里的每个人都在盯着他们的 10 个智能体之类的东西工作,他们什么时候去和同事讨论这些事情?或者如果他们自己都不理解代码,他们怎么向公司里的其他人解释代码是做什么的,并一起构建?我认为没有人真正在大规模上很好地解决这个问题。我听过一个播客,是关于 Anthropic 的协作团队的,至少我记得的是,播客问,嘿,你们是怎么把它做得这么好的?他们说,是的,团队里的每个人都编写了自己的版本,然后我们选了最好的一个。我当时想,天哪。如果那是我们在团队合作中能做到的最好的,你知道,这有点——我的意思是,从某种意义上说这很酷,因为它就像集成方法应用到人身上,你可以探索更多。但它也确实消除了所有的团队合作。那时每个人都孤立地工作,只是在重复。是的。所以我很好奇我们将如何解决这个问题。
And it's hard—I mean, another really hard social problem I think right now is how do you do teamwork? That's something I also think about as a manager. If everyone on my team is staring at their 10 agents or something working on some stuff, when are they going to go and talk to their colleagues about the stuff? Or if they don't understand the code themselves, how are they going to explain to the other people in the company what the code does and build together? And I don't think anyone has really solved it well at scale. I heard a podcast also with the Anthropic co-working team, and at least the way I remember it was that the podcast asked, hey, how did you make this so good? And they were like, yeah, everyone on the team coded their own version and then we picked the best one. And I was like, damn. If that's the best we can do in teamwork, you know, it's kind of—I mean, in a way it's cool because it's like the ensemble method just applied to people, and you can explore so much more. But it also does remove all of the teamwork. At that point it's just everyone working in isolation, duplicating. Yeah. So I'm curious how we're going to solve that.
这是性能与能力的问题,我想我们集体在打赌,性能才是最重要的。
And it's performance versus competence, and I guess we're kind of making a bet collectively that performance is all that matters.
托马斯,能请你上节目真是太棒了。
Thomas, it's been great to have you on the show.
是的,非常有趣。谢谢你,蒂姆。
Yeah, it's been so much fun. Thank you, Tim.