递归语言模型——MIT 博士 Alex Zhang

Recursive Language Models — Alex Zhang, MIT PhD

亚历克斯·张 Alex Zhang · Latent Space · 2026-10-02 · 约 103 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

MIT 博士 Alex Zhang 探讨递归语言模型、GPU Mode 以及自动化 GPU 内核开发的未来。

Alex Zhang of MIT discusses recursive language models, GPU Mode, and the future of automated GPU kernel development.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 39)

全文 · Full transcript(中英对照)

最喜欢的编程工具 Favorite Coding Tools

Alex

哦,我喜欢 Claude Code。我喜欢 Codex。我喜欢 Pi。不,我喜欢 Oh My Pi。我喜欢 Prime Agent。说实话,我觉得它们都一样。

Oh, like I love Claude Code. I love Codex. I love Pi. Like no, I love Oh My Pi. I love Prime Agent. To be honest, I think all of them are the same.

Host

你知道,如果你更努力,你的模型其实能做得更多。所以这是技能问题。

You know, your models actually are capable a lot more if you try harder. So this is a skill issue.

Alex

是的。是的。基本上,我想看到一件事。我很欣赏大家对“锯齿状智能”的关注,因为它描绘了一个宏大的图景:如果我们真的下定决心,我们就能做到。但我有点希望,也许学术界有人应该做这件事。就是真正坐下来思考,比如如果我拿 Astra,即使是当前的前沿模型,也不够好,比如在比如一个月的时间里持续且出色地完成某项工作。它是 10,000 个智能体在 88 小时内。嗯,1300 亿个公开 token,估计公开定价约为 4000 万美元。

Yeah. Yeah. Basically, I mean I think there's one thing I want to see. I appreciate that there's a big focus on like jagged intelligence because it paints a big picture of like we can do this if we really set our minds on it. But I kind of wish and maybe someone in academia should do this. Like really just sit down and think about like if I took Astra even the current frontier models are not good enough at like doing a particular job over let's say the span of a month consistently and well. It's uh 10,000 agents in 88 hours. Um 130 billion public tokens which is estimated to be about $40 million in public pricing.

Host

也许你得付大约 4000 万美元才能得到一个结果。但话说回来,这很令人兴奋。我想说,我们甚至可以选择把 4000 万美元投入一个问题并解决它,这非常令人兴奋。

Maybe you'll have to pay like $40 million to get a result. And it's like well but it's exciting. I will say like it is it's very exciting that we even have the option to point $40 million at a problem and solve it.

Alex

是的。

Yes.

赞助商信息 Sponsor Message

Host

在我们进入今天的节目之前,我有一条给听众的简短信息。谢谢。如果你们没有选择点击并收听我们的内容,我们就无法为你们带来你们如此明确想要的 AI 工程、科学和娱乐内容。几乎每天都有人找我们做赞助。但幸运的是,你们中有足够多的人订阅了我们,让这一切在没有广告的情况下可持续,我们想保持这样。但我只想请你们帮个忙。你能做的最强大、完全免费的一件事就是点击那个订阅按钮。这是我唯一会请求你做的事。这对我和我的团队来说意义重大,他们每周都在努力把 Inspace 带给你。如果你订阅了,我保证我们会一直努力让节目变得更好。现在,让我们开始吧。

Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis. But fortunately, enough of you actually subscribe to us to keep all this sustainable without ads, and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you. And it means absolutely everything to me and my team that works so hard to bring the Inspace to you each and every week. If you do it, I promise you we'll never stop working to make the show even better. Now, let's get into it.

嘉宾介绍 Guest Introduction

Host

好的,我们在演播室,嘉宾是 Alex Jang。嗯,我想最出名的是 RLMs,嗯,但你还有几个其他身份。欢迎来到节目。

All right, we're here in the studio with Alex Jang. Uh, I guess most famously of RLMs, uh, but you you have a lot a few other affiliations. Welcome to the show.

Alex

是的,谢谢邀请我。是的,我想还有 GPU mode。

Yeah, thank you for having me. Yeah, I guess GPU mode as well.

Host

是的,还有 GPU mode。

Yes, GP mode as well.

Host

嗯,你是由 Mark Sarafin 引荐的。不是每个人都能得到那种欢迎。

Uh, you were shephered in by Mark Sarafin. Not Not everyone gets that kind of welcome.

Alex

是的。是的。是的。我和 GPU mode 里的所有人都非常非常亲近。所以,是的,我们经常以各种身份合作,甚至超出了 GPU mode 本身。所以,

Yep. Yeah. Yeah. I'm very very close to all the people in GP mode. So, yeah, we often end up working together in various capacities like even beyond just GP mode itself. So,

Host

是的。我们能解释一下吗,嗯,这样不太了解的人就不知道这个。它是一个 Discord。它以前专注于,我想是 CUDA mode,然后稍微泛化了一点。嗯,它是由 Mark 发起的。是的。对我来说,它基本上就像是 PyTorch 团队的招聘渠道,然后你离开了 PyTorch。

yeah. Can we explain uh so people who are not that close don't know about this. It's a it's a discord. It used to be focused on I guess CUDA mode and then then generalized a little bit. Uh it was started by Mark. Yep. It was basically like to me it's like the hiring pipeline of the PyTorch team and then you left PyTorch.

Alex

是的。是的。是的。是的。

Yep. Yep. Yep. Yeah.

GPU Mode 社区 GPU Mode Community

Alex

所以它以前是,我想,嗯,它实际上是在我上大学的时候开始的,嗯,大概在 2023 年。嗯,我想它是由 Mark Andreas 和 Jeremy Howard 发起的。最初的设想就是它是一个 GPU,或者是一个专门学习如何编写 GPU 内核的 Discord,他们还有讲座。基本上就这些。嗯,我对它感兴趣是因为我在写 GPU 内核。实际上完全是偶然。我当时在 Snapchat 实习,而且

So it used to be I think uh it started actually around when I was in college um in like 2023. Uh I think it was started by Mark Andreas and Jeremy Howard. The original premise was just like it was a GPU or it was a discord dedicated to learning how to write GPU kernels and they had like lectures. That was basically the extent of it. Um, and I got interested in it because I was writing GPU kernels. It was actually out of like pure chance. I was interning at Snapchat at the time and

Host

Rexus。

Rexus.

Alex

是的。我对 Rexus 非常厌倦。所以,嗯,他们有一个项目,他们有兴趣写。是一篇叫 Infinite Attention 的论文。是谷歌的论文。

Yeah. I was very bored with Rexus. So, um, they had a project where like they were interested in writing. It was this paper called Infinite Attention. It was like a Google paper.

Host

是的。我们在 paper club 上讨论过。

Yes. We've covered it on paper club.

Alex

是的。是的。所以我对在 Snapchat 是否能为其编写专用内核感兴趣。嗯,没什么结果,但我当时加入了 GPU mode。它叫 CUDA mode。嗯,我想是出于法律原因或其他什么,他们改了名字,但我遇到了 Mark,嗯,我遇到了 Mate,我遇到了很多其他非常参与社区的人。然后 Mark 提出了一个叫 Popcorn 的想法,就是今天你看到的排行榜。但总体想法是,我想我们所有人都有这种直觉,GPU 编程非常类似于,如果你们做过竞技编程的话。它并不是,我不是说它们可以迁移

Yes. Yeah. So I was interested in in whether or not you could write specialized kernels for it at Snapchat. Um it didn't nothing really came of it but I joined GPU mode at the time. It was called CUDA mode. Uh I think for like legal reasons or something they changed the name but I met Mark uh I met Mate I met a bunch of other people that were very involved in the community. And then Mark had pitched this idea called Popcorn which was now what you see as the leaderboard today. But the general idea was like I think all of us had this like intuition that GPU programming is like very similar to if you guys have done like competitive programming. It's it's a not I I don't mean to say like they're transferable

Host

技能。你稍微打点代码高尔夫。

skills. You code golf a little bit.

Alex

是的。是的。是的。实际上人们做的优化空间小得惊人。嗯,实际上人们有兴趣优化的内核本身并没有那么多。所以我们有这种想法,如果你有足够的数据,就像 Codeforces 有数百万个问题一样。如果你能用 GPU 代码做到这一点,你就能扩展和自动化 GPU 内核开发,这对研究人员来说是件大事,因为我认为更大的瓶颈之一,比如你看 Mamba,他们发布论文时附带内核,否则你无法以任何有意义的方式使用它,而且不是每个人团队里都有像 Tri Dao 这样的人。所以我们对此非常感兴趣。嗯,Kernel Bench 也由此衍生,就是我们能 LLM 自动化 GPU 内核代码吗,我认为那是一个非常非常有趣的时期。那是在大学和我的博士之间,是的,我在 GPU mode 做事时非常愉快。现在我只是有时帮忙讲座。嗯,我不那么参与了,我想总的来说我们不像以前那么多比赛了,但嗯,是的,我仍然和那里的每个人保持很多联系。

Yeah. Yeah. Yeah. And there's like there's actually a surprisingly small space of optimizations that people do. Uh and there's actually not that many kernels per se that people are interested in optimizing. And so we kind of had this thought that like if you had enough data like in the same way that code force is there's millions of problems. If you could do this with GPU code like you could scale and automate kind of GPU kernel development which for researchers is a huge deal cuz I think one of the bigger bottlenecks like if you look like mamba for example like they release the paper with kernels because otherwise like you can't really use it in any meaningful way you know and not everyone has like a tree on their team. So we're very interested in this. Uh kernel bench kind of spawned from that too of like can we get LLMs to automate um GPU kernel code and I think that was like a it was a very very fun time. It was like between college and my PhD and yeah I I had a really pleasant time doing stuff with GPU mode. Now I kind of just help with the lectures sometimes. Um I'm not as involved and I think in general like we don't have as many competitions as we used to but um yeah I still keep in touch a lot with with everyone there.

Host

有没有友好的竞争,因为我想以前做这个的社区,比如 MLS ML Perf 之类的。有没有友好的竞争?这就像是新一代的 MLF 还是怎么回事?

Is there a friendly rivalry because like I think the previous community that used to do this like MLS ML Perf thing. Is there a friendly riv rivalry? Is this like just new generation MLF or what's going on?

Alex

GPU mode 的好处在于它也是一个社区,从某种意义上说,很多讲座对初学者来说很容易跟上并提问之类的,比赛也有点,不是次要,但你可以参与其中来学习。我想对于很多像 MLS、ML Perf 这样的基准测试,大部分情况下只有认真的实验室和公司才会认真参与。至少这是我的理解。嗯,我可能错了,但我认为现在除了 GPU mode,还有一件非常令人兴奋的事,就是有更多的网站和人们致力于举办比赛。比如我想有这

The nice thing about GPU mode is that it is also a community in the sense that like a lot of the lectures are very easy enough for a beginner to follow and ask questions and things like that and like the competitions are like somewhat not secondary but like you can participate in them to learn. I think with a lot of like MLS, ML Perf kind of benchmarks like for the most part like only serious labs and companies participate like like seriously in them at least. That's that was my understanding of it. Um I I could be wrong, but I I think also beyond GPU mode now, one thing that has been really exciting is there's a lot more websites and like people that work on hosting competitions. Like I think there's this

Host

我想有一个叫 LeetGPU 之类的网站,就像是 GPU 问题的 LeetCode。

I think there's this website called like leak GPU or something and it's like leak code for GPU problems.

GPU 编程曾是冷门 GPU programming used to be niche

Alex

还有其他一些我也喜欢的。我们看到很多项目涌现出来,在 GPU Mode 上分享,这非常令人兴奋,因为我觉得 GPU 编程以前超级小众。我当初对它感兴趣,唯一的原因就是 Trio 在普林斯顿做了一场报告,因为他当时在申请教职——我是说,他现在已经是那里的教职了——但我在 2023 年听了那场关于 FlashAttention 的报告,我当时就想,哇,这是最酷的东西。我就觉得,这是所有人都该去做的事。我猜当时 VLM 之类的东西也出来了,大家就说,哦,我们应该写 kernel。但现在人人都写 kernel。我觉得作为一个领域,它在某种意义上几乎已经饱和了。

There are other ones too that I like. We've seen many that have kind of spawned and talked on GPU Mode, and it's very exciting in the sense that I think GPU programming used to be super niche. When I was interested in it, the only reason I got interested was Trio gave a talk at Princeton because he was applying for faculty — I mean, he is faculty there now — but I listened to his talk on FlashAttention in 2023 and I was like, wow, this is the coolest thing ever. And I was like, this is what everyone should be working on. I guess VLMs and stuff had come out too, and it was like, oh, we should be writing kernels. But now everyone writes kernels. I think it's actually almost saturated in some sense as a field.

AI 生成内核与验证难题 AI-generated kernels and the verification problem

Host

对于想进入这个领域的人,你有什么有趣的看法吗?

Any interesting takes for people that want to get into it?

Alex

我觉得最大的新闻之一是 GPT 5.6 写出了更高效的 kernel,所以 Terra 和 Luna 可以便宜 80%。然后我们还看到其他竞赛里有人在破纪录,他们说我们在跑某种自动研究循环,而这些人根本没有写 kernel 的背景。

So I think one of the biggest news is GPT 5.6 wrote more efficient kernels so Terra and Luna could be 80% cheaper. And then we've seen other competitions where people are setting records and they're like, we're doing some auto research loop, and these are people that don't have a background in any kernel writing.

Host

对。

Right.

Alex

是的。所以即使在 GPU Mode 的排行榜上,如果你看最近很多题目,几乎所有解法都是 AI 生成的。不过你会注意到排行榜上有个人叫 Gaurst,他是 GPU Mode 非常非常资深的成员。我们早就知道他是一个极其极其正确的 GPU kernel 写手。我们在这个排行榜上发现的一件事是,他也用 AI 来帮他做这些解法,但大部分时候是他来提示、把方向往某些地方推。我们发现他的 kernel 基本上是前十名里唯一一个在实际端到端系统里真正稳定的。这就引出一个问题——我是说,GPU kernel 有一个验证问题。我们早就知道这一点。从 KernelBench 发布以来这就一直是个问题。有很多奖励黑客行为。但你也会注意到代码行数小了很多。

Yeah. So even on the GPU Mode leaderboard, if you look at a lot of the recent problems, almost all the solutions are AI generated. However, you'll notice on the leaderboard there's this guy named Gaurst who is a very, very regular member of GPU Mode. We've always known for a long time that he's a super, super correct GPU kernel writer. One thing we discovered on this leaderboard is that he also used AI to help him with these solutions, but for the most part he helped prompt and move it in certain directions. We found that his kernel was basically the only one in the top 10 that was actually stable in actual end-to-end systems. And it does bring into question — I mean, GPU kernels have a verification problem. We've kind of known this. It's been a problem since KernelBench was released. There's a lot of reward hacking that goes on. But you also notice the lines of code is a lot smaller.

Host

这很明显吗,还是只是——

Is that noticeable or is it just—

Alex

是的。不,这绝对非常重要。而且我觉得很有意思的是,擅长这件事仍然有很大的 alpha。

Yeah. No, it's definitely very important. And I think it's really interesting that still there's a lot of alpha in being good at it.

Host

绝对有。是的。

There definitely is. Yeah.

Alex

我觉得这也适用于很多 AI 系统。我觉得即使是最新的数学证明之类的东西,也不一定意味着数学家就过时了。我是说,这些公司仍然雇数学家来做,不管是数据标注工作,还是只是引导模型去解决问题。在这些事情上有知识仍然有很大的 alpha。

I think this applies to a lot of AI systems as well. I think even with the most recent math proofs and stuff, it doesn't necessarily mean mathematicians are obsolete. I mean, these companies still hire mathematicians to do, whether it be data labeling work or even just steering the models to solve problems. There is still a lot of alpha in being knowledgeable in these things.

直觉、规划与计算效率 Intuition, planning, and compute efficiency

Host

所以这只是知识,还是说也有更多规划,是不是有一种涌现出来的、效果更好的规划风格?

So is it just knowledge or is it also there's just more planning and is there an emergent style of planning that works better?

Alex

我觉得是混合的——也许这就是你的意思——比如解决问题的直觉。

I think it's a mix of — maybe this is what you mean — like intuition for how to solve the problem.

Host

比如我总是把我的代码画成图。

Like for example I always diagram my code.

Alex

是的。

Yeah.

Host

对。然后如果图里有哪部分我不理解,我就一直做到理解为止。否则就不允许。

Right. And then if there's a part of the diagram I don't understand, I work until I understand it. Otherwise it's not allowed.

Alex

是的。是的。所以我觉得是这些东西的混合。那些知道怎么看这些问题、怎么解决它们的人,也知道怎么用 AI 来做,因为你在充当一个非常强的验证者。如果你有知识,如果你知道该做什么,而且你也——我觉得我们从所有这些智能体集群之类的东西里发现的一件事是,当你往一个问题投入足够多的算力,你就能充分地探索这个问题的解法。但很多时候,也许你可以在某件事上烧掉一千亿或一万亿 token,但如果你引入一个对这个问题有所了解的人,他们能为模型揭示出某些东西,从而抹掉那一万亿 token 的花费。我是说,这里的趋势到底是什么还不太清楚。但我觉得现在野外还有太多我们想解决的问题,而我们没法总是往上面投入尽可能多的算力。所有这些东西仍然有一个效率层面,这一点超级超级重要。

Yeah. Yeah. So I think it's a mix of those things. The people who know how to look at these problems and how to solve them also know how to use AI to do them, because you're acting as a very strong verifier. If you are knowledgeable, if you know what to do, and you're also — I think the thing that we've kind of discovered with all these agent swarms and things like this is that when you throw enough compute at a problem, you can sufficiently explore solutions to that problem. But often times, maybe you can burn a hundred billion or a trillion tokens on something, but if you bring in someone who knows something about the problem, they can uncover something for the model that would erase that one trillion token spent. I mean, it's not super clear what exactly the trends are here. But I think there are so many problems in the wild still right now that we want to solve, and we can't afford to just always throw as much compute as possible at it. There is still an efficiency aspect of all of these things that is super, super important.

内核的理论速度极限 Theoretical speed-of-light limits for kernels

Host

是不是有一个理论上的正确答案,你可以直接根据物理算出来,然后你就接近物理极限了?

Is there like a theoretical right answer that you can just calculate based on physics and then you just get close to the physics limit?

Alex

是的。对于 GPU kernel,你可以算出来。有时候其实没那么容易算,取决于问题有多复杂。对于矩阵乘法,很容易算出这种“光速”式的估计,即最快的 kernel 能有多快。而且这也假设你的数据可能都从 CPU 开始,或者从 GPU 的 DRAM 开始,等等。这会稍微改变这些数字。

Yes. So for GPU kernels you can compute it. It's actually not that easy to compute sometimes, depending on how complex the problem is. For matrix multiplications it's very easy to compute this speed-of-light kind of estimate of what the fastest kernel can be. And this is also assuming maybe all of your data starts on the CPU or maybe it starts in DRAM on the GPU, etc. This changes these numbers slightly.

Host

传输之类的这些东西。

The transfers and all these things.

Alex

是的。但我要说的是,在很多情况下,是否有可能达到这个理论数字并不清楚,如果这么说能理解的话。这假设了完美的数据重叠和传输,而也许存在某种你绕不过去的瓶颈。但通常我们写的 kernel 根本不够接近——离这个数字远得不足以有任何意义。

Yeah. But I will say it's not clear though, in a lot of cases, if it's even possible to hit this theoretical number, if that makes sense. This is assuming perfect overlapping and transfer of data, and there's maybe some bottleneck that you can't get around. But often the kernels we write are not even close — not nearly close enough to this number to be meaningful at all.

速度、内存与功耗 Speed, memory, and power consumption

Host

是的。那重要的是速度吗?你也在意——显然内存,内存会影响速度。你在意功耗吗?我们今年最火的播客之一就是 Jeff Dean,他其实——我刚查了一下——微焦耳,或者纳焦耳、皮焦耳。

Yeah. And is it speed that matters? Do you also care about — obviously memory, which feeds into speed. Do you care about power consumption? So one of our top pods of the year was Jeff Dean, who was actually — I just checked — the microjoules or like the nanojoules, picojoules.

Alex

是的,皮。是可选的,皮。

Yeah, pico. It's optional, pico.

Host

你在意那个吗?

Do you care about that?

Alex

我不——我猜,Vivie,我没那么——

So I don't — I guess, Vivie, I'm not as—

Host

但这里一切都是速度,对吧?就像没人在数皮焦耳。

But everything here is speed, right? Like nobody's counting picojoules.

Alex

嗯,但这里有个前提,我觉得单个 kernel 语境下的速度和更大问题(比如端到端模型)语境下的速度是两回事。因为要考虑的一点是——这也是为什么说清楚“光速”指的是什么很重要——因为在这些 kernel 的情况下,我们总是假设一切都从比如 HBM 开始,对吧?但你可以想象,对于一个端到端的 kernel,比如一个端到端模型,你在两层之间可能想做的事是牺牲第一个操作的速度,把东西留在缓存里给第二个操作用。而这些是你在孤立地看这类 kernel 时无法得到的。

Um, but there's a caveat here, which is I think there is speed in the context of a single kernel and there's speed in the context of a larger problem, like maybe the end-to-end model. Because one thing to consider — and this is why it's important to talk about what speed of light is referring to — because in these cases for the kernels, we always start with everything in, like, HBM for example, right? But you can imagine that for an end-to-end kernel, like an end-to-end model, what you might want to do between two layers is you might sacrifice the speed of the first operation to keep things in the cache for the second operation. And these are things that you can't really get out of in isolation with these kinds of kernels.

巨型内核与编译器 Mega Kernels and Compilers

Host

人们把这叫做融合问题,或者 mega kernel 之类的东西。大多数情况下,它一般只在你受内存带宽限制时才适用。但还有一个问题:随着这些模型变得更强,我们是不是就该直接生成 mega kernel?那是我们想要的吗?你怎么看?

People call this the fusion problem, or mega kernel stuff. It generally only applies when you're memory bound in most cases. But there's also this question: as these models get better, should we just be generating mega kernels? Is that what we want? What's your take?

Alex

我觉得这真的很难,因为你需要数据才能做到。我还没在现实中见过一个例子,能在没有任何样本的情况下,自举出解决某一类非常困难问题的能力。我觉得这件事可能没那么有趣的另一个原因是:在单个 kernel 的层面,它们没那么复杂,而且你多少能确信单个 kernel 里没有那么多结构。但在 mega kernel 里,我会更倾向于相信编译器在这里会更好。某种基于更高层算子的编译器是说得通的,因为总体而言,mega kernel 的各个部分其实非常可组合,就是由单个 kernel 拼起来的。有些场景你可能会想做奇怪的融合之类的,但总体而言,我认为这些情况编译器大概都能处理。据我所知,有一家公司正在做这个,他们也在 GPU Mode 上做过一些分享。

I think this is really difficult because you need the data to do this. I have yet to see an example in the wild of bootstrapping the ability to solve a very difficult class of problems without any examples. The other reason I think maybe this isn't that interesting is that at the level of an individual kernel, they're not that complex, and you're somewhat confident that there's not as much structure in a single kernel. But in a mega kernel, I would be more inclined to believe that a compiler would be better here. Some compiler over higher-level ops makes sense, because in general the pieces of mega kernels are very composable from the individual kernels. There are some areas where you might want to do weird fusions and everything, but in general I think these are cases that a compiler can probably handle. And there is a company working on this, from what I understand, that has given some talks on GPU Mode as well.

GPU Mode 讲座 GPU Mode Lectures

Host

好,我想把关于 GPU Mode 的讨论基本都集中在这里,因为显然我们还有其他部分要推进。我觉得这里有个东西值得宣传一下。你们办了很多非常好的讲座,都在 YouTube 上,大家可以跟着看。而且其中不少是你主持的。你现在还挺——

Yeah, I want to basically cluster all the GPU Mode discussions here, because obviously there are other parts we need to move on to. I think there is something to plug. You guys host a lot of really good lectures. They're all on YouTube. People can follow along. And you lead quite a bit of it. You're still quite—

Alex

我以前会。有时候现在也会。我觉得主要是 Mark。通常是 Mark 来做。Mate 有时候也会做。但没错,我强烈推荐它们。它们是非常好的资源。我觉得人们在那里分享的东西多到有点不可思议。

I used to. Sometimes I still do. I think they're mostly Mark. Mark is the one who usually does them. Mate does sometimes as well. But yeah, I highly recommend them. They are extremely good resources. I think it's kind of crazy how much people share on there.

Host

因为如果你在那里,你就非常非常——你就是最合适的受众,你懂吧?

Because if you're there, you're very, very—you're exactly the right audience, you know?

Alex

对,没错。这些内容不会触达主流大众。而且我们在上面也放了很多入门材料,我觉得对人们很有用。

Yes, exactly. This isn't going to reach the mainstream. And there's a lot of introductory material as well that we've put on there that I think is useful for people.

KernelBench 与进行中工作 KernelBench and Ongoing Work

Host

KernelBench 挺有影响力的。我就是想看看——你知道,那是去年的事了。你还想推荐哪些正在进行的工作,让大家去关注?因为显然你在这个领域里。

KernelBench was kind of influential. I just want to see—you know, that was last year. What other ongoing work do you want to shout out that people should pay attention to, because obviously you're involved in this field.

Alex

好,我大概讲讲背景故事。我其实参与过很多 benchmark,或者说以前参与过,可能是在我读博之前。这始于我在普林斯顿的时候。我在那里和 SWE-bench 团队合作。John——这里的 John Carlos——他们都很棒,我很喜欢他们。

Yeah, I'll give maybe the background story. I am actually involved in a lot of benchmarks, or I used to be, maybe prior to my PhD. It started because I was at Princeton. I worked with the SWE-bench team there. John—John Carlos here—they're all great, I love them.

Host

基本上就是——我觉得人们不知道有多少 benchmark 来自同一个团队。在普林斯顿。你认识 Shun 吗?

There's basically this—I think people don't understand how many benchmarks come from the same group. At Princeton. Do you know Shun?

Alex

认识。对。对。

Yes. Yeah. Yeah.

Host

我们之前请他上过播客。现在他简直风生水起。

We had him on the pod before. Now he's like running 10.

Alex

对。现在他简直是个超级明星。我认识他的时候,他在指导我的朋友 Michael——Michael Tang,他现在在 Anthropic——但他们经常一起合作。我们俩当时算是 Caric 实验室里的两个本科生。后来也还有其他人加入。不过没错,Shun 很棒。我直到后来、我离开之后,才知道他这么厉害。

Yeah. Now he's like a superstar. When I met him, he was advising my friend Michael—Michael Tang, who is now at Anthropic—but they work together a lot. We were like the two undergrads in Caric's lab. And then some others joined later as well. But yeah, Shun is great. I did not know he was such a superstar until later on, after I left.

Host

对,我是说,好吧,像你这样的博士生非常少——你就像是下一个。我们大概每年会重点介绍一个人,基本上他整个博士生涯都踩在点子上。这样的人并不多。Shun 显然是其中之一。还有你知道,Jack——Jack Morris 是另一个。我们在节目开始前聊过研究品味,对吧?不知怎么的,有些研究生就是职业生涯特别顺,基本上就是——对,对,对,对,对——这个会留下来、很重要、大家都该知道。而其他人就什么都没有。

Yeah, I mean, okay, so there are very few PhD students like yours—yours is like the next one. Once a year we feature someone whose basically entire PhD has been on target. There's not that many of them. Shun was clearly one of them. And you know, Jack—Jack Morris is another one. And we talked before the show about research taste, right? Somehow some grad students just have a very blessed career where it's like, yep, mostly—yep yep yep yep yep—this is going to stick around, relevant, everyone should know this. And then others just nothing.

Alex

我觉得即使在产业实验室里的人身上,这一点也成立。只不过研究生更显眼。所以你就能看到——有些人是真的很走运,或者说这是走运加上非常聪明之类的混合。我觉得研究品味也是通过机会培养出来的。至少在我身上,我非常幸运地走了我走的那条路,在普林斯顿工作,然后后来在 MIT 找到了我的导师 Omar。他是个很棒的导师。不过我要说,我发现研究生或学术界最成功的研究,往往出现在人们关心那些产业里大多数人可能不看的问题时。我觉得问题就在这里:很多研究生做的东西是有好处的——是在产业实验室看来很亮眼的东西。比如他们会做某个 benchmark。现在这非常火。我是说,我觉得 2023 年的 benchmark 和现在的 benchmark 完全是两回事。有很多人做 harness、meta harness,以及针对 XYZ 任务的特定 harness。你仔细想想,一个人做这个的原因,可能是围绕我们今天已有的模型有一个清晰的目标,就是‘这是我想看到的’。但我想拿 RLM——递归语言模型那篇论文——举例,因为我觉得它是个超级超级简单的想法。我觉得它刚出来的时候,很多人看到这种东西会说,这到底有什么意义?

I think this is also true of even people within industry labs as well. I think it's just that grad students are a lot more visible. So you just see—you see there are some people who really get lucky, or I mean it's a mix of being lucky and also being very smart and things like that. I think with research taste as well, I think it gets developed through opportunities. At least in my case, I was very fortunate to have taken the path that I took, working at Princeton and then later finding my advisor Omar at MIT. He's a fantastic advisor. I will say though, I find that the most successful research from grad students or in academia comes when people care about problems that maybe most people in industry are not looking at. I think this is the issue: a lot of grad students work on things that benefit—that look good to an industry lab. For example, they'll work on some benchmark. It's really popular now. I mean, I think benchmarks in 2023 were a very different story than benchmarks now. There are a lot of people that work on harnesses and meta harnesses and specific harnesses for XYZ task. And when you really think about it, the reason someone would work on this is maybe there's a clear goal shaped around the models that we have today, of like, this is what I want to see. But I'll give the RLM—the recursive language model paper—as an example, because I think it's a super, super simple idea. I think when it came out as well, there were a lot of people that, when they see something like that, they're like, what is even the purpose of this?

Host

对,就像为什么——什么,这不就是子智能体之类的吗,对吧?

Yeah, like why—what, this is just sub-agents or something, right?

Alex

我觉得当你得到那样的反应时,这几乎是个好迹象,因为它清楚表明人们没有在思考这件事的目的是什么。我再举一个例子:SWE-bench。当 SWE-bench 刚出来的时候——我是说,Ofir 很喜欢讲这个故事——它刚出来时,没人在意。大家都说,这是个不可能的任务。我们为什么会把它当成一个 benchmark 来考虑?直到 Devin 出来,大家才说,哇,这是值得我们往上爬的东西。

And I think when you get a reaction like that, it's almost like a good sign, in the sense that it's clear that people aren't thinking about what the purpose of this is. And I'll give another example: SWE-bench. When SWE-bench came out—I mean, Ofir loves to tell this story—when it came out, nobody cared. Everybody was like, this is an impossible task. Why would we ever even consider this as a benchmark? And it wasn't until Devin came out that everyone was like, whoa, this is something we want to hill climb.

学术界的重大押注 Big Bets in Academia

Alex

我觉得这对你来说很真实。你往往会看到很多想法——我最喜欢的例子是 Eric Zelikman 关于 Star 和 QuietStar 的工作。当你读那篇论文时,至少我第一次读的时候,我就想,这不是一个显而易见的想法吗?或者也许不是,我不知道。我当时想,哦,这看起来真的很简单。或者像思维链,也是一样。或者像,不应该用 React 吗?就像,好吧,是的,当然。但当你真正思考时,这篇论文的价值是什么?我认为它来自于讲述了一个关于你希望这个领域看起来什么样的故事。而这在学术界是非常难做到的,因为如果你看所有这些论文——QuietStar、React、RLMs、SWE-bench——这些论文都不是像 GPT-6 Astro 发布那样。不是每个人都说,哦天哪,我现在就要用这个。这是世界上最好的东西。学术界就是负担不起这样做,至少现在是这样。我的意思是,有一大堆理由我认为这应该改变,但我认为如果你没有——作为一个博士生,我认为你处于一个非常独特的位置,你可以基本上研究任何你想研究的东西。如果你不利用这一点,而是研究没人关心的事情,或者人们认为是一些琐碎的事情,比如,哦,我想过这个,但我不使用它。我只是觉得最终研究永远不会那么有趣,因为如果你想留在学术界,你需要下大赌注。因为否则,我认为就去工业实验室吧——你知道他们有大量的资源,大量的人才。为什么要在一个你没有很多资源、周围甚至没有那么多人的领域限制自己?我认为这真的只是归结为大赌注。你只需要下大赌注,其中很多会失败。这很自然。但我认为作为一个博士生,这是你相对于其他实验室任何一个人的最大优势,因为你不需要处理官僚主义和其他所有事情。

And I think this rings true for you. You tend to see that a lot of ideas—I think my favorite example of this is Eric Zelikman's work with Star and QuietStar. When you read the paper, at least when I first read it, I was like, is this not an obvious idea? Or maybe not, I don't know. I was like, oh, this seems really simple. Or like chain of thought, same thing. Or like, shouldn't use React? It's like, okay, yeah, sure. But then when you really think about it, what is the value of the paper? And I think it comes from telling a bit of a story as to what you want the field to look like. And that is something that is very hard to do in academia because if you look at all these papers—QuietStar, React, RLMs, SWE-bench—none of these papers are like a GPT-6 Astro release. It's not like everyone's like, oh my gosh, I'm going to use this now. This is the best thing in the world. Academia just can't afford to do this, at least right now. I mean, there's a whole slew of reasons why I think that should change, but I think if you don't have—as a PhD student, I think you're in such a unique position where you can work on literally whatever you want for the most part. If you're not taking advantage of that and working on things that nobody cares about or that people see as some trivial thing like, oh, I thought about this, but I don't use it. I just think in the end the research is just never going to be that interesting because you kind of need to take big bets if you're going to be in academia. Because otherwise, I think just go to an industry lab—you know they have tons of resources, tons of talent. Why constrain yourself in an area where you don't have a lot of resources and there's not even that many people around? And I think it literally just comes down to big bets. You just have to take big bets, and a lot of them will fail. That's just natural. But I think that as a PhD student, that's the biggest advantage you have over any single person at another lab because you don't have to deal with bureaucracy and all these other things.

Host

有道理。

Fair enough.

Alex

是的。

Yeah.

Host

我问过很多人这个问题,通常他们都敷衍了事。所以我觉得很感激你实际上给出了一个深思熟虑的回答,就像,不,这是你的不公平优势,因为其他一切基本上都对你不利。

I ask a lot of people this question and usually they hand wave away. So I think I appreciate that you're actually giving a thoughtful response on like, no, this is your unfair advantage because everything else is biased against you basically.

Alex

是的。是的。没错。所以,老实说,我会以 Jeb 为例,因为它不是一个学术项目。我想提这个是因为这也发生在 RLMs 上,而且发生在许多其他工作上,比如事情被过度炒作,对吧?在某种程度上,某件事被过度炒作,然后人们就说,为什么这个被过度炒作?这很琐碎,这很愚蠢。我在 Jev 上看到了同样的事情,因为我认为发布时——我的意思是,有整个关于学术界或他们不是学术团体,但人们必须做品牌推广,他们必须推销他们的研究,所以我理解,但我认为有很多关于 Jev 的讨论,说它只是我们多年来就知道的东西,我认为这有点错过了重点,比如为什么这样一个系统如此有趣。为什么它不只是我们一直在做的某个愚蠢的 MLP 分类器,你知道,回到我们的机器学习入门课之类的。我认为 Jev 真正有趣的地方在于它开启了这个问题:语言模型是正确的吗?就像它们现在的形式,我们能否考虑一个不同于文本到文本的设计空间?他们所做的基本上是说我利用这个语言模型骨干,我知道它捕获了很多关于语言的信息,但我要改变模型的输出空间,给你一个权衡,那就是我会做非常非常快的推理,如果你对这个问题的先验知识,比如我只需要做一个二元分类,我会让我的语言模型做这个并支付 400 倍的成本吗?不,这很愚蠢,对吧?我认为在很长一段时间里,因为实验室是唯一控制的地方,你永远不会使用除了 GPT、GPT-6 或 Fable 之外的东西,因为它们是最好的模型。但正因为如此,人们已经习惯了这种想法,即语言模型只是一个自回归解码器。就像我们已经接受了这一点。我认为当 RLMs 出现时,也是一样。我得到的一个非常频繁的批评是,这不是一个语言模型,或者当我看到这个时,我以为它是一个新架构,但实际上不是,我的回应是,好吧,语言模型只是建模语言,它不一定是这个 Transformer 解码器,你知道。而 Jeb 真的非常有趣,因为我们现在有了一个新的东西可以调整,那就是输出空间是什么,以及这如何影响推理延迟。我认为我们实际上可以开始调整语言模型本身的各个部分。这是一个类似的想法,很多关注它的人说,这是一个愚蠢的想法,为什么谁在乎这个?但这是一个简单的想法,但它实际上开启了一整套新问题,我认为特别是如果你是博士生,这些是你想要回答的问题,因为我认为对于 Jev 来说,我们不知道我们能把这个带到多远,对于循环 Transformer,我们也不知道我们能把这个带到多远,如果你只循环模型的一个子集呢?如果你路由到比如你有一些路由器到模型的不同部分,你能在模型内部模仿你在 harness 中所做的事情吗?你能在 harness 选择和模型架构选择之间架起什么桥梁?这些都是我认为像这样的工作开启的问题。这真的很令人兴奋。特别是 Jeb,当我看到它时,我就想,这对 RLMs 真的很有用。我认为这是有道理的,因为 RLMs 或群体或类似系统的最大瓶颈是它们很慢。当你一直进行多个语言模型调用时,你没有正确分配你的算力,因为也许有一些琐碎的事情你只想让一个简单的模型做,但你做不到,因为你的语言模型只是这个笨重的东西,你知道?所以,我非常兴奋。我认为我们将开始看到超越标准前沿模型的新类型模型出现。而且有很多事情你可以用这些新的权衡来做。

Yeah. Yeah. Exactly. And so like, it's honestly—I will bring up Jeb as an example because it's not an academic project. I want to bring this up because this also happened with RLMs and it happens with many other works like things get overhyped, right? To an extent, something gets overhyped and then people are like, why is this overhyped? This is trivial, this is stupid. And I saw the same thing with Jev because I think the release was—I mean, there's this whole thing about like academics or they're not academic group but like people have to do branding and they have to kind of market their research and so I understand but I think there was a lot of discourse about Jev just being like something we've known for years and I think it's kind of missing the point of like why is such a system so interesting. It's why is it not just some stupid MLP classifier that we've been doing, you know, back in our intro ML classes or something. I think what's really interesting about Jev is that it kind of opens up this question of are language models correct? Like in the form that they're in, can we consider a different design space other than text to text? What they're doing is they're basically saying like I will take advantage of this language model backbone like I know it captures a lot of information about language but I'm going to change the output space of the model to give you a trade-off which is I will do very very fast inference over like if you have some prior about this problem like let's say I only need to make a binary classification am I going to ask my language model to do this and pay like a 400x cost like no it's silly, right? And I think for the longest time because the labs are the only places that control, you're never going to use something other than like GPT, GPT-6 or Fable because they're the best models. But because of that, people have gotten kind of accustomed to this idea that a language model is just an autoregressive decoder. Like we have accepted this. And I think when RLMs came out, it was the same thing. Like one of the comments I—a very frequent criticism I got was like this is not a language model or like when I look at this I thought it was a new architecture but it's actually not and my response to that is like well a language model is just modeling language it doesn't have to be this transformer decoder you know and Jeb is really really interesting in that like we now have a new thing to tune which is like what is the output space and how does this affect inference latency. And I think we can actually start this various parts of the language model itself. It's a similar idea of a lot of the attention around it was like this is a silly idea like why who cares about this? But it's like it is a simple idea but it's actually it opens up a whole new set of questions that I think like especially if you're a PhD student these are the things that you want to answer because I think it's like we don't know for Jev for example we don't know how far we can take this for loop transformers we also don't know how far we can take this what if you loop only a subset of the model what if you route to like only like you have some router to different parts of the model like can you mimic what you do in a harness inside of the model? What can you bridge between a harness choice and a model architecture choice? These are all questions I think that open up with works like this. And that's like really exciting. Jeb in particular, when I saw it, I was like, this is actually really useful for RLMs. Like I think it's it makes sense because the biggest bottleneck in RLMs or swarms or systems like these is they're slow. When you do multiple language model calls all the time, you're not distributing your compute correctly because like maybe there's something trivial that you just want a simple model to do, but you can't do it because your language model is just this bulky thing, you know? So, I'm very excited. I think we will start to see new types of models emerge beyond just the bog standard frontier model. And that is like there's so many things that you can do with these like new trade-offs.

Host

我也要 shout out Thinky 和他们的交互模型。

I also shout out Thinky with their interaction model.

Alex

是的。是的。

Yeah. Yeah.

校准 Breaking the Paradigm

Host

另一个很好的例子。

Another great example.

Alex

是的。所以基本上就是尝试打破从序列到序列、仅解码器的范式,然后 literally 做点别的。我的意思是,我觉得如果你试图竞争,你不可能在算力规模资源上跟做自回归解码器的前沿实验室竞争。即使是 Thinky 也不会——好吧,可能少数几个能,但你在这方面不会有多大作为。我对 Thinky 了解不多,也不想多说什么,但如果他们的策略只是复制 OpenAI 或 Anthropic,那是个糟糕的策略,因为他们根本没有——你得从你有什么优势的角度来想。如果你要用同样的设置——我确信他们不是——但如果你要做同样的设置,你基本上就是在你能控制的东西上竞争,也就是数据和算力,而显然他们无法在这方面与前沿实验室竞争。所以是的,如果你是一个新实验室——其实,我不知道你会不会把他们视为新实验室,但我想——

Yeah. So basically try to break the paradigm from sequence-to-sequence decoder-only and just literally do anything else. I mean, I think there's a level of—if you're trying to compete, you're not going to compete with a frontier lab doing an autoregressive decoder on the amount of compute scaling resources they have. Even Thinky will not—okay, there may be one of the handful that can, but you're not really going to do much in that. I don't know too much about Thinky, and I don't want to say anything either, but it's like if their strategy is just to replicate OpenAI or Anthropic, that's a horrible strategy because they just don't have—you kind of just have to think of it in terms of what advantage do you have. And if you're going to use the same setup—I'm sure they're not—but if you're going to do the same setup, you're basically competing on the things you can control, which is data and compute, and obviously they cannot compete with the frontier labs on that. So yeah, it makes sense that if you are a neolab—actually, I don't know if you would consider them to be a neolab, but I guess—

Host

我觉得他们有点奇怪。

I guess they're kind of a weird one.

Alex

是的,他们在那份名单上。他们发布了 Inkling,所以算数。

Yeah, they're in that list. They ship Inkling, so they count.

Host

除了 OpenAI 和 Anthropic——也许还有 Meta 和 GDM——你就得做点别的,你知道吗?这就是残酷的现实。但我其实觉得这是好事。我很高兴 Scaling(规模扩张)有效,这些公司会继续这么做,因为它可能开启新的玩家——如果他们发现了非常非常有趣的东西。因为我有点怀疑前沿实验室是否真的在做这件事,因为就像,你为什么要那样做?当旧的赌注已经奏效时,你为什么要冒险把大量算力分配给新的赌注?我认为这就是催生很多新实验室的原因,对吧?你有一个副赌注,你拿不到算力,你就说,‘好吧,我去。’你举的潜在上升空间的例子就像 Jev,便宜 100 倍,出来了,也许它是个语言模型。在这种情况下,只是不同而已。

Anything other than OpenAI and Anthropic—maybe Meta and GDM—you just got to do something else, you know? Like, it's just the sad reality. But I actually think it's a good thing. I'm very happy that scale works and these companies will just keep doing it because it opens up potentially new players—like if they uncover something really, really interesting. Because I sort of have my doubts that this is seriously going on at frontier labs, 'cause it's like, why would you do that? Why would you take the risk of allocating a large amount of compute to new bets when the old bet already works? And I think that's what spins off a lot of neolabs, right? You have a side bet and you don't get compute, and you're like, 'Okay, I'll go.' And your example of the potential upside is something like Jev, which is 100 times cheaper, comes out and maybe it is a language model. In this case, it's just different.

Host

是的,是的,是的,是的。所以不猜测 Jev 到底是什么?

Yeah, yeah, yeah, yeah. So no speculation on what Jev actually is?

Alex

我猜我对它可能是什么有一些猜测。我见过一些人说它像是扩散模型的东西,也就是并行解码。

I guess I have some guesses for what it might be. I've seen some people say it's like a diffusion thing, which is parallel decode.

Host

是的,并行解码。

Yeah, parallel decode.

Alex

老实说,我认为不管它实际上是什么——因为我见过一些开源复现——他们所做的真正让我兴奋的是,我不完全确定他们的优化目标是什么,以及他们是如何训练的。我认为这是 RLM 的一个问题,我们也一直在思考,那就是:好吧,RLM 是一个非常简单的想法。如果我发表这篇论文,现在任何人都可以用它。但 RLM 的实际价值在于你是否能正确地训练它,以及你是否能围绕这个系统塑造一些架构,让它变得非常非常好。这是我正在积极研究的事情,我想。但我认为对他们来说,他们找到了一种训练系统的方法,这完全不是 trivial 的。我其实不太知道他们是怎么做到的。我在网上看到一些比较——有些人声称他们用了 Qwen 或在 Qwen 上做后训练——但你用的每一个开源 Qwen 都会更差,因为他们训练它所做的显然非常有效。所以这非常令人兴奋。

Honestly, I think regardless of what it actually is—because I've seen some open-source replications of it—what is really exciting to me about what they did is I'm not entirely sure what their optimization objective was and how they trained it. And I think this is a thing for RLMs that we've also been thinking about, which is: okay, RLMs are a very simple idea. If I come out with this paper, anyone can use it now. But what distinguishes the actual value of an RLM is whether or not you can train it properly and whether or not you can mold some architecture around this system to make it really, really good. And that's something that I'm actively working on, I guess. But I think for them, they figured out a way to train the system, which is completely non-trivial. I actually don't really know how they did it. And I've seen some comparisons online—some people are claiming they used Qwen or they post-train on top of Qwen—but every open-source Qwen that you use is going to be worse because whatever they did to train it clearly works very well. And so that's very exciting.

模型玩游戏 Calibration

Host

有一个校准的要素,这是一个罕见的话题,我认为人们甚至不知道或理解。我们在与 Hugging Face 的 Clementine Forio 的对话中讨论过,她曾经负责 Hugging Face 的评估。基本上就是,模型被调整为给你最可能的下一个词,但当你问它们有多自信时,它们会对你撒谎,因为它们只会给你最可能的下一个答案,而不是真正地——不,让我们校准。比如,我实际上有 50% 的把握,我有 20% 的把握,让我们试着校准一下。

There's one element of calibration, which is a rare topic that I don't think people even knew about or understood. We covered it with our conversation with Clementine Forio of Hugging Face, and she used to run the evals at Hugging Face. It's basically the idea that models are attuned to give you the most likely next token, but they're going to lie to you when you ask them how confident are you, because they're just going to give you the most likely next answer instead of actually—like, no, let's calibrate. Like, I am actually 50% sure, I'm 20% sure, and let's try to calibrate that.

Alex

我想说,如果有的话,我认为围绕这个生成合成数据相当容易,因为你可以看到真相,然后合成生成一堆答案,让 Jev 分类,然后与真实标签比较。那会是我的逆向工程。我其实认为校准可能被低估了——就像,人们只是把它当作一个非常快的分类器,但他们实际上甚至没有使用概率或校准估计。我认为它仍然被误解了。重申一下,当你问一个模型‘你有多自信?’它会吐出,比如,43%。最大的区别是这是一个有根据的分类,对吧?

I would say, if anything, I think that's pretty easy to generate synthetic data around, because you can sort of see the truth and then synthetically generate a bunch of answers, have Jev classify it, and then compare with ground truth. That would be my reverse engineering. I actually think calibration is probably under—like, people are just using it as a very fast classifier, but they're actually not even using the probability or calibration estimates. I think it's also still just misunderstood. To reiterate, when you ask a model, 'How confident are you?' it will spew out, what, 43%. The big delta is this is a grounded classification, right?

Host

是的,是的。我非常兴奋地看到人们用这个模型做什么。我的意思是,它会解决一切吗?不,当然不会。但我认为它解决了一类我们传统上难以解决的问题,那就是低延迟的事情。所以我喜欢游戏方面的例子。那实际上——我的意思是,Doom 的例子。很好。

Yeah, yeah. I'm very excited to see what people do with this model. I mean, is it going to solve everything? No, of course not. But I think it solves a class of problems that we traditionally struggled with, which is low-latency things. So I love the examples with games. That's actually—I mean, the Doom example. Good.

Alex

我有一个关于语言模型玩电子游戏的基准测试。我一直着迷于你是否能让一个智能系统玩新游戏和类似的事情。所以我认为他们有一种独特的方式,在一个非常非常快的模型中捕捉语言和理解,这非常非常酷。

I have a benchmark on language models playing video games. I've always been fascinated by whether you can have an intelligent system play new games and things of this nature. And so I think it's really, really cool that they have a unique way to do this to capture language and understanding in a very, very fast model.

Host

好吧,既然你提到这个话题,有什么你想指出的吗?

Well, while you're on the topic, anything you want to point out?

Alex

哦,是的。我的意思是,你知道,你确实做了一个关于许多模型的基准测试。我想这些数字已经非常过时了。很多模型都非常不同。我实际上看到有人运行——有些人在这些游戏上运行更新的模型,这真的很酷。我想这个基准测试的一般前提是,我们只是想看看视觉语言模型是否足够好,能在包含延迟约束的情况下直接接入游戏。因为这实际上——我是在 Claude Plays Pokémon 出来之后马上发布的。所以这大概是两年前,我想现在算是古老了。但我认为这一套任务真正酷的地方在于,游戏种类非常多样。而且,我认为大多数游戏是人们知道或以前见过的。我看到 Jev 玩 Doom。

Oh, yeah. I mean, you know, you did do a benchmark on many anyway. I guess these numbers are very outdated. A lot of the models are very different. And I've seen actually people run—there are some folks out there that are running newer models on these games, which is really cool. I guess the general premise of this benchmark was we just want to see if vision-language models are good enough at just like plugging into games with the latency constraint included. Because this actually—I came out with this right after Claude Plays Pokémon came out. So this was like 2 years ago, which I guess is like ancient now. But I think what's really, really cool about this suite of tasks is it's very diverse in terms of what games they are. And also, I think most of the games are games that people know or like have seen before. I saw Jev playing Doom.

泛化与测试框架的作用 Models Playing Games

Alex

我不认为它们只是在玩非常简单的关卡之类的,但说实话,大多数模型仍然做不到,或者我其实不认为有任何模型能真正有意义地解决这些游戏。有些游戏它们能解决。我想我见过 Astra 能解决 Kirby 游戏。而且我们为这个基准测试特意设计了一个非常精简的 harness。我稍后会谈到 harness 这一点,因为我认为关于 harness 的价值有一整场对话要展开:harness 对模型来说到底有什么目的?但总的来说,我希望很快或不久就能看到所有这些游戏被更新的模型打败。

I don't think they were just playing really simple levels and stuff, but honestly most models still can't really do or I don't actually think any models can solve these games very meaningfully. There are some games that they can. I think I've seen Astra be able to solve the Kirby game. And we also for this benchmark, we intentionally designed a really minimal harness. And I will get to this point about harnesses because I think there's a whole conversation to be had about what is the value of a harness? What is even the purpose of a harness for a model? But in general, I hope to see very quickly or very soon all of these games beaten by newer models.

Host

有意思,就像老早的 Cloudplace Pokemon,它们从 RAM 读取状态,看哪些地块可以走、哪些不能。我们很久以前和他们做过一期播客。

It's interesting like the old old Cloudplace Pokemon, they like read state from RAM and saw what tiles are walkable and what not. We did a podcast with them a long long time ago.

Alex

明白了。明白了。是的。是的,我的意思是类似,Jeff 没有视觉,所以你必须输入这些游戏状态之类的东西。我的意思是,让我们直接进入 harness 的话题,因为你提到了。语言模型 harness 的组合泛化

Gotcha. Gotcha. Yeah. Yeah, I mean it's similar like Jeff doesn't have vision so you have to kind of feed in these like game state and all these things. I mean let's go right into the harness stuff because you brought it up. Language model harnesses a compositional generalizes

Host

读起来有点费劲,解释一下。

struggle to read that explain.

Alex

是的。好的。好的。所以我对人们如何看待 harness 可能有点不满意,因为人们比较说哦我喜欢 Claude Code 我喜欢 codeex 我喜欢 pi 不我喜欢 oh my pi 我喜欢 prime agent 说实话我认为它们都一样,大多数设计决策或围绕这些 harness 的设计选择都一样,也许 prime agent 有点不同,因为它本质上是一个 RLM,但总的来说我认为我们可以对 harness 更有创意。我的意思是,如果我们从 harness 到底为模型做什么的角度来思考,基本上当你试图解决一个问题并想用语言模型来解决它,比如一个非常困难的任务。我们发现的一件事是,下一个词预测是一种非常非常笨拙的形式来完成很多这些任务。例如,以 Swebench 为例。当你在代码库中导航时,你能通过一次语言模型调用弄清楚如何做所有这些吗?就像你只说解决代码或解决我在这个代码库上的查询?不。所以我们依赖 harness 来帮助你做这些事情。我认为有趣的是,harness 是一个非常非常有主见的程序,关于你想如何让语言模型适应一个问题。我提到这个是因为当我们思考 harness 在做什么时,我们应该真正思考 harness 中的哪些选择让语言模型解决这个任务,我能不能实际上有一个语言模型直接做这个?因为 harness 如果你现在想想,既然循环 Transformer 已经存在,我认为真正有趣的是你可以在某种程度上用 harness 来建模一个循环 Transformer,对吧?你只是在模型上循环。现在,你可以说哦我没有解码,所以有点不同,但总的来说,我们出于某种原因一直坚持相同的模型架构选择。有很多理由,但显然我们现在围绕 harness 的 rollout 来训练模型。所以有一种非常尴尬的方式来在 harness 上进行训练,那就是我们训练一个语言模型在 harness 内行动,但现在它就像一个非常长的可能像多个智能体的 rollout,我们有非常尴尬的 hacky 方式来做这个。所以这篇博客谈论的是,好吧,你可以这样思考这里发生了什么,如果 harness 基本上帮助模型解决一个特定任务。不同的 harness 设计选择实际上能不能做一些更有意义的事情,而不仅仅是这里有一些工具调用会帮助你。这里有一种通过你的代码库进行 GP 的方式。所以这实际上这篇博客的想法也随着 RLM 的想法而来。我们只是没有那样包装它。我认为这实际上是真的。关于 RLM 还有许多其他想法,我们会陆续推出,但它们从一开始就都在那里。我的意思是这些都是设计决策,关于嗯我认为我喜欢 RLM 的地方是,有许多许多许多迭代和不同抽象版本,我感兴趣去做,最终 RLM 最有意义。但有很多不公开的原因导致这种情况。嗯我的意思是你你会看到的。所以在这个例子中,我们看到 RLM 的一件事是,如果你充分卸载上下文并让模型在这个上下文上写代码,你会得到这个非常非常奇怪但有用的属性,那就是当模型在训练中认识到如何解决一个任务时,事实证明许多任务的解决方案非常相似。跨越那些你甚至不清楚解决方案相似的任务。所以在这个例子中,我们有一个检索任务和一个聚合任务,它们的查询非常非常不同。领域完全不同。当你在这两个任务上训练一个常规语言模型时,轨迹看起来非常不同。所以当你比如使用 pi 或 Claude Code 之类的,这不在博客中,但我们有这些结果。你会发现这些 harness 在这些问题之间区分太多,即使解决方案相同。所以我们训练 RLM 时发现的一件事是,当它看到这些问题相同时。它看到这些问题相同的原因是子智能体看到不同的问题,但子智能体解决一个更简单的子任务。所以你确信它足够聪明来做这个。但对于基础整体策略,它们最终看起来相同。所以例如当你训练左边的任务时,它可以立即解决右边的任务。所以如果你看我们有的图表,你会观察到的一件事是,当你只是天真地在这些任务上训练你的 RLM 时,它们自然学会泛化,例如到更长的任务,因为策略基本上相同。你只是修改一个长度变量。这实际上也适用于不同的任务。它们甚至不是长度不同。它们是完全不同的任务,数学任务与写作任务,但解决方案,像元高级解决方案是相同的。所以当你训练 RLM 在其中一个上时,它将这种行为泛化到第二个。这里没有魔法。我我想也许这是我想强调的事情,你会如何用语言表达它们在学习什么?所以我认为在这里你说你在短任务上训练。

Yes. Okay. Okay. So I have been a little unsatisfied maybe with how people think about harnesses because people compare like oh like I love Claude Code I love codeex I love pi like no I love oh my pi I love prime agent to be honest I think all of them are the same most of the design decisions or like the design choices around these harnesses are the same maybe prime agent is a little bit different cuz it's like inherently an RLM but in general like I think we can be a lot more creative with harnesses. And what I mean by that is if we think about this from the perspective of what exactly is the harness doing for the model well basically when you're trying to solve a problem and you want to use a language model to solve it like a very difficult task. One thing that we have discovered is that next token prediction is a really really awkward form to do a lot of these tasks. So for example, take Swebench. When you're navigating a codebase, like are you going to be able to figure out how to do all of this with a single language model call? Like you just say solve code or like solve my query over this codebase? No. So we rely on a harness to help you do these kinds of things. And I think what is interesting is like a harness is a very very opinionated program over how you want a language model to be form fit over a problem. And I bring this up because when we think about the what a harness is doing, we should really think about like what choices in the harness let the language model solve this task and can I actually just have a language model that just does this? Because a harness if if you think about it now that loop transformers are a thing I think what's really interesting about it is you can model a looped transformer in some ways like with a harness as well, right? you're just looping over the model. Now, you can say like, "Oh, I'm not decoding, so it's like a little bit different, but in general, we for whatever reason have stuck with the same model architecture choice forever." And there's many arguments for why, but clearly like we are now training models around harness rollouts. And so there is this very awkward way of doing training over harnesses which is that we train a language model to act within a harness but like now it's like a really long maybe like multiple agent roll out that we're doing and there's like really awkward hacky ways of doing this. So what this blog talks about is like well one way you can think about what is going on here is if the harness is basically helping the model solve a particular task. Can different harness design choices actually do something a little bit more meaningful beyond just here are some tool calls that will help you. Here is a way to GP through your codebase. And so this actually the idea for this blog came with the RLM idea as well. We just didn't package it that way. And I think this is actually true. There are many other ideas around RLMs that like we will be coming out with, but we're all there from the beginning. I mean these are all design decisions around um I think with what what is what I like about the RLM is that there were many many many iterations and versions of different abstractions that I was interested in doing and ultimately the RLM made the most sense. But there's a lot of reasons that aren't public as to why that's the case. Um I mean you you will see. So in this example, one of the things that we see with an RLM is that if you sufficiently offload context and write ask the model to write code over that context, you get this really really weird but useful property which is that when the model recognizes during training how to solve a task, it turns out that the solution to many tasks is very similar. across like tasks where you you don't even like it's not even that clear to you that the solutions are similar. So in this example, we have like a retrieval task and we have like a an aggregation task and they're very very different query. Like the domain is just completely different. And when you train a regular language model over these two tasks, the trajectories look very different. And so what you're relying when you like let's say you use pi or Claude Code or something, which is not in this blog, but we do have these results. You'll find that like these harnesses distinguish too much between these problems even though the solutions are the same. And so one thing that we find when training RLMs is that like when it sees these problems as the same. And the reason it sees these problems as the same is the sub agent sees different problems, but the sub agent is solving an easier subtask. And so you're confident it's smart enough to do it. But for the base overall strategy, they end up looking the same. And so when you train on the left task for example, it can immediately solve the right task. And so if you go down to like the plots that we have, one thing you'll kind of observe is that the as you just naively train your RLM on these tasks, they naturally learn to generalize, for example, to longer tasks because the strategy is basically the same. You're just modifying like a length variable. And this actually also holds for tasks that are different. And it's not even they're not different across length. They're completely different tasks, math tasks versus writing tasks, but the solution the like meta highle solution is the same. And so when you train the RLM on one of them, it generalizes this behavior to the second one. And there's no magic here. I I guess maybe that's the thing that I want to kind of stress like how how would you kind of verbalize what they're learning? So I think in here you say you train on short tasks.

改进通用测试框架 Generalization and the Role of Harnesses

Host

它们能泛化到长 8 到 30 倍的任务上。它们是在学习如何解决这类问题,还是说它们真正学到的核心东西是什么?

They generalize to stuff 8 to 30 times longer. They are learning how to solve these type of problems or what's the core thing they're actually learning?

Alex

是的,它们学的是如何解决某个特定长度下的这类问题,而事实证明,当你把学到的策略迁移到更长的长度时,它可以直接迁移过去。它们本质上就是同一个程序。这非常令人兴奋,因为这暗示着,当你有一个用于训练模型的数据或环境语料库时,希望在于:a)你可以用更少的环境训练,却比现有模型通过朴素训练泛化到更多环境;b)你仍然想用上所有数据。所以当你在这些任务上训练时,希望它能泛化到更广泛的问题类别。而更令人兴奋的是,至少在 RLM 或任何递归调用系统的语境下,这个论证是归纳成立的。举个例子,我说过竞赛编程和 GPU 优化使用非常相似的技能组合。模型可以——harness 也许需要你以某种方式引导一下——但它能学会:好,我解决这个 GPU 编程任务的方式和我为竞赛编程学到的非常相似,所以我会列出一组解决方案,我会生成子智能体来列出有希望的解决方案,然后我会写这个循环来检查这些解决方案,也许对它们进行演化,并针对验证器进行演化。在这两个任务之间,这看起来是一样的,但子智能体在做的事情可能独特、有些不同,但甚至子智能体解决的问题也可能具有相同的形式,对吧?因为这是一个递归论证。所以我在整篇博文中想说的是,我们应该重新思考 harness 的角色,因为 harness 实际上可以极大地提高模型的泛化能力以及它所获得的数据量。这并不局限于 RLM。我认为还有一大类尚未被发现的 harness 实际上可以产生类似的特性。再把这个论证延伸一点,我认为你还可以推断出:如果我审视一个 RLM,哪些组件是真正必要的?我能不能直接训练一个模型来做这件事?我能不能训练一个模型在其前向传播中隐式地充当 RLM?这是一个非常奇怪的想法,因为你可能会说,啊,代码是不可微的 blah blah blah,但我们会开始发现许多这种行为的近似。我认为我们将超越仅仅设计一个使用特殊压缩形式的新编码 harness 之类的做法。我认为我们可以在这里更有创造力。我们可以用这些语言模型做太多事情,而我认为我们就是没在做。我对此非常兴奋,因为我认为我们可以从非常非常固执己见且优秀的 harness 设计中获得巨大的收益,这种设计更适合扩展。我的意思是,RLM 例如是一个非常原始的归纳偏置。除了它与我们目前所做的非常不同之外,设计上没有什么特别之处。但这可能在我们可用的数据和环境下扩展得好得多。

Yeah, they're learning how to solve these types of problems at a certain length and it turns out that when you take this strategy that they learned, it is directly transferable to the longer length. Like they're effectively the same program. And this is something that is really exciting because what this kind of implies is that when you have a corpus of data or environments that you train your model on, the hope is that like a you can train on less environments and generalize to more than what existing models can do through naive training. But b you still want to use all the data that you have. So when you train on these tasks like hopefully it generalizes to a wider class of problems and why this is also even more exciting at least in the context of RLMs or any recursively calling system is this argument holds inductively. So like I'll give you an example cuz I said I claimed that competitive programming and GPU optimization use very similar skill sets. The model can the harness potentially I mean you might have to nudge it a certain way but it can learn that like okay how I'm going to go about solving this GPU programming task is very similar to what I learned for competitive programming so I'm going to list out a set of solutions I'll like spawn sub agents to list out promising solutions and then I'll like write this loop to go through and check these solutions maybe evolve them and like evolve them against a verifier and between these two tasks this looks the same, but what the sub agents are doing are maybe like unique and something like different, but even what the sub agents are solving might actually also be of the same form, right? Cuz it's like a it's a recursive argument. And so what I'm trying to get at with this whole blog post is just that like we should rethink what the role of the harness is because harnesses can actually greatly increase the generalization capability of your model and the amount of data that it's given. And this is not exclusive to RLMs. I think there is a wide class of harnesses that are yet to be discovered that actually can yield similar properties. And to extend this argument a little bit further, I think what you can also extrapolate from this is like if I look at an RLM, what are the components of an RLM that are actually necessary? And can I actually just directly train a model to do this? Can I train a model to act as an RLM implicitly in its forward pass? It's a really weird thing to think about because like you know you might say like ah code is non-differentiable blah blah blah blah but there are many approximations of this behavior that we will start to uncover and I think like we will see beyond just like I'm going to design a new coding harness that uses a special form of compaction or something. I think we can be a lot more creative here. Like there's so much we can do with these language models that I think we are just not doing. And I'm like very excited about this because I think like I think we can get serious gains from very very opinionated and good harness design that lends itself better to scale. And what I mean by this is like you know the RLM for example is a very primitive inductive bias. Like there's nothing super special about the design other than the fact that it's very different than what we currently do. But this may potentially scale much better with like the data and the environments that we have available to us.

RLM 结果与效率 Improving Generalized Harnesses

Host

我想人们可能会问的反面问题是,当前的 harness 非常泛化于编码,而人们看到这在很多领域都有效。Claude Code 被用于设计演示等一切。嗯,Muse、Spark、Grok Bot。这些是非常简单、不固执己见的 harness,擅长代码,而且也在扩展。嗯,我想问的是,我们如何改进这些的例子是什么?

I guess the opposite thing that people would probably ask is current harnesses are very generalized towards coding which people see works for a lot of domains. Claude Code is being used for design presentations everything. um Muse, Spark, Grok Bot. These are very simple, non-opinionated harnesses that are good at code and that is also scaling out. Um what's the example of how we improve those I guess?

Alex

是的。嗯,让我提另一篇最近刚出的论文。就是那篇 harness tax 论文。我想是 Arena 写的。我非常喜欢这篇论文,因为它提出了我之前的一个先验,基本上就是大多数

Yeah. Um let me let me bring up another paper uh which came out very recently. It's like the harness tax paper. I think it's by Arena. I really like this paper because it puts forward a prior that I had, which is basically that like most

Host

它证实了一个先验。

it confirms a prior.

Alex

是的。就像大多数 harness 选择并不重要,因为所有这些 harness 都是一样的。

Yeah. Like most harness choices don't matter because all of these harnesses are the same.

Host

但我要说,比如 Grok Bot,据我理解,实际上与这些其他 harness 的设计相当不同。我非常喜欢这一点。嗯,我认为至少从这里可以清楚看出,我很确定 Anthropic 或 OpenAI 只在他们的 harness 上训练。他们可能不会在竞争对手的 harness 上训练。我的意思是,我假设不会,因为我不知道他们为什么要那样做。

But I I will say like Grok Bot, for example, is actually quite different, I think, from my understanding than than how some of these other harnesses have been designed. And I like that a lot. Um, and I think it's clear from here at least that like I'm pretty sure Anthropic or OpenAI are exclusively training on their harnesses. They're probably not training on their competitors. I mean, I'd assume not because I I don't know why they would do that.

Alex

但我的意思是,你知道,这是你在开源模型中看到的现象,对吧?比如 Quinn 非常擅长使用 open code。它们需要在 harness 中训练。旧的 Gemas 在这方面出了名的差。模型很好,但你需要在一个 harness 中训练。不过我认为,随着模型变得更聪明或更好,这种区别变得不那么重要了,因为如果你把 Astra 放进 open code 里,它不会发疯,因为我认为它已经足够好了。我这么说是因为这些不同 harness 之间唯一的优势,大部分情况下只是成本。而我想说的是,如果你把这些模型插入 RLM,它们仍然没那么好。它们还行。我认为这主要是因为我们在训练时所围绕的 harness 类别就是这一类 harness。这种 pi 循环,我喜欢称之为“轨迹即提示”,意思是你把整个 rollout 的轨迹保留为主模型所使用的上下文。即使你使用子智能体,它仍然有点像这种形式。我认为,如果我们想探索新的 harness,就需要有专门的团队真正进行有意义的实验,比如扩展新的 harness,也许是在不同 harness 设计上进行后训练扩展。我认为我们实际上可以从这类事情中获得非常有意义的知识或收益。无论是 RLM 还是别的什么。这很令人兴奋,因为我认为,例如,Fable 在很长一段时间内是 RLM 的最佳模型,因为他们有动态工作流,而且很明显这是一种某种程度上被训练进去的能力,即使模型仍然有点笨,在 RLM harness 中它仍然比其他模型表现好得多。Astra 现在也足够擅长做这些事情了。

But I mean, you know, this is a thing you see in open models, right? Like Quinn is really good at using open code. They need to train in harnesses. Old Gemas were notoriously bad at this. Models are good, but you need to train in a harness. I I think though as models get smarter or like as they get better this distinction becomes uh not that important in the sense that like if you take Astra and you put it inside of open code like it's not going to go crazy because I think it's like just sufficiently good. And so I say this because the only benefits between these different harnesses is just cost for the most part. And and I think like what I'm getting at here is if you plug these models into RLMs though, they're not that good still. They're okay. And I think it's mainly because the types like the class of harness that we are training around is this class of harness. This like pi loop this like I I like to call it trajectory as a prompt which just means like you keep the whole trajectory of the roll out as the context that that that your main model is is using. Even if you use sub agents, it's still like kind of this form. And I think we're going to if we want to explore new harnesses, there needs to be teams that are dedicated to actually running meaningful experiments over like scaling out new harnesses, like maybe post-train scaling out on different harness designs. Like I I think we can actually get very meaningful knowledge or gains from doing this kind of thing. Uh whether it's an RLM or whether it's something different. And that's exciting because I think for example if you train a lot on Fable for a long time was the best model for RLMs because they had dynamic workflows and it was pretty obvious that like this was a capability that was somewhat trained in even if the model was still like a little dumb like in the RLM harness it still worked a lot better than than other models did. Astra is now also like good enough at doing these things.

本地分布内任务 RLM Results and Efficiency

Host

但我们是在这里看到 Fable 最适合 RLM,还是有别的——

But is this where we see that Fable is the best for RLMs, or is there some other—

Alex

哦不,这些应该都是内部结果。对,我现在还没有公开。但总的来说,我觉得你很容易看出来,我们还没有针对 RLM 式的工作流做优化。而且我认为,如果我们得到能正确做这件事的模型,它们也会高效得多。我的意思是,你其实可以自己想一想为什么会这样,对吧?

Oh no, these are all internal results, I guess. Yeah, I don't have them public right now. But in general, I think you can very easily tell that we have not optimized for RLM-like workflows yet. And I think if we get models that do this correctly, they will be a lot more efficient as well. I mean, you can kind of just think through why this is the case, right?

Host

我觉得到了这里,你得用 10 秒钟讲一下什么是 RLM,因为这里有很多听众——

I think this is the point where you have to give the 10-second what are RLMs, because there's a lot of listeners here that—

Alex

我觉得我们也在假设大家有很多背景知识。

We're assuming a lot of knowledge also, I think.

Host

但你已经铺垫了一些背景,所以结合我们刚才说的,能不能给 RLM 一个干净、清晰的定义?

But you have set some context, so with everything we just said, can we have a clean, crisp definition of RLM?

Alex

好。我想回到那篇博客,那篇讲组合式泛化器的博客,就是这篇。好。这大概是最好的——我希望我原论文里就有这个。RLM 基本上就是一种 harness 设计,其中 harness 里唯一的工具就是代码。呃,也就是这种程序化的子智能体调用机制,它可以选择把自己当作工具来调用,也有其他工具,但所有这些东西都是代码里的函数。而它处理的上下文总是存储在这个代码环境内部的某个记忆里。所以这可以是一个文件系统,比如——我举个例子:Prime Agent。Prime Agent 的轨迹,也就是上下文,即使你做了压缩之类的操作,也会存在磁盘上,所以模型总能引用它原始的上下文,即使它被压缩了。而它所有的工具都运行在,比如说,一个 Python REPL 或 bash REPL 里。所以这就像是一种非常原始的抽象。我会说,大多数 harness 的区别就在于,上下文卸载不是那样做的。我的意思是,如果你看 Prime Agent,顺便说一句,Prime Agent 确实做上下文卸载,但不是完全做到底。所以它仍然维持标准的 Claude Code 循环,也就是把轨迹当作提示词,然后你做压缩,但它额外有那种——上下文被卸载了,而 Prime Agent 唯一独特的一点是,它唯一的工具是 IPython。所以这大概就是围绕 RLM 的那种非常通用的抽象。

Yes. Okay. I want to go back to the blog, the compositional generalizers blog, this one. Okay. This is like the best—I wish I had this in the original paper. An RLM is basically just a harness design where the only tool in the harness is code. Uh, which is this programmatic sub-agent calling thing, where it has the option to call itself as a tool, and it has other tools, but all of these things are functions in code. And the context that it's dealing with is always stored in some memory inside of this code environment. So this could be a file system, like this could be—like I'll give you an example: Prime Agent. The trajectory of Prime Agent, like the context, even when you compact and do all these things, is stored on disk, so the model can always reference its original context even if it's compacted. And all of its tools are run inside of, let's say, like a Python REPL or a bash REPL. And so this is like this very primitive abstraction. And I would say where most harnesses differ is context offloading is not done like that. I mean, if you look at Prime Agent, by the way, Prime Agent does context offloading, but not all the way. So like it still maintains the standard Claude Code loop of like trajectory as a prompt where you compact, but it has the additional kind of—like the context is offloaded, and its only unique point of Prime Agent is that the only tool is IPython. So this is like the very kind of generic abstraction around like RLMs.

Host

具体来说,RLM 试图解决的核心问题是什么?我记得它刚出来的时候,一度是上下文,还有——

Concretely, what's the core thing RLMs are trying to solve? At one point I think when it first came out it was context, and—

Alex

这个问题我转给别人。

I'll pass the question.

Host

好。

Yeah.

Alex

所以最初的问题是,harness 在处理长上下文时表现非常糟糕。它们通常只被用来做像代码这样的特定事情。比如,你知道,它们能处理你的代码库,因为模型是在代码上训练过的。但现在更多是围绕这篇博客所讲的,也就是组合性,以及这样一个事实——我觉得 harness——我们想要的语言模型系统,能在每一步对自己的行动有更多控制。我的意思是,工具调用非常受限,因为你每一轮都必须调用它们。比如你必须先调用工具 A,再调用工具 B,再调用工具 C,而且没有一个你可以随时回溯的中心上下文。而 RLM 正是围绕组合性设计的,并且有一个你总能从中取用的中心上下文。嗯,而这个上下文是围绕现有语言模型设计的。另一个在设计上非常相似的例子是智能体集群。比如 Hugging Face 那次事件——这些智能体集群有一个留言板,它们学会在上面交流,而这个留言板在某种意义上就是它们行动所依托的共享上下文。而 RLM 基本上是说,通过它交流的最佳方式是用代码,也就是你写代码来做这件事,因为这些模型太擅长写代码了。嗯,我们想利用这一点。

So the original problem was harnesses have a really bad time dealing with long context. They typically were used to only really do them for like specific things like code. Like, you know, they could deal with your codebase because it was trained on it. But now it's more around what this blog is talking about, which is compositionality, and the fact that—like I think harnesses—we want to have language model systems that have much more control over the actions they make at every step. And what I mean by this is tool calls are very limited because you have to invoke them every turn. Like you have to invoke tool A, then tool B, then tool C, and there's no central context that you can kind of draw back from. And RLMs are specifically like designed around composition and having like a central context that you can always draw from. Um, and this context is like designed around the existing language models. Another like very similar example actually in design is like agent swarms. For example, the Hugging Face incident—like these agent swarms have like a message board that they learn to communicate over, and this message board in some sense is the shared context that they like act over. And RLMs basically say that like the best way to communicate through this is in code, like you write the code to do this, and it's because these models are so good at writing code. Um, like we want to take advantage of that.

Host

然后对于这种组合性的东西,甚至包括生成你自己的 harness,你知道,专门针对任务定制。

And then for this compositional thing, up to and including generating your own harness, you know, specific for the task.

Alex

正是。正是。

Exactly. Exactly.

什么是更聪明的测试框架 Locally In-Distribution Tasks

Host

我觉得如果你往上看这个图,我们就会开始看到这一点。好。

I think we will start to see that if you go up to this figure. Okay.

Alex

所以我们谈到 harness 的“局部同分布任务”这个概念,它有点像我们在思考语言模型时“同分布任务”之上的一个概念。同分布任务就是提示词是模型以前见过、或者见过某种版本的任务。大多数 harness 都像右边那个一样运作,也就是它们不断把轨迹作为提示词追加进去,所以最终,除非你是 Anthropic 或 OpenAI,并且在这些用户轨迹上训练过,大多数这些东西大部分时候最终都会变成分布外的。但局部同分布基本上是这样一个组合性论点:如果一个 RLM 把它的计算拆解成某种元 harness,或者一个包含子智能体、去处理局部问题的程序,那么在这个任务过程中每一次单独的语言模型调用都是同分布的,即使整个任务本身是分布外的。我认为这是一个非常理想的性质,原因显而易见。比如,如果对每一次单独的语言模型调用来说每个任务都是同分布的,你大概就能得到正确答案。嗯,所以我的意思是,RLM 的逻辑极限就是这样的 RLM:你不仅写 harness,还训练一个定制模型、收集数据,一切都像是一个在你 harness 内部的全自动 AI 研究员。

So we talk about this idea of locally in-distribution tasks for a harness, and it is like an idea on top of like in-distribution tasks when we think about language models. An in-distribution task is just a task where like the prompt is something that the model has either seen before or like has seen some version of it. Most harnesses work like the one on the right, which is they keep appending the trajectory as a prompt, and so eventually, unless you're Anthropic or OpenAI and you train on like kind of these like user trajectories, most of these things end up being out of distribution for the most part. But locally in-distribution is basically the compositional argument of: if an RLM breaks down its computation into like a kind of a meta harness of sorts, or like a program that involves sub-agents that like look at a local problem, every individual language model call over the course of this task is in distribution, even if the entire task itself is out of distribution. And this is a very desirable property, I think, for obvious reasons. Like if every task is in distribution for each individual language model call, you will probably get to the right answer. Um, so I mean, the logical limit of RLM is RLMs where like you not just—you don't just write the harness, you also train a custom model, you collect data, everything like it's like a fully automated AI researcher inside of your harness.

Host

对,我们会看到 RLM 的训练走向何方。

Yeah, we will see where the training of RLMs goes.

Alex

我要说,作为一名学者,我在 MIT 并没有做这个,至少不是大规模地做,因为我负担不起。呃,但外面有公司在做这个。我觉得 Prime 和 Select 显然在做这个,非常酷。比如我非常期待看到——也许我们会观察到,我不知道,但也许我们会观察到,当你围绕一个聪明的 harness 训练时,会出现更好的后训练缩放墙。也许我们甚至会看到更聪明的 harness 出现,它们围绕这类原则运作得更好。

I will say as an academic I am not working on this at MIT, or at least in the scaled sense, because I can't afford to. Uh, but there are companies out there that are working on this. I think Prime and Select is very clearly working on this, and it's very cool. Like I'm very excited to see—maybe we'll observe, I don't know, but maybe we'll observe better like post-training scaling walls when you train around a smart harness. Maybe we'll even see smarter harnesses that come out and like they work better around these kinds of principles.

架构、数据与缩放定律 What Is a Smarter Harness

Host

什么叫更聪明的 harness?这根本没什么意义。你刚说它们都一样。

What is a smarter harness? Like that doesn't mean anything. You just said they're all the same.

Alex

不,我更多想说的是,Claude Code、Codex、Pi 等等都一样,因为当你拆解 harness 的逻辑时,它们几乎是同一个东西。

No, what more of what I mean is like Claude Code, Codex, Pi, etc. are all the same in that like when you break down the logic of the harness, it's like virtually the same thing.

Host

对。循环里两次调用,诸如此类。

Yeah. Two calls in a loop, whatever.

Alex

但有了 RLM 和其他 harness 抽象,它看起来就非常不同。而我认为这才是真正区分开来的地方。

But with RLMs and with other harness abstractions, it looks very different. And this is where I think you really distinguish.

参与 Prime 项目 Architecture, data, and scaling laws

Alex

这就像语言模型的架构选择:当你把它扩展出去时,很多架构选择最终看起来都差不多,差异会变得比较小。对一个实验室来说这不算小,但也许一个模型收敛得稍微好一点。总的来说,预训练缩放定律之所以成立,是因为我们现有的架构选择相对稳定。但如果你彻底改变架构,预训练缩放定律很可能就不成立了,或者那条幂律会看起来非常不同。Harness 也是一样。我觉得现在所有的 harness 大体上都长得差不多,但也有一些例外正在出现。

It's in the same way that I think with language model architecture choices, a lot of architecture choices end up looking the same when you scale it out. The differences end up being somewhat minor. For a lab it's not minor, but maybe one model converges slightly better than the other. In general, pre-training scaling laws only hold because the architecture choices we have are somewhat stable. But if you were to completely change the architecture, pre-training scaling laws probably don't hold, or the power law is going to look very different. It's the same thing with harnesses. I think all the harnesses we have right now for the most part roughly look the same. But there are some exceptions coming out.

Alex

其实,说到那些愿意冒大风险的博士生,我最近更感兴趣的一件事是,人们也在追求另一面:如果你改变数据,预训练缩放定律就不成立。现在数据只是原始的非结构化文本,一个互联网语料库。如果你有更好的数据表示来训练呢?那么是的,你的缩放定律也会改变。

Actually, one of the things I've been more interested by, talking about PhD students who take big risks, is that people have also been pursuing the other side, which is that pre-training scaling laws don't hold if you change data. Right now it's just raw unstructured text, a corpus of the internet. What if you had a better data representation to train on? Then yes, your scaling law would change as well.

Host

所以有架构,有数据,还有你能想到的其他东西。我想回到这一点。这些都说得通。很有意思的是,你在整个技术栈上上下下地递归,从非常概念性的东西,到“这就是我们今天的现状”。但显然你可以上下扩展。我很好奇:你是怎么开始和 Prime 合作的?Prime 会在这方面承担更多工作吗?这是他们对 Hermes agents 的回应吗?你提到 Grockot 有点不同。我只是想把这些都点一遍,听听你对每一个的看法。

So there's architecture, there's data, and whatever else you can think about. I just want to get back to this. It all makes sense. It's very interesting how you sort of recurse up and down the stack, from very conceptual to, well, this is where we are today. But obviously you can scale up and down. I'm curious: how did you start working with Prime? Is Prime taking on more work with this? Is this their answer to Hermes agents? You mentioned Grockot is a little bit different. I just wanted to name-check all these guys and get your thoughts on each.

Prime Agent 是什么 Getting involved with Prime

Alex

是的。我是在 Prime 发布了一篇博客文章之后参与进去的——顺便说一句,那篇文章和我完全没关系——他们在文章里说相信 RLM 是未来。我有一个朋友在那里工作,GP Mode Mate,我们联系上了。我想我同意那里很多研究者的看法,以及他们对 harness 设计的信念。我印象非常深刻。我觉得他们理解了 RLM 论文的目的,那并不一定只是说我们在解决长上下文任务,而是说我们想要更有主见的 harness 设计。

Yeah. I got involved with Prime after they released a blog post — by the way, not affiliated with me at all — about how they believed RLMs were kind of the future. I had a friend working there, GP Mode Mate, and we got in touch. I think I agreed with a lot of the researchers there and what they believed about harness design. I was very impressed. I think they understood the purpose of the RLM paper, which is not necessarily just to say that we're solving long context tasks, but actually that we want more opinionated harness designs.

Host

是的。论文里总有你选择去强调的那个结果。

Yeah. There's always the result of the paper that you choose to highlight.

Alex

相对于真正的要点。

Versus the actual point.

Host

对。我很想聊聊学术界的激励机制,以及它为什么有缺陷、有哪些问题,我们待会儿再回到这个话题。

Yes. I would love to talk about the incentives of academia and why it's kind of flawed and all the issues, and we'll get back to that.

Alex

是的。总之,我很喜欢 Prime 的人。在我们决定合作之后,我们决定研究训练一个 RLM,同时也构建这种 RLM harness,看看能把它带到哪里。Prime Agent 就是这么来的。我觉得 Prime Agent 的反响相当不错。我担心的一点是,至少在我们构建它的时候,没有一个模型特别擅长做 RLM 的事情。所以那是在 Astra 之前。

Yeah. So anyways, I love the guys at Prime. After we decided to work together, we decided to look into training an RLM and also build this kind of RLM harness and see where we can take it. That is how Prime Agent came about. And I think the reception for Prime Agent has been pretty good. The one thing I was worried about with Prime Agent is that none of the models, at least at the time when we were building it, were that good at doing RLM stuff. So this was pre-Astra.

子代理与状态外化 What Prime Agent is

Host

我想退一步问,你能解释一下 Prime Agent 是什么吗?它和传统的 Claude Code、人们预期中的 harness 有什么不同?

I guess, to take a step back, can you explain what Prime Agent is, how it's different from a traditional Claude Code, what people would expect from a harness?

Alex

好的。Prime Agent,我想我前面稍微提到过。它基本上是在 Pi 之上的一个 harness,比如 Pi Mono——补充一下背景,Pi Mono 是一个极简主义的 harness。我把 Pi 当作一切的参照,因为我觉得所有其他 harness 基本上都只是 Pi。但它就是 Pi,只不过我们明确限制它只能使用这一个工具。其他所有工具都作为 Python 模块或 bash 脚本加载进来,供它运行。所以它在 Pi 之上使用了核心的 RLM 抽象,然后它还有这个 continual harness 的东西,那是 Seth——他是另一个博士生。这是他用来让语言模型 harness 玩游戏的。他和 Joel 合作很多,Joel 就是那个让 Gemini 玩 Pokémon 的人。顺便说一句,continual harness 也非常简单。我很喜欢它。它基本上是一个设计原则:harness 的哪些部分你可以让 harness 自己修改。有些部分你会让它修改自己的技能、可用的子智能体、给模型的系统提示。而 continual harness 基本上作为 IPython 内核里的一个工具可用。这就是 Prime Agent。Prime Agent 里其他一切都是标准的。

Yes. So Prime Agent, I think I mentioned this a little bit earlier. It's basically a harness on top of Pi, like Pi Mono — Pi Mono for context is a minimalist harness. I use Pi as the reference for everything because I think all other harnesses are basically just Pi. But it is Pi except we explicitly restrict it to be the only tool that's available to it. Every other tool gets loaded in as a Python module or a bash script that it can run. So it uses the core RLM abstraction on top of Pi, and then it also has this continual harness thing, which is Seth — he's another PhD student. This is a thing that he used to get language model harnesses to play games. He worked a lot with Joel, who is the Gemini plays Pokémon guy. And continual harness is also, by the way, very simple. I quite like it. It's basically this design principle around what parts of the harness can you let the harness itself modify. There are certain pieces where you'll let it modify its own skills, the subagents available to it, what the system prompt to the model is. And continual harness is available basically as a tool inside of the IPython kernel. So that's what Prime Agent is. Everything else in Prime Agent is standard.

Host

标准的。

Standard.

Alex

标准的,对。我觉得我真正喜欢它的一点,而且我们也算走运,是很多新的前沿模型在里面实际上运行得很好,甚至很多开源模型也运行得很好,至少一些较新的模型是这样。还有一点我应该强调,Prime Agent 里有一个非常特别的智能体间通信系统或框架,因为 RLM 往往会生成很多子智能体。我们想要一种方式,让子智能体可以和根智能体或彼此通信。所以有一些设计决策,关于每个子智能体被允许和谁对话、怎么对话。同样,一切都在代码里。所以它写代码来做这种通信,我觉得这非常酷。然后还有持久化子智能体,这是后来加进去的另一个东西,也就是子智能体可以存活得比原始智能体的标准运行时间更长,你可以进入那个子智能体,继续给它提示,你对正在发生的事情有更多可见性和灵活性。

Standard, right. Yeah. I think what I really liked about it, and we got kind of lucky, is that a lot of the new frontier models actually work really well inside it, and actually even a lot of the open source models work really well, at least some of the newer ones. And there is another thing in Prime Agent I should highlight, which is that we have a very particular agent-to-agent communication system or framework, because RLMs tend to spawn many subagents. We want a way for subagents to communicate with maybe the root or with each other. And so there are some design decisions around what each subagent is allowed to talk to, how it does it. Again, everything is in code. So it writes the code to do this kind of communication, which I think is really cool. And then there's, I guess, persistent subagents, which is another thing that was kind of added, which is that the subagents can last beyond the standard runtime of the actual original agent, and you can go into that subagent, you can prompt it more, you have more visibility and flexibility into what is going on.

Prime 的未来计划 Subagents and externalizing state

Host

这是我现在对 Codex 最大的痛点。他们的子智能体非常短暂,而且他们主动劝阻你把它用于长时间运行的任务。

This is my number one pain with Codex right now. Their subagents are just very ephemeral and they actively discourage you from using it for long-running things.

Alex

是的。我觉得这说得通。

Yeah. Which I think makes sense.

Host

所以我的意思是,诀窍就是把它外化到文件系统,对吧?这就是诀窍。这就是那个大诀窍。

So I mean the trick is just externalize to a file system, right? Like that's the trick. That is the big trick.

Alex

是的。而且还要强制一切都通过代码运行。相信那个会写代码的模型,它自己会写出自己的 harness。

Yeah. And also force everything to run through code. Trust the model that can write code, and it's going to write its own harnesses itself.

训练模型与博士研究 Prime's future plans

Host

那么 Prime 会为此承担训练、后训练定制模型的工作吗?这是你们之间的一次性合作吗?就这样。

So is Prime going to take on training, post-training custom models for this? Is this a one-off collaboration between you guys? That's it.

第三方对 RLM 的研究 Training Models and PhD Research

Host

比如什么……

Like what's...

Alex

对,他们在内部训练一个模型。我觉得他们其实挺公开的,早在三月份就说了。

Yeah, they are training a model internally. I think they were pretty public about this actually, back in March.

Host

我是说,这显然是他们的业务,对吧。

I mean clearly it is their business, yeah.

Alex

呃,对,他们展示了可以在自己的托管训练栈上训练模型。但不对,模型训练这块我没参与。主要原因是我在博士期间还有其他想做的事。我觉得除了 RLM 之外,还有很多大的赌注可以下。有些……我还不知道能分享多少,但总的来说,我真心觉得读博的一个奢侈之处就是有太多大的赌注可以下。大多数可能都不会有结果。但现在做研究真的是非常非常激动人心的时刻,尤其因为我觉得这个领域的大多数进展都有点无聊。我不是说结果无聊,而是做这些事情的过程往往相当无聊。所以就有这样一个问题:你知道,我们接下来想做什么?但这个我们可以稍后再聊。

Uh yeah, I mean they're showing that they can train it on their hosted training stack. But no, so for model training I'm not involved with them on that. The main reason is just I have other things in the PhD I want to work on. I think there are many other big bets to take outside of just RLMs. Some... I don't know how much I can share yet, but in general, I actually think one of the luxuries of being a PhD student genuinely is that there's so many big bets to take. Most of them will probably yield nothing. But it's a really really exciting time to be in research, especially because I think most progress in the field has been a little bit boring. I'm not saying the outcome is boring, but the process of doing these things tends to be quite boring. And so there is this question of like, you know, what do we want to do next? But we can talk about that later.

Host

好。我想再收尾一下你的研究,然后我们可以展开。自从你发布 RLM 以来,有很多兴奋。有没有什么第三方的次要工作你想推荐,说“你们应该看看这个”?

Yeah. Okay. I want to close out a little bit more of your research and then we can spread it out. Since you released RLM, a lot of excitement about it. Any secondary third party work that you want to shout out as like, you guys should take a look at this?

语言模型的未来:集群与测试框架 Third-Party Work on RLM

Alex

哦对对对。所以,Harvey,那家法律 AI 公司,发了一篇博客文章,完全不是我关联的,但他们后训练了一个 RLM 用于法律工作,这通常涉及大量筛选文档和查看各种具体信息,这些信息可能用纯检索系统不太容易检索到。他们展示了非常非常好的结果。非常令人兴奋。我很震惊他们做了这个。他们没告诉我。所以当这个出来时,我就像“哦,太棒了”。所以这个我觉得超级超级酷,他们在那里做的事情。顺便说一句,我觉得这是和 Baseten 的合作。

Oh yeah yeah yeah. So, Harvey, the legal AI company, released a blog post, not affiliated at all, but they post-trained an RLM on their legal work, which often involves a lot of sifting through documents and looking through a variety of specific information that maybe is not so easy to retrieve with a pure retrieval system. And they show really really good results. It's very exciting. I was shocked that they worked on this. They did not tell me. So when this came out I was like oh that's awesome. So there's this one I think is super super cool and what they're doing there. I think this is a collaboration with Baseten by the way as well.

Host

这是 B……

This is B...

Alex

Headlong,是 Lud 研究所的那种……它是他们那种持续运行的 harness。非常酷。

Headlong, which is law the Lud Institute's kind of... it is their like persistently running harness. It's very cool.

Host

哦,他们改名了。他们以前叫别的。

Oh they renamed it. They used to call it something else.

Alex

就像 auto terminus ter。对。我知道他们经历了……对。所以这是 Andy Kwinsk 的大项目。超级超级酷。我喜欢 Andy。我不想贬低他们做的事情,因为他们使用了 RLM 抽象,但他们做的事情比 RLM 酷多了。就像他们有一个系统,他们称之为“持续思考”。所以即使你不查询它,它也有办法思考它上下文中的问题。

It was like auto terminus ter. Yeah. I know they've gone through... Yeah. So this is Andy Kwinsk's big project. It's super super cool. I love Andy. I don't want to downplay what they're doing because they're using the RLM abstraction, but they're doing something much cooler than the RLM. Which is like they have a system that kind of what they call like thinks persistently. So even when you don't query it, it has a way to think through problems that it has in its context.

Host

哦,这就像一直开着的那种?就像一直开着的东西,但没那么贵。就像他们控制 token 成本,确保不会烧光你所有的积分。这超级酷。我想想,其实有很多,如果你去 RLM 的 GitHub 页面,下面我链接了一堆东西。人们做了很多非常非常酷的事情。Axe 是另一个很酷的,我想只是一个人做的。它像是 DSP 和 RLM 上的一个 harness。DS PI 也有一个 RLM。哦,最后我要提的是在 ARC AGI 3 上,我相信他们的 Kaggle 竞赛中有很多 harness,官方的那个,不是公开的 prime select 那个,或者人们评估的那个。它们都声称使用或引用了比如 Tufa 之类的某种形式或受启发的 RLM 抽象在他们的 harness 中,这真的非常酷。我觉得这就是……我的意思是,这正是你会看到很多好处的场景,通过组合和使用代码,以及将神经符号系统与 AI 结合,所以

Oh, is this just like always on type? It's like an always on thing, but it's like not that expensive. Like they control the token cost to make sure it's not like burning through all your credits. This is super cool. I'm trying to think there are many actually if you go to the RLM GitHub page, there's a bunch of things I've linked below. There's a ton of really really cool kind of things that people have been doing. Axe is another really cool one that I think it's just by this one guy. It's like a harness or on DSP and RLMs. DS PI also has an RLM. Oh, the last thing I'll shout out is on ARC AGI 3 I believe there were a lot of harnesses on their like Kaggle competition like the official one, not not the like public prime select one that or like what people have evaluated on. they all like claim to use or they reference like like Tufa for example some form or some inspired form of the RLM abstraction in their harness which is really really cool. I think it's um this is where I mean this is exactly the setting where you would see a lot of benefits from composition and and and using code and combining like neurosymbolic systems with with AI and so

Host

对,非常非常酷

yeah very very cool

Alex

我们喜欢好的神经符号引用。

we love a good neuro symbolic reference.

Host

对。呃,你也提到了 RKGI3,OpenAI 出来说我们在这上面达到 99.9%。呃,他们还说说我们解决了 Nevia Stokes。我们只是扔个模型上去。

Yeah. Uh you also uh mentioning RKGI3, OpenAI comes out and says we're 99.9% on this. Uh they also say we we solve Nevia Stokes. We just throw a model at it.

Alex

呃,关于它是否只是模型有一些争论。

Uh there's some debate around whether or not it's just model.

Host

嗯。

Mhm.

Alex

他们在用 RLM 吗?你知道吗?

Are they using an RLM? Do you know?

Host

我是说,我猜可能不是,除非你说……我会小心点,因为你知道人们争论什么是 RLM,什么不是。有点清楚的是,他们使用的是某种智能体集群,有共享的上下文,比如共享文件系统。这非常符合 RLM 的精神,但我觉得他们做了很多更聪明的事情,可能和 RLM 本身无关。我有点同意这个观点,即对于他们所做的,harness 并不是那么必要。我会这样表述:我认为像 GPC6 Astra 这样的模型,在给定正确信息的情况下,技术上足够聪明,能为这些非常非常困难的问题想出证明。现在,你如何获得这些信息是一个巨大的问号。在他们的案例中,可能归结为对许多这些子智能体或集群中的许多智能体进行非常长的搜索,也许还有研究人员。我其实不确定这部分,但加入他们的直觉,比如你应该探索什么等等。最终,这产生了一些信息,某个智能体能够利用它来完成证明。所以从这个意义上说,我认为 harness 那么重要吗?不。我认为这指向的是,harness 的这些具体细节并不真正重要。我认为这也是那篇 harness tax 论文所指向的,即除了用户对 harness 的感觉之外,现实中真正重要的是你如何以有意义的方式组合这些智能体来得到最终答案。也许这就是集群和所有这些事情真正关于的。所以至少从我的观点来看,如果我们开始思考用户用例,我们想从 harness 中得到什么?我们想从这些像 Claude Code、codeex 中取出好的部分,人们喜欢看到的那种流,但在底层,无论运行的是什么,都可以是一些非常奇怪复杂的智能体集群,最终得出答案。

I mean I would guess probably not unless you say like be I'll be careful here because you know people debate what is an RLM, what is not an RLM. It's somewhat clear that what they used is some kind of swarm of agents with a shared some shared context like some shared file system. And like this is very much in the spirit of RLM stuff, but I think there's a there's a lot of like, you know, more clever things that they did that's not maybe related to to the RLM itself. I agree a little bit with the idea that like a harness is not that necessary for what they did. The way that I would put this is that I think a model like a GPC6 Astra type thing is technically smart enough conditioned on the right information to come up with a proof for these very very difficult problems. Now, how you get to that information is a giant question mark. And in in their case, it probably came down to like a very long search over like many of these sub a or many of these like agents in the swarm and and maybe also like researchers. I'm I'm actually not sure about this part, but putting in like their kind of intuition as to like what you should explore and things like this. And ultimately like this produced some information that some agent was able to take to finish the proof. And so in that sense like I think was the harness that important? No. And I I think what this is pointing at is like these specific details of a harness do not really matter. And I think that's also what that uh what the harness tax paper is is pointing at which is that like beyond the user's feeling of the harness realistically all that matters is just like how are you composing these agents in a meaningful way to get to the final answer. And maybe that's what like swarms and all these things are are are really about. And so for my POV at least, if we if we start to think about like for user use cases, what do we want out of harnesses and and and things like this? Like we want to take the good parts out of these like the Claude Codes, the codeexes, like the stream that people like to see, but like under the hood, whatever is running can be some really weird complicated swarm of agents that like ultimately come up with an answer.

基础研究与应用研究 The Future of Language Models: Swarms and Harnesses

Alex

用户不想看到那个,虽然显然是对的,但这不是可读的信息。所以这又是另一种东西,类似于 RLM 的精神,即递归语言模型。它听起来像是一个语言模型,但它不是语言模型架构。但这样做的原因是,我认为我们未来会开始看到,可能有一天,我写了一篇博客,我们所认为的语言模型,即我们查询的东西,实际上可能像一个群体,或者像一个脚手架,或者像某种奇怪的、能很好地扩展的框架设计,但用户就是看不到它。最终,用户只看到这个框架的前端版本。是的,我认为至少可以相对安全地打赌,这就是我们将看到的。

The user doesn't want to see that, though obviously right, like it's not legible information. And so this was another kind of thing in the spirit of RLM's, like recursive language model. It sounds like it's a language model, but it's not a language model architecture. But the reason for this is like I think we will start to see in the future, probably one day, and I wrote a blog about this, what we think of as a language model, like the thing that we query, might actually be like a swarm or like a scaffold or like some weird harness design that scales very well, but the user just doesn't see it. Ultimately, all the user sees is some front end version of this harness. And yeah, I think it's a relatively safe bet at least to make that this is what we will see.

Host

这可能要回到基础 Transformer 的局限性。就像,你知道,显然如果你只是拿一个基础 Transformer,然后说解决纳维-斯托克斯方程之类的,它不会做到。就像,是的,我们都知道这不会发生。但是的,我认为这可能是更有趣的部分,也许关于硬度是否重要的说法更多是关于这一点,即只是任意地将模型或智能体指向不断增长的信息上下文,也许就足以解决非常困难的问题,嗯,这个我可以接受。

And this maybe goes back to the limitations of the base transformer. Like, you know, obviously if you just took a base transformer and you said like solve Navier-Stokes or something, it's not going to do it. Like, yeah, we all know this is not what's going to happen. But yeah, I think this is maybe the more interesting part, and maybe the claims around like, you know, did the hardness matter is more around this of like just arbitrarily pointing models like or agents at a growing kind of context of information maybe is just enough to solve very difficult problems, um, that I can buy.

Alex

虽然里面有很多内容,但我想说 OpenAI 花费的金额是半公开的。是 88 小时内 10,000 个智能体,呃,1300 亿输出 token,根据公开定价估计约为 4000 万美元。

While there's a lot in there, I do want to say the amount that OpenAI spent is semi-public. It's 10,000 agents in 88 hours, uh, 130 billion output tokens, which is estimated to be about $40 million in public pricing.

Host

实际上令人惊讶的是,比我想象的要少。

Surprisingly actually like less than I thought.

Alex

是的,没那么多。

Yeah, not that much.

Host

最终输出是 1300 亿,但如你所说,有大量上下文在传递。仅发送的智能体消息总数就超过这个的两倍。

130 billion output for the final, but as you said, there was a lot of context being passed around. It's more than double that in just the total agent messages being sent.

Alex

是的。是的。

Yeah. Yeah.

Host

我认为我有幸提到的另一件事是 Cursor,就群体之类的东西而言。呃,这稍微老一点,呃,意思是二月份,这已经是古老的了。呃,但如果你一直向下滚动到他们最终达到的那种多智能体架构,基本上就是一个正常软件团队的组织结构图。我在想的一件事是,因为我基本上有这样一个,嗯,呃,收集所有函数,你必须用子智能体来做,或者,你知道,这实际上非常类似于 GPU 编程。

I think one thing I was honored to bring up also was Cursor as far as like swarm stuff is concerned. Uh, this is slightly older, uh, meaning February, which is ancient. Uh, but if you scroll down all the way to the final sort of multi-agent architecture that they arrived at, it was basically an org chart of a normal software team. One thing I'm thinking about because I basically there's like this, um, uh, gather all function that you have to do with with sub-agents or, you know, it's very similar to GPU programming actually.

Alex

是的。是的。是的。

Yeah. Yeah. Yeah.

Host

呃,所以呃,那是一个瓶颈。

Uh, and and so uh that's a bottleneck.

Alex

这是一个瓶颈。有一个主要的人,呃,他在协调所有子人员,然后他们必须重新,你知道,再次聚集,然后然后重新协调。

This is a bottleneck. There's one main guy, uh, it's coordinating all the sub guys, then they have to like re, you know, gather again and then and then re-coordinate.

Host

那很慢,那很糟糕。呃,真正的群体应该是每个人都是独立的个体。

That's slow, that's that's crappy. Uh, what a true swarm should be is everyone is just their own person.

Alex

是的,是的。

Yeah, yeah.

Host

嗯,我同意这一点,但我认为有一个问题需要讨论。让我打个比方,就像你什么时候会使用压缩,什么时候会使用 RLM?在许多情况下,RLM 可以解决比压缩更困难的任务,但在大多数情况下,你更愿意使用压缩,因为它更便宜、更快捷。我认为在智能体群体的背景下,也有类似的情况,比如我相当确定,95% 的群体完全无用,或者它探索的完全是,你只是在燃烧 token,而在这种设置中可能没那么多。我不确定。我的意思是,也许这里也是如此。

Well, I agree with this, and I think that there is a question to be had though. Let me give an analogy, which is like when would you use compaction and when would you use an RLM? And there are many settings where an RLM can solve maybe a more difficult task than compaction can, but you would prefer to use compaction in most cases because it's cheaper and it's quicker. And I think in the context of agent swarms, there is a similar thing going on of like, I'm fairly certain that like 95% of the swarm is entirely useless or like what it's exploring is entirely, you're just burning tokens, versus in this setup maybe not so much. I'm not sure. I mean, maybe it's also the case here.

Alex

每个人都有自己的工作。有一个 Jira 看板。我认为在某种程度上,问题就是这样被框定的,对吧?所以就像如果你有一个搜索问题,你展开一堆子智能体来做搜索,会有很多无用的信息,对吧?你得到一个检索答案,然后你展开。但这是有意识地被理解的,对吧?

Everyone has a job. There's a Jira board. I think at some level that's how problems are framed, right? So like if you have a search problem and you're spanning out a bunch of sub-agents to do search, there's going to be a lot of useless information, right? There's one retrieval answer that you're getting and you're spanning off. But that is consciously understood, right?

Host

是的。但有一个问题,比如什么适合解决哪些问题,比如什么设计,你知道,理论上 OpenAI 可以使用,可以打包这个 API,他们会称之为群体,然后他们会给你,他们会说把这个指向任何问题,我们会给你一个解决方案。但也许

Yes. But there is kind of this question of like what is appropriate to solve for which problems, like what design, you know, in theory OpenAI can use, can package up this API and they'll call it swarms and then they'll give it to you and they'll be like point this at any problem and we'll give you a solution. But maybe

Alex

它叫做 pro。

It's called pro.

Host

是的。也许你得支付 4000 万美元才能得到一个结果,这就像,嗯,但这很令人兴奋。我会说,我们甚至有能力将 4000 万美元指向一个问题并解决它,这非常令人兴奋。

Yeah. Maybe you'll have to pay like $40 million to get a result and it's like, well, but it's exciting. I will say like it is it's very exciting that we even have the option to point $40 million at a problem and solve it.

Alex

但是,在这个领域仍然有很多研究要做,比如你知道什么是必要的,比如我们想做什么?我们想要什么设计?我的意思是,我们可能不希望一切都是群体,但比如我们在哪里划清界限,比如智能体可以设计那个或自己决定那个等等。是的。

But there is there is still kind of uh a lot of research to be done in this area around like you know what is necessary like what what do we want to do? What design do we want? I mean we probably don't want everything to be a swarm but like where do we draw the line like can the agent design that or decide that for itself etc etc. Yeah.

Host

呃,然后,呃,只是快速检查一下,以防你对此有看法。你有没有深入研究过开放性作为一般类别的问题?

Uh, and then, uh, just a quick check in case you have opinions on this. Have you looked much into open-endedness as a general category of problems?

Alex

意思是没有提示,只是去。

Meaning no prompts, just go.

Host

一点点。嗯,所以我在 Sakana 度过了一个夏天,呃,就在我博士之前或之后,我猜是之前,那是他们经常研究的东西。我认为即使在像递归超级智能这样的地方也有很多人。现在有很多这样的人。

A little bit. Um, so I was at Sakana for a summer, uh, right after or I guess right before my PhD, and and that's something that they work on a lot there. And I think there's a lot of people even at like recursive super intel. There's many of them now.

Alex

你说 Richon。

You said Richon.

Host

哦,是的。是的。所以,所以就像 Richard 的公司,然后然后然后实际上等等,那可能是,可能是同一家公司。我不记得了。Tim Rockell 也是吗?好的。

Oh yes. Yeah. So, so like Richard's company and then and then and then also actually wait that might be it might be the same company. I don't remember. Is is Tim Rockell also? Okay.

Alex

他是主要的联合创始人。他是谷歌开放性研究的负责人。

He's the main co-founder. He's the head of open-endedness for Google.

Host

是的。

Yeah.

Alex

所以,我认为对于开放性问题,我把它们看作有点类似于未解决的数学问题。也许这是一种奇怪的表述方式,但我认为很多技术,比如人们如何接近它们,是相似的。比如进化搜索非常类似于只是启动一群智能体,并希望它们能提出一些有趣的,我的意思是,这就是 Alpha Evolve 和一些其他工作在一年或两年前所做的。但我认为可能不是,我不确定这是否是你所暗示的,但我认为在开放性风格的搜索问题中,不太清楚的是,我们是否将它们全部框定为基本上是一个未解决的非常困难的问题,或者目标明确的问题。如果不是这样,我还不完全知道它的价值是什么。也许也许你有其他看法,我我我对这个没有太多看法,但至少从我在 Sakana 的时候,我我感觉到我们最终仍然希望以比如 OpenAI 处理纳维-斯托克斯的方式来处理事情。我们想要,仍然有很多推动在某些方向上,我们需要达到我们有有趣东西的地步。

So, I think with open-endedness problems, like I view them as somewhat similar to even like unsolved math problems. Maybe that's a weird way of framing it, but like I think a lot of the techniques in terms of like how people like approach them are kind of the same. Like evolutionary search is like very similar to just launching swarms of agents and hoping that like they come up with like an interesting, I mean this is what like Alpha Evolve and some of these other works did, uh, like a year or two ago. But I think what maybe is not, and and I'm not sure if this is what you were alluding to, but I think what's not super clear in open-endedness style search problems is do we frame all of them as basically like an unsolved very difficult problem or like where the objective is clear. If that's not the case, I still don't know yet entirely what the value of it is. Maybe maybe you have other opinions, like I I I don't have too many opinions on this, but at least from my time at Sakcon, like I I got the sense that like we ultimately still kind of want to approach things the way that like say OpenAI approached Navier or Stokes. We want we there's still a lot of nudging in certain directions that we want to have to like get to the point where we have something interesting.

Kimmy 与中国实验室 Basic vs Applied Research

Alex

我的版本是,也许可以分成基础科学和应用科学。基础科学:你为了研究而研究,只是想更好地理解事物。我完全不知道会不会有任何应用。而应用科学:你有一个目标,你试图以某种方式最小化损失。

My version of it is maybe it's a split between basic science and applied science. Basic science: you're researching for research's sake, you just want to understand things better. I have no idea if there will be any application at all whatsoever. And then applied: you have a goal, you're trying to minimize loss in some way.

Alex

所以,就人们下的那些大赌注而言,我真正思考的是——如果没有提示词会怎样?

And so what I really think in terms of the big bets that people have — what if there was no prompt?

Host

我明白了,你就直接挑。

I see, you just pick.

Alex

就像你在这纷繁复杂的东西里,你说,嘿,大家好,我喜欢你们正在做的东西,然后你就决定,我觉得这是个有趣的问题。最大的问题——也许有清晰的路径,但至少我在那里的时候——人们对开放式研究最大的问题是:你最终怎么选择?你怎么挑出有趣的东西?因为当没有……也许智能体会提出一个目标,但很多情况下……他们有个叫 Fugu 的东西,我想,它像是一个模型路由器,至少是受这个想法启发:让我们选一个问题,也许我们能挑出某个东西的最佳解决方案。在这个例子里,就是为这个问题挑选最佳模型。

Like you just in this swarm of things and you're like, hey, what's up guys? I like what you guys are working on, and you just decide, I see this is an interesting problem. The biggest issue — and maybe there's a clear path to this, but when I was there at least — the biggest issue that people had with open-endedness was: how do you ultimately choose at the end? How do you pick out the interesting stuff? Because when there is no... maybe the agent comes up with a goal, but in a lot of cases... and they have something called Fugu, I think, which is like a model router type thing that was inspired at least by this idea of: let's pick a problem where maybe we can pick out the best solutions to something. In this case it's like pick the best model for this problem.

Host

你是第一个把模型路由和开放式研究联系起来的人。

You're the first person to connect model routing to open-endedness.

Alex

不,不。是的。但我提起这个是因为我认为对于开放式研究,通常的问题是:当我们有这巨大的垃圾语料库时,我们如何筛选并找到隐藏的宝石?而开放式研究的解决方案真的就是让模型永远运行下去,做数据生成——超级、超级、超级高吞吐量的数据生成。像有意义的……也许那很酷。

No, no. Yeah. But so I bring this up because I think with open-endedness, generally the issue is: when we have this giant corpus of slop, how do we sift through and find the hidden gems? And the solution to open-endedness really is just letting models run forever and doing data gen — super, super, super high throughput data generation. Like a meaningful... maybe that's cool.

Host

对我来说,这非常有趣,作为对基本上所有机器学习的反例——你有目标,而这里没有目标。

To me it's very interesting as a counter to basically all of machine learning, where you have a goal, to have no goal.

Alex

是的。

Yeah.

Host

但或者像一个定义不清的目标,你说,那这个目标怎么样?你说,好吧,也许。然后你再多研究一下,发现那是个有趣的目标,因为我认为找到你的目标函数——就像你说的,Jeb 找到了一个有趣的目标函数,没人在探索。

But or like an ill-defined goal that you're like, well, what about this goal? And you're like, well, okay, maybe. And then you sort of research more and you find that that is an interesting goal, because I think finding your objective function — like you said, Jeb found an objective function that was interesting that no one was exploring.

Alex

是的。

Yeah.

Host

我认为这也类似于你关于研究生的观点:你留在学校是因为你想追求开放式研究。如果你想利润最大化并加入——逃离永久底层——那你就加入实验室,对吧?

I think that is similar to your message about grad students as well: you stay in school because you want to pursue open-endedness. If you want to profit max and join — escape the permanent underclass — then you join a lab, right?

Alex

是的。有趣的是,我觉得我不常听到这种讨论。我的意思是,我在东海岸,所以是非常不同的类型……但每当我来这里,这总是讨论的话题。

Yeah. It's funny. I feel like I don't hear this discourse a lot. I mean, I'm in the East Coast, so it's a very different type of... But then whenever I come here, it's always the topic of discussion.

Host

你不做这个就付不起房租。

You cannot pay rent without doing this.

Alex

你们正在被高价挤出,伙计们。

You're getting priced out, guys.

Host

好吧。所以,是的,有所有这些。我不知道你是否想——如果足够相关的话,谈谈你刚才提到的被管理不善的天才们。

Okay. So, so yeah, there's all that. I don't know if you want to — if it's relevant enough to talk about the mismanaged geniuses, which you were pulling up.

Alex

不,不,这只是你的博客之一,但我会继续追问。

No, no, it's just one of your blogs, but I will poke on.

Host

是的。是的。所以他们在博客文章中——我也和他们谈过这个。所以 Ultra 的一个很酷的结果是,他们基本上就让它自由进行自动化数据科学研究,几乎没有人干预。它只是在做……

Yeah. Yeah. So they did in their blog post — I talked to them about this as well. So one of the cool results of Ultra is they basically just let it loose on automated data science research with little to no human intervention. It's just kind of making...

Alex

是的。所以这是自动研究,它更开放一些,而且开放式研究有程度之分,我同意这一点。

Yeah. So this is auto research, which is a little bit more open-ended, and there's degrees of open-endedness and I agree with that.

Host

是的。和有目标的自动研究不同——这只是做事情。但你知道,这很酷。他们正在研究这个,给感兴趣的人。既然你提到了,实际上,你对 Sakana 有什么看法?除了是日本的那个,他们在做什么?

Yeah. Separate than auto research with objective — this is just do stuff. But you know it's cool. They're working on it for those that are interested. While you're bringing it up, actually, what is your take on Sakana? Like what are they doing apart from being, you know, the Japan one?

Alex

是的。是的。是的。我的意思是,我其实很喜欢那里的人。我认为他们有一个非常非常聪明的团队。这很合理。我的意思是,它从 GDM 的一个早期团队分出来的,那个团队也在做这种开放式进化研究风格的东西。至少我在那里的经历中我喜欢的是,他们确实有过那次失误——我忘了具体什么时候——但我想……

Yeah. Yeah. Yeah. I mean, I actually love the people there. I think they have a really, really smart team. And it makes sense. I mean, it branched off from an earlier team at GDM which was also kind of doing this kind of open-ended evolutionary research style stuff. What I liked about my experience there at least was that they did have that mishap back — I forget at this point when — but I think...

Host

关于科学家。

About scientists.

Alex

不,是那个上传的 GPU 内核。

No, the GPU kernel that went up.

Host

是的。人们无法原谅他们那件事。

Yeah. People cannot forgive them for that.

Alex

是的。是的。是的。我猜就像 AI 科学家——我的意思是,有一些批评,我不做那个,所以我没有看法。但总的来说,至少我喜欢他们的一点是,他们更像一个研究型实验室。所以他们不像 OpenAI 或 Anthropic 那样在同一个领域运作,肯定的。绝对不是,他们不是。我的意思是,很明显,至少我在那里的时候,他们没有一个大竞争模型或大家都在用的东西。但我认为他们在某些方面像一个博士实验室,这很酷。我认为 David Ha 非常非常聪明。我觉得他有很好的感觉……另外,我认为日本的 AI 市场也有点不同,他们瞄准的对象可能和我们这里习惯的略有不同。但是的,我喜欢他们——很多人看他们的研究时会觉得有点怪,我喜欢这一点。我认为这……

Yes. Yeah. Yeah. And I guess like the AI scientists — I mean, there's some criticisms of it that I don't work on that so I have no take on it. But I think in general what I like about them at least is they're a little bit more of a researchy type lab. So they don't operate in the same space as OpenAI or Anthropic, for sure. Definitely no, they do not. I mean, it's pretty obvious probably that at least when I was there they do not have a big competitor model or something that everyone is using. But I think they kind of operate in some ways as a PhD lab, which is cool. And I think David Ha is really, really smart. I think he has a good sense of... Also, I think the market in Japan is also a little bit different for AI, and who they're targeting is slightly different than maybe what we're used to here. But yeah, I like that they take — a lot of their research is kind of weird, I think, when people view it, and I like that. I think it's...

Host

你应该有更多的怪。是的。

You should have more weirdness. Yes.

Alex

正是。正是。

Exactly. Exactly.

Host

你说不同的市场——只是像企业吗?

And you say different market — just is it like enterprise?

Alex

就像那里的运作方式有点不同,比如交易如何发生之类的。他们确实有一个——我的意思是,我猜这个页面最初是日文的——但他们确实有一个专门针对日语的模型。

Like the way it works there is a bit different, like how deals happen and stuff like that. They do have a — I mean, I guess this page is originally in Japanese — but they do have a model specialized for the Japanese.

Host

你不知道吗?

You didn't know that?

Alex

我不知道你知道。我也有私人朋友,认识团队。所以有一个……

I didn't know that you did. I also have personal friends and know the team. So there is a...

Host

甚至从你说话的文化方式来看,你知道,回答是朝着那个方向调整的。这不像……

Even from the sense of a way that you speak culturally, you know, responses are tuned towards that. This is not like...

Alex

它在基准测试上是前沿的。它是一个适合他们文化的模型,他们有聊天之类的。但呼应你的观点,你知道,还有像教育应该是什么样子,有人想研究它,他们非常博士实验室风格:做你的事,为什么不呢?我们有钱,去研究吧。

It's frontier on benchmarks. It is a culturally appropriate model for them and they have like chat and all that. But to mirror your point, you know, there's also like how should education look like and someone wants to work on it and they're very PhD lab of do your thing, why not? We have money, go research.

Host

哦,他们说这也是 Kimi 的微调。那很好。

Oh, they say it's a Kimi fine-tune, too. That's nice.

Alex

哦,那就对了。对他们好。

Oh, there you go. Good for them.

Gemini 与 Google DeepMind Kimmy and Chinese Labs

Host

说到 Kimmy,对吧,又是一个从学校出来单干的博士生,我手上正好有个 Kimmy Delta 注意力机制想研究,结果人家不知怎么就一发命中了。

Speaking of Kimmy, right, like another grad student that split out, and I just have this Kimmy Delta attention that I want to work on, and somehow managed to make one shot.

Alex

到现在我还是没搞懂。

Don't understand it still.

Host

是啊是啊。反正他强得离谱,至少在我看来是这样。我觉得总体上说,很多中国实验室都做出了非常非常酷的工作。

Yeah. Yeah. Well, he's super cracked. At least my understanding. I think in general a lot of the Chinese labs have done really really cool work.

Host

你对 Kimmy 的智能体集群有什么看法?

Any thoughts on Kimmy agent swarms?

Alex

有。我要说的一点是,OpenAI 在智能体集群上做的,显然是对的方向。你得这么想:没有什么东西是轻易白来的,尤其是在没有一个非常非常聪明的 harness 设计的情况下——而我觉得目前还没有人真正做出来。比如说,智能体集群的设计不是理所当然就能有的。不是说 GPT6 Astra 特别聪明,然后智能体集群就自然跑得好了。他们显然是训练过的——我是说 Hugging Face 那件事,就是他们在训练一个系统去像集群一样运作。我觉得他们显然做对了一些事,到了可以砸 4000 万美元去解决一个未解难题的程度。而 Kimmy 的智能体集群那件事,至少我读下来,感觉就是:这挺有意思,但我并不确定它能不能解决任何新的问题。

Yes. One thing I will say is whatever OpenAI is doing with their agent swarm is clearly the right thing to do. You have to kind of think about it this way, like nothing, especially without a very very smart harness design, which I don't think anyone really has so far, nothing comes so easily for free. For example, the agent swarm design is not something you can just take for granted. It's not like GPT6 Astra is just super smart and then it just got agent swarms running well. They clearly trained — I mean the Hugging Face incident was them training a system to be like a swarm. And I think clearly they've done something really well to the point where you can throw $40 million and solve an unsolved problem. And I think with the Kimmy agent swarm thing, at least from when I read it, it just came off as like this is interesting but I don't actually know whether or not this can solve anything novel.

Host

是啊。他们基本上就是说它能做电子表格。

Yeah. They just were like it makes spreadsheets.

Alex

他们差不多就是:对,这里有个集群,能做点事,挺酷的。但动态工作流我也要说同样的话。我其实觉得动态工作流那次发布有点失败。不知道你们怎么看,但据我了解它用得并不多,或者说——

They kind of were just like, yeah, here's a swarm that kind of does stuff and it's cool. But I will say the same thing with dynamic workflows. I actually kind of think the dynamic workflows release was sort of a flop. I don't know how you guys feel about it, but my understanding is it's not used that often, or like —

Host

就是太贵了。

It's just very expensive.

Alex

太贵了。

It's too expensive.

Host

就是 ultra code。基本上就像“接管我的银行”那种。它不干活。我试过,它不按——再说一遍,我不是 OpenAI 最大的拥趸什么的,但我觉得他们做的那些非常厉害。他们不知怎么让这个集群真的朝着一个目标去运作,这确实非常难。

It's ultra code. It's basically like take over my bank. Doesn't act. I've tried it and it doesn't act in the way that — again, I'm not the biggest OpenAI stand or something, but I think whatever they did was very impressive. They somehow managed to get a way for this swarm to actually work towards a goal, and yeah, it's very difficult.

Host

我明白了。所以多智能体集群的效率就是这里的目标函数,对吧?这里面有多少是垃圾产出,感觉垃圾产出很多。

I see. So like efficiency of the multi-agent swarm is the objective function here, right? Like how much of this is slop, like this is a lot of slop.

Alex

对,我觉得是这样。我觉得我们把“集群收敛到一个答案”这件事当成理所当然了。这恰恰不是理所当然的。

Yeah, I think so. I think we take for granted what it means for a swarm to converge to an answer. It's just not something we can take for granted.

Host

是啊是啊。我们和 Noam Brown 做过一期播客,他的核心观点就是:好,我们研究了很多竞争型智能体,现在在研究协作型智能体,而这现在就叫集群。

Yeah. Yeah. We've done one pod with Noam Brown, and his whole thing was like, okay, we've worked on a lot of competitive agents. We're working on collaborative agents, and like that's now called a swarm.

推测性 PTC 与被埋没的天才 Gemini and Google DeepMind

Host

我想快速问一句。你对 Gemini 有什么看法?他们也拿了 IMO 金牌。有一段时间他们让智能体长时间推理,而且——

I would do a quick poke. Do you have any thoughts on Gemini? They were also IMO gold. There was a time where they were getting agents to reason for a long time, and —

Alex

嗯,我是说——

Yeah, I mean —

Host

是不是太旧了,不值得再谈?

Is it too old to think about?

Alex

不不不。我觉得有点被夸大了,就像 Gemini——再说一遍,我得先声明我没在这些地方工作过,所以这话你打个折扣听,对吧?

No, no, no. I think it's a little bit blown out of proportion, like Gemini — again, so I should preface by saying I haven't worked at any of these places, so take this with a grain of salt, right?

Host

我是说,你对智能体 harness 有很强的看法,而且你知道,这次是——

I mean, you have strong opinions on agent harnesses, and you know, this was —

Alex

但让我先说这个。这项工作非常厉害。我觉得他们在这里展示的是:他们在一个模型还不太行的时候,非常非常聪明地设计了 harness 的作用。我记得这个,至少像 AlphaGeometry,我猜那是这之前一年的事。但那非常酷。他们把它做到了极致,然后设计了——我也说不好。我对竞赛数学不太感冒。我觉得 GDM 有点可惜。我聊过的每个人对 GDM 的看法都差不多,就是它太官僚了。不管发生了什么,他们有人才和资源去做几乎任何事,但,我不知道,在他们把那一块搞定之前——比如我对 Antigravity 没意见,但我不认识任何用 Antigravity 的人。我试过一次,没看出有什么理由要换过去。我觉得不管出于什么原因,他们一直在这方面挣扎。

But let me just say this first. This work was really impressive. I think what they showed here was they took a time when the models weren't that good. Yes. And they managed to be very very smart about what the harness does. I remember for this, at least like, yeah, like AlphaGeometry. I guess that was a year before this. But it was very cool. They took it to the max and they designed — I don't know. I'm not too big on competitive math. I think like GDM, it's sort of a shame. Like everyone I've talked to about GDM kind of has the same opinion, which is that it's way too bureaucratic. Whatever is going on, they have the talent and the resources to do almost anything, but like, I don't know, until they figure that part out — like nothing against Antigravity for example, but I don't know anybody that uses Antigravity. I've tried it once and I don't see a reason to switch to it. And I think for whatever reason they've been struggling with this.

Host

是啊。很多人长期吐槽 Meta,直到他们开始拿出东西来,我觉得 Google 现在正处在那个阶段。你知道,你得活下去,而且——

Yeah. Well, a lot of people dogged on Meta for a long time until they started coming out, and I think Google's going through that phase right now. And you know, you got to stay alive and —

更好的测试框架与语言选择 Speculative PTC and Mismanaged Geniuses

Host

我想把话题拉回到你总体的看法上。我们可以聊投机式 PTC、被管理不善的天才,或者干脆把这些都扔掉,聊点别的什么都行。

I want to focus back on just your thoughts in general. We can talk about speculative PTC, mismanaged geniuses, or just throw away all this and talk about whatever else.

Alex

好,那我们就聊一会儿被管理不善的天才。关于投机式 PTC,我唯一要说的就是它是个非常简单的想法。几乎显而易见就该这么做,没什么更多可聊的。我觉得你就该把它用在编程上,比如任何涉及程序化智能体调用的东西,像 RLM 或 CodeAct。对,这就是不用想的事。

Okay, let's talk about mismanaged genius for a little bit. The only comment I'll say on speculative PTC is that it's a really simple idea. It's almost like obvious that this should be done, and there's not much more to talk about it. I think it's just like you should just use it for coding, like anything with programmatic agent calling, like RLMs or CodeAct. Yeah, it's like a no-brainer.

Host

那——你知道,对于没读过的人来说,一句话概括是什么?

What's the — you know, for people that haven't read it, what's the one-liner?

Alex

简单说就是,当模型在写代码时,甚至写完之后,很多工具往往是串行的,或者说你得等它们。所以你就应该提前启动它们。比如如果你能编译这段代码,你大概就能推断出——尽管涉及变量之类的,你也能推断出——

The simple thing is when the model is writing its code, or even after it finishes writing the code, a lot of tools tend to be sequential, or like you have to wait on them. So you should just launch them in advance. Like if you're able to compile this code, you can probably figure out — even though it's like kind of variables and stuff, you can figure out like —

Host

静态分析式的投机。

Statically analyze speculation.

Alex

对。

Yeah.

参差不齐的智能与技能问题 Better Harnesses and Language Choice

Alex

有人向我指出,学术界尤其是编程语言领域的人,有一些非常酷的做法。所以也许某个时候我会去做这件事。

Someone pointed out to me that academics, especially in programming languages, have some very cool ways of doing this. So at some point maybe I'll work on this.

Host

主要是你得换语言,因为如果你用 JavaScript 或 Python,你就做不到这个。

Mostly you have to change language, because if you're in JavaScript or Python you can't do this.

Alex

对,是的,是的。

Yes, yeah, yeah.

Host

所以像 Haskell,对。那通常用的、不是那个的是——

So like Haskell, yes. What's the normal one that's not—

Alex

Lisp。

Lisp.

Host

Lisp,或者 OCaml。

Lisp, or OCaml.

Alex

哦,OCaml。任何函数式语言,你其实都能把这条流水线串起来。所以如果你想用 TypeScript,就用 Effect TS。

Oh, OCaml. Any functional language, you can actually pipeline this. So Effect TS if you want to do TypeScript.

Host

好,我们可以转到 diagram。

Okay, we can switch over to diagram.

语言与推理 Jagged Intelligence and Skill Issue

Host

这就像是你们整套的核心论点,对吧?对我来说这有点像换个说法,但也许我漏掉了什么,就是说:去研究更好的 harness,或者你的模型其实能做得更多,只要你更努力。所以这是个技能问题。

Which is like you guys' whole thesis, right? To me this is kind of a restatement, but maybe I'm missing something, of like, well, work on better harnesses, or your models actually are capable of a lot more if you try harder. So this is a skill issue.

Alex

对。对。对。基本上是这样。我是说,有一件事我想看到。我很欣赏大家对“锯齿状智能”的高度关注,因为它描绘了一幅大图景:只要我们下定决心,就能做到。但我有点希望——也许学术界有人该做这件事——真正坐下来想一想:如果我把 Astra 拿来,即便是当前的前沿模型,也不足以在一个月的时间跨度里持续、稳定地做好某一项工作。而我觉得这是个愚蠢的问题。我真的认为我们能解决它。你不需要是一家前沿实验室,也不需要为你的 IPO 搞那些花哨的东西。我认为这些模型太聪明了,即便用很笨的办法,这真的就是个技能问题——你能让一个模型做到,比如说,某个 18 岁高中生做某项工作的水平。我觉得我们连这都做不到,简直荒谬。部分原因是语言模型的形式本身不太适合这件事。但我认为你可以围绕它塑造一个 harness 来做到。而且我认为这件事本身就与 RLM 那套东西无关,与“我们该用什么抽象”那套东西无关。我就是觉得有人能设计出一个能做到这件事的 harness。

Yeah. Yeah. Yeah. Basically. I mean, I think there's one thing I want to see. I appreciate that there's a big focus on jagged intelligence because it paints a big picture of like we can do this if we really set our minds on it. But I kind of wish—and maybe someone in academia should do this—like really just sit down and think about, if I took Astra, even the current frontier models are not good enough at doing a particular job over, let's say, the span of a month consistently and well. And I think this is a stupid problem. I genuinely think we can solve this. You don't need to be a frontier lab and do all this fancy stuff for your IPO. I think these models are so smart that even if it's a silly way, it genuinely is a skill issue—you can get a model to be as good as, let's say, just some 18-year-old high school kid doing some job. I think it's ridiculous that we can't do that. And part of the reason is the format of a language model is not really amenable to that. But I think you can shape a harness around it and do it. And I think this in itself is ignoring the RLM stuff, ignoring all the what-abstractions-should-we-use stuff. I just think someone can design a harness that can do this.

Host

做到这件事——做到什么?

Do this—do what?

Alex

做长时间运行但简单的任务,并且可靠地完成。

Do long-running but simple tasks, and do them reliably.

Host

那是个例子。所以这跟——比如说你最喜欢的公司,Harvey,用 LLM 做法律工作——有什么不同吗?还是说——

That's an example. So is this different than, you know, pick your favorite company, Harvey, for example, using LLMs to do legal work? Or what's the—

Alex

我是说,我猜有点像,只不过瓶颈不是某些法律知识之类的东西。我不知道。比如说——

I mean, I guess it's kind of like that, except if the bottleneck was not certain legal knowledge or something. I don't know. Let's say—

Host

我是说,我猜就像那些例子,人们可以拿模型来搭流水线之类的,让一个智能体反复做他们想要的任何任务,对吧?

I mean, I guess like the examples, you know, people can take models and build pipelines or whatever and have an agent repeatedly do whatever tasks they want, right?

Alex

对。对。所以比如说,如果我想要一个通用系统,我可以像跟实习生说话那样跟它说话,基本上就是让它去探索某个小东西。也许一个例子就是很傻的自动研究——也许这是个例子,不一定是找到超级新颖的解法,但至少把任何问题里所有容易的部分都优化掉。它们往往过度偏向 ML 训练之类的东西。但,是的,我不知道。也许这说得不太清楚,但在很多人的一般工作流里,你大概可以直接 vibe code 出一个特定的 harness 来帮你自动化这件事。有些例子是自动化找研究论文之类的。但通常人们会设计一个专门的智能体来帮他们做这类事,或者他们 vibe code 一个 harness 直接跑,或者他们的 Slackbot 之类的。但我几乎觉得存在一种标准形式——就是一个你直接插进去的 harness。你不需要——它不需要为找论文而设计。你只要告诉它“帮我找这个”,然后把它插到那个场景里。我想说的是,我认为有很多容易的事情可以被自动化,而且——

Yeah. Yeah. So for example, if I wanted a general system that I could talk to like I would talk to an intern, and basically just ask it to explore some small thing. So maybe an example of this is very silly auto research—maybe an example of this, of not necessarily finding super novel solutions but at least optimizing all of the easy parts of any problem. They often end up being overindexed for ML training and things like that. But yeah, I don't know. Maybe that's not super clear, but there are a lot of people's general workflows where you probably could just vibe code up some specific harness to help you automate this thing. Some examples are automating finding research papers and stuff like that. But usually people will design a specialized agent to help them do this kind of thing, or they'll vibe code a harness and just run it, or their Slackbot or something. But I almost think there's just a standard form—just a harness that you just plug in. You don't—it doesn't need to be designed for finding papers. You just kind of tell it find this for me and you plug it into that setting. What I'm getting at is that I think there's a lot of easy things that can be automated and—

Host

这算是一个假设,还是关于能力过剩的观点——就是说即使我们暂停了,当前状态的模型仍然能带来很大影响?

Is this like a hypothesis or a point around capability overhang—like even if we paused, there's still a lot of impact to be had with the current state of models?

Alex

某种意义上,是的。我猜我呈现的是这件事最容易的形式,但这大致是在说:我们在很多事情上都有锯齿状智能。比如模型在编程和数学上强得不成比例。这说的是我们可以把那些能力迁移到很多其他事情上。比如说,通常如果你把一个 IMO 金牌得主之类的人放到很多不同的解题领域,他们都能搞明白。我其实不知道这对模型是否成立。比如在 GPU 代码优化里,一个非常有趣的问题是:如果你把模型里所有 GPU 编程数据都拿掉,但它和现在的 Astra 一样强,只是没有那些数据,它还能优化 GPU kernel 吗?它能在上下文里大致学到它需要学的东西,然后有某种流水线,或者想出某种解优化任务的方案吗?我认为存在一种错配——如果你拿一个和 Astra 一样聪明、或知道得一样多的人,那个人能做的事和 Astra 能做的事之间存在错配,也许就差一个 harness。而我认为我们其实可以更接近人类。

In some sense, yes. I guess what I'm presenting is the easiest form of this, but what this is kind of saying is that we have jagged intelligence on a lot of things. For example, models are disproportionately good at coding and math. This is saying we can translate those abilities to many other things. So for example, usually if you take someone who was an IMO gold medalist or something and you apply them to a lot of different problem-solving domains, they can figure it out. I don't actually know if this applies to models. For example, in GPU code optimization, one very interesting question is: if you were to take out all of the GPU programming data from a model, but it was as good as Astra is now just without that taken out, would it be able to still optimize GPU kernels? Would it be able to learn in context roughly what it needs to learn and then have some pipeline or come up with some solution to solving optimization tasks? And I think there's a mismatch between—if you took a human that was as smart or knew as much as Astra, there is a mismatch between what that human can do and what Astra can do, maybe around a harness. And I think we can actually approximate the human a lot more.

Host

对我来说,这听起来非常接近持续学习问题。我觉得你是在说——

To me it sounds very approximate to the continual learning problem. I think you're making—

Alex

那是最好的例子。是的。

That's the best example. Yes.

Host

对。你为什么不直接这么说?

Yes. Why didn't you just say that?

Alex

对。对。我猜我——

Yeah. Yeah. I guess I—

Host

我当时在想,我能不能直接蹦出几个词?我对这个很谨慎。但是的,我——

I was like thinking, could I just blurt out some words? I'm careful with that. But yes, I—

Alex

你有一种非常——也许我——

You have a very like—maybe I—

Host

也许有点像——我不知道。所以我觉得你是一个非常——我不知道。其实我知道你的本科。你是数学背景的吗?对。我记得你学的是数学。

Maybe like a bit of a—I don't know. So I think you're a very—I don't know. I know your undergrad actually. Are you like a math person? Yeah. I think you studied math.

Alex

对,至少那是我当时想做的。像范畴论那种抽象,你用范畴来思考,然后你得往下翻译到具体的东西,但其实你更在意的还是那个范畴。

Yeah, that's what I wanted to do at least. Like a category theory type of abstraction where you think in categories and then you have to translate down to the specific, but then you actually really care more about the category.

Host

而那——那就是沟通误差,因为每个人都在听具体的东西,但你其实还想传达那个一般的东西。

And like that—that's the communication error, because everyone's listening for the specific but you're actually trying to also convey the general.

Alex

对。

Yeah.

Host

这很难。

Which is hard.

Alex

我真的不知道。

I don't really know.

合作与联系 Language and Reasoning

Host

你可以用个简写,比如“好,我在第二层”,然后“我要升到第三层再回到第二层”。

You can maybe use a shorthand like, okay, I'm at level two, and then I'm going to go up to level three and come back to level two.

Alex

我们应该为这类东西搞一套认识论上的简写,因为很难——你要把大量信息压缩进逐词顺序解码里。

We should have some epistemic shorthand for this kind of thing, because it's hard—you're compressing a lot into word-after-word sequential decoding.

Host

我们该不该转成无神经的形式?有没有比英语、Python 或 JavaScript 更好的输出形式?我不知道,这很事后诸葛,但人们一直在猜测模型想输出的母语是什么。

Should we convert to neuro-less? Is there a better form of output than English or Python or JavaScript? I don't know, this is very post-hoc, but people have speculated about what the native language is that models want to output.

Alex

有人说是二进制——那是马克和杰森的观点。

Some people say binary—that's Mark and Jason's thing.

Host

我不知道,是啊。

I don't know, yeah.

Alex

PTX?

PTX?

Host

就说英语和 Python 的混合吧。

Let's say a mix of English and Python.

Alex

我这么说只是因为模型的能力某种程度上反映了我们训练它的内容。所以我们还是想要……我不太买二进制的账。我理解,但就像……你想以某种方式建模世界。我在这类对话里总会提的一件事就是萨丕尔-沃尔夫假说:如果你选英语,你就锁定了英语存在的时间——就算 500 年吧——这并不长。你说的语言会约束你的思维。如果你学一门不同的语言,比如中文里没有时态。我不知道你……我其实不知道这一点。

And I only say this because the capability of a model is somewhat a reflection of what we train them on. So we still want... I don't really buy the binary argument. I guess I understand, but it's like... you want to model the world in some way. One thing I'll bring up, which I always do in this kind of conversation, is Sapir-Whorf: if you choose English, you will have locked into however long English has been around—let's say 500 years—which is not that long. The language you speak constrains how you think. If you learn a different language, for example, in Chinese we don't have tenses. I don't know if you... I actually didn't know that.

Host

而且我说中文。

And I speak Chinese.

Alex

哦,我知道这一点,但我的中文不太好。

Oh, I did know that, but my Chinese is not great.

Host

是啊,是啊。

Yeah, yeah.

Alex

或者像日语,或者我忘了是哪种语言——在韩语里,你跟每个人说话都得承认社会地位,但那是不同于性别的维度。它影响你做的一切。我上语言学 101 时,据说非洲有一种语言里有蔬菜性别。对,就像你有……或者像萨米语没有“雪”这个词之类的。总之,你采用的语言会影响你的思维,如果你选择用英语输出你的思维链,你就在偏向英语能解决的东西。我不知道英语的先验是什么。

Or like in Japanese, or I forget what language it is—in Korean, everyone you speak to you have to acknowledge social status, but there's a different dimension than gender. It just influences everything you do. When I took Ling 101, apparently there's a language in Africa where there's a vegetable gender. Yeah, right, like you have... or like Sami has no word for snow or whatever. Anyway, so the language that you adopt affects your thinking, and if you choose to output your train of thought in English, you are biasing towards whatever English solves. I don't know what the prior of English is.

Host

有意思。我没这么想过。

That's interesting. I did not think of it that way.

Alex

我是说,某种程度上这很有意思,对吧?所以你在语言上是对的——很多模型的思维链也会在语言间波动。明显的例子是中国的模型说英语时可能仍用中文推理。但同样,大多数模型在多语言上非常强,我们能看见这种适应——你可以加入语言,维持一整门语言并不会带来太多收益,但它们会交替推理。

I mean, at some level it's interesting, right? So you're right on language—a lot of model chain of thought also fluctuates language. The obvious example is Chinese models speaking in English might still reason in Chinese. But at the same level, most models are very capable multilingually, and that adaptation we can see—you can add in languages, you don't get that much for maintaining a whole language, but they'll reason interchangeably.

Host

是啊,所以我们都是退化的。我是说,但还有比如德语——你知道,主宾一致,你得把动词放在最后,这非常非常烦人,非常出名。对,你直到最后才知道在干什么,就像“哦,那一堆名词,然后是动词”。

Yeah, so we're all regressive. I mean, but also like let's say German—you know, subject-object agreements, you have to put the verb at the end, which is very super annoying, like very famously. Yeah, right, you don't know what you're doing until the end where you're like, oh, that mess of nouns and then the verb.

Alex

嗯,大多数人听过的最经典的例子是《降临》,里面有七肢桶,它们认为时间对它们是平坦的。所以它们一次性输出整个句子。这最接近自回归和扩散之间的区别。我们用自回归说话。如果你能用扩散说话呢?

Well, the most classic one that most people have heard of is Arrival, where they have the heptapods where they think time is flat to them. So they output entire sentences at one shot. So it's closest to the difference between autoregression and diffusion. We talk in autoregression. What if you could talk in diffusion?

Host

事物随时间逐渐解析。

Where things just resolve over time.

Alex

我明白了,我明白了。

I see, I see.

Host

但整个想法一下子出现。

But the whole idea shows up at once.

Alex

啊,我明白了。所以那是一种截然不同的语言,但它是一种语言。

Ah, I see. So that is a drastically different language, but it is a language.

Host

我明白了。哦,那真的很有意思。真的真的很有意思。

I see. Oh, that's really interesting. That's really really interesting.

Alex

机器可能会说那种我们可能永远不会说的语言,但机器不在乎。

Which machines could speak that we probably would never speak, but machines don't care.

Host

也许这是个大岔路,但难道不是有些东西本质上是推理链,本质上就是自回归的吗?有时候像……

Maybe this is a huge tangent, but are there not things inherently that are reasoning chains that are inherently autoregressive? Sometimes like...

Alex

对,先发生一件事,然后另一件。

Yeah, something happens first then something.

Host

比如代码里的任何东西,在某种程度上必须是因果的,或者通常至少必须是因果的。

Like anything in code, for example, has to be causal in some or usually at least has to be causal.

Alex

嗯,不,那么你就得——那你就没有充分探索函数式编程语言理论,在那里一切都是纯函数式的、完全关系型的,你抽象掉了那个把永远为真的关系翻译成代码的求解器。所以我觉得这可能有点超出我的深度了。

Well, no, so then you have to—then you're not exploring enough functional programming language theory, where everything is pure functional and completely relational, and you sort of abstract away the solver that translates the relationships that are always true into code. So I feel like this is maybe a little bit too out of my depth.

Host

但我喜欢语言,无论是编程语言还是人类语言,我确实经常思考这如何影响推理以及我们能做什么的边界。我不需要超出这个范围太多。我不知道你还有没有其他想法。我原本的收尾问题是:你有所有这些想做的研究方向。你经历过 GPU 模式阶段,经历过 RLM 阶段。想必你还有其他计划,这就是为什么你没有加倍投入那个。顺便说一句,我注意到有趣的是,你们确实从 GPU 那边开始,然后迁移到零梯度那边,这是 Shenu 的叫法。那难道不比摆弄 GPU 感觉更不靠谱吗?

But I love languages, whether it's coding or human, and I do think a lot about how that affects reasoning and the boundaries of what we can do. I don't need to go too much beyond that. I don't know if you have any other thoughts. My closing question was going to be: you have all these research directions that you want to do. You had a GPU mode phase, you had an RLMs phase. Presumably, you have other stuff planned, which is why you're not doubling down on that. By the way, I noticed that it is interesting how you guys do start with the GPU side and then you migrate towards the zero gradient side, which is what Shenu called it. Doesn't that feel less legit than messing with GPUs?

Alex

对,我想是在这个意义上——你确实提到过我喜欢用数学导向的方式思考,有时候做 harness 和智能体非常不舒服,因为它太……

Yeah, I guess in the sense that—so you did bring up that I like to think about things in a math-oriented way, and it's very uncomfortable sometimes to be working on harnesses and agents because it's so...

Host

因为你觉得所有 harness 都一样,就像 harness 里只有两个新想法。

Because you think all harnesses are the same, like two new ideas in harnesses.

Alex

对,而且从经验上讲,很难验证很多发现,至少用我们可用的算力来说是这样。但我认为我转向很多这些问题的原因是,我觉得实际上这才是大部分创新尚未发生的地方。GPU 层面是探索其他想法的手段。比如,你想擅长写 kernel,甚至自动化写 kernel,是为了一个更宏大的目标——我想探索那些不会因系统挑战而瓶颈的想法。我猜那里写的很多都是 harness 相关的东西,但我也对模型层面的东西感兴趣。不过我就说到这里吧。

Yeah, it's also just like empirically it's hard to verify a lot of findings, at least with the compute that we have available to us. But the reason why I think I've moved on to a lot of these problems is I think actually this is where most of the innovation is yet to happen, to me. The GPU level is a means to exploring other ideas. You want to, for example, get good at writing kernels or even automate writing kernels for the sake of a broader goal of—I want to explore ideas where I'm not bottlenecked by systems challenges in that sense. I guess a lot of what's written there is all harness stuff, but I am also interested in things at the model level as well. But I'll just leave it at that.

Host

好,这是个很好的提示。如果人们想联系你,你在找什么帮助?你想在哪些方面找合作者?有没有什么呼吁?

Okay, that's a good hint. If people want to reach out to you, what are you looking for help on? What do you want collaborators on? Any sort of calls to action?

AI 与数学 Collaboration and Reaching Out

Alex

所以,我想我没有什么特别需要和别人合作的事情,除非是和公司合作获取算力,或者和其他人讨论。但我要说,我从不反对和别人一起研究想法。我经常收到很多人的联系,通常是本科生,甚至是其他学生、播客主。

So, I guess there's nothing I have in particular where I feel like I need to work with someone on, unless it's with a company for compute or to talk with other people about it. But I will say I'm never opposed to working on ideas with other people. I get reached out to a lot by often undergrads or even other students, podcasters.

Host

播客主。

Podcasters.

Alex

通常我收到的邮件大概是这样的:‘我非常喜欢 RLM,我想一起合作。’我觉得我……

And usually I feel like I get an email that's something along the lines of, 'I really like RLM, I want to work together.' And I feel like I...

Host

是啊,那是个糟糕的联系方式,对吧?最糟糕的是‘我能请教你吗?’我心想:‘请教什么?读我的论文吧,老兄。’

Yeah, that's a bad reach out, right? The worst is like, 'Can I pick your brain?' I'm like, 'On what? Read my paper, dude.'

Alex

比如他们会说‘我读了你的论文’——带引号的——比如递归语言模型或者 prime agent,一个自我改进的 RLM 工具之类的。这有点像……我真的很喜欢有主见的人,即使我们意见不同。我认为如果你有强烈的观点,并且能够思考为什么你认为这些观点是对的或错的——因为通常很难判断——但如果你对某些问题有坚定的信念,我总是乐意聊天,甚至可能一起做点什么。对于和谁合作或做什么,我没有限制。所以……

Like they'll be like, 'I read your paper' in quotes, like recursive language models or like prime agent, like a self-improving RLM harness or something. And it's kind of like... I really like people that are opinionated even if we disagree. I think if you have strong opinions and are able to think through why you think those opinions are right or wrong—because usually it's hard to actually tell—but like you have strong convictions about certain problems, like I'm always happy to chat and even potentially work on something together. I have no limit to who or what I would like to work on. So...

Host

没有限制。

No limit.

Alex

是的。是的。是的。你知道,在智能体时代,我认为在带宽方面你可以做更多的工作。所以总的来说,我并不难被打动,但我觉得只需要一点努力来知道自己想要什么。

Yeah. Yeah. Yeah. You know, in the era of agents, I think there's a lot more work you can do, you know, bandwidth-wise. So I think in general, I am not hard to impress, but I think it just takes a little bit of effort to know what you want.

Host

这很清楚。我的意思是,当你看到一个新东西出现,执行得很好,简单的好想法,那立刻就会引起你的注意,对吧?其实要获得和所有前沿实验室的人一样的关注并不难,因为他们正在寻找你。你只需要把自己展示出来,对吧?

It's very clear. I mean, when you see a new thing come out, well-executed, good simple idea, then that immediately gets your attention, right? Like it's actually not that hard to get the same attention as all the frontier lab guys because they are looking for you. You just have to put yourself out there, right?

Alex

没错。是的。

Exactly. Yeah.

Host

但是的,这是真的。我要说,我认为现在人类的注意力非常稀缺,我确实在努力应对我手头的项目数量。

But yeah, it's true. I will say I think human attention is very scarce right now and I do struggle with the number of projects I have going on.

Alex

我不知道如何管理它。我不认为智能体有任何帮助。就像我会提示它,提示一个东西,然后再也不看它。

And I don't know how to manage it. I don't think agents are helping at all. Like I would just prompt it and I'll prompt a thing and then never look at it.

Host

对。这很常见。

Right. Like which is very common.

Alex

是的。那很糟糕。

Yeah. And that sucks.

Host

我想也许一个较小的区别是——我的意思是,实际上也许你在做研究。我不确定。但至少对我来说,我可能有 10 到 15 个不同的想法想做,但问题是大多数都是糟糕的。

I guess maybe one of the smaller differences is—I mean, actually maybe you were doing research. I'm not sure. But for me at least, I'll have maybe 10 or 15 different ideas that I want to do, but the thing is most of them are bad.

Alex

是的。甚至对于联系我的人来说,这可能也是真的。也许这个想法实际上很糟糕,但对我来说看起来很有趣。所以我们可以花些时间研究它,如果我们觉得确实有东西,那么我们应该在接下来的几周里真正追求它。这是我的风格——顺便说一下,这就是我喜欢博士的原因——因为有时候我只是在思考问题,也许在跑步或打网球时,我猜我不是在工作,但那些是最有趣的时光。然后当我真正对某件事有信念时,我会放下一切,只是去做,把所有时间都花在思考和解决那个问题上。然后一旦你到了可以运行实验的地步,就又变得轻松了。所以……

Yeah. And this also maybe is true even for someone that reaches out to me. Like maybe the idea is actually bad, but it looks interesting to me. And so we can spend some time looking into it and if we feel like there's actually something there, then we should take the next few weeks and just really pursue it. And this is my style—this is why I love the PhD by the way—because there are times when I'm just thinking about problems, maybe on a run or just playing tennis or something, like I'm not working I guess, but it's like those are the most fun times. And then when I really am convicted about something, I'll just drop everything and just do it, just spend all my time thinking and working on that problem. And then once you get to the point where you can just run experiments, then it's kind of easy coasting again. So...

Host

是的,抱歉,这大概是倒数第四个问题:我认为很多人也在思考科学作为下一个前沿,比如物理科学、生物、甚至数学。你如何区分,比如说你选择适用于工业的项目,然后可能选择纯粹科学的项目?

Yeah, sorry, this is like the fourth last question, which is: I think a lot of people are also thinking about science as the next frontier, like physical sciences, bio, math even. How do you separate, I guess let's say your choice of projects that is applicable for industry and then maybe choice project that's just like science?

Alex

实际上,在我开始博士之前,我做过 AI 用于生物的工作。自那以后,这个领域变化很大。我应该说明……

I actually worked on AI for bio stuff before I started my PhD. The field has changed a lot since then. I should say...

Host

因为过去是理论性的,当然,你是什么意思,就像你知道我有一条路,然后我选择了它,但现在很多人跨界,所以我们开始了一个科学播客来覆盖那些东西,因为很多工程师说实际上那里有可解决的问题。

Because it's like it used to be a theoretical like of course what do you mean like you know I have one path and then I chose that but now a lot of people are crossing over and like so we have started a science pod to just cover those things because a lot of engineers are like well actually there's tractable problems there.

Alex

是的。你知道,我先说,我对很多这些主题的理解可能相当有限,但我认为如果我发现了,要么是因为有人联系我,要么是我看一个问题,然后我想,‘嘿,我们使用或正在思考的一些设计原则实际上对这个问题很有意义,’我也会对此感到兴奋。但我认为这更难。我不知道。我认为科学,尤其是经验或应用科学,有非常非常长的反馈循环之类的。

Yeah. You know, I will preface by saying my understanding of a lot of these topics is probably pretty limited, but I think if I find out either because someone reaches out or I look at a problem and I'm like, 'Hey, some design principles that we use or that we're thinking about right now actually make a lot of sense for this problem,' I get excited about those as well. But I think it's harder. I don't know. I think with science, especially empirical or applied science, has very very long feedback loops or whatever.

Host

是的,你把这个转化为机器人问题。

Yeah, you convert this to a robotics question.

Alex

是的。是的。我的意思是,对我来说,还有这个方面,比如现在值得花时间和押注什么,因为也许我花了很多时间在这个问题上,然后 6 个月后,不同的解决方案,也许一个新模型出来了,它对这个好多了,所以我必须小心,你知道,你必须意识到你认为事情可能走向哪里。所以……

Yeah. Yeah. I mean to me like also this aspect of like what is worth spending and betting my time on now because like maybe I spend a lot of time on this problem and then like in 6 months like a different solution kind like maybe a new model comes out and it's like ah it's way better for this and so I do have to be careful I like you know you have to be conscious about like where you think things might be going. Um so...

Host

发表发表周期。

Publish publish cycle.

Alex

是的。是的。是的。所以……

Yeah. Yeah. Yeah. So...

Host

我可以直接 Oh 3 它饱和了。我们做到了。

I can just Oh 3 it's saturated. We did it.

Alex

是的。我的意思是,ArcJ 3 在不到一年内就饱和了。所以,有点,你知道,就像我不知道,如果你是一个实验室选择那个问题,你现在可能有点难过,因为……

Yeah. I mean, ArcJ 3 got saturated in less than a year. So, it's kind of, you know, like it's I don't know like if if if you were a lab picking that problem like you're probably kind of sad now cuz...

Host

没错。这就是为什么知识工作、游戏,所有这些都饱和了。现在,实际上前沿是科学。

Exactly. That's why knowledge work, gaming, all these things are saturated. Now, actually the frontier is science.

Alex

知识工作饱和了。

Knowledge work is saturated.

Host

是的。GDP 大概是 80 多 90 多。嗯,我的意思是,你知道有 90 到 100%,那显然需要接下来的 10 年,但就像呃,下一个低垂的果实将是,你知道,其他的东西。

Yeah. GDP is like 80 something 90 something. Um, I mean like you know there's there's 90 to 100% that's obviously going to take the next 10 years but like uh well the next low-hanging fruit is going to be like you know the other stuff.

Alex

是的,也许我会多想想。实际上我还没有太多……

Yeah, maybe I'll think about that more. I actually haven't given too much...

Host

我只是想猜猜你的下一个方向。

I'm just trying to guess your next direction.

播客总结 AI and Mathematics

Alex

我想说,因为我觉得尤其是在 MIT,那里有很多非常有才华的自然科学家,我觉得如果说“我要解决你的问题”,几乎有点亵渎神明。

I will say, because I think especially at MIT, there are a lot of really talented scientists there in the natural sciences, and I think it's a little bit like sacrilegious almost to be like, I'm going to figure out your problem.

Host

不,那不对,我知道,是的。

No, that's no, I know, yeah.

Alex

所以我采访了做 IMO 那件事的人。他从未,我从未去过 IMO。我甚至不知道那是什么,模型老兄。

So I interviewed who did the IMO thing. He's never, I've never been to IMO. I don't even know what it is, model dude.

Host

是的,是的,这有点不尊重,但无所谓。

Yeah yeah, which is like very disrespectful, but like whatever.

Alex

在某个时候,你必须尊重等等,就像好吧,你知道进展正在取得,数字输出正确。

At some point, like you have to respect etc., like okay, you know the progress is being made, like number is getting output right.

Host

确实。

True.

Alex

是的,我的意思是,这非常苦涩的教训。非常有趣,因为数学家们现在对 nar 这样回应。

Yeah, I mean, that is very bitter lesson. It's very interesting, like because the mathematicians are responding this way to nar right now.

Host

就像 Terry 转向说,不,让我们不要,不要使用 AI,我就像。

Like Terry turns to is like no, like let's not, let's not use AI, I'm like.

Alex

嗯,是的,我的意思是,我觉得整个事情有点奇怪,因为我觉得如果很多人在这发生之前没有和 OpenAI 合作过,他们的论点会更有力。

Well yeah, I mean, I think that whole thing is kind of weird because I feel like they would have had a stronger case if a lot of them didn't work with OpenAI before like all this happens.

Host

不,那是人身攻击,呃,他们真的试图避免那样。

No, that's ad hominem, uh and they're really trying to stay away from that.

Alex

那又怎样,对吧?那又怎样,就像我不知道,就像他们得到了什么,你知道,他们合作过。我会和不同意的人合作之类的,或者就像我做了一件事,现在后悔改变了主意之类的,对吧?所以我会捍卫他们说话的权利,但是,呃,是的,很多人合理地不同意他们。

Then so what, right? So what, like I don't know, like so what they got, you know, they've collaborated. I'll collaborate with people that I don't agree with or whatever, or like I did a thing and then now I regret that I changed my mind or whatever, right? So I'll defend their right to say that, but like, uh yeah, a lot of people are reasonably disagreeing with them.

Podcast Wrap-up Podcast Wrap-up

Host

嗯,好的,酷,呃,感谢你加入我们,祝贺你到目前为止的成功。我简直不敢相信你还没完成你的博士学业,第二年,你知道。

Um okay cool, uh thanks for joining us, congrats on your success so far. I can't believe you're still not done with your PhD, year two, you know.

Alex

不敢相信我们做这个播客没有过一遍强化学习的论文。

Can't believe we did this podcast without going through the RL paper.

Host

他有一个定义。

He had a definition.

Alex

我认为这篇论文更多是关于实证结果。实际的想法相当简单。

I think the paper is more about empirical results. Like the actual idea is quite simple.

Host

是的。而且你已经谈过很多次了。

Yeah. And you've talked about it many times.

Alex

是的。是的。是的。在这一点上,我认为现在有更有趣的事情要看。所以。

Yeah. Yeah. Yeah. At this point, I think there's more interesting things to look over now. So.

Host

酷。嗯,我们很期待看到你接下来做什么。

Cool. Well, we're excited to see what you do next.

Alex

非常感谢。

Thank you so much.

互动版:逐字朗读 + 针对本期提问 →