From Transformer Inventor to Continuous Thought Machine: A New Path in AI
打开互动全文版(中英对照 + 朗读 + 问答)→一位 Transformer 共同发明者讨论离开过度饱和领域,探索连续思维机器等新架构,以及 AI 研究中研究自由的重要性。
A Transformer co-inventor discusses leaving the oversaturated space to explore new architectures like the Continuous Thought Machine, and the importance of research freedom in AI.
尽管我参与了 Transformer 的发明,幸运的是,没有人比我研究它们更久,可能除了其他七位作者。所以今年早些时候,我决定大幅减少专门针对 Transformer 的研究,因为我觉得这个领域已经过度饱和了。并不是说 Transformer 没有更多有趣的事情可做,而是我要利用这个机会做一些不同的事情,真正增加我在研究中的探索力度。
Despite the fact that I was involved in inventing the Transformer, luckily no one's been working on them as long as I have, with maybe the exception of the other seven authors. So I actually made the decision earlier this year that I'm going to drastically reduce the amount of research that I'm doing specifically on the Transformer because of the feeling that it's an oversaturated space. It's not that there are no more interesting things to be done with them. I'm going to make use of the opportunity to do something different, to actually turn up the amount of exploration that I'm doing in my research.
我们刚刚发布了连续思维机。它是今年 Euro 2025 的亮点。你应该关注它,因为它具有原生的自适应计算能力。这是一种构建循环模型的新方法,使用更高层次的神经元概念和同步作为表示,让我们通过受生物学和自然启发的方式,以更人性化的方式解决问题。
We just released the continuous thought machine. It's a spotlight at Euro 2025 this year. You should care about it because it has native adaptive compute. It's a new way of building a recurrent model that uses higher level concepts for neurons and synchronization as a representation that lets us solve problems in ways that seem more human by being biologically and nature inspired.
在 Transformer 时代,AI 研究的氛围实际上非常不同,因为现在感觉类似的事情不太可能发生了,因为我们拥有的自由度降低了。Transformer 是非常自下而上的。并不是某个人有一个宏大的计划从上而下地告诉我们这就是我们应该研究的东西。而是一群人在午餐时交谈,思考当前的问题以及如何解决它们,并且有自由花上几个月的时间专门尝试这个想法,最终产生了这个新架构。
The atmosphere in AI research was actually quite different back during the Transformer years because it doesn't feel like something similar could actually happen right now, due to the reduced amount of freedom that we have. The Transformer was very bottom-up. It's not that somebody had this grand plan that came down from on high that this is what we should be working on. It was a bunch of people talking over lunch, thinking about what the current problems are and how to solve them, and having the freedom to have literally months to dedicate to just trying this idea and having this new architecture fall out.
我们已经花费了数亿美元。最大的基于进化的搜索大概只有几万次。我们有这么多算力。如果你扩大这些搜索算法的规模,会发生什么?我相信当有人最终下定决心真正扩大这些进化类生命实验的规模时,你会发现一些有趣的东西。但我在一个人们全力投入这一种技术的环境中提出了这个想法,却无人问津。所以现在我有自己的公司,可以追求这些方向。
We've spent hundreds of millions of dollars. The biggest sort of evolution-based search is probably in the tens of thousands. We have all this compute. What happens if you scale up these search algorithms? I'm sure you'll find something interesting when someone eventually does bite that bullet and really scale up these evolutionary sort of life experiments. But I pitched it in an environment where people were just going all in on this one technology. I got zero interest. So now I have my own company and I can pursue those directions.
观众知道我是 Kenneth Stanley 思想的超级粉丝。他的书《为什么伟大不能被计划》改变了我的人生,简直不可思议。他说的就是我们需要让人们追随自己的兴趣梯度,不受目标和委员会等的束缚。因为这就是我们进行认知觅食的方式。当太多议程混杂其中时,你最终会得到一团灰色的浆糊,无法发现有趣的新奇和多样性。我想这基本上就是你们公司 Sakana 的论点,即拥抱这些想法。
The audience will know I'm a huge fan of Kenneth Stanley's ideas. So his book, Why Greatness Cannot Be Planned, changed my life. It was absolutely insane. And what he was speaking to is that we need to allow people to follow their own gradient of interest unfettered by objectives and committees and so on. Because that is how we do epistemic foraging. When you have too many agendas involved in the mix, you kind of end up with a gray goo and you don't discover interesting novelty and diversity. And I suppose that's basically the thesis of your company, Sakana, to lean into those ideas.
是的,完全正确。在公司里,我们是那本书的超级粉丝。实际上,我们希望下周能邀请他来公司演讲。这是我们内部经常讨论的一种哲学。我们有这本书的副本,包括最近的日文译本。你知道,作为联合创始人之一,我的主要工作之一,也是我必须持续为公司做的事情,就是确保保护研究人员目前拥有的自由,因为拥有资源去做这件事确实是一种特权。而且不可避免地,正如我所见,随着公司成长,越来越多的压力会涌入,压缩自由。但我认为,因为我们如此坚信这种哲学,我希望我们能尽可能长久地给予人们现在所拥有的全部研究自由。
Yes, exactly. At the company, we're a massive fan of that book. We're hoping to have him come and talk at our company next week, actually. And it's a philosophy that we do talk about internally. We have copies of the books including the recent Japanese translation. As you know, one of the co-founders, one of my main jobs, one of the main things that I have to keep doing for this company is making sure that we protect the freedom that the researchers currently have, because it's a privilege really that we have the resources to be able to do that. And inevitably, as I've seen happen, as the company grows, more and more pressure comes in and it narrows the freedom. But I think because we believe in this philosophy so strongly, I'm hoping that we can give people all the research freedom that we do now for as long as possible.
那么,随着公司成熟,那些限制自由的过程是什么?我是说,你会如何描述?
And what are those processes that curtail freedom as a company matures? I mean, how would you describe that?
这个行业从未有过如此多的兴趣、人才、资源和资金,这很好,但不幸的是,这增加了人们为了与其他从业者竞争、试图从这项技术中获取价值并赚钱的压力。我认为这就是现实。作为一家初创公司,你会有兴奋感和尝试新事物的感觉。在最初阶段,你有一些缓冲时间,所以你有自由去尝试不同的事情。但不可避免地,人们开始要求投资回报,或者期望你推出一些产品。这不幸地降低了研究人员的创造力,因为发表论文或创造对产品有用的技术的压力增加了,所以自主感开始下降。但我确实告诉人们,当他们开始为公司工作时,我希望他们研究自己认为有趣和重要的事情,我是认真的。
It's great that there's never been so much interest and people and talent and resources and money in the industry, but unfortunately that just increases the amount of pressure people have in order to compete with all the other people working on it and trying to get the value out of this technology and making money. And I think that's what just happens. As a startup, you have a feeling of excitement and trying something new. And right at the beginning, you have a bit of a runway. So you have the freedom to try different things. But inevitably, people are starting to ask for returns on their investments or they're expecting you to churn out some product. And this just unfortunately reduces the creativity that researchers have because the pressures to publish or the pressure to create technology that's actually useful for the products that we have goes up, and so the feeling of autonomy I think starts to go down. But I literally tell people when they start working for the company, I want you to work on what you think is interesting and important, and I mean it.
在 YouTube 上有一个现象叫观众捕获,我认为可能还有一个现象叫技术捕获。在谷歌早期,它是相当开放的,而 Transformer 现在是所有 AI 技术的普遍骨干,你参与其中是一项巨大的成就。但 OpenAI 也有类似的故事。他们现在开始看到所有这些商业化的机会。他们将变成 LinkedIn、应用平台、搜索平台、社交网络。我想这可能会发生在你们身上,可能性很大,尤其是考虑到我们今天要讨论的新论文——连续思维机。
In YouTube there's a phenomenon called audience capture, and I think there might be a phenomenon called technology capture. In the early days of Google it was quite open-ended, and I mean Transformers is now the ubiquitous backbone of all AI technology and it's a huge achievement that you're involved in. But there's a similar story with OpenAI. They're now starting to see all of these commercialization opportunities. They're going to become LinkedIn, an application platform, a search platform, a social network. And I guess this could happen to you guys, that there's a very strong chance, especially with your new paper that we're going to talk about today, this continuous thought machines.
它可能是一项革命性技术,但随后它的商业化路径就会变得显而易见。压力就是这样来的。我喜欢“受众捕获”这个类比。我认为大型语言模型确实存在某种程度的捕获,对吧?它们效果太好了,以至于每个人都想研究它们。我真的很担心我们现在陷入了某种局部最优,我们需要努力摆脱它。
It could be a revolutionary technology, but then it will become obvious how it could be commercialized. And that's how those pressures come in. I like the audience capture analogy. I think there's definitely been some kind of capture by large language models, right? They worked so well that everyone wanted to work on them. And I'm really worried that we're kind of stuck in this local minimum now, and we sort of need to try to escape it.
我们谈到了 Transformer,但我想谈谈 Transformer 之前的一段时间,因为我觉得很有启发性。当然,Transformer 之前的主要技术是循环神经网络,对吧?当时也有类似的感觉。当 RNN 出现,我们发现了这种新的序列到序列学习,那也是一个巨大的突破,翻译质量大幅提升,语音识别质量大幅提升。当时也有类似的感觉:好吧,我们找到了这项技术,只需要完善它。那时,我最喜欢的任务甚至是字符级语言建模。所以每当有新的基于 RNN 的字符级语言建模论文出来,我都会很兴奋。我会想快速阅读论文,看看他们是如何改进的。但论文总是对同一架构的微小修改,比如 LSTM 和 GRU,也许用单位矩阵初始化以便使用 ReLU 函数,或者把门放在不同的位置,或者以稍微不同的方式堆叠它们,或者有向上和侧向的门控。我记得我最喜欢的一个是层次化 LSTM,它实际上会决定是否计算不同层。如果你在维基百科上训练,观察它决定计算或不计算的结构,看起来模型确实捕捉到了句子的结构。我以前很喜欢这类东西。但改进总是像 1.26 比特每字符、1.25、1.24。这样的结果就可以发表,令人兴奋。但在 Transformer 之后,我后来加入的团队首次将非常深的 Transformer 模型(仅解码器)应用于语言建模,我们立即得到了大约 1.1 的结果。这个结果太好了,以至于有人走到我们桌前礼貌地说:“我觉得你算错了,是奈特而不是比特每字符吧?”我们说:“不,不,不,这确实是正确的数字。”后来让我震惊的是,突然间,所有这些研究——而且是很优秀的研究——突然变得完全多余了。
So, we spoke about the transformers, but there's a time just before the transformers that I'd like to talk about because I think it's quite illustrative. So, of course, the main technology before transformers was recurrent neural networks, right? And there was a similar feeling, right? When recurrent neural networks came in and we discovered this new sort of sequence-to-sequence learning, that was also a massive breakthrough, right? The translation quality went up massively, voice recognition quality went up massively. And there was this similar sort of feeling then of like, okay, we found the technology and we just need to perfect this technology. Back then, even my favorite task was character-level language modeling. So every time a new RNN-based character-level language modeling paper came out, I got quite excited. I'd want to quickly read the paper, like, how did they get the improvements? But the papers were always these slight modifications on the same architecture, right? It was LSTMs and GRUs, maybe initializing with the identity matrix so that you could use the ReLU function, or maybe if you put the gate in a different place, or if you layer them in a slightly different way, or if you had gating going upwards as well as sideways. And I remember one of my favorites was this hierarchical LSTM where it would actually decide to compute or not compute the different layers. And if you trained on Wikipedia and looked at the structure of when it decided to compute or not compute, it kind of looked like the structure of the sentences were actually being picked up by the model. And I used to love that sort of stuff. But the improvements were always like 1.26 bits per character, 1.25 bits per character, 1.24. That was a result that was publishable, that was exciting. But then after the transformer, the team that I went on to afterwards, we applied for the first time very deep transformer models, decoder-only transformer models, to language modeling and we immediately got something like 1.1. So something that was so good that people actually came to our desk and politely told us, 'I think you made an error, a calculation. Do you think it's nats not bits per character?' And we're like, 'No, no, no, it really is the correct number.' What struck me later is that all of a sudden, all of that research—and to be clear, very good research—was suddenly made completely redundant.
是的。
Yes.
所有那些对 RNN 的无尽排列组合突然看起来像是浪费时间。我们现在也处于类似的情况,很多论文只是采用相同的架构,进行无休止的不同调整,比如归一化层放在哪里,训练方式略有不同。我们可能正在以完全相同的方式浪费时间。我个人认为我们还没有结束。我不认为这是最终的架构,我们只需要继续扩大规模。某个时候会有突破,然后我们会再次清楚地意识到我们现在正在浪费大量时间。
All of those endless permutations to RNNs were suddenly seemingly a waste of time. We're kind of in the situation right now where a lot of the papers are just taking the same architecture and making these endless amount of different tweaks, like where to put the normalization layer and slightly different ways of training them. And we might be wasting the time in exactly the same way. I personally don't think we're done. I don't think that this is the final architecture and we just need to keep scaling up. There's some breakthrough that will occur at some point, and then it will once again become obvious that we're kind of wasting a lot of time right now.
是的。所以我们成了自己成功的受害者,陷入了这个吸引盆。有很多吸引盆。Sarah Hooker 谈到过硬件彩票,这是一种架构彩票。这让我想起了农业革命:这种相变发生了,所有拥有这些必要技能的人——这些多样化的生存技能——都消亡了。这实际上相当矛盾,因为我们需要这些技能来迈出下一步。
Yeah. So we are a victim of our own success and this basin of attraction. There are so many basins of attraction. Sarah Hooker spoke about the hardware lottery, and this is a kind of architecture lottery. It actually made me think of the agricultural revolution, which is that this kind of phase change happened and all of the folks that had these skills that were so necessary, these diverse skills for living and surviving, they died out. And that's actually quite paradoxical because we need those skills to take the next step.
所以我们现在处于这个阶段。我们有“基础模型”这个术语,暗示你可以用基础模型做任何事情。在企业界,过去我们有数据科学家、机器学习工程师,甚至在中型企业里做这些架构调整。而现在我们只有 AI 工程师,他们只做提示工程之类的。所以你是说,我们需要多样化的基本技能来思考新的解决方案和架构,但这些技能正在消亡。我想我不同意这一点。我认为问题在于我们有很多非常有才华、非常有创造力的研究人员,但他们没有发挥自己的才能。例如,如果你在学术界,有发表论文的压力。有发表论文的压力时,你会想:“好吧,我有一个很酷的想法,但它可能行不通。可能太奇怪了。可能很难被接受,因为我需要更多地推销这个想法。或者我可以只尝试这个新的位置嵌入。”问题在于,当前学术界和公司里的环境并没有真正给人们自由去做他们可能想做的研究。
And so we're now in this regime. We've got the term 'foundation model', and the implication is that you can do anything with a foundation model. In the corporate world, we used to have data scientists, ML engineers doing these architectural tweaks even in midsize enterprise. And now we just have AI engineers who are just doing prompt engineering and so on. So you're saying that the fundamental skills that we need to be diverse, to think of new solutions and new architectures, they're dying out. I think I'm going to disagree with that. I think the problem is we have plenty of very talented, very creative researchers out there, but they're not using their talents. For example, if you're in academia, there's pressure to publish. And if there's pressure to publish, you think to yourself, 'Okay, I have this really cool idea, but it might not work. It might be too weird. It might be difficult to get it accepted because I have to sell the idea more. Or I can just try this new positional embedding.' The problem is that the current environment both in academia and in companies are not actually giving people the freedom that they need to do the research that they probably want to do.
我的意思是,还有一件有趣的事,即使有很棒的新研究,我和 Seb Hoger 聊过,他有所有这些新的架构想法,但 OpenAI 没有实施它们。我的意思是,谷歌在做这个扩散语言模型,很酷。我想听听你的看法,为什么会这样。现在有一些哲学观点在流传,比如通用表征的概念,认为存在通用模式,而 Transformer 的表征类似于大脑中的表征。这导致了这样一种想法:我们不需要使用不同的架构,因为只要我们有更大的规模和更多的算力,条条大路通罗马。那我们为什么还要费心去做不同的事情呢?
I mean, there's also this interesting thing that even in spite of great new research, I mean I was speaking to Seb Hoger and he's got all of these new architectural ideas and OpenAI aren't implementing them. I mean Google are doing this diffusion language model which is quite cool. And I'd like to know your opinion on why that is. So there's a few philosophies floating around like this concept of a universal representation that there are universal patterns and the transformer representations resemble those in the brain. And it's rather led to this idea of, well, we don't need to use different architectures because if we just have more scale and more compute, then all roads lead to Rome. So why would we bother doing it any differently?
实际上有更好的,对吧?研究中已经有一些架构被证明比 Transformer 效果更好。但好得还不够,不足以让整个行业放弃这样一个成熟的架构,因为你熟悉它。你知道如何训练它。你知道它如何工作。你知道内部机制。你知道如何微调它们。你所有的软件都已经为训练 Transformer 设置好了。
There's actually better, right? There is actually already architectures that have been shown in the research to work better than transformers. Okay. But not better enough in order to move the entire industry away from such an established architecture where you're familiar with it. You know how to train it. You know how it works. You know how the internals work. You know how to fine-tune them. You have all this software is already set up for training transformers.
所以如果你想推动行业摆脱它,仅仅更好是不够的。它必须明显碾压式地更好。Transformer 就比 RNN 好那么多。你只需把它应用到一个新问题上,训练速度快得多,准确率高得多,你不得不迁移。我认为深度学习革命也是另一个例子。当时有很多怀疑者,人们早在那个时候就在推动神经网络,而其他人说:‘不,我们认为符号方法会更好。’但后来他们证明它好到无法忽视。这个事实让寻找下一个东西变得更加困难。这就是那个引力中心,总是把你拉回:‘好吧,但 Transformer 已经足够好了。是的,你在这里搞了一个很酷的小架构,看起来准确率更高,但 OpenAI 那边把它放大十倍就打败了它。所以我们就继续这样吧。’
So if you want to move the industry away from that, being better is not good enough. It has to be obviously crushingly better. Transformers were that much better over RNNs. You just applied it to a new problem and it was so much faster to train and you got such higher accuracy that you just had to move. And I think the deep learning revolution was also another example of that. You had plenty of skeptics and people were pushing neural networks even back then and people were saying, 'No, we think symbolic stuff will work better.' But then they demonstrated it as being so much better that you couldn't ignore it. This fact makes finding the next thing even harder. That's the gravitational pull always pulling you back to, 'Okay, but a transformer is good enough. And yeah, you made a cool little architecture over here that looks like it's got better accuracy, but OpenAI over here just made it 10 times bigger and it beats that. So let's just keep going.'
我还想补充一个可能的原因,你知道,我很喜欢那篇关于碎片化纠缠表征的论文。存在一个捷径学习问题,我认为这里有点海市蜃楼。这些语言模型可能存在一些我们尚未完全意识到的问题。而且我们还看到,我们开始劣化这个架构。我们知道我们需要用于推理的自适应计算。我们知道我们需要不确定性量化之类的东西。而我们正在做的是把这些东西附加在上面,而不是拥有一个内在就能完成所有这些我们知道需要的事情的架构。
May I also submit that there could be an additional reason which is, you know, I love that fractured entangled representations paper. There's this shortcut learning problem and I think there's a little bit of a mirage going on here. There might be problems with these language models that we're not fully aware of. And there's also this thing that we're seeing that we are starting to bastardize the architecture. So we know we need to have adaptive computation for reasoning. We know we want things like uncertainty quantification. And what we're doing is we're bolting these things on top rather than having an architecture which intrinsically does all of these things that we know we need.
是的。我认为我们的连续思维机器就是试图更直接地解决这些问题,卢克稍后会告诉你更多。当前技术仍然有些不太对劲。我认为一个流行的说法是‘锯齿状智能’。你可以问一个 LLM 某个问题,它能解决一个博士级别的问题,然后下一句话它就能说出一些明显错误的东西,令人震惊。我认为这实际上反映了当前架构中可能存在的根本性问题。尽管它们很惊人,但当前技术实际上太好了。这是另一个难以摆脱它们的原因。它们在以下意义上太好了。你提到了我们有这些基础模型。这没问题,所以我们有了可以用它们做任何事的基础。是的,我认为当前的神经网络非常强大,如果你有足够的耐心、足够的算力和足够的数据,你可以让它们做任何事。但我不一定认为它们‘想要’这样做。我们有点在强迫它们。它们是通用函数逼近器,但我认为可能存在一类函数逼近器,它们更愿意以人类表征事物的方式来表征事物。所以有一篇相当冷门的论文是我对此的典型例子。它叫‘智能矩阵指数化’。我认为它实际上被拒稿了。所以你可能可以投影图一,但有一张图展示了它解决经典的螺旋数据集——需要分离螺旋中的两个类别。它展示了经典 ReLU 多层感知器和 tanh 多层感知器的决策边界。你可以看到它们都解决了问题。从技术上讲,它们都解决了问题,因为它们正确分类了所有点,并在这个非常简单的数据集上得到了很好的测试分数。然后他们展示了他们在论文中构建的 M 层的决策边界,它是一个螺旋。该层将螺旋表示为螺旋。如果数据是螺旋,我们难道不应该把它表示为螺旋吗?然后如果你回头看螺旋和经典 ReLU 多层感知器的决策边界,很明显你只有这些微小的分段线性分隔。这就是我的意思。是的,如果你训练这些东西足够多,并把这些小的分段线性边界推来推去,它可以拟合螺旋并获得高准确率。但当我看到那些图像时,我并没有感觉到 ReLU 版本真正理解它是一个螺旋。而当你把它表示为螺旋时,它实际上能正确外推,因为螺旋会继续向外延伸。
Yeah. And I think our continuous thought machine is an attempt at addressing those more directly, which Luke will be able to tell you more about later. There's something still not quite right with the current technology. I think the phrase that's becoming popular is 'jagged intelligence'. The fact that you can ask an LLM something and it can solve literally a PhD level problem, and then in the next sentence it can say something so clearly obviously wrong that it's jarring. And I think this is actually a reflection of something probably quite fundamentally wrong with the current architecture. As amazing as they are, the current technology is actually too good. Another reason why it's difficult to move away from them. So they're too good in the following sense. And you spoke about the fact that we have these foundation models. That's okay, so we have the foundation that we can do anything with them. Yes, I think current neural networks are so powerful that if you have enough patience and enough compute and enough data, you can make them do anything. But I don't necessarily think that they want to. We're sort of forcing them. They're universal function approximators, but I think there are probably a space of function approximators that will more want to represent things in the way that a human represents them. So there's actually quite an obscure paper that is my poster child for this. It's called 'Intelligence Matrix Exponentiation'. And I think it was actually rejected. So you can probably project the image of figure one, but there's an image of it solving the classical spiral dataset of needing to separate the two classes in the spiral. And it has the decision boundary for both a classic ReLU multi-layer perceptron and a tanh multi-layer perceptron. And you can see they both solve it. Technically, they both solve the problem because they classify all the points correctly and get a very good test score on this very simple dataset. And then they show you the decision boundary for the M-layer that they built in this paper and it's a spiral. The layer represented the spiral as a spiral. Shouldn't we, if the data is a spiral, represent it as a spiral? And then if you look back at the decision boundaries for the spiral and the classic ReLU multi-layer perceptron, it's clear that you just have these tiny little piecewise linear separations. And that's what I mean. Yes, if you train these things enough and you push these little piecewise linear boundaries around enough, it can fit the spiral and get a high accuracy. But there's no feeling when I look at those images that the ReLU version actually understands that it is a spiral. And when you represent it as a spiral, it actually extrapolates correctly because the spiral just keeps going out.
你触及了一个迷人的点,因为我们之前讨论了对自适应性和自适应计算的需求。我深受 Randall Beer 的神经网络样条理论的启发,我们多次邀请过他。你可以看看 TensorFlow 游乐场。你可以看看当你在螺旋流形上使用 ReLU 网络时会发生什么。你可能会认为这些东西基本上就是一个局部敏感哈希表,因为它们划分空间并可以预测螺旋流形。但我们想要做一些更不同的事情。这也与‘冒名顶替者’问题有关,因为仅仅追踪螺旋流形而不延续模式,这两者有很大区别。所以从冒名顶替者的角度来看,仅仅追踪模式并不是抽象地或建设性地学习它。如果我们建设性地学习它,就像你在论文中提到的复杂化、抽象构建块,并且你可以进行自适应计算,你就理解了螺旋。这意味着通过自适应计算,你可以延续螺旋,然后你可以更新模型的权重,使其具有自适应性,因为这对智能至关重要。所以我们知道我们需要能够做这些事情的模型。但出于某种原因,它们如此复杂,几乎比一个自适应智能系统更好,因为它们告诉我们我们想听的话。它们看起来如此智能,但我们知道它们缺少这些基本属性。
You're touching on something fascinating there because we were talking about the need for adaptivity and adaptive computation. I'm really inspired by Randall Beer's spline theory of neural networks and we've had him on many times. You can look on the TensorFlow playground. You can look what happens when you have a ReLU network on this spiral manifold. And you'd be forgiven for thinking that these things are basically a locality sensitive hashing table, because they partition the space and they can predict the spiral manifold. But we want to do something a little bit more different than that. And it also comes into this impostor thing because just tracing the spiral manifold but not continuing the pattern, there's a big difference between that. So from an impostor perspective, just tracing the pattern is not learning it abstractly or constructively. If we learned it constructively, as you speak about in your paper this complexification, the abstract building blocks and you can do adaptive computation, you understand the spiral. That means that with adaptive computation, you can continue the spiral and then you can update the model's weights so it has adaptivity because that's so important for intelligence. So we know that we need models that can do these things. But for some reason they're so sophisticated, they're almost better than an adaptive intelligent system because they tell us exactly what we want to hear. They seem so intelligent, but we know that they're missing these fundamental properties.
当我看到视频生成模型时,我仍然相当怀疑。你知道,我们经历了一个阶段,你可以通过某人手上的手指数量来检测它们。是的,通过更多的数据、更多的算力、更好的训练技巧,好吧,它们屈服了。现在它们通常确实有五根手指。
I'm still fairly skeptical when I see video generation models. You know, we went through a phase where you could detect them because of the number of fingers on somebody's hand. And yes, with more data, with more compute, with better training tricks, okay, they submit. And now they usually do have five fingers.
但我们到底解决了问题,还是只是用蛮力强迫神经网络知道它有五根手指,而实际上可能有更好的表示空间?说我们应该像螺旋一样表示螺旋,这几乎荒谬到有争议。但你知道,如果它能普遍做到这一点,比如像人类一样表示手,那么数手上有几根手指可能就容易多了。不幸的是,它们效果太好了。不幸的是,Scaling(规模扩张)效果太好了,因为人们太容易把这些扫到地毯下。
But did we fix the problem or did we just use more brute force to force the neural network to know it's five fingers, where something that actually had a much better kind of representation space? It's almost mad that it's controversial to say that we should represent a spiral like a spiral. But, you know, something that could do that generally, that if it represented a human hand the way that maybe I represent a human hand, then maybe it would be much easier to count how many fingers are on a hand. It's unfortunate that they work so well. It's unfortunate that scaling works so well because it's too easy for people to just sweep these problems under the carpet.
你们可能创造了今年最好的论文。这可能是带我们进入下一步的创新。你在欧洲也获得了焦点奖?
You guys have possibly created what I think might be the best paper of the year. This could actually be the innovation which takes us to the next step. And you got the spotlight in Europe as well?
是的。
Yeah.
今年,恭喜。我认为这证明了这篇论文有多棒。
This year and congratulations on that. So I think that's testament to how amazing this paper is.
CTM(连续思维机器)其实离我们陷入的局部最小值并不远。我们并没有找到全新的技术。我们只是采用了一个简单的生物启发想法,即神经元同步,甚至不一定是生物学上合理的。大脑并不是所有神经元都连接在一起以同步。但这是我想鼓励人们做的研究。推销它很容易。我们从未担心被抢先,那种压力完全消失了。所以我们没有急于推出这个想法,因为觉得可能别人也在做同样的事。我认为我们能获得焦点奖是因为我们打磨出了一篇精致的论文。我们花时间做好科学,得到想要的基线,尝试所有任务。鼓励研究者多冒险,尝试这些更投机、长期的想法,可悲的是我觉得这并不难推销。我想让 CTM 成为它有效的典范。它有点冒险,我们不知道会不会发现有趣的东西,但第一次尝试就找到了,并成为一篇成功的论文。
The CTM, the Continuous Thought Machine, is actually not that far outside of the local minimum that we're stuck in. Right? It's not as if we went and found this completely new technology. We took a simple biologically inspired idea, right, of the fact that neurons synchronize, and not even necessarily in a biologically plausible way. Brains don't literally have all their neurons wired together in a way that they work out their synchronization. But it's the sort of research that I want to encourage people to do. And the way to sell it is quite easy. I think at no point did we have to worry about being scooped, right? That stress was taken away from us completely. So there was no pressure to rush out with this idea because we thought, well, there's probably somebody else working on exactly this. And I think the reason we were able to get a spotlight is because we were able to create such a polished paper. We took the time to do the science properly, to get the baselines that we wanted and do all the tasks that we wanted to try. Encouraging researchers to take a little bit more of a risk, to try these slightly more speculative long-term ideas, I think the sad thing is I don't think it's necessarily a very difficult thing to sell. And I want to have the CTM as a poster child of it works, right? It was a bit of a risk. We didn't know if we were going to find something interesting, but it was our first shot and we did find something interesting and it became a successful paper.
如果我们找到一个能获取知识、设计新架构、做你所说的开放式科学的系统,你能看到未来进步主要由模型本身驱动吗?
If we do find a system which can acquire knowledge, design new architectures, do the open-ended type of science that you're speaking to, can you see a future where at some point the locus of progress will be mostly driven by the models themselves?
我想是的。至于是否完全取代我们,我反复思考。强大的算法正在帮助我们做研究。我认为它可能只是更强大的版本。我们发布的 AI 科学家展示了端到端的过程:从给系统一个研究论文的想法开始,然后放手让它运行。思考想法、写代码、运行代码、收集结果、写论文。我们最近甚至让一篇 100% AI 生成的论文被一个研讨会接收。但我们这样做是为了展示在真实系统中能做到。我希望它更互动:先给一个想法,然后它带回更多想法,和我讨论,再去写代码。我想看代码并检查,然后讨论结果。这就是我设想的近期未来,或者我想如何用 AI 做研究。
I think so. Whether or not that's going to replace us completely, I go back and forth on. Powerful algorithms are helping us do research, right? And I think it might just end up being a more powerful version of that. So I know the AI scientist that we released, we showed that you could actually go end to end, right? Go from seeding the system with an idea for a research paper and then just take your hands off and just let it go. Think about the idea, write the code, run the code, collect the results, and write the paper. To the point that we were actually able to get a 100% AI generated paper accepted to a workshop recently, right? But I think we did that to show that you could do it as a sort of demonstration in a real system. I think I would want it to be much more interactive. I would want to be able to seed with an idea and then have it come back with more ideas, have a discussion with me, then go away to write the code. I want to look at the code and check it and then discuss the results as they're coming out. So that's the sort of near-term future that I would envision or how I would like to do research with an AI.
你能反思一下吗?是因为你觉得我们需要监督,因为模型还不理解?你知道路径依赖的想法。我们需要监督,因为我们有路径依赖,可以引导语言模型的生成。也许未来语言模型自己会理解得更好。但还有输出维度:我们希望产生扩展人类兴趣的产物。我们希望它与人类相关。
And could you introspect on that? Is it because you feel we need supervision because the models don't yet understand? You know there's this path dependence idea. So we need to do supervision because we have the path dependence so we can guide the generation of the language models. Maybe in the future the language models will just understand better themselves. But there's also the output dimension which is that we want to produce artifacts that extend the fogyny of human interest. We want it to be human relevant.
是的。我认为更多是在初始想法中,可能无法精确描述你想要什么。就像我带实习生时,不能直接说有个疯狂想法,解释完就让他们自己干四个月。需要来回沟通,因为我有特定的探索方向,需要不断引导他们朝我最初的想法走。所以基本上就是这样。你有深刻的理解,有丰富的来源、历史和路径依赖,这意味着你能采取创造性的、直觉的步骤,尊重那种模糊性。他们尊重你拥有的深层抽象理解,而实习生还没有。
Yeah. I think it's more that in that initial seed idea it's probably impossible to actually describe exactly what you want. It's exactly the same with, you know, when I have an intern. I can't just have an intern come into the company and I go, I have this mad idea and then just explain it to them and then just leave them alone for 4 months. There's a back and forth because I have a particular idea that I want to explore and I need to keep steering them in the direction that I had in my mind originally. So I think it's more like that basically. You have such a deep understanding. So you have this rich provenance and history and path dependence and that means you can take creative steps, intuitive steps for you respect the fogyny. They respect all of this deep abstract understanding that you have and interns don't yet have that.
但也许未来的 AI 模型会拥有那种理解。
But maybe AI models in the future will have that.
当然。如果它们到了我的输入变得有害的地步,那就会那样。就像国际象棋,曾经人机结合能打败纯引擎,但现在不是了。加入人类反而让机器更差。
Yeah sure. If they get to the point where my inputs becomes detrimental then yeah that'll be a thing. It's kind of like chess, right? There was a point at which chess engine and human fusion actually beat chess engines. That's not true anymore, right? Adding a human into the mix actually makes the bots worse.
哦,有意思。我之前不知道。
Oh, interesting. I wasn't aware of that.
是的。所以当那天到来时,AI 科学家该怎么做是一个更广泛的讨论。
Yeah. So what to do when that day comes for AI scientists is a broader discussion.
我觉得现在是时候更详细地讨论这篇论文了。这个连续思维机器你之前提到过。卢克,首先介绍一下你自己,然后给我们讲讲这个。
I think now is a good segue to talk about this paper in a little bit more detail. So this Continuous Thought Machine you were just pointing to it before. Luke, first of all, introduce yourself and set this thing up for us.
我叫卢克,是 Sakana AI 的研究科学家,主要研究方向就是这个连续思维机器。整个团队花了大约八个月时间。我做了很多工作,但也有很多人在不同领域做不同部分。我觉得一篇论文八个月的周期在现在的 AI 研究中似乎有点长。但回到技术点,我们称之为连续思维机器。
My name is Luke. I am a research scientist at Sakana AI and my primary sector of research is this Continuous Thought Machine. It took us somewhere in the region of about eight months working on this project with the whole team. I did a lot of the work but we also had a lot of people in different areas and doing different parts of it. I think an 8-month life cycle for a paper seems a bit long for AI research at the moment. But yes, to the actual technical points of the paper. So we call it Continuous Thought Machine.
它最初有个不同的名字。我们之前叫它异步思维机,但每次人们问异步是什么意思时,都变得有点困惑。所以连续思维机基本上依赖于三个创新点。第一个是我们所谓的内部思维维度,这并不完全是新东西。它在概念上与潜在推理的思想相关。本质上是在一个序列维度上应用算力。当你开始在这个领域和框架中思考想法和问题时,你会开始理解许多看似智能的问题解决方案往往具有序列性质。例如,我们在连续思维机上测试的主要任务之一就是迷宫求解任务。对于深度学习来说,解迷宫是相当简单的。如果你让任务对机器容易,就很容易做到。其中一种方法是给神经网络(比如卷积神经网络)一张迷宫图像,它输出一张同样大小的图像,其中没有路径的地方是 0,有路径的地方是 1。有一些非常出色的工作展示了如何仔细训练这些网络并基本上无限扩展它们。这很迷人,也是一个非常有趣的解决方案。然而,当你抛开这种方法,问一个更人性化的方式来解决这个问题时,它就变成了一个序列问题。你必须说向上、向右、向上、向左等等,从起点到终点画出一条路线。当你约束这个简单的问题空间,并要求机器学习系统像那样解决它时,结果发现它变得更具挑战性。所以这成了我们 CTM 的 Hello World 问题,应用内部序列思维维度就是我们解决它的方式。另外两个创新点我们可以谈谈。我们重新思考了神经元应该是什么。在这个领域有很多优秀的研究,特别是在认知神经科学中,探索生物系统中神经元的工作方式。然后我们看看另一边,深度学习神经元的工作方式,典型的例子是 ReLU。它在某种意义上要么关闭要么开启。这种对大脑神经元的高度抽象感觉有点短视。所以我们处理这个问题时说,好吧,在逐个神经元的基础上,让这个神经元本身成为一个小模型。这最终在构建系统动态方面做了很多有趣的工作。第三个创新点如前所述,我们有一个内部维度,思考在其中发生。我们问:表征是什么?生物系统在思考时的表征是什么?仅仅是神经元在任何给定时刻的状态吗?那能捕捉到一个想法吗?如果我可以有争议地使用思考和想法这些词,我的哲学是:不,不能。想法的概念是随时间存在的。那么,如何用工程语言来捕捉它呢?我们不测量循环模型的状态,而是测量神经元如何成对地与其他神经元同步。这为我们可以用这种表征做的大量事情打开了大门。
It originally had a different name. We called it asynchronous thought machines before, but every single time people asked us what the asynchronous part it became a bit confusing. So continuous thought machines basically depends on three novelties. The first one is having what we call an internal thought dimension, and this is not necessarily something new. It's related conceptually to the ideas of latent reasoning. And it's essentially applying compute in a sequential dimension. And when you start thinking about ideas and problems in this domain and in this framework, you start understanding that many problems that look like intelligent solutions are often solutions that have a sequential nature. So for instance, one of the primary tasks that we tested in the continuous thought machines was this maze solving task. And solving mazes for deep learning is quite trivial. It's really easy to do if you make the task easy for machines. And one of the ways to do this is you give an image of a maze to a neural network like a convolutional neural network and it outputs an image same size of the maze and it's zeros where there isn't a path and ones where there is a path. There's some really brilliant work showing how you can train these in a careful way and scale them up essentially indefinitely. And this is fascinating and a really interesting idea of how to solve this. However, when you take that approach out of the picture and you ask what is a more human way to solve this problem, it becomes a sequential problem. You have to say well go up, go right, go up, go left, whatever the case may be to trace a route from start to finish. And when you constrain that simple problem space and you ask a machine learning system to solve it like that, it turns out to actually get much more challenging. So this became our hello world problem for the CTM and applying an internal sequential thought dimension to this is how we went about solving this. Two other novelties that we can touch on and talk about. We sort of rethought the idea of what neurons should be. There is a lot of excellent research in this world, in cognitive neuroscience particularly, exploring how neurons work in biological systems. And then we get on the other side of the scale how deep learning neurons work, which the quintessential example is a ReLU. It's off or on in a sense. And this very high level abstraction of neurons in the brains feels a little bit myopic. So we approached this problem and said well let's on a neuron by neuron basis let this neuron be a little model itself. And this ended up doing a lot of interesting work on how to build dynamics in the system. The third novelty here is as I said before we have this internal dimension over which thinking happens. We ask the question, well, what is the representation? What is the representation for a biological system when it's thinking? Is it just the state of the neurons at any given time? Does that capture a thought, if you wish, if I can be controversial and use the term thinking and thought and my philosophy with this is no, it doesn't. That the concept of a thought is something that exists over time. So, how do we capture that in engineering speak? We instead of measuring the states of the model that is recurrent, we measure how it synchronizes how neurons synchronize in pairs along with other neurons. And this opens up the door to a huge array of things that we can do with this type of representation.
你刚才谈到了推理的序列性质,我来唱个反调。我是说,有篇 Anthropic 的生物学论文,他们谈到了规划和思考,他们说这个东西在提前规划,因为我认为你的系统实际上可以说是在做规划,它在计算上确实不同,你能解释一下吗?
You were talking about this sort of sequential nature of reasoning and devil's advocate. I mean there was that Anthropic biology paper and they were talking about planning and thinking and they were saying that this thing is planning ahead because I think your system actually we can say does planning it's actually different computationally can you explain that?
是的,我认为从图灵机的角度来看计算边界非常有趣,因为能够写入磁带、从磁带读取然后再次写入以实现图灵完备系统的概念显然是一个改变了世界的惊人想法。我认为 Transformer 与我们用 CTM 试图做的事情之间的主要区别在于,CTM 思考的过程,我们可以应用那个内部过程来分解问题。所以问题本身可能有一个单一的解决方案,你可以一次性完成。就像我解释的迷宫,你可以一次性处理,但某些问题的表述方式使得一次性解决变得指数级困难。在迷宫任务中,一个很好的例子是,如果你试图一次性预测路径上的 100 步或 200 步,我们训练的任何模型,甚至我们的模型都做不到。我们需要构建一个自动课程系统,模型先预测第一步,当它能预测第一步时,我们再开始训练它预测第二步、第三步和第四步。由此产生的行为变得有趣。我喜欢做研究的方式,也是我鼓励与我一起工作的人做研究的方式,是理解模型的行为。我们现在已经到了一个地步,我们构建的模型明显具有智能,其方式不断让我们惊讶,而将其分解为一组指标甚至一个有限的性能指标对我来说可能不是正确的方法。理解这些模型在系统中以特定方式训练时所采取的行为和行动,似乎更能揭示引擎盖下实际发生的事情。
Yes, I think the boundary in terms of computation from a Turing machine perspective if you wish is really interesting because the notion of being able to write your tape, read from a tape and then write again to be a Turing complete system is obviously an incredible idea that has completely changed the world. And I think the primary difference with let's talk about transformers versus what we're trying to do with the CTM is that the process that the CTM thinks in, we can apply that process, that internal process to breaking down a problem. So the problem itself can be a single there is a single solution to this problem and you could do that in one shot. You could as I explained with the maze you could just process that in one shot but there are certain phrasings of problems that are real problems that doing so becomes exponentially more challenging. So in the maze task, a really good example is that if you try to predict 100, 200 steps down the path in one shot, no models that we could train, not even our model could do that. And we needed to actually build an autocurriculum system where the model first predicted the first step and then when it could predict the first step, then we started training it on the second and third and fourth step. And the sort of resultant behavior of this is where it gets interesting. One of the ways that I like to do research and that I encourage people who work with me to do research is understand the behavior of a model. We're getting to a point now where the models that we build are demonstrably intelligent in ways that keep surprising us and breaking that down into a single set of metrics or even a finite single metric about performance seems maybe not to be the right way to do it for me. And understanding the behavior and the actions that those models take when you put them in a system and train them in a certain way seems to reveal more about what's actually going on under the hood.
非常酷。我想我没注意到这一点。所以你是在做固定数量的步骤,所以有一个上下文窗口,你说你把它设在大约 100 步?
Very cool. And I think I didn't pick up on this. So you're doing a fixed number of steps so you have like a context window and did you say that you've set that around 100 steps?
对于迷宫任务,模型在每个步骤都观察完整的图像。CTM 会吸收观察完整的图像,为了论证,这些图像可以是语言的 token、语言模型的输出,这些输入可以是模型需要排序的数字等等,它应该对数据不可知,这就是我们试图构建它的方式。但在迷宫任务中,模型可以连续地观察数据,无论它在哪里,它可以同时看到整个图像,但它使用注意力从数据中检索信息,并且它有大约 100 步可以思考。我们所做的是,在某个时刻,模型解决了迷宫中的三步。所以它说:我要向上、向上、向右。
So for the maze task, the model always observes the full image at every step. The CTM will absorb observe the full image for argument sake those images could be tokens from a language, the output of a language model, those inputs could be numbers that the model has to sort whatever the case may be it should be agnostic to data that's how we've tried to build it. But in the maze task, the model can continuously just observe the data no matter where it can look at the whole image simultaneously but it uses attention to retrieve information from the data and it has let's call it 100 steps that it can think through. And what we do is we pick up at some point the model solves three steps through the maze. So it says I'm going to go up, up, and right.
然后它是正确的。但接着它走错了方向。此时我们停止监督,只训练它解决第四步——比它原本能做的多一步。实践中我们做五步,但原理不变。这样做就形成了一个自我引导机制。我想直觉敏锐的听众会明白这如何扩展到其他领域,比如语言预测中的多步预测等序列任务。
And then it's correct. But then it makes the wrong turn. And at that point, we stop supervision. We only train it to solve the fourth step. So one more than what it could. In practice, we do it five, but the principle holds. And when you do that, it's a self-bootstrapping mechanism. And I think the intuitive listener will understand how that extends to other domains, other sequential domains for instance like language prediction, many tokens ahead, that sort of thing.
我对自适应计算这个概念很感兴趣。第一个问题是,性能对步数有多敏感?第二个问题是,能否有任意数量的步数——比如基于不确定性或某种标准,你可以用更少的步数?最后一个问题是,能否有任意或无限数量的步数?
So I'm really interested in this idea of adaptive computation. So I guess the first question is how sensitive was the performance to the number of steps and then the next question would be could you have an arbitrary number of steps which means that you know perhaps based on uncertainty or some kind of criterion you could do fewer steps and then the final question is could you have potentially like an arbitrary or unbounded number of steps?
非常好的问题。我先回答关于步数敏感性的不确定性部分。一个很好的例子是我们在 ImageNet 分类上训练模型,损失函数很简单:我们运行 50 步,选取两个不同的点——第一个是性能最佳点(损失最低),第二个是最确定点(置信度最高)。这两个点给出 0 到 49 之间的两个索引,我们在两个点都应用交叉熵,损失取这两个交叉熵的平均值。这样会诱导出这样的行为:简单样本几乎在一两步内解决,而困难样本自然需要更多思考,使模型能够自然地利用所有可用时间,而无需强制。
Yeah, really super question. I think that I'll answer the uncertainty question first about the sensitivity to steps. So a very good example of this is we just trained the model on ImageNet classification and our loss function is quite simple. What we do is we run it for, for example, 50 steps and we pick up two distinct points. The first one is where is it performing the best, i.e. where is the loss the lowest, and the second one is where is it most sure or where is it most certain, and those give us two indices between 0 and 49 inclusive, and we apply cross entropy at both of those points; we just make the loss the average of the cross entropy at those points. So what this does is it induces a behavior where easy examples are solved almost immediately in one or two steps, whereas more challenging examples will naturally take more thinking, and it enables the model to use the full breadth of time that it has available to it just in a natural fashion without having to force it to happen.
你们决定将每个神经元建模为 MLP,这非常有趣。请谈谈这一点,还有同步的概念——你们用内积来确定参数同步的程度,这随时间展开作为驱动力。能更详细解释一下吗?
So you've decided to model every neuron as an MLP which is really fascinating. Talk about that but also there's this notion of synchronization and I think you use the inner product to determine the extent to which the parameters are synchronized and this kind of unfolds over time as the driving force. Can you explain that in a bit more detail?
当然。我认为先解释论文中所谓的神经元级模型(NLM)是个好主意,因为它与此相关。想象一个循环系统,它是一个状态向量,每一步都在更新。我们跟踪这个状态向量,它随时间展开,对于系统中的每个神经元(第 i 个神经元),我们有一个展开的时间序列——它是连续的(实际上是离散的,但值是连续的)。这些时间序列定义了所谓的随时间变化的激活。同步很简单,就是测量两个时间序列的点积。因此,在一个有 d 个神经元的系统中,有 d 的平方除以 2 个不同的同步对。神经元 1 和神经元 2 可以通过它们的同步方式关联,神经元 1 和神经元 3 也是如此。神经元级模型通过接收有限的历史(比如固定窗口的神经元激活)来工作,它们不是直接使用原始激活,而是利用这段历史作为信息来处理单个输出激活,这就是从所谓的预激活到后激活的过程。这看起来可能有些随意,它对性能有帮助吗?事实证明是的,但这并不是万能的解决方案,也不是我们的目标。我们的目标是尝试做一些生物上合理的事情。在生物学(大脑如何在生物基质中实现事物)和深度学习(高度并行化、学习极快、可反向传播等所有让我们走到今天的美好特性)之间找到一条线,在这条线上我们可以汲取一些生物学灵感,但仍然用深度学习来训练。事实证明,神经元级模型是一个很好的中间方案。同步的概念应用于这些神经元级模型的输出之上。
Absolutely. I think it's a good point to explain the neuron-level models as we call them in the paper, or NLM, first because it ties into this. So you can imagine a recurrent system is a state vector that is being updated from step to step. We track that state vector and that state vector unfolds, and for each individual neuron, each i-th neuron in the system, we have an unfolding time series. It's a continuous time series. Well, it's discrete, but it's a continuous value. And those time series define what we call the activations over time. And synchronization is quite simply just measuring the dot product between two of these time series. So you have a system of d neurons and essentially you have d over two squared different synchronization pairs. So neuron one can be related to neuron two by how they synchronize, and neuron one can also be related to neuron three, etc. The neuron-level models function by taking in a finite history like a fixed window of neuron activations coming in, and instead of just being a raw activation, they use that history as information to process a single activation out, and that is what moves from what we call pre-activations to post-activations. And the principle here is that this might seem rather arbitrary and does it help for performance? Turns out it does, but that's not really the catch-all solution here. That's not what we're after. What we're after here is trying to do something biologically plausible. Find the line somewhere between biology, which is how the brain implements things in the biological substrate that we have, versus deep learning, which is highly parallelizable, super fast to learn, backpropable, all of the nice properties that have got us this far, and find a line somewhere where we can take some sprinkling of biological inspiration but still train it with deep learning. And it turns out that neuron-level models is a nice interim that we can do this with. The concept of synchronization is applied on top of the outputs of those neuron-level models.
关于扩展,我认为时间复杂度相对于同步矩阵的维度是二次的,对吧?在你们的论文中,你们提到用子采样来提升性能,但这如何影响稳定性?这样做有什么代价吗?
On the scaling, I think the time complexity is quadratic with respect to the dimension of the synchronization matrix, right? And in your paper you were talking about subsampling to improve the performance, but how did that affect the stability? Were there any things that cost you doing that?
好问题。关于稳定性,我们发现了一件有趣的事——这也是我们在论文实验中贯穿始终的感受:无论我们在什么任务上尝试,它似乎都能在各种超参数设置下正常工作。而通常使用 RNN 和 LSTM 等循环模型进行时间反向传播时,你会遇到问题:运行很多内部时间步后,学习似乎会崩溃。但因为我们使用了同步,它在某种意义上触及了所有时间步的所有神经元,所以非常有助于梯度传播。一个可能与你问的同步问题有点偏离的有趣点是:我们有一个 d 个神经元的系统,如前所述,有 d 的平方除以 2 种可能的组合。这实质上意味着系统的底层状态或表示比仅取那 d 个神经元要大得多。至于这对下游计算、性能以及我们能做什么意味着什么,正是我们目前积极探索的方向。
Yeah, it's a neat question. I think in terms of stability, what we found was kind of fun, and this was a sentiment that we had throughout the experiments that we ran with this paper: it tended, no matter what we tried it on, it just kind of worked with all spreads of hyperparameters. And the problems that you have with backprop through time typically with recurrence models like RNNs and LSTMs — it's a challenge and you run for many internal ticks with the RNNs or the LSTMs and the learning seems to break down — but the fact that we use synchronization in some sense touches all of the neurons through all of the time, so it really helps with gradient propagation. A nice interesting point that's maybe a bit oblique to what you asked about synchronization is we have a system of d neurons and like I said earlier, there are d over two squared possible combinations. This essentially means that our underlying state or underlying representation to the system is quite a lot larger than what you would get with just taking those d neurons. And as to what that means in terms of downstream computation and performance and the things that we can do with this is what we're actively exploring right now.
你们用了指数衰减率。
You guys used an exponential decay rate.
系统随时间展开。如果任意两个神经元之间的同步依赖于相同的时间尺度,那可能会过于受限。例如,大脑中的有些神经元在很长的时间尺度上放电,有些则在很短的时间尺度上。它们一起放电的方式会影响其他神经元并使其放电。但生物大脑中的一切都在不同的时间尺度上发生,这就是为什么不同思考状态对应不同脑电波。除此之外,我们在连续思维机中使用指数衰减,可以允许非常急剧的衰减,从而表明这两个配对在一起的神经元,真正重要的只是它们此刻如何一起放电。但如果衰减非常缓慢,那么本质上捕捉的是这些神经元在极长时间内的全局放电模式。
You have the system that unfolds over time. It would be maybe a little bit too constrained if the synchronization between any two neurons depended on the same time scale. So for instance, there are neurons in your brain that are firing over very long time scales and very short time scales. The way that they fire together impacts other neurons and causes those neurons to fire. But everything in biological brains happens at diverse time scales. It's why we have different brain waves for different thinking states, for instance. But beside that point, what we do with the exponential decay in the continuous thought machines is it allows us for a very sharp decay to say that these two neurons that are pairing together, what only really matters is how they fire together right now. Right? But if we had a very long and slow decay, essentially that's capturing a global sense of how those neurons are firing over an extremely long period of time.
所以这本质上是一种捕捉不同神经元如何快速同步放电、而其他神经元缓慢放电或不放电的方式。这让我之前提到的那个表示空间——D 除以 2 的平方表示空间——变得更加丰富,我们可以通过更精细地调整计算这些表示的方式来丰富这个空间。我们昨天聊过,Luke,当人们把 Transformer 应用到像 ARC 挑战这类需要推理的任务时,需要做很多特定领域的 hack。比如去年挑战的获胜者使用了深度优先搜索采样,有些人也在尝试用语言表示或 DSL。部分原因与语言的可达性有关,对吧?语言非常稠密,这意味着你可以单调地增加。但如果我没理解错,你的系统在推理、离散和稀疏领域以及样本效率方面可能有一些有趣的特性,因为我们想构建一个能在 ARC 挑战这类任务上表现良好的系统。你能用简单的语言解释一下,为什么你认为这种架构可能比 Transformer 好得多?
So this was essentially a way of capturing this idea of how different neurons could fire together very quickly and other neurons could fire together very slowly or not at all. And this lets that representation space that I spoke about, that D over 2 squared representation space, become more rich, and we can enrich that space with more subtle tweaks to how we compute those representations. We were speaking about this yesterday, Luke, that when folks apply transformers to things like the ARC challenge or things that need reasoning, we need to do lots of domain-specific hacks. So the architects who were the winners of last year's challenge, they did depth-first search sampling, and some folks have been experimenting with using language representations or using DSLs. And some part of this is to do with the reachability of language, right? Language is quite dense, which means you can kind of monotonically increase. But if I understand correctly, your system might have some interesting properties for reasoning and for discrete and sparse domains, and also for sample efficiency, because we want to build a system that can actually do well on things like the ARC challenge. But can you explain in simple terms why you think this architecture could be significantly better than transformers for doing those things?
我认为过去几年语言模型文献中真正令人着迷的工作,很多都与一个可以称为新缩放维度有关。在某种意义上,我把思维链推理视为向系统添加更多算力的一种方式。这显然只是它真正含义的一小部分。但我认为这在某种意义上是一个相当深刻的突破。现在我们试图做的是让那个推理组件完全内部化,但仍然以某种顺序方式运行。我认为这相当重要。你之前提到 Gemini 的扩散语言建模,我认为现在有很多不同的方向在探索这个。我确实认为,连续思维机结合同步和多层次时间表示的概念,在那个空间上提供了其他人尚未探索的灵活性。那个空间的丰富性——能够投影下一步来解决 ARC 挑战,以及接下来的 100 步、200 步,将其分解成一个过程,然后模型可以在其高维潜在空间中快速搜索这个过程——这感觉像是一个很好的方法。
I think a lot of the really fascinating work in the last few years that I found fascinating in the literature of language models has been related to what one can actually call a new scaling dimension. I in some sense see chain-of-thought reasoning as a way of adding more compute to a system. That's obviously just one small part of what that really is and what that really means. But I think it's quite a profound breakthrough in some sense. Now what we're trying to do is have that reasoning component be entirely internal yet still running in some sort of sequential manner. And I think that that's rather important. And you spoke earlier about Gemini's diffusion language modeling, and I think that there are a lot of different directions that are exploring this right now. I do think that the continuous thought machine with the ideas of synchronization and multi-hierarchical temporal representations gives a certain flexibility on that space that other people are not yet exploring, and that richness of that space being able to project the next step to solve the ARC challenge and the next 100, the next 200 steps to be able to break that down into a process that a model can then very quickly search that process in its high-dimensional latent space becomes something that feels like a good approach to take.
你认为这种架构和 Alex Graves 的神经图灵机之间有关系吗?
Do you see any relationship between this architecture and Alex Graves' neural Turing machine?
是的,这很有趣。我确实认为。我认为使用神经图灵机最困难的部分之一是写入和读取内存的概念,因为它是一个离散动作。这有其自身的挑战。而且,我不会说连续思维机肯定接近图灵完备,但它的概念是在潜在空间中进行推理,并让那个空间以丰富的方式展开以适应不同的任务。这实际上引出了我觉得很有趣的一点,想和你分享。再考虑一下 ImageNet 任务或任何分类任务。这是一个很好的测试平台。有很多图像非常简单,也有很多图像非常困难。当我们训练比如 ViT 或 CNN 来做这个任务时,它必须将所有推理嵌套在同一个空间中。它必须将决策过程——从一只非常明显的猫到一些复杂奇怪的、代表性不足的类别——全部嵌套在那个系统、那个数据集中,并且必须以并行的方式嵌套,直到最后一层然后分类。我认为将其分解,在不同的时间点你可以说“我完成了,可以停止”与“我完成了,可以停止”,这让你能够将一个数据集或任务自然地分割成从简单到困难的组成部分。而且我们知道课程学习和这种连续意义上的学习似乎是个好主意。这是人类学习的方式。如果我们能在架构上实现这一点,并让模型自然地表现出这种特性,这似乎值得探索。我不确定你是否了解模型校准以及神经网络往往校准不良的问题。
Yes, that's really interesting. I do. I think that one of the most challenging parts about working with a neural Turing machine is the concept of writing to memory and reading from memory because it is a discrete action. And that has its own challenges associated with it. And yes, I wouldn't go so far as to say that the continuous thought machine is definitively nearing Turing completeness, but the notion of doing reasoning in a space that is latent and letting that space unfold in a way that is rich towards a different set of tasks. And this actually brings me to a point that I find quite interesting that I'd like to share with you. Consider again the ImageNet task or any sort of classification task. It's a nice test bed. There are many images that are really easy and many images that are really difficult. When we train, for instance, a ViT or a CNN to do this task, it has to nest all of that reasoning in the same space. It has to put all of its decision-making process for a very simple obvious cat versus some complex weird underrepresented class in that system, in that dataset, and it has to nest it all in parallel in a way that we get to the last layer and then we classify. I think breaking that down where you have different points in time where you can say 'now I'm done, I can stop' versus 'now I'm done, I can stop' lets you take a dataset or take a task and actually naturally segment it into its easy to difficult components. And I think we know that curriculum learning and learning in this continuous sense again seems to be a good idea. It's how humans learn. And if we can get at that architecturally and just have that fall out in a model, again, this seems like something worth exploring. I'm not sure if you know much about model calibration and how neural networks tend to be poorly calibrated.
哦,请讲,Tommy。这是一个有点老的发现,但如果你训练一个神经网络足够久,它拟合得非常好,正则化也做得非常好,你会发现模型是未校准的,这本质上意味着它对某些错误类别非常自信,而对某些正确类别却不自信。对于一个完美校准的模型,你希望如果它预测某个类别正确的概率是 50%,那么 50% 的情况下它应该正确,以此类推。所以一个校准良好的模型,如果它预测是猫的概率是 0.9,那么 90% 的情况下它应该是正确的。而实际上,大多数训练足够久的模型都会变得校准不良。有很多事后技巧来解决这个问题。我们在训练后测量了 CTM 的校准度,发现它几乎完美校准,这又是一个有力的证据,表明这似乎是一种更好的方法。这类研究的特点是,我们并没有刻意去创建一个校准良好的模型,对吧?我们甚至没有试图创建一个能够进行某种自适应计算时间的模型,对吧?我非常喜欢 Alex Graves 那篇关于自适应计算时间的论文,是吗?但那篇论文中有大量的超参数扫描,因为在那篇论文中,他需要对计算量施加一个损失。
Oh, go for it, Tommy. It's a bit of an old finding, but if you train a neural network for long enough and it fits really really well and you've regularized it really really well, you'll find that the model is uncalibrated, which essentially means that it is very certain about some components where some classes where it's wrong and uncertain for some classes where it's correct. Essentially what you want for a perfectly calibrated model is if it predicts a probability that this is in the correct class with 50%, 50% of the time you want it to be correct about that class, and so on and so forth. So a well-calibrated model if it's predicting a probability of 0.9 that it is a cat, then 90% of the time it should be correct. And it actually turns out that most models that you train for long enough get poorly calibrated. And there are loads of post-hoc tricks to fixing this. We measured the calibration of the CTM after training and it was nearly perfectly calibrated, which is again a little bit of a smoking gun that this actually seems to be probably a better way to do things. The flavor of this kind of research is such that we didn't actually go out and actually try to create a very well-calibrated model, right? And we didn't even try to create a model that was necessarily going to be able to do some kind of adaptive computation time, right? I was a very big fan of the paper on adaptive computation time by Alex Graves, was it? But that paper had a massive amount of hyperparameter sweeps in it because in that paper he needed to have a loss on the amount of computation that was being done.
因为任何时候你尝试做某种自适应计算时间的研究,你都在对抗神经网络是贪婪的这一事实,对吧?因为显然,获得最低损失的方式是使用你所能获得的所有算力。
Because anytime you try to do some sort of adaptive computation time research, what you're fighting is the fact that neural networks are greedy, right? Because obviously the way to get the lowest loss is to use all the computation that you have access to.
所以,除非你有一个额外的损失函数,带有惩罚项说‘好了,你不允许使用所有算力’,并且非常小心地平衡损失,那篇论文中的模型才会出现有趣的动态计算时间行为。但看到连续思维机(Continuous Thought Machine)时,真的很令人欣慰,因为我们设置损失函数的方式(Luke 之前描述过),自适应计算时间似乎自然而然地就出现了。所以我认为研究更应该是这样。
So unless you had an extra loss that had a penalty saying, 'Okay, you're not allowed to use all the computation,' and very carefully balanced loss, that's when you actually got the interesting dynamic computation time behavior falling out of the model in that paper. But it was really gratifying to see with the Continuous Thought Machine that because of the way we set up the loss that Luke described earlier, adaptive computation times seem to just fall out naturally. So that's more the way I think research should go.
没错,因为我们实际上并没有一个具体的目标,比如要解决某个特定问题,或者要发明什么东西。更多的是我们有了这个有趣的架构,然后我们只是顺着‘有趣性’的梯度走。
Okay, because we don't actually have a specific goal, like a specific problem we're trying to fix, or something we're trying to invent. It's more that we have this interesting architecture and we're just following the gradients of interestingness.
是的。关于这一点,我认为你的论文最令人兴奋的地方在于,我们谈到了路径依赖性,以及这种逐步构建的理解,这种复杂化的过程。我的意思是,这或许很契合世界模型的主题,也契合主动推理——我特意给‘主动推理’加了引号,因为它不是卡尔·弗里斯顿的主动推理,也许叫自适应推理之类的——但我们想要构建能够持续学习、能够更新参数、最重要的是能够构建路径依赖性理解的智能体。因为这完全不同于仅仅理解事物是什么。你如何到达那里才是非常重要的。而这种架构可能允许这些智能体使用这种算法探索空间中的轨迹,找到最佳轨迹,并实际构建一种按照关节来划分世界的理解。
Yes. And on that point, I think maybe the most exciting thing about your paper is, you know, we were talking about path dependence and having this understanding which is built step by step, this process of complexification. And I mean maybe this is apropos in the theme of world models in general and also active inference—and I say active inference in big quotes because it's not Karl Friston's active inference, maybe adaptive inference or something like that—but we want to build agents that can continue to learn, that can update their parameters, and most importantly can construct path-dependent understanding. Because that's completely different to just understanding what the thing is. It's how you got there that is very important. And this architecture potentially allows these agents using this algorithm to explore trajectories in spaces, find the best trajectories, and actually construct an understanding which carves the world up by the joints.
是的,这是一个非常巧妙的视角。我实际上没有那样想过,但没错,我认为当你考虑模糊问题时,这种立场变得非常有趣,因为用一种方式划分世界和用另一种方式划分世界,性能是一样的。
Yeah, that's a really neat perspective. I haven't actually thought about it like that, but yes, I think that particular stance becomes really interesting when you think about ambiguous problems, because carving the world up in one way is as performant as carving it up in another way.
是的。你知道,语言模型中的幻觉可能是在以某种精细的方式划分世界,但在我们衡量‘这是幻觉’的标准下,它并不高效,而实际上那并不是真的。但在另一条路径上,通过自回归生成词元来划分世界,你会得到不同的世界划分。而能够训练一个模型,让它隐式地意识到自己实际上是在以不同的方式划分世界,并探索那些方式、那些划分的下降,正是我们所追求的。我认为这是一种非常令人兴奋的方法,采取‘让我们把这个问题分解成可解决的小部分,并学会那样做’的立场,以及我们如何以自然的方式做到这一点,而不需要太多技巧。
Yeah. You know, perhaps the hallucination in language models is carving the world up in some fine way, but it's just not performant in our measure of 'this is hallucination' and actually that's not true. But in some other trace down the path of wanting to carve the world up through an autoregressive generation of tokens, you end up in a different carve-up of that world. And being able to train a model that can be implicitly aware of the fact that it is actually carving up the world in a different way, and explore those manners, those descents down the carve-up, is something that we're after. And I think it's quite an exciting approach to try to take a stance of 'let's break up this problem into small solvable parts and learn to do it like that' and how can we do this in a natural way without too many hacks.
是的,这是我一直在思考的事情,因为尽管我很喜欢乔莱的智力衡量标准,但对他来说,适应新奇就是得到正确答案,而你给出那个答案的原因非常重要。在机器学习中,我们有这个问题:我们提出这种成本函数,反而导致了捷径问题。但你知道,我们可以构建一个符号系统,我们可以做老式人工智能,然后说‘好吧,我们需要做这种有原则的知识构建,保持语义。’但我们没有那样做。我们在做一个混合系统。但必须有一种自然的推理方式,尽管最终目标是这个成本函数,但由于我们遍历这些开放空间的方式,我们实际上可以在机制上更有信心,我们正在做与世界对齐的推理。
Yeah, it's something I've been thinking about because, as much as I love Chollet's measure of intelligence, for him adapting to novelty is getting the right answer, and the reason why you gave that answer is very, very important. And in machine learning, we have this problem that we come up with this kind of cost function that rather leads to this shortcut problem. But you know, we could just build a symbolic system, we could be GOFAI, and we could say, 'Okay, we need to do this principled kind of construction of knowledge maintaining semantics.' Well, we're not doing that. We're doing a hybrid system. But there must be some natural way of doing reasoning where, in spite of the end objective being this cost function, because of the way that we traversed these open-ended spaces, we can actually have more confidence mechanistically that we're doing reasoning which is aligned to the world.
我认为这是看待这一特定研究方向的一个很好的方式,而且显然我们不是唯一这样想的人,也不是唯一试图这样做的人。我们拥有的是一个适合它的架构,而且令人惊讶的是——这又不是目标。做这种类型的研究不是目标。能够把世界分解成这些我们可以以自然方式推理的小块也不是目标。相反,我们所做的是尊重大脑,尊重自然,然后说,‘好吧,如果我们构建这些受启发的东西,实际上会发生什么?会出现哪些不同的解决问题的方式?’然后当这些不同的解决问题的方式出现时,我们可以开始问哪些大的哲学和基于智能的问题?这就是我们现在所处的位置。所以有时可能会感觉,尤其对我而言,问题太多,而回答这些问题的人手太少。但我认为有趣、令人兴奋且鼓舞人心的事情是,我可以尝试鼓励其他年轻研究人员的是,做你充满热情的事情,想办法构建你在乎的东西,然后看看它会带来什么。看看它打开了哪些门,以及如何更深入地探索那些领域。
I think that's a great way of seeing this particular avenue of research, and I think that obviously we're not the only people thinking like this, and we're not the only ones trying to do this. What we have is an architecture that's amenable to it, and surprisingly so—it wasn't, again, the goal. It's not the goal to do this type of research. It's not the goal to be able to break the world down into these small chunks that we can actually reason over in a way that seems natural. Instead, what we did was pay respect to the brain, pay respect to nature, and say, 'Well, if we build these inspired things, what actually happens? What different ways of approaching a problem emerge?' And then when those different ways of approaching a problem emerge, what big philosophical and intelligence-based questions can we then start to ask? And that's where we're at right now. So it might feel at times, especially for me, too many questions and too few hands to answer those questions. But I think the fun and exciting thing and the encouraging thing that I can try to encourage other younger researchers out there is that, you know, do what you're passionate about and figure out how to build the things that you care about, and then see what that does. See what doors that opens up and see how to explore deeper into those domains.
我们昨天讨论过这个,对吧?你可以把语言看作一种迷宫。
We were talking about this yesterday, weren't we? That you can think of language as being a kind of maze.
是的。比如,有什么阻止我们采用这种架构并用它构建下一代语言模型呢?我的意思是,老实说,正如你所知,这是我目前正在积极尝试探索的事情。而且,我认为迷宫任务在加入模糊性时变得非常有趣,当有多种方法可以解决迷宫时。老实说,这我还没试过,也许下周我应该试试。但基本上,你可以想象一个智能体,或者这里的 CTM,观察迷宫并走一条轨迹。令人惊讶的是,我们看到了这个——在我们最近更新的 arXiv 论文中,也就是最终定稿版本,我们添加了一个额外的补充部分,不在主要技术报告中。那个补充部分基本上是说,‘嘿,我们看到了这些酷炫的事情发生’,我们列出了我认为 14 件在研究过程中发生的趣事,这些显然没有写进论文,但我们希望人们知道这些奇怪的事情。这就是其中一件奇怪的事:我们在训练过程中观察发生了什么。在训练的某个时候,也许在训练进行到一半时,我们可以看到模型会先走迷宫的一条路径,然后突然意识到‘哦不,该死,我错了’,然后回溯,再走另一条路径。但最终它变得非常好,并且在这个过程中进行了一些分布式学习,因为它有一个多头注意力机制。
Yes. Like, what is to stop us from taking this architecture and building the next generation language model with it? I mean, that's honestly, as you know, something that I am actively trying to explore right now. And yeah, I think the maze task gets really interesting when you add ambiguity to it, when there are many ways to solve the maze. And honestly, this isn't something I've tried yet, and maybe it's something I should try next week. But it's essentially you can imagine an agent, or the CTM in this case, observing the maze and taking a trajectory. And surprisingly, we saw this—we have a section in our recently updated paper on arXiv, the final camera-ready version of this paper, where we added an extra supplementary section that is not in the main technical report. And that supplementary section is basically, 'Hey, we saw this cool stuff happen,' and we list, I think, 14 different interesting things that happened while we were doing the research that obviously didn't make it into the paper, but we wanted people to know about these strange things that happened. And this is one of the strange things: we watched during training what was happening. And at some time during training, maybe halfway through the training run, we could see what the model would do is it would start going one path in the maze, and then suddenly it would realize, 'Oh no, damn, I'm wrong,' and would backtrack and then take another path. But eventually it gets really good, and it does some sort of distributed learning in this because it's got an attention mechanism with multiple heads.
所以它实际上可以很好地找出如何做到这一点并改进其解决方案。但在学习的早期阶段,它会探索多条路径,然后返回并回溯。我们有一组非常迷人的实验也展示了——实际上我们在网上有一些补充材料展示这一点——而且我不太确定这意味着什么,这有点像深刻的哲学问题,但如果你试图解决一个迷宫却没有足够的时间,结果发现有一种更快的算法可以做到。当我看到这一点时,我震惊了。所以如果我们限制模型的思考时间,但仍然让它尝试解决一个长迷宫,它不会逐段追踪迷宫,而是快速跳转到大致需要的位置,然后向后追踪并填充那条路径,然后再向前跳,跳过顶部并向后追踪那一段,然后再跳,它表现出这种迷人的跳跃行为,这是基于系统的约束。再说一次,这只是我们观察到的一个现象,从深层意义上讲这意味着什么,它与给模型思考时间与不给思考时间的关系,以及思考时间是否足够?会发生什么?当你以这种方式约束模型时,它会学习哪些不同的算法?我觉得这非常迷人,是一个值得探索的有趣问题。它是否告诉我们一些关于人类如何思考的信息?它是否告诉我们一些关于我们在受约束与开放环境下如何思考的信息?在这方面可以提出很多有趣的问题。
So it can actually just figure out how to do this pretty well and refine its solution. But sometime early on in the learning it descends multiple paths and comes back and backtracks. We have a really fascinating set of experiments that also showed—and this we actually have some supplementary material online showing this—where, and I don't really know what this says, it's kind of a deep philosophical thing, but if you're trying to solve a maze but you don't have enough time, turns out that there's a faster algorithm to do it. And this blew my mind when I saw it. So if we constrain the amount of thinking time that the model has but still get it to try solve a long maze, instead of tracing out that maze, what it does is it quickly jumps ahead to approximately where it needs to be and it traces backwards and it fills in that path backwards, and then it jumps forward again, leapfrogs over the top and traces that section backwards, and then leapfrogs, and it does this fascinating leapfrogging behavior that is based on the constraint of the system. And again, you know, this is just an observation we made, and what that means in a deep sense and how it's related to giving a model time to think versus not, and is it enough time to think? What happens? What different algorithms does the model learn when you constrain it in this way? I find that quite fascinating and an interesting thing to explore. Does it tell us something about how humans think? Does it tell us something about how we think under constrained settings versus open-ended settings? There are a number of cool questions you can ask on this front.
你们两位都是群体方法和集体智能的忠实粉丝,因为我们既可以向上扩展也可以向外扩展,那么向外扩展不仅意味着简单的并行化,还包括并行模型之间的某种权重共享等,这可能会带来什么?
You guys are both huge fans of population methods and collective intelligence, and because we can scale this thing up and we can scale it out, what would it mean to scale this thing out not only just in a kind of trivial parallelization but in terms of having some kind of weight sharing between parallel models and so on? What would that give you potentially?
这是一个有趣的研究领域。所以我们团队正在积极探索的一个方向是记忆的概念,特别是长期记忆,以及这对这类系统意味着什么。例如,可以设计一个实验:将一些智能体放入迷宫,让它们尝试解决这个迷宫,但不是像我们论文中那样做,而是在一个非常受限的环境中,智能体只能看到周围大约 5x5 的区域,我们给智能体一些保存和检索记忆的机制。任务就是解决迷宫,找到出口,模型需要学习如何构建记忆,以便能够回到之前见过的地方,知道‘我上次做错了’,然后选择不同的路线。然后你可以看到并行智能体在同一个迷宫中共享记忆结构,当它们都能访问那个记忆结构并拥有一个共享的全局记忆——几乎像文化记忆一样——时,会发生什么,通过许多智能体尝试使用这个记忆系统来解决这个全局任务。我确实认为记忆将是未来人工智能发展的一个关键要素。
This is a fun area of research. So one of the active things that we're trying to explore in our team is concepts of memory, long-term memory, and what this means for a system like this. So an experiment that one can construct, for instance, is to put some agents in a maze and let them try to solve this maze, not how we did it in the paper but in a very constrained setting where an agent can only see maybe a 5x5 region around it, and we give that agent some mechanism for saving and retrieving memories. The task, if you wish, is to solve that maze, find your way to the end, and the model needs to learn how to construct memory such that it can get back to a point where it's seen before and know 'I did the wrong thing last time' and go a different route. And you can then see this with parallel agents in the same maze with a shared memory structure, and see what actually happens when you can all access that memory structure and have a shared global—almost like a cultural memory—that we can access and solve this global task by having many agents trying to use this memory system. And I do think that memory is going to be a very key element to what we need to do in the future for AI in general.
刚才提到了推理这个话题,我认为最近我们在推理方面取得了很大进展,对吧?因为这实际上是人们正在研究的主要方向之一。我们最近发布了一个名为 Sudoku Bench 的数据集,我很高兴看到它几周前在你们的播客中被自然地提及。
So the subject of reasoning came up just a second ago, and I think there's a perception that recently we made a lot of progress in reasoning, right? Because it's actually one of the main things that I think people are working on. We released a dataset recently called Sudoku Bench, and I was actually quite happy to see it come up organically on your podcast a few weeks ago.
克里斯·摩尔,对吧?
Chris Moore, right?
是的。所以我想跟你聊聊这个基准测试,因为我觉得我在推广它时遇到了一些问题,因为它表面上听起来不是特别有趣,因为数独给人一种已经被解决的感觉,对吧?所以,一堆数独题对推理能有多大的意义?确实。我们说的不是普通数独,而是变体数独。变体数独通常是在普通数独的基础上——也就是在行、列和宫格中填入数字 1 到 9——再加上任意额外的规则。而且它们都是手工制作的,约束条件极其不同。这些约束实际上需要非常强的自然语言理解能力。例如,数据集中有一个谜题,它用自然语言告诉你谜题的约束条件,然后说,‘哦,顺便说一下,那个描述中的某个数字是错的。’对吧?所以你必须能够在开始解谜之前就对规则本身进行元推理。还有其他谜题,比如在数独上叠加了一个迷宫,老鼠必须沿着一条通往奶酪的路径走出迷宫,但路径上有约束,比如数字以及它们的和。很难真正描述这些变体数独有多么多样化,我认为它们如此多样化,以至于如果有人能击败我们的基准测试,他们必然已经创建了一个极其强大的推理系统。目前,最好的模型只能达到大约 15%的正确率,而且仅限于数据集中最简单、最小的数独题。我们即将发布一篇关于 GPT-5 性能的博客文章,性能确实有提升,但它仍然完全无法解决人类能解决的谜题。我非常喜欢这个数据集,实际上也是我创建它的最初催化剂,是因为安德烈·卡帕西说过一句话:‘好吧,我们从互联网上获得了所有这些数据,但如果你想要 AGI,你真正想要的不是人类创造的所有文本,而是他们在创造文本时脑海中的思维痕迹,对吧?如果你能从中学习,那么你就会得到真正强大的东西。’我心想,好吧,那些数据一定存在于某个地方。我首先想到的是哲学,比如有一种哲学流派,你只是写下你的想法而不加思考,就像意识流。我以为那可能行得通。但后来当我不再想这件事,在闲暇时间,我观看了一个名为 Cracking the Cryptic 的 YouTube 频道。
Yes. So, I wanted to tell you a little bit about this benchmark because I think I've been having a little bit of issue promoting it because it doesn't on the surface sound particularly interesting, because Sudoku has a sort of feeling that it's already been solved, right? So, how interesting can a collection of Sudokus be for reasoning? Exactly. We're not talking about normal Sudokus. We're talking about variant Sudokus. And what variant Sudokus are, are usually normal Sudokus, right? So put the numbers one to nine in the row, the column, and the box, but then literally any additional rules on top of that. And they're all handcrafted. They all have extremely different constraints. Constraints that actually require very strong natural language understanding. So for example, there's one puzzle in the dataset where it tells you the constraints of the puzzle in natural language and then says, 'Oh, by the way, one of the numbers in that description is wrong.' Right? So you have to be able to meta reason about the rules themselves even before you start solving the puzzle. There are other puzzles where you have a maze overlaid on the Sudoku and the rat has to work out a way through the maze by following a path to the cheese. But then there are constraints on the path that it takes, like what numbers and what they can be add up to. It's difficult to really describe how varied these variant Sudokus are, and I think they're so varied that if anyone was actually able to beat our benchmark, they would necessarily have to have created an extremely powerful reasoning system. Right now, the best models get around 15%, but only on the very simplest and the very smallest Sudoku puzzles in the set. We're going to be putting out a blog post about GPT-5's performance, and it is a jump, but it's still completely unable to solve puzzles which humans can solve. And what I really like about this dataset, and actually was the catalyst for me creating it in the first place, was that there was a quote by Andrej Karpathy saying, 'Okay, so we have all this data from the internet, but what you really want, if you wanted AGI, you wouldn't want all of the text that humans have ever created; you would actually want the thought traces in their head as they were creating the text, right? If you could actually learn from that, then you would get something really powerful.' And I thought to myself, well, that data must exist somewhere. My first thought was maybe philosophy, like there's a type of philosophy where you just write down your thoughts without thinking, like stream of consciousness. I thought maybe that could work. But then when I wasn't thinking about it and I was, you know, in my leisure time, I was watching a YouTube channel called Cracking the Cryptic.
是的。
Yes.
那两位英国绅士会为你解决这些极其困难的数独谜题。没错。有时他们的视频长达四个小时,他们是专业人士——这就是他们的工作。
Where these two British gentlemen will solve these extremely difficult Sudoku puzzles for you. Right. Sometimes their videos are four hours long, and they're professionals—this is their job.
我意识到完美的是,他们极其详细地告诉你他们解决这些特定谜题所用的推理过程。所以我们征得他们的同意,拿走了他们所有的视频,这些视频代表了数千小时非常高质量的人类推理,比如思维痕迹,我们抓取并用于模仿学习。我们确实在内部尝试过。结果发现我做得有点太好了,创造了一个非常难的基准。所以我们仍在努力让那东西工作,如果我们取得一些成功,我们会发布它。我想真正强调这个推理基准确实与众不同。你不仅得到非常扎实的东西,比如你确切知道它对错,所以你可以尽情做强化学习,而且你不能轻易泛化。每个谜题都是手工精心设计的,在规则上有一个新的独特转折,称为“突破”,你必须理解它。而现在,尽管我们取得了所有进展,当前的 AI 模型无法实现那个飞跃。它们找不到这些突破。它们会退回到,好吧,我试试不,我试试五,我试试六,我试试七。推理变得非常无聊,完全不像我们从那个 YouTube 频道开源的那些转录中看到的那样。所以我只想提出挑战,这是一个非常难的基准,我认为在这个基准上的进展将真正意味着 AI 的整体进步。
And what was perfect I realized is they tell you in agonizing detail exactly what reasoning they used to solve those particular puzzles. So we with their permission took all of their videos which represents thousands of hours of very high quality human reasoning like thought traces and scraped them and made that available for imitation learning. We did try to do this internally. Turns out that I did a little bit too much of a good job of really creating a very difficult benchmark. So we're still trying to get that stuff working and we'll publish it if we have some success. I want to really sell the fact that this reasoning benchmark really is different. Not only do you get something that's super grounded, like you know exactly if it's right or wrong, so you can do RL to your heart's content, but you can't generalize very easily. Each puzzle is deliberately designed by hand to have a new and unique twist on the rules called a breakin that you have to understand. And right now, despite all the progress we've made, the current AI models can't take that leap. They can't find these breakins. They'll fall back to, okay, I'll try no, I'll try five, I'll try six, I'll try seven. The reasoning becomes really boring and nothing like what you see in the transcripts that we've open sourced from this YouTube channel. So I just want to put the challenge out there that this is a really difficult benchmark and I think progress on this benchmark will really mean progress in AI generally.
你能反思一下,在看了这个 Cracking the Cryptic YouTube 频道之后,这些模式有多多样化吗?因为克里斯对我说,哦,你知道这些人,他们上 Discord 服务器,得到这些创意疯狂的想法,我着迷了。也许我只是理想主义,但我喜欢这个想法,即存在一个知识的演绎闭包。有一棵巨大的推理树,我们都拥有树的不同部分,深度不同。所以越聪明、知识越丰富,你在树中走得越深。但在这种理想化的形式中,有一棵树,所有知识都源自或发源于这些抽象原则。原则上我们可以构建推理引擎,它们可以从第一原理推理,而且可能在计算上不可约简。所以你必须执行所有步骤。感觉因为不拥有整棵树,我们需要做的是四处摸索。我们四处寻找乐高积木。哦,那是一个好乐高积木。我可以把它应用到这个问题上。也许这就是我们现在在 AI 中需要做的,就是尽可能多地获取这棵树。但我们能一直做到最底层吗?
Could you reflect a bit after watching this Cracking the Cryptic YouTube channel? How diverse were the patterns? Because Chris was saying to me, oh you know these guys they go on Discord servers, they get these creative crazy ideas and I'm obsessed. Maybe I'm just being idealistic, but I love this idea of there being a deductive closure of knowledge. That there's this big tree of reasoning and we're all in possession of different parts of the tree to different depths. So the smarter and the more knowledgeable you are, the deeper down the tree you go. But in this idealized form, there is one tree and all knowledge kind of originates or emanates from these abstract principles. And we could in principle build reasoning engines that could just reason from first principles and it might be computationally irreducible. So you have to perform all of the steps. And it feels like because we're not in possession of the full tree, what we need to do is kind of fish around. We fish around to find Lego blocks. Oh, that's a good Lego block. I can apply that to this problem. And maybe that's just what we need to do in AI for the time being is we need to just acquire as much of the tree as possible. But could we just do it all the way down?
是的,迷人的问题。那棵树可能非常庞大。当人类解决这些谜题时,他们肯定在实时学习并发现树的新部分。这有点像元任务,因为不仅仅是推理,你在推理关于推理。我认为我们现在在 AI 中没有这个。因为如果你看视频,他们会说类似的话,‘好吧,这看起来像一个 parask,或者这是一个集合论问题,或者,你知道,也许我应该拿出我的路径工具并追踪这个。’当然,专业人士确实像你说的那样,在头脑中已经有了大量推理乐高积木的集合。所以他们会识别出,哦,那种规则通常需要这种乐高积木。观察他们多么擅长直觉地知道,而像我这样解决得不多的人需要花很多时间四处寻找,比如好吧,也许我应该试试这个,或者也许我试试那个,这真的很迷人。但即使他们也不完美。所以你可以看到他们采取某种推理并开始构建。好吧,也许我们应该这样解决,然后发现这没有足够消除歧义,然后回溯,然后走另一条路。同样,我们在当前 AI 尝试解决这个基准时没有看到这一点。树非常大,我猜树中许多主题之间的系统发育距离非常大。所以很难跳跃,我认为这就是为什么作为集体智能我们合作得这么好,因为我们实际上找到了跳跃到树的不同部分的方法。我认为这可能就是为什么我们试图应用于此的强化学习算法目前不起作用,因为为了学习如何获得这些突破,理解解决这些谜题所需的细微推理,你必须采样它们。而这是一个如此稀有的空间,需要如此特定的推理才能达到特定的突破,这种技术不起作用。社区里肯定有一种感觉,好吧,这就是你现在解决问题的方式。我们有强化学习,是的,我们可以让这些语言模型做我们想做的事。但它不适用于这个数据集。
Yeah, fascinating question. That tree is probably massive. And as a human is solving these puzzles, they're definitely learning in real time and discovering new parts of this tree. And it's sort of a meta task, because it's not just reasoning, you're reasoning about the reasoning. And I don't think we have that in AI right now. Because if you watch the videos, they'll say something like, 'Okay, this looks like a parask or this is a set theoretic problem or, you know, maybe I should get my path tool out and trace this around.' And of course the professionals they do have this already massive collection of reasoning Lego blocks as you say in their head. So they'll recognize okay that type of rule usually needs this kind of Lego block. It's actually fascinating to watch how good they are at just intuitively knowing where someone like me who hasn't solved as many needs to spend a lot of time looking around like okay maybe I should try this or maybe I try this one. But even they're not perfect. So you can watch them take a certain kind of reasoning and start building up. Okay, maybe we should solve it like this and then go and know that doesn't disambiguate it enough and then backtrack and then go down another path. Again, something that we do not see current AI doing when they're trying to solve this benchmark. The tree is very big and I guess the phylogenetic distance between many of these motifs in the tree is just so large. So it's so difficult to jump between and I think that's why as a collective intelligence we work so well together because we actually find ways to jump to different parts of the tree. And I think that's probably why the RL, the current state of the RL algorithms that we're trying to apply to this just isn't working because in order to learn how to get these breakthroughs to understand what the sort of nuance reasoning is to get these puzzles, you have to sample them. And it's such a rare space, it's such a specific kind of reasoning that's required to get to the specific breakthrough that this kind of technique doesn't work. And there's definitely a feeling in the community like, okay, this is how you just solve things now. Like we have RL, yes, we can get these language models to do what we want. It doesn't work for this dataset.
伙计们,非常荣幸邀请你们上节目。在我们结束之前,你们在招聘吗?因为我们有很棒的机器学习工程师和科学家观众,我认为为 Sakana 工作会是梦想的工作。
Guys, it's been an absolute honor having you on the show. Just before we go, are you hiring? Because we've got a great audience of ML engineers and scientists and I think working for Sakana would be the dream job.
你太客气了。是的,我们确实在招聘,正如我在这次采访中早些时候说的,我真诚地希望给人们尽可能多的研究自由。我愿意打这个赌。我认为会有非常有趣的东西从中产生。而且我认为我们已经看到了很多有趣的东西从中产生。所以如果你想做你认为有趣和重要的工作,来日本吧。
That's very kind of you. Yes, we are definitely hiring and as I said earlier in this interview, I honestly want to give people as much research freedom as possible. I'm willing to make that bet. I think things that are very interesting will come out of this. And I think we've already seen plenty of interesting things coming out of this. So if you want to work on what you think is interesting and important, come to Japan.
而日本恰好是世界上最文明的文化。这可能是千载难逢的机会,各位。所以联系吧,伙计们。说真的,非常感谢你们。非常荣幸邀请你们两位上节目。
And Japan just happens to be the most civilized culture in the world. It might be the opportunity of a lifetime, folks. So get in touch, guys. Seriously, thank you so much. It's been an honor having you both on the show.
非常感谢。非常感谢。这很棒。
Thank you very much. Thank you so much. It's been great.