Nested Learning & AI Sleep: Towards Continual Learning with Ali Behrouz
打开互动全文版(中英对照 + 朗读 + 问答)→Ali Behrouz 讨论其开创性的嵌套学习和语言模型需要睡眠的研究,探索受生物学启发的架构如何实现 AI 的持续学习和记忆巩固。
Ali Behrouz discusses his groundbreaking work on Nested Learning and Language Models Need Sleep, exploring how biologically inspired architectures enable continual learning and memory consolidation in AI.
大家好,欢迎回到《认知革命》。今天,我很高兴能与 Ali Behrouz 进行对话。他是康奈尔大学的研究生、谷歌的研究员,也是《嵌套学习》的作者。这期节目是几个月前录制的,虽然我通常认为 AI 内容经不起时间考验,但这次与 Ali 的对话是个例外。他的工作是我在追求真正持续学习的新机器学习架构领域中所见过的最具启发性和潜在变革性的成果之一。这当然是当今最重要的能力进步之一。可以说,它是当前模型与能够像人类一样加入并贡献于人类团队的数字 AGI 之间的主要差距。而 Ali 正在用一种既受生物学启发又技术优雅的方法推进前沿。他的重磅论文《嵌套学习》被 Jeff Dean 本人誉为可能范式转变的先兆,提出了一种简单策略:通过以不同频率更新系统的不同部分,让模型在持续适应当前上下文的同时保留核心知识,就像人类管理从工作记忆到长期记忆的多时间尺度记忆一样。他的最新工作《语言模型需要睡眠:学习自我修改和巩固记忆》——我实际上是在这次录制中第一次现场听到,现在终于完全公开——从人类在睡眠中巩固记忆和从梦境中学习的方式中汲取灵感,引入了一种新的离线模式,模型通过蒸馏将新知识从高频更新层转移到更慢演化的层,并通过生成和训练基于近期经验的合成数据来学习概念之间的新抽象和连接。除了这些架构的细节——像许多 AI 创新一样,我觉得它们既极其令人兴奋又有点可怕——我们还讨论了 Scaling 如何可能从堆叠更多层转向嵌套更多频率更新率;Ali 如何将机器学习系统的所有组件理解为压缩给定上下文流的联想记忆形式;为什么这让他称深度学习架构为幻觉,以及他如何通过开发能够学习更新规则并优于 Adam 和 Muon 的表达性优化器来操作化这一概念性见解。我们还讨论了注意力机制如何可以被理解为无限频率更新模块,以及为什么 Ali 认为注意力层将因此永远成为 AI 系统的固定组成部分。我们涵盖了实证结果,显示 Ali 的新架构在标准指标上与 Transformer 竞争有效,同时在困难任务上表现更优,例如从多达 1000 万个 token 的上下文中有效回忆信息,以及同时学习翻译多种从未见过的语言。最后,我们讨论了为什么 Ali 认为持续学习既是机会也是隐私和对齐的巨大风险;人机关系可能如何演变;以及为什么 Ali 谨慎乐观地认为,基于我们与它们的交互而随时间演变的模型既能更有效地满足个人需求,也能导致一个更加多样化和希望稳定的 AI 生态系统。对我来说,底线是:尽管关于当前架构是否能扩展到 AGI 及超越的争论和猜测很多,但很有可能概念性突破会在我们甚至设法回答这个问题之前就使其变得无关紧要。Transformer 显然改变了世界,但它们不是历史的终结。尽管跟上 AI 发展很困难,但任何想了解未来走向的人都不能忽视像 Ali 这样的新研究方向。所以,话不多说,希望你喜欢这次与 brilliant Ali Behrouz 一起的深度探索,预览那些以越来越像人类的方式持续学习的 AI 系统。《认知革命》由 Mercury 赞助,这是一家超过 30 万家雄心勃勃的公司和个人信任的金融科技公司,用于管理他们的财务。在过去的几个月里,我在个人 AI 基础设施方面取得了巨大进步。现在,我在 Mac mini 上运行着 Claude Code 和 OpenClaw 的高上下文实例,它们能做的事情令人惊叹。然而,在开始使用 Mercury 之前,我没有很好的方式让它们支付东西。我不想让它们无限制地访问我的钱,但我的旧银行没有给我其他选择。有了 Mercury,我可以创建任意数量的虚拟卡,每张卡都有自己的每日、每周或每月消费限额,并且我可以将任何卡锁定到单一购买类别甚至单一商家。现在,我有一张卡可以让我的智能体用来购买我们家的杂货,而且只能买杂货。我可以随时创建另一张卡,用于给智能体一个可能需要购买的随机一次性项目。这实际上只是 Mercury 的 AI 友好产品的开始。你的银行提供 API 密钥、MCP 或 CLI 工具吗?如果没有,请访问 mercury.com 查看 Mercury。Mercury 是一家金融科技公司,不是 FDIC 保险的银行。银行服务由 Choice Financial Group 和 Column N/A 提供,成员 FDIC。感谢 Mercury 对《认知革命》的支持。现在,节目开始。Ali Behrouz,《嵌套学习》和新的《语言模型需要睡眠:学习自我修改和巩固记忆》的作者,欢迎回到《认知革命》。
Hello, and welcome back to the Cognitive Revolution. Today, I'm excited to share a conversation with Ali Behrouz, grad student at Cornell, researcher at Google, and author of Nested Learning. This episode was recorded a few months back, and while I normally believe that AI content does not age well, this conversation with Ali is an exception. His work is some of the most inspired and potentially transformative that I've seen anywhere in the quest for new machine learning architectures that are capable of genuine continual learning. This, of course, is one of the most important capability advances on the horizon today. Arguably, it is the main gap between today's models and a digital AGI that would be capable of joining and contributing to human teams just as humans do. And Ali is advancing the frontier with an approach that is both biologically inspired and technically elegant. His blockbuster paper, Nested Learning, which has been touted as a harbinger of a possible paradigm shift by no less than Jeff Dean, develops a simple strategy that allows models to rapidly adapt to their current context on an ongoing basis while preserving core knowledge by updating different parts of the system at different frequencies, much like humans manage memory on multiple timescales from working memory to long-term memory. His latest work, Language Models Need Sleep, Learning to Self-Modify and Consolidate Memories, which I actually heard about live for the first time on this recording, and which has now finally become fully public, takes inspiration from how humans consolidate memories and learn from dreams while sleeping, introducing a new offline mode in which models transfer new knowledge from their high-frequency update layers to their more slowly evolving layers via distillation, and also learn new abstractions and connections between concepts by generating and training on synthetic data derived from their recent experiences. In addition to the details of these architectures, which, like so many AI innovations, I find both extremely exciting and a bit scary, we also discuss how scaling for performance may shift from stacking more layers to nesting more frequency update rates. How Ali understands all components of machine learning systems as forms of associative memory that compress a given context flow. Why this leads him to call deep learning architectures an illusion, and how he's operationalized this conceptual insight by developing expressive optimizers that learn update rules and are capable of outperforming both Adam and Muon. We also discuss how the attention mechanism can be understood as an infinite frequency update module, and why Ali expects that attentional layers will therefore remain fixtures of AI systems indefinitely. We covered the empirical results showing that Ali's new architectures compete effectively with transformers on standard measures, while also outperforming them on hard tasks, such as effectively recalling information from up to 10 million tokens of context, and also learning to translate multiple previously unseen languages at the same time. Finally, we discuss why Ali sees continual learning as both an opportunity and a huge risk for privacy and alignment. How human-AI relationships might evolve, and why Ali is cautiously optimistic that models that evolve over time based on our interactions with them could both serve our individual needs more effectively, and also lead to a more diverse and hopefully stable AI ecosystem overall. The bottom line for me is that for all the debate and speculation about whether or not current architectures can scale to AGI and beyond, there is a very good chance that conceptual breakthroughs will render that question moot before we even manage to answer it. Transformers have changed the world, clearly, but they aren't the end of history. And as tough as it is to keep up with AI developments, anyone who wants to get a handle on where things are going from here can't afford blind spots when it comes to new research directions like Ali's. And so, without further ado, I hope you enjoy this deep dive preview of AI systems that learn on an ongoing basis in increasingly human-like ways with the brilliant Ali Behrouz. The Cognitive Revolution is brought to you by Mercury, the fintech that more than 300,000 ambitious companies and individuals trust to run their finances. Over the last few months, I have made tremendous strides with my personal AI infrastructure. Today, I've got high context instances of both Claude code and open claw running on a Mac mini, and it's amazing what they can do. However, until getting started with Mercury, I didn't have a great way for them to pay for things. I didn't want to give them unrestricted access to my money, but my old bank didn't give me any other options. With Mercury, I can create as many virtual cards as I want, each with its own daily, weekly, or monthly spending limit, and I can lock any card to a single category of purchase or even a single merchant. Now, I have a card that my agent can use to buy our family's groceries, and only our groceries. And I can create another anytime I want to give an agent a random one-off project that might require making a purchase. This is honestly just the start of Mercury's AI-friendly offerings. Does your bank offer API keys, an MCP, or a CLI tool? If not, check out Mercury at mercury.com. Mercury is a fintech company, not an FDIC-insured bank. Banking services provided through Choice Financial Group and Column N/A, members FDIC. Thank you to Mercury for supporting the Cognitive Revolution. And now, on with the show. Ali Behrouz, author of Nested Learning and the new Language Models Need Sleep, Learning to Self-Modify and Consolidate Memory. Welcome back to the Cognitive Revolution.
非常感谢你邀请我。我很感激。
Thank you very much for having me. I appreciate it.
我对今天的对话感到超级兴奋。感谢你愿意抽出时间回来,更深入地探讨你的工作。我认为这非常迷人,确实是我最近看到的最具启发性的工作之一。据我理解,你的方法很大一部分是观察人类认知的构成,识别我们作为人类所做的那些似乎非常重要且对我们成功运作于世界至关重要的东西,然后试图找出 AI 版本可能是什么样子,并开始开发架构或系统设计——甚至可能比架构更抽象——使这些能力在 AI 系统中成为可能。这些想法有些效果非常好,而且当我真正花时间深入理解它们时,它们感觉如此优雅和正确,这让我非常震惊。
I am super excited about this conversation today. I appreciate you for being willing to take some time and come back and do a deeper dive into your work. I think it is super fascinating and genuinely some of the most inspired work that I have seen in recent times. And a big part of your method, as I understand it, is looking at what human cognition consists of and identifying things that we as humans are doing that seem like they're quite important and really critical to our successful function in the world, and then trying to figure out kind of what an AI version of that might look like, and then starting to develop the architectures or system designs, maybe even more abstractly than architectures, that start to make those capabilities possible in AI systems. And it's really striking both like how well some of these ideas have worked, and also striking to me how elegant they feel and and how kind of right they seem as I really take time to dig in and understand them.
那么第一个问题,宏观来看:你如何看待自己正在做的事情?显然,你识别出了当前系统的不足。你设想通过开发新架构来解锁什么?
So, first question, big picture: how do you think about what you are trying to do? Obviously, you identify gaps in current systems. How do you conceive of what you are trying to unlock with the new architectures you're developing?
从大脑中汲取灵感。对我来说,这对不同的人意味着不同的事情。我真的很喜欢从大脑和进化中获取灵感。主要原因是,我认为大脑有大量数据以自然的方式进行训练。所以,我们现在看到的是一个非常复杂的生物大脑的进化版本。因此,它是一个很好的灵感来源。但当我这么说时,我并不是说要复制大脑并完全做大脑所做的事情,因为大多数时候我们并不知道那到底是什么。关于大脑如何工作,有不同层次的理解。第一个层次是我们知道它有效。第二个层次是大脑中有一些模块,每个模块负责不同的部分——记忆等——这大致就是过程。当我们想从大脑中获取灵感时,难点在于我们关注哪个粒度层次。如果你过于深入细节,会有两个问题:一是过度拟合到一种特定的智能形式,二是我们实际上并不知道大脑是如何做那件特定事情的。在我所做的所有工作中,例如在 Titans 和 Nessen 学习上,我们可以看到模型在实际应用中面临挑战。那么问题就是,这个特定的挑战是否被人类解决了,人类能否简单地做些什么来应对它?然后第二个问题是他们如何做到这一点。例如,在 Titan 中,我们讨论了惊喜度量,以及记忆应该如何分解为短期和长期记忆。但大脑并不是通过梯度下降来理解惊喜的;那只是高层次的理解。如果我们把它作为灵感来源,那很好。但如果我们深入细节,我们可能会面临挑战,因为我们对大脑的理解随时间变化,我们可能会过度拟合到某个特定的设计选择。
Getting inspired from the brain. For me, it means different things for different people. I really like to get inspired from the brain and generally evolution. The main reason is that I think it has a lot of data to train itself in a natural way of training. So, one thing we can see now is a very evolved version of a very complicated biological brain. So, it's a great source of inspiration. But when I say that, I don't mean I want to replicate the brain and fully do what the brain does, because most of the time we don't know what that actually is. There are different levels of understanding about how the brain works. The first level is that we know it works. The second level is that there are some modules in the brain, each responsible for different parts—memory, etc.—and that's generally the process. The hard part when we want to get inspired from the brain is at what level of granularity we focus. If you go too much into the details, there are two issues: one is overfitting to a specific form of intelligence, and another is that we don't actually know how the brain does that specific thing. In all the works I have done, for example, on Titans and also on Nessen learning, we can see that models face challenges in real-world applications. Then the question is whether that specific challenge is solved by humans, and can humans simply do something to address it? Then the second question is how they can do that. For example, in Titan, we discussed a surprise metric, and how memory should be decomposed into short-term and long-term memory. But the brain is not exactly performing gradient descent to understand surprise; that's just a high-level understanding. If we keep that as a source of inspiration, it's great. But if we go into more details, we might face challenges because our understanding of the brain changes over time, and we might overfit to one specific design choice.
关于必要学习,我认为当前模型缺少的一枚硬币有两个部分。一是它们如何适应所处的环境和上下文。二是模型如何理解新知识并随时间将其融入参数,从而避免灾难性遗忘——即它们训练过的特定任务被遗忘。这两点很重要,当前模型面临很多挑战。如果你有一个非常大的模型,它需要随时间更新。这就是为什么所有大语言模型都有知识截止日期。例如,如果你问 ChatGPT 特定信息,并说不允许使用任何工具,你可能会看到知识截止,这很难克服。如果你想持续更新模型,有两个巨大挑战:灾难性遗忘和效率。参数很多,你无法持续更新所有参数。有一些解决方案,比如监督微调(SFT)或强化学习,但模型仍然可能面临灾难性遗忘。此外,在某个时候,你需要将上下文中的所有知识转移到实际参数中。如果你只是不断总结词元,并在词元空间中做所有事情,主要问题是你将超出大语言模型的上下文长度。所以,当前大语言模型范式的主要问题是它们无法持续学习并随时间获取新知识和技能。此外,它们在理解世界的不同抽象层次方面也有限。科学中的一切都是以最简单的方式解释世界。这就是我们学习的方式——我们不想保留所有东西;我们想理解潜在模式。这种压缩过程以及从数据中理解不同抽象层次是当前大语言模型所欠缺的。
About the necessary learning, I think one coin that is missing in current models is about two parts. One is how they can adapt to the environment and context they are in. Another is how the model can understand new knowledge and incorporate it into their parameters over time, so they can avoid catastrophic forgetting—meaning that a specific task they have trained on is forgotten. These two are important things, and current models face a lot of challenges. If you have a very large model, it needs to be updated over time. That's why you can see a knowledge cutoff for all LLMs. For example, if you ask ChatGPT about specific information and say you are not allowed to use any tools, you might see a knowledge cutoff, which is challenging to overcome. If you want to keep updating the models, there are two huge challenges: catastrophic forgetting and efficiency. With many parameters, you cannot keep updating all of them. There are some solutions, like supervised fine-tuning (SFT) or RL, but the model can still face catastrophic forgetting. Also, at some point, you need to transfer all the knowledge in your context to the actual parameters. If you just keep summarizing tokens and do everything in token space, the main issue is that you will exceed the context length of LLMs. So, the main issue with the current LLM paradigm is that they cannot continually learn and obtain new knowledge and skills over time. Also, they are limited in understanding different levels of abstraction of the world. Everything in science is about explaining the world in the simplest way possible. That's how we learn—we don't want to keep everything; we want to understand underlying patterns. This compression process and understanding different levels of abstraction from data is something current LLMs fall short on.
那么,我想从几个不同角度进一步探讨你的直觉。我当然感受到了——而且我认为大多数用户确实感受到了——你强调的那些问题。
So, a couple of different angles I want to probe your intuition a bit more on this point. Certainly I have felt—and I think most users honestly have felt—the problems you're highlighting.
从某种意义上说,我觉得自己相对于当今 AI 的最大优势在于这种持续的连贯性和相当稳定的身份认同。比如,我早上知道自己是谁,大致记得昨天想做什么,而且基本能接着干。我可能能从日常做的事情中学到更多,但至少能学到一些东西并吸收进去。显然,当前的模型做不到这一点,这对它们来说是一个很大的弱点。这就是为什么当初 Mamba 论文出来时我那么兴奋,因为我想,‘哇,这东西看起来能和 Transformer 竞争,但它有一个固定大小的记忆空间。’显然我们不能让记忆空间呈二次方增长到无穷大,所以我们必须有某种大小有限的东西来工作。那是一个显著的进步,而且你也在 Mamba 架构上做了一些工作。Titan 更是如此:这是另一种思考固定大小记忆模块的方式,它在运行时通过梯度下降更新。这似乎很可能让它比 Mamba 架构更强大,但类似的结构——一个固定大小的东西,能保留所需并逐渐舍弃不需要的——似乎非常重要。我想知道你是否从另一个角度来考虑。到目前为止,我们说的是它不能做什么,我们能做什么,它不能做什么。另一种思考方式是,我们希望我们的 AI 变成什么样?今天,我们大多拥有需要被唤醒的聊天机器人。它们要么需要我们发送消息,要么我们越来越多地使用 cron 作业和其他触发器来让 AI 醒来并做点什么。但如果这些事情没有发生,它们就是惰性的。它们只是坐在那里,直到有人呼叫它们。你有没有一种感觉或愿景,认为你理想的未来 AI 会有所不同?它看起来更像一个拥有 AI 优势的另一个人,还是一个修补了弱点的 LLM?当你梦想着 2030 年每天与你密切合作的 AI 时,你设想的是什么?
And in some ways I feel like the biggest advantage that I have relative to an AI today is this kind of ongoing coherence and reasonably stable identity. Like I know who I am in the morning and I kind of know what I was trying to do yesterday and I can mostly pick up where I left off. I probably could learn a lot more from the things that I do on a daily basis, but I at least learn some stuff and take it on board. Obviously the current models don't really do that and that is a big weakness for them. That's why I was so excited about the original Mamba paper when that came out because I was just like, 'Wow, here's something that seems like it's competitive with Transformers, but it has a fixed-size memory space.' Obviously we can't grow the memory space quadratically to infinity, so we're going to have to have something that's bounded in size that can work. That was a notable step and you've done some work with the Mamba architecture as well. Titan is even more so: here's another way to think about having a fixed-size memory module that was updated with gradient descent at runtime. That seems like it would likely make it even more powerful than the Mamba architecture, but a similar kind of structure of a fixed-size thing that can keep what it needs and gradually let go of what it doesn't seems super important. I wonder if you come at it from the other angle. So far we've said here's something it can't do, we can do this, it can't do that. Another way to think about it is what do we want our AIs to be like? Today we have mostly chatbots that need something to wake them up. They either have to go send them a message or increasingly we have cron jobs and other triggers that get the AI to wake up and do something. But if those things don't happen, they're inert. They just sit there until somebody calls their number. Do you have a sense or a vision of what your ideal future AI would be like that's different than that? Does it look more like another person but with AI advantages, or does it look like an LLM with weaknesses patched? When you dream of a 2030 AI that you're working closely with on a daily basis, what do you envision?
回答这个问题有不同的方面。从技术角度来看,我认为过去 40 年的整个研究,大部分都建立在一个范式上:我们有预训练或一般的训练阶段,以及测试阶段。但关键是,如果我们有持续学习,我们可以看到最近有很多关于持续学习的研究,关于如何做到这一点,以及如何克服许多挑战。但关键是,一个真正的持续学习者没有测试和训练时间。所以,如果我们在任何设计选择中听到这个名字,可能意味着它不是真正的持续学习者,因为没有测试,没有训练时间。所以问题是,这对模型来说是一个统一的过程吗?我个人认为我们仍然需要至少两个阶段。它的工作方式是,我们应该有一个模型活跃的阶段。所以它主动接收信息,无论是通过用户查询,还是通过视觉模型、世界模型或任何类似的东西。但关键是,模型接收一些信息,通常对输入数据进行一些计算,此时它是活跃的。但另一方面,还有另一个阶段,模型不必等待输入数据。它可能不接收任何输入数据。它完全与外部世界隔绝。但问题是,即使在那个时候,模型应该是静态的,不执行任何计算,还是模型需要开始思考某个过程,思考其参数内部的数据等等?所以我认为我们可以将过程分为两部分,正如我提到的。一个是活跃阶段,另一个是另一个阶段。我们可能可以称之为睡眠时间,因为没有输入,但人工大脑或模型本身仍在尝试执行一些计算。所以我认为这是定义持续学习方向上不同阶段的一个好方法。然后我认为一个好的模型是一个在两方面都表现很好的模型。它应该正确地接收信息,编码、处理并以最佳方式理解它。另一方面,当它进入睡眠时间时,它也应该开始处理之前学到的东西,并用于自我改进。所以这就是我从技术角度认为理想模型应该做的。
There are different aspects to answer this question. From the technical point of view, I think generally the entire research, most part of the research in the past 40 years, is built on a paradigm that says we have a pre-training or a generally training phase, and we have a test phase. But the point is, if we have continual learning, we can see that there are a lot of recent studies about continual learning, how we can do that, and how we can overcome a lot of challenges. But the point is, a true continual learner doesn't have a test and train time. So potentially, if we hear this name in any design choices, potentially it means that it's not a true continual learner, because there is no test, there is no train time. So the question is, is it a uniform process for the model or not? My personal opinion is that we still need at least two phases. How it works is that we should have one phase where the model is active. So it actively receives information, whether through the user query, for example, or through vision models, world models, or anything similar. But the point is, the model receives some information and generally it performs some computation on the input data and it's active at that point. But on the other hand, there is another phase where the model does not have to wait for input data. It might not receive any input data. It's completely locked from the world outside of it. But the question is, even at that time, should the model be static without performing any computation, or does the model need to start thinking about some process, thinking about the data that it has inside its parameters, and so on? So I think we can break the process as I mentioned into two parts. One is the active phase and another is another phase. Potentially we can call it a sleep time because there is no input, but still the artificial brain or generally the model itself is trying to perform some computation. So I think that's a good way of defining different phases in this direction of continual learning. And then I think a good model is a model that performs very well on both sides. It should receive the information properly, encode it, process it, and understand it in the best way possible. And on the other hand, when it goes to sleep time, it should also start processing what it has learned before and use that for self-improvement. So that's what I think an ideal model should do from that technical point of view.
嘿,我们稍后继续采访,先听一段赞助商信息。今天的节目由 Anthropic 提供,他们是 Claude 和 Claude Code 的开发者。在过去的几个月里,Claude 帮助我构建并完善了一个个人深度上下文数据库,现在包含了我过去整整 5 年的所有电子邮件、Slack 消息、推文、跨平台私信、视频通话和播客转录。在此基础上,我们还添加了描述我与数百个联系人、组织和想法关系的摘要文章。现在有了这个,几乎没有什么 Claude 帮不上忙的。在报税季,我让 Claude 帮我整理。它浏览了我的收件箱,找到了我所有 10 份兼职工作的 1099 表格,并为我建立了一份关于我的支出和捐赠的全面报告。对于我的天使投资,Claude 现在可以根据我与创始人的通话和电子邮件往来,以我的风险基金要求的精确格式起草投资备忘录。当有人需要帮忙时,Claude 通常能做得和我一样好。最近,一位朋友联系我,问我是否认识适合他正在招聘的职位的人。起初我没想到任何人,但后来我想到问 Claude,果然,它找到了两个很好的候选人。Claude 是那些不满足于“足够好”的人们的 AI。
Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Anthropic, makers of Claude and Claude Code. Over the last few months, Claude has helped me build and refine a personal deep context database that now contains all of my emails, Slack messages, tweets, DMs across platforms, video calls, and podcast transcripts going back a full 5 years. On top of that, we've now layered summary articles describing my relationship with hundreds of contacts, organizations, and ideas. And now that this exists, there's almost nothing that Claude can't help with. For tax season, I asked Claude to help me get organized. It went through my inbox, tracked down 1099s for all 10 of my part-time jobs, and built me a comprehensive report on my expenses and donations. For my angel investing, Claude can now draft investment memos in exactly the form that my venture fund requires, based on the calls I've had and the emails I've exchanged with the founders. And when someone needs a favor, Claude can often do it as well as I can. Recently, a friend reached out to ask if I know anyone who might be a fit for a role that he is currently hiring for. Initially, nobody came to mind, but then I thought to ask Claude, and sure enough, it identified two great leads. Claude is the AI for minds that don't stop at good enough.
它是一个真正理解你整个工作流程并与你一同思考的协作者。无论你是在深夜调试代码,还是制定下一个商业策略,Claude 都能扩展你的思维,帮你解决重要问题。因此,对于值得解决的问题,请访问 claude.ai/tcr 开始使用 Claude。网址是 claude.ai/tcr,并查看 Claude Pro,它包含今天节目中提到的所有功能。再次提醒,网址是 claude.ai/tcr。但另一方面,我认为存在很多挑战。我们现在知道的模型非常大。所以,即使是一篇介绍新 LLM 或架构的简单学术论文,也需要在拥有数十亿参数(10 亿、20 亿等)的模型上进行实验。一般来说,如果你想随时间更新一个非常大的模型,它需要大量的算力。并且需要一些技术来使这个过程成为可能。总的来说,当我们思考这个方向时,正是关于嵌套学习的一些想法开始出现的时候。因为一般来说,如果你考虑嵌套学习,我们可以看到在每个时间步,我们不必更新所有内容。我们只需要更新所有参数的一小部分。所以这可能是克服效率挑战的一种方式。因此,如果我想总结我想说的,我认为一个理想的模型应该有两个阶段。首先,它应该是一个持续学习者。它应该与世界互动。而且,它应该有两个阶段:一个是关于非常活跃的过程,另一个是关于自我改进以及如何巩固记忆、如何理解知识、如何连接看似无关的不同事物等等,类似于我们大脑在睡眠时所做的事情。
It's the collaborator that actually understands your entire workflow and thinks with you. Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. So, for problems worth solving, get started with Claude at claude.ai/tcr. That's claude.ai/tcr, and check out Claude Pro, which includes all of the features mentioned in today's episode. Once more, that's claude.ai/tcr. But then the other hand, I think there are a lot of challenges. The models that we know right now are really large. So, even a simple academic paper presenting a new LLM or architecture needs to perform experiments on models with billions of parameters, 1 billion, 2 billion, or something like that. And generally, a very large model requires a lot of computation if you want to keep it updated over time. And it needs some techniques to make this process possible. And generally, for us, when we were thinking about this direction, it was a time when some ideas about nested learning started. Because generally, if you think about nested learning, we can see that at each time step, we don't have to update everything. We just need to update a small subset of all parameters. So that potentially is a way to overcome the challenges about efficiency. So, if I want to summarize what I wanted to say, I think an ideal model should have two phases. One, generally, it should be a continual learner. It should interact with the world. And also, it should have two phases: one is about a very active process, and another one is about self-improvement and how it can consolidate its memory, how it can understand the knowledge, how it can connect different things that seem to be irrelevant, and so on and so forth, similar to what our brain does when we go to sleep.
是的,这只是我的个人观点,可能完全错误,但我认为我们不应该过多关注人类能做什么,而应该更关注人类希望从 AI 那里得到什么。我认为我们不想创造与人类非常相似的东西。我的意思是,那也是一个非常有趣的方向,但我个人对此并不太感兴趣。我认为人类的设计已经很棒了,但另一方面,对于 AI 来说,我们需要能够理解我们需求的 AI 模型。例如,我认为在当前 LLM 的范式中,你可以看到当它们拥有更多功能后,它们变得非常出色。比如 ChatGPT 现在有了记忆,Claude 也有一些功能,Gemini 等等。现在它们更能理解你的需求。当你说“给我写这张图”时,它们可以简单地为你写出来。所以我认为这总体上是一个非常有前途的方向,因为我们不想复制人类。定义智能有很多不同的方式,它不必与人类智能相同。所以我认为我们再次可以从人类那里获得灵感,但我们需要理解为什么想要获得灵感。我们是想复制人类能做的事情,还是想从自然中获得灵感以理解一些底层规则?例如,一个极端的例子是我们不能穿越时间。所以如果我提出一个想法说我的 AI 模型试图打破世界的因果律,那可能是一个错误的方向,因为它打破了我们世界的一些自然规则。或者让我再举一个例子。当我们谈论“LLM 需要睡眠”这个标题时,并不意味着 LLM 真的需要去睡觉和休息。它意味着从人脑中,似乎有一个非常普遍的规则:它有两个阶段——学习,然后是另一个处理、巩固记忆和发现接收数据之间潜在模式的阶段。所以这是来自大脑的一个非常高层面的灵感。所以,简而言之,我认为我们不想复制人类能做的事情,也不想拥有人类智能,但另一方面,我们想要一种新的智能形式,它非常合适且设计良好,能够理解人类需求,从而帮助人们完成许多没有 LLM 时会面临挑战的事情。
Yeah, that's just my personal opinion and might be completely wrong, but I think we shouldn't focus too much on what human can do and instead focus more on what human wants from AI. I think we don't want to create something that is very similar to us. I mean, that's also a very interesting direction, but I personally am not really interested in that. I think we have a great design for humans, but on the other hand, for the AI side, I think we need AI models that are capable of understanding what we want. For example, I think in the current paradigm of LLMs, you can see that they are great after some time when we have more features. For example, ChatGPT has memory now, and Claude has some features, Gemini and everything. Now they are more capable of understanding what you want. When you say, for example, 'write this image for me,' they can simply write that for you. So I think that's a very promising direction in general because we don't want to replicate humans. There are many different ways to define intelligence, and it doesn't have to be the same as human intelligence. So I think here again, we can get inspired from humans, but we need to understand why we want to get inspired. Do we want to get inspired because we want to replicate what humans can do, or do we want to get inspired from nature to understand some underlying rules? For example, one extreme example is we cannot travel through time. So if I come up with an idea that says my AI model is trying to break causality in the world, then that might be a wrong direction because it breaks some natural rules about our world. Or let me give you another example. When we talk about the title 'LLM needs a sleep,' it doesn't mean that LLM literally needs to go to sleep and rest. It means that from the human brain, there seems to be a very general rule that it has two phases: learning and then another phase of processing, consolidating memory, and finding underlying patterns between received data. So that's a very high-level inspiration from the brain. So yeah, in short, I think we don't want to replicate what humans can do, and we don't want to have human intelligence, but on the other hand, we want to have a new form of intelligence that is really proper and designed in a good way to understand human needs, so it can help people do a lot of things that they might face challenges without LLMs.
是的,它们在某些方面已经超越了人类,所以它们平衡我们弱点的机会是不可思议的。你刚才说的一句话我想抓住的是“定义智能的多种方式”。也许我可以给你一个我对嵌套学习的高层概述。我认为在论文中,你们非常强调展示某些等价性,比如我们今天做事的方式是你们正在开发的更通用框架的一个特例。在我看来,嵌套学习范式的核心思想,以及我认为对于一个像我这样简单的人来说最令人兴奋的是,相当长一段时间以来,我们通过堆叠越来越多的层并让它们变得更大,实现了模型越来越强的表达能力。这效果非常好。我们已经能够将这个范式推得非常远。但这有点疯狂,对吧?我们只是把这一层反复堆叠,你知道,无论 80 层还是 120 层深,或者多少层,就这样了。感觉一个更成熟的解决方案应该比这更复杂,对吧?但这就是我们实现这种表达能力的方式,有时也会用到术语“计算深度”。而嵌套学习所做的是带来一种不同的方式,以实现更高水平的表达能力或计算深度,那就是堆叠的不是层而是级别。级别与层的区别在于,层是相同的东西,或者它们可以交替。显然,我们也有这种交错架构。但这些是信息从一层传递到下一层时按顺序进行的。但前向传播就像逐层通过所有层到达终点,就是这样。
Yeah, certainly they're already superhuman in some ways and so the opportunity for them to balance out our weaknesses is incredible. One phrase that you said there that I wanted to latch onto is 'multiple ways to define intelligence.' And maybe I'll just give you my high-level pitch for what nested learning is. I think in the paper there is a big emphasis on showing certain equivalences where you're like the way that we're doing things today is sort of a special case of a more general framework that you're developing. The core idea as I see it in the nested learning paradigm and what I think is potentially most exciting for a simple person like myself is that for quite some time now, we have achieved greater and greater expressivity of models by stacking more and more layers and just making them bigger. And that has worked remarkably well. We've been able to push that paradigm incredibly far. It's kind of crazy though, right? That we just have this one layer stacked over and over again, you know, whatever 80 or 120 layers deep or however many layers, and that's kind of it. It feels like a more mature solution should be somehow more elaborate than that, right? But that's been the way that we've achieved this expressivity, or sometimes the term 'computational depth' is thrown around. And what nested learning is doing is bringing a different way to the table to achieve higher levels of expressivity or higher levels of computational depth, and that is by stacking not layers but levels. And what differentiates a level from a layer is that a layer is like the same thing or they can alternate. Obviously, we have these sort of interleave architectures too. But these are things that are sort of in sequence as information passes from one layer to the next. But a forward pass is like passed through all the layers one by one and get to the end, and that's kind of the thing.
层级范式带来的不同之处在于,不同层级可以有不同的更新频率。这样一来,整个系统的某些部分可以更持久,而另一些部分则可以更激进地、近乎实时地更新。这显然更符合我们人类的特点,对吧?我们不是一个每次都以完全依赖方式处理信息的静态事物。我们的状态在很大程度上取决于刚刚经历的事情,但只是在一定程度上,对吧?比如,我的情绪或当前想法反映了今天早些时候发生的事情,但我对世界的宏观看法从今早到现在并没有改变。所以,显然存在某种不同信念、不同表征、不同回路构成的层级结构,有些更新得非常快,有些则非常慢,而且它们显然被整合在一起协同工作。我们在机器学习中还没看到这一点,除了少数非常边缘的实验案例。而现在,你通过这种嵌套学习范式真正开始证明,不仅能让它工作,而且正如我们将看到的结果,还能让它与 Transformer 竞争,甚至似乎拥有一些新的优势,我称之为微技能优势。通过某些早期诊断,你可以看到:‘哦,这东西能做一些 Transformer 做不到的、性质不同的事情,同时在通用困惑度评分上也略胜一筹。’那么,你对这种高层总结有什么反应?另外,我也很想听听你对计算深度或表达性这个概念的看法。我在某种程度上想把它类比为 G 因子。你知道,人们常说 AGI 中的 G 代表通用性。在人类智商语境中也有 G 因子,这是一种无形的特质,衡量你在广泛领域的能力。这同样指向通用性。也许在机器学习中,G 就是损失函数,或者存在某种基本的等价关系,但也许不是。我不太确定。所以我非常想知道你是怎么想的。显然我们在触及某种东西,而且已经看到了巨大进步,但那个东西到底是什么,我真的很想听听你的直觉。
What the levels paradigm brings to it that's different is that different levels can have different update frequencies. And with that, you now have the possibility for some parts of the overall system to be much more durable, and some parts to be updating much more radically in something much closer to real time. And that obviously feels like much more aligned to what we are, right? We're not like one static thing that processes information in a fully dependent way each time. We have, you know, we're very much our state in any given time is very much contingent on what we just experienced, but only to a degree, right? Like, my mood or what is currently on my mind is a reflection of what happened earlier today, but my big picture views about the world, they didn't change from this morning until now. So, there's clearly some sort of hierarchy of different kinds of beliefs, different kinds of representations, different kinds of circuits that we have, which are updated in some cases very quickly, in other cases very slowly, and obviously they are integrated together and work together. And we just haven't seen that in machine learning, except maybe in a few very far-flung experimental cases. And now you're really starting to show that with this nested learning paradigm, not only can you make it work, but as we'll get into with results, you can make it work in a way that is competitive with transformers and even seems to have some of these new or what I call micro skill advantages, where you can see with these certain early diagnostics that, 'Oh, this can do something that's qualitatively different than what a transformer can do, even as it also outperforms it a bit in terms of general perplexity type scoring.' So, how would you react to that kind of high-level summary? And then I also really would be interested to get your take on what is this concept of computational depth or expressivity? I'm tempted in some ways to make an analogy to just like the G factor. You know, people talk about obviously the G in AGI is the generality. There's also G in the context of human IQ, which is the G factor, this sort of intangible something that's like how capable you are across a very wide range of things. Again, that's just getting at generality. Maybe in machine learning it's as simple as G is sort of loss or there's maybe some fundamental equivalence there, but maybe not. I don't really know. So, I'm very interested in how you think about what that is. Clearly, we're getting at something, but we've seen huge progress, but what is that something is another thing I really would love to get your intuition on.
我们在嵌套学习上研究了非常非常久。可能超过一年半。我个人遇到的主要问题——也是我们团队在论文作者之后经常讨论的——就是很难把我们想实现的东西形式化。因为即使在数学公式里,也很难用正式的方式写出我们在这个范式中到底想做什么。经过反复讨论,我们最终提出了这个具体框架:存在一种更新频率,让每个模块在等待时能有时间做点事情。我认为这是描述为什么需要多种频率的最佳方式。另一方面,我们需要一种知识迁移方法,将知识从一个层级传递到另一个层级。所以,当慢网络等待快网络的所有计算和信息处理时,慢网络应该能从中获益,因为我们为快网络的计算付出了成本。这种优势来自于快网络和慢网络之间良好的知识迁移,无论是从低层到高层还是反之。其核心思想是:当有一个更新非常快的网络时,我可以利用它的快速计算为慢速计算侧提供一些东西。这样,慢网络或较慢的层级可以专注于数据的高层知识抽象,而更新更频繁的网络则可以专注于快速适应和处理高分辨率数据。总的来说,这就是我们描述这个框架的方式:我认为这两方面必须同时存在,相互补充。一方面是更新频率,另一方面是层级之间的知识迁移。
We were working on nested learning for a very, very long time. Potentially, I mean, even more than a year and a half or something. The main issue that I personally had and we'll be discussing a lot in the group after the authors of the work was it was really, really hard to formalize what we wanted to deliver. Because, you know, even in the mathematical formulation, it was really hard to write it in a formal way and say what we want exactly to do in this paradigm. And so after some back and forth discussions, I think we came up with this specific framework saying that there is a frequency of updates to somehow give time to each module to do something when it's waiting. I think that's the best way of describing why we need to have multiple frequencies. And on the other hand, we need to have a knowledge transfer method to somehow transfer the knowledge from one level to another level. So, when the slow network is waiting for all the computation and information processing of the fast network, then there should be some advantages for the slow network that we are paying all these costs for doing some computation by the fast network. And that advantage comes when we have a good way of knowledge transfer between the fast network and the slow one, which is like lower level or to the higher. So, the idea there is that when we have a network that is updating really fast, then I can use that fast computation to give something to a slow computation side. So, the slow network or the slower level can somehow focus on the more high-level knowledge abstraction of the data, and then the more frequently updated network can somehow focus on fast adaptation and how they can process higher resolution data. So, generally that was the way we could somehow describe this framework, saying that I think these two sides need to be there to complement each other. One is the frequency of update and another one is the knowledge transfer between the levels.
对于第二个问题,我认为从某个特定角度来看,当前的模型非常高效。为什么高效?因为从 LLM 能提供的东西来看,它们非常廉价。例如,如果你想在某些任务上用人类来匹配它们的能力,成本可能会非常不同。所以从这个角度,我们可以说计算以及当前的 LLM 范式非常高效。我们可以为模型中的每个参数或人工神经元执行更多的计算。这有助于实现不同的事情。其中之一是帮助我们拥有更智能的模型——虽然这是个很主观的术语。但你知道,当有更多计算时,似乎就在进行某种内部思考。所以,一个基于 Transformer 的简单 LLM:当 token 到来时,我对它进行一些计算,通过所有层,然后预测下一个 token。这就是它的工作方式。但现在假设对于某个特定 token,我不只是做一次简单的计算,而是根据过去的数据进行更多的内部计算,或者一般性地组合或混合数据等等。在这种情况下,我可以看到质量可能会提升,这个特定设计选择的下一个 token 预测质量会提升。为什么呢?因为它可以被解释为一种内部思考过程。所以现在,模型中的每个特定参数都在执行多次计算。这不再是一个参数、一步计算,而是一个参数、多步计算。这就是一个优势。
For the second question, I think in my opinion the current models are very efficient from one specific point of view. Why are they efficient? Because you can see that from what we can get from LLMs, they are very cheap. For example, if you want to match their power in some tasks with humans, then potentially the cost would be very, very different. So from that perspective, we can say that computation and generally the current LLM paradigms are very efficient. And we can perform more computation per each parameter or artificial neuron that we have in our model. And it can help us for different things. One is that it can help us to have a smarter model. It's a very subjective term to describe this. But you know, when we have more computation, it seems that we are performing some internal thinking. So, a simple LLM based on transformers: when a token comes, I will do some computation on the token, pass it through all the layers, and then predict the next token. That's how it works. But now let's assume that for a specific token, instead of just a simple pass of computation, I also perform more internal computation with respect to the past data or generally combine or mix the data and something like that. In that case, I can see that the quality can potentially go up, the quality of next token prediction for this specific design choice can go up. And why is that? Because it can be interpreted as a form of internal thinking process. So, now for each specific parameter that I have in my model, it is performing several computations. So, it's not just one parameter, one step of computation. It's one parameter, a couple of steps of computation. So, that's one advantage.
另一个是关于记忆视角和这些模型的适应性。所以,当我们有一个能快速适应上下文的模型时,它就有可能学习上下文。这就是我们在“大规模学习”中试图传达的主要信息之一——我们所知道的一切在某种程度上都是上下文学习的一种形式。总的来说,我认为人类语言中创造新词是件好事,但另一方面,如果同一个概念我们创造太多词汇,只会让人困惑,甚至可能产生误导。所以,我认为我们应该创造新词来区分不同概念,但对于一个特定概念,我们需要坚持使用一个特定的词,以避免误导。从这个角度看,我们可以说一切只是上下文学习的一种形式。我们已经知道什么是上下文学习,现在只需要理解我们一直在做的事情如何成为上下文学习的一种形式。所以我们开始展示,例如,反向传播是上下文学习的一种形式,是一种联想记忆。当它是联想记忆时,我们就可以说模型的通用预训练阶段是上下文学习的一种形式。或者当我们谈到注意力机制或 RNN 的上下文时,同样也是上下文学习的一种形式。所以,当我们执行梯度计算时——我们可以基于梯度下降或其他优化过程定义任何 RNN——这意味着我们正在对当前发生的上下文进行学习。我认为这就是 Nessel 学习试图解决的两个主要问题:一是每个神经元获得更多计算,二是适应性和持续学习方面。
Another one is about memory perspective and adoption of these models. So, when we have a model that is adapting to the context very fast, then potentially that model can learn context. So, that's somehow one of the main messages that we try to deliver in the massive learning, which was everything that we know of somehow is a form of in-context learning. So, generally I think it's a great thing in human language that we create new words, but on the other hand, if we just create a lot of words for the same concept, it can just make us confused or it can be misleading somehow. So, I think we should create new words to differentiate different concepts, but if we have a specific concept, we need to stick to one specific word that we have for that concept to avoid misleading or anything like that. So, from that perspective, we realize that we can say everything is just a form of in-context learning. So, we already know what is in-context learning. Now, we just need to understand how what we have been already doing is a form of in-context learning. And so, that was the part we started to show that, for example, back propagation is a form of in-context learning. It is a form of associative memory. And when it's a form of associative memory, then we can say that the general pre-training phase of the model is a form of in-context learning. Or when we go to the context of attention or RNN, so on and so forth, again, it's a form of in-context learning. So, when we perform gradients, which we can define any RNN based on gradient descent or other form of optimization process, when we can do that, it means that we are doing some learning on the context that is happening right now. So, I think these two are the main things that Nessel learning is trying to address. One is about generally more computation per neuron, and another one is about adaptability and continual learning side.
你能更具体地描述一下吗?比如各个层级的相对大小、结构、上下文窗口或长度、频率是怎样的?请用非常黑白分明的方式为我们描绘出来。
Can you just describe in more specific detail, like what are the relative sizes of the levels, what are the structures of the levels, what are the context windows or lengths of the different levels, what are the frequencies? Just like map the thing out for us in kind of very black and white terms.
好的,我们从 Transformer 结构开始。在 Transformer 中,我们有注意力模块和 MLP 模块。预训练时,注意力会关注不同的上下文,试图组合所有 token,每个 token 关注它之前的所有 token,等等。然后有一个 MLP 模块,负责长期记忆。模型预训练完成后,MLP 模块就固定了,不再改变,它包含了预训练期间压缩的所有信息。而注意力在推理时负责当前上下文,MLP 模块负责长期记忆和通用的世界知识。现在,我们简单扩展这个想法:保留注意力,但用多个 MLP 模块代替单个 MLP 模块,每个模块以不同频率更新。这样做的好处是,注意力让你快速适应上下文,它非常强大,像完美记忆一样缓存一切。但另一方面,你可能想要多级记忆,这就是我们定义连续记忆系统的部分。所以,我们用多个 MLP 模块代替一个 MLP 模块。第一个 MLP 模块更新非常快,这可能导致灾难性遗忘,因为它可能会忘记几千个 token 之前的信息。但关键在于,其他 MLP 模块尚未更新,第一个模块遗忘的知识仍然存在于它们的参数中。因此,当我们通过所有层进行反向传播时,知识可以恢复。这帮助我们实现时间上的循环过程:第一个 MLP 模块可能忘记某些东西,但如果被遗忘的特定数据样本或技能很重要,它可以通过其他尚未更新的 MLP 模块恢复,因为这些模块仍然保留着相关知识。这是扩展整个 Transformer 模块的一种非常简单的方式。我们称这种变体为 Hop Attention,它是注意力加多个 MLP 模块的组合。我们还有另一种变体,即实际的 Hop 架构。我们想说的是,注意力虽然是完美记忆,能缓存一切、可扩展、速度快,但它的更新频率是无限的——这意味着注意力不知道 token 之间的时间依赖关系,需要位置编码,即使有位置编码,它也不擅长处理需要顺序推理的任务。我们的想法是用另一种将键映射到值的联想记忆来替代注意力。一种潜在的架构是 Titan,我们可以简单地用 Titan 加上连续系统来替代。但论文的前几节我们讨论过,如果只有简单的线性更新过程(如 Titan 每个块内部的更新),它可能比自引用过程弱。什么是自引用过程?梯度下降或一般的反向传播就是一种自引用过程。自引用过程的想法是学习如何学习,如何学习如何学习,等等,有很多层次。但在计算上,实现所有这些层次是不可行的。
Yeah, so we start from the transformer structure. In transformer we have attention block and then MLP block. So, what is happening there is that in the pre-training we have different contexts that attention attends to; the attention side is trying to combine all the tokens and each token attends to all the tokens before that in the context, and so on and so forth. And then there is an MLP block. And that MLP block is responsible for long-term memory. Now, when the model is pre-trained, the MLP block is fixed. It's not changing anymore. So it has all the information compressed during the pre-training. And then we have attention. So, at inference time, attention is responsible for the context it is getting, and the MLP block is responsible for the long-term memory and generally the very general knowledge of the world. Now, let's just simply extend this idea. A simple extension of this idea is that we can simply keep attention, and then instead of just one MLP block we have multiple MLP blocks. Each of them are updated with different frequency. So, now what is going on there is that why it's helpful is that when you have attention, you have fast adaptation to the context. Attention is very powerful. It's like a perfect memory. It caches everything and so it's great. On the other hand, you might want to have multiple levels of memory. And that's the part we define continuum memory system. So, instead of one block of MLP, we have multiple blocks of MLP. And now you have your first MLP block. It is updated very fast. So, what would happen in that case? In that case, the first updating process of this MLP block can cause catastrophic forgetting, because this MLP block can simply forget the information that it gets, for example, a couple of thousand tokens ago. But the point is, since the other MLP blocks have not updated so far, the knowledge that is forgotten by the first MLP block is still in their parameters. So, when we perform back propagation through all these layers, then the knowledge can come back. So, it provides and it helps us to have a loop process in time. The first MLP block can forget something, but if that specific data sample or skill that is forgotten is important, then this can come back through the other MLP blocks, which have not been updated so far, and they still have the knowledge about that specific skill or data sample. So, that's a very simple way of extending the whole transformer block. And so, we call this variant Hop Attention. It's a combination of attention plus multiple MLP blocks, and that's what we called Hop Attention. Now, we have another variant, which is the actual Hop architecture. What we are saying is that attention, as I mentioned, is a perfect memory. It can cache everything, it can scale, and it's fast, and so on and so forth. This is great. But the point is it's still the update of attention has infinite frequency. What does it mean? It means that attention doesn't know anything about the temporal dependency of all the tokens. And it needs something like positional encoding, or even with the help of positional encoding, attention is not a great model for tasks that are sequential, tasks that require sequential reasoning or something like that. What was our idea? Our idea was to replace attention with another associative memory that maps keys to values. So, that's what attention does. And now, we want to replace it with another module that tries to map keys to values. And so, one potential architecture here is Titan. So, we can simply just replace Titan and have Titan plus continuum system. And so, that's a simple idea. But the point is, in initial sections of the paper, we discussed that if you have a simple linear process of updates, which is what is happening inside each chunk of Titan update, this process can somehow be weaker than the case where we have self-referential process. So, what is self-referential process? Gradient descent or generally backpropagation is a form of self-referential process. So, what is the idea in the self-referential process? And the idea there is that we want to learn how to learn and how to learn and how to learn and how... There are a lot of levels of how to learn how to learn. And so, computationally, it's infeasible to implement all those levels of how to learn how to learn and so on and so forth.
Schwitalla 等人提出了一个想法,即自指模型。当我们有一个键值记忆时,让模型自指的一种方式是让模型生成自己的值。具体来说,假设我们想在记忆中记住某件事,比如一个特定的事件,你想记住它,比如一个特定的词。我们的大脑中存在联想记忆,我们试图将这个特定的词映射到我们已经知道的另一个概念上,从而记住它。你肯定见过这样的情况:你生成一个特定的词,然后把它映射到你已经知道的另一个词上,这样你就能记住它。我们生成要映射键的值,从而也能记住键。自指过程是一个非常通用的概念,但在这篇论文中,一个具体的设计选择是:联想记忆的值由其自身参数生成。也就是说,模型自己生成自己的值,然后尝试将键映射到值。这个过程可能是完全顺序的,你无法以简单的方式并行化它,因此它完全理解了数据中的因果关系。在需要顺序思考、顺序推理等任务中,我们可以预期这个模型比简单的注意力机制表现更好,因为简单的注意力机制根本无法顺序理解数据的因果关系。这就是我们的想法。在 Hope 架构中,我们用自修改 Titan 替换了 Titan,这正是一个生成自己值函数(联想记忆的值)的 Titan 模块。这就是 Hope 架构:自修改 Titan 加上连续记忆系统。
So, there is one idea by Schwitalla et al. They had this idea of a self-referential model. One way to make a model self-referential when we have a key-value memory is that the model generates its own value. What is happening there? Let's say that in our memory we want to memorize something. There is a specific event happening and you want to memorize it, say a specific word. In our brain, we have associative memory and we try to map this specific word to another concept we already knew, so we can memorize it. You have seen cases where you generate a specific word and map it to another word you already knew, so you can remember that. We generate the value that we want to map our keys to, so we can memorize the key as well. For the self-referential process, it's a very general concept, but in this paper, a specific design choice is that the value of the associative memory is generated by its own parameters. So, the model itself generates its own value and then tries to map keys to values. Potentially, this process is fully sequential. You cannot parallelize it in a simple format, so it has a full understanding of the causality in our data. Potentially, in tasks that require sequential thinking, sequential reasoning, or anything like that, we can expect this model to work better than simple attention because simple attention doesn't have the ability to sequentially understand the causality of the data at all. That was our idea. In the Hope architecture, we replace Titan with self-modifying Titan, which is exactly a Titan module where it generates its own value function—the value of the associative memory. So, that's the Hope architecture: self-modifying Titan plus continuous memory system.
我想再花点时间聊聊你说的“生成自己的值”是什么意思。我比较熟悉 Transformer 架构,我们有 K、Q、V 向量,对吧?训练过程会随时间修改所有这些向量。一般的启发式理解是:对于每个 token,查询向量表示这个 token 在找什么,键向量帮助表示其他 token 能提供什么,并对齐它们,找到匹配的地方,也就是相关性。然后第三个,值向量,会带来概念,这些概念被送入下游层,确保我们有正确的激活以继续处理。但所有这些向量都是学习得到的,对吧?K、Q、V 都是学习得到的。所以我不太清楚你说的模型学习自己的值是什么意思,因为 Transformer 不也学习自己的值向量吗?我不太清楚你在这里做的区分。我们假设大家至少对 Transformer 有一般了解,知道它是怎么工作的。
I think I want to spend one more beat on what you mean when you say generating its own value because I'm kind of like, okay, I know the transformer architecture pretty well. We've got these K, Q, and V vectors, right? And the training process modifies all of those over time. The general heuristic is that for each token, there's the query vector that indicates what this token is looking for, the key vector that helps indicate what other tokens have to offer and aligns those, finding where there's a match, basically where there's relevance. And then the third one, the value, brings up the concepts that then get fed into the downstream layers to ensure we have the right activations for continued processing. But all of those are learned, right? All K, Q, and V are learned. So, I'm not entirely clear on what you mean when you say that the model learns its own values because doesn't the transformer sort of learn its own value vector as well? I'm not quite clear on the distinction you're making there. Let's assume people are at least generally familiar with the transformer and know how that goes.
通常在 Transformer 中,更准确地说在 softmax 注意力中,我们会对 Q、K、V 进行投影,然后输出进入注意力机制。注意力机制对 Q、K、V 的投影没有任何控制。但当我提到模型或一般的联想记忆试图生成自己的值,然后将键映射到值时,我指的是类似梯度下降的过程。如果我们回想梯度下降,我们有前一个权重状态 W_t 减去损失函数的梯度,等于下一个权重状态。如果你观察这个过程,你可以用链式法则分解梯度,写成关于输出的梯度乘以输入数据。现在你可以看到这变成了联想记忆的形式,非常类似于线性注意力,因为它是 W_{t+1} = W_t - K,其中 K 是 X_t,然后 V 是关于输出的梯度。所以这个过程非常类似于线性注意力或任何线性循环模型,但有趣的是它与线性注意力不同,因为如果你看值(关于输出的梯度),这个梯度是 W_t 的函数,是当前权重状态的函数。所以基本上,联想记忆中的值分量不是来自这个循环公式之前的另一个分量,而是每次由这个循环过程生成的。这就是自指过程的本质。在一个非常简单的自修改 Titan 版本中,如果我们想把它设计成一个简单的 Titan,过程是:我有 X,然后进行 QKV 投影,投影成 QKV,然后将它们全部传入 Titan 模块。这是一个简单的 Titan 模块。但如果我想要一个自修改 Titan,那么所有这些 Q、K、V 投影的参数都在 Titan 模块内部进行优化。所以基本上,模型有能力以某种方式修改自己的更新规则,并生成自己的值用于进一步的记忆。这就是两者的主要区别。
Generally in transformer, or more accurately softmax attention, what happens is that we have a projection of Q, K, V, and then the output goes to attention. So, attention doesn't have any control over the Q, K, V projections. But when I say that the model or generally the associative memory tries to generate its own value and then map keys to the value, I mean something like gradient descent. If we recall gradient descent, we have something like the previous state of the weights, W_t, minus the gradient of the loss function, which equals the next state of the weights. If you look at this process, you can break the gradients using the chain rule and write it as gradients with respect to the outputs times the input data. Now you can see that this takes the form of associative memory, very similar to linear attention, because it's W_{t+1} = W_t - K, where K is X_t, and then V, which is the gradient with respect to the outputs. So, this process is very similar to linear attention or any linear recurrence model, but the interesting part is it is different from linear attention because if you look at the value, which is the gradient with respect to the outputs, this gradient is a function of W_t, a function of the current state of the weights. So basically, the value component in the associative memory is not from another component before this recurrence formula, but it is generated by this recurrent process every time. That's what is going on with this self-referential process. In a very simple version of self-modifying Titan, if we want to design this as a simple Titan, what happens is that I have this X and then QKV projection, I project it into QKV and then pass all of them to the Titan module. That's a simple Titan module. But if I want to have a self-modifying Titan, then all of these parameters of the projection of Q, K, and V are optimized inside the Titan module. So basically, the model has the control of somehow modifying its own updating rule and somehow generate its own value for further memory. That's the main difference between these two.
是的,我认为大家需要抓住的关键短语,包括我自己,是“修改自己的更新规则”。这实际上也是 Mamba 架构的一个主题。Mamba 架构的作者之前做了很多状态空间模型的工作,而 Mamba 的一个重大突破是:状态在每个时间步的更新方式现在变成了输入的函数。这增加了表达能力——使状态本身的更新依赖于它当时接收的输入,从而解锁了更好的性能。这里似乎也有类似的情况,你说我们不想让最终的值输出——基本上,我们不想过早计算值向量。我们希望它稍后出现,并且比传统的 softmax 注意力更依赖于输入、更依赖于历史。
Yeah, I think the key phrase for people to latch onto there, starting with myself, is modifying its own update rule. This is definitely a theme in general with the Mamba architecture as well. The authors of the Mamba architecture had done a bunch of previous state-space model work, and the big unlock with Mamba specifically was that the way in which the state is going to be updated at each time step now became a function of inputs. That increased expressivity—the potential for expressivity of making the actual update of the state itself dependent on the input that it's receiving at that time unlocked better performance. There's something very similar going on here, it seems like, where you're saying we want to make the final value output not—we don't want to calculate the value vector too early, basically. We want that to come a little bit later and be more input-dependent, more history-dependent than it has traditionally been in softmax attention.
这个直觉对吗?还是说其中还缺少些什么?
Is that a good intuition or is there something still missing from that intuition?
我认为这个直觉非常完美。通常,这种值的投影也在模块内部进行更新。这一点非常重要,因为它有助于适应性。一般来说,模型本身对上下文也非常适应。因此,对于每一个到来的 token,模型都在尝试学习一些东西。所以,它从关联记忆的值项中生成值的方式,与它更新记忆的方式完全相同。因此,生成值也是一个非常自适应的过程。
I think it's a perfect intuition. Generally, this projection of the value is also updating inside the module. That's a very important point because it helps for the adaptability. Generally, the model itself is also very adaptive to the context. And so, from every token that comes, the model is trying to learn something. And so, the way that it generates the value from the value term for the associative memory is exactly the same as the way it updates its memory. So, it's a very adaptive process to also generate the value.
那么,你能带我们走一遍单个时间步吗?也许我们可以先用注意力机制,再用 Titans 来演示。并突出其中的微小差异,但也要拉远视角看全局。所以,我们现在有了一个新的基础模块,对吧?如果我们先做注意力机制部分,我们有一个注意力机制,然后有多个 MLP 按更新频率从最快到最慢依次排列。然后这个模块被堆叠成层。对吗?老实说,我有点困惑,关于没有训练测试区分这一点,因为仍然存在一个训练过程,你只是拿一堆数据跑一遍模型,对吧?所以,从研究者的角度来看,也许从模型的角度来看,区别不大,但从研究者的角度来看,你仍然坐在那里运行一个过程,拿一堆数据让模型从中学习。这有点像批量处理,而不是用户每小时都在参与。那时没有其他事情发生,对吧?那么,它有什么不同呢?你知道,当我为 Transformer 做这件事时,我确实有一些很好的并行化优势。我有兴趣回头理解当前的硬件范式在多大程度上与你这里的一些东西配合良好,以及基本的循环性在多大程度上可能带来挑战,但先把这个放一边。今天,我可以并行地跑一堆 token。我们可以累积所有这些梯度,然后应用梯度,然后进入下一个时间戳,继续这样做,我对信息如何流动有很好的直觉。我可以在脑海中可视化前向传播,然后可视化反向传播的向后传递,逐渐更新所有权重。新架构的过程有何不同?与我们更熟悉的范式相比,核心差异是什么?
So, can you take us through a single time step? And maybe we can do this with the attention help and then the Titans help. And just highlight the little difference there, but let's also zoom out to the big picture. So, we've got now a new fundamental block, right? If we do the attention help thing first, we've got an attention mechanism, and then we've got multiple MLPs arranged in sequence from fastest update to slowest update frequency. And then that block gets stacked into layers. Correct? I'm a little bit confused to be honest on the sort of there's no training test distinction because there is still some training process where you're just taking a bunch of data and running it through the thing, right? So, from the researcher's perspective, maybe from the model's perspective, there's not so much of a distinction, but from the researcher perspective, you are still sitting there running a process that takes a bunch of data and has the model learn from that. Which is kind of a bulk process that's not like a user is engaging hourly. There's nothing outside of that process happening at that time, right? So, how is it different? You know, when I do this for a transformer, I have some of these really nice parallelization benefits. I am interested to come back to understand to what degree the current hardware paradigm plays nicely with some of the stuff you've got here and to what degree the fundamental recurrence may present challenges, but bracket that for a second. Today, I can run a bunch of tokens through the thing in parallel. We can accumulate all these gradients and then we sort of apply the gradients and then we have the next time stamp and we kind of keep doing that and I have a pretty good intuition for how information flows. I can visualize the forward pass in my mind and then I can visualize the backward pass of back propagation going through and gradually updating all the weights. How does the procedure with the new architecture vary? What are the core things that are different from the paradigm that we're more used to?
我认为主要的区别在于我提到的更新方面。所以,对于酷注意力机制,我认为它与当前范式非常相似,架构也与 Transformer 非常相似。它实际上是 Transformer 架构。我们只是用多个 MLP 块替换了单个 MLP 块。因此,我认为当我们进行推理时,主要区别在于对于每个 MLP 块,我们需要跟踪当前状态。是时候更新这个 MLP 块了,还是它还有时间被更新?如果是后者,那么我们就使用该特定 MLP 的最后更新状态来进行推理。如果还没有,我的意思是,如果到了更新的时候,那么我们首先通过所有反向传播的东西和当前块中到目前为止看到的所有 token 来更新它。然后,当权重更新后,我们再进行推理。从研究的角度来看,可能很难去掉这一部分。我认为即使从研究的角度来看,也许最好说我们有评估时间,而不是没有评估时间。因为似乎我们总是在更新,模型随着时间的推移不断更新,没有训练时间和测试时间。但关键是,似乎在一段特定时间内,我们不进行某些评估,我们等待一段时间,然后开始在不同的下游任务上评估模型,就像其他任何事情一样。我认为从模型的角度来看,它不知道自己是处于测试时间还是训练时间,因为一切都一样,这是一个非常独特、统一的过程。但从我们这边来看,我们是否想要评估模型并测量其在特定任务上的准确性等等,这显然很重要。所以,回到 Hope 架构,这通常适用于我提到的 Hope Transformer 或 Hope 注意力模型。但当我们进入实际的 Hope 架构时,一切又都一样了。一切都非常相似。唯一的区别是注意力机制被自修改 Titan 取代。而对于自修改 Titan,一切又都与 Titan 非常相似。所以所有的改变都只是在模型设计内部,从更高层次来看,推理非常相似。上下文或文档进入自修改 Titan,然后对于每个 token,我们得到一个输出,然后进入 MLP 块。我们有多个 MLP 块,等等。所以一切都与当前范式非常相似。
I think generally the main difference comes from the update side that I mentioned. So, for the cool attention, I think it's very similar to the current paradigm and the architecture is very similar to transformers. It's actual transformer architecture. We just replace the MLP block with multiple MLP blocks. And so, I think when we want to do inference, then the main difference comes from the fact that for each of the MLP blocks, we need to track where we are. Is it the time that we want to update the MLP block, or it's still has time to get updated? If it's the later case, then we use the last updated state of that specific MLP for doing the inference. If it has not, I mean, if it's the time to get updated, then we first update it through all the back propagation stuff and all of the tokens that we have seen so far in the current chunk. And then, when the weight is updated, then we perform the inference. From the research point of view, it might be a little bit hard to remove this part. I think even from the research point of view, it might be better to say that we have evaluation time and not evaluation time. Because it seems that we are always like update when the model is always gets updated over time, and there's no training time and test time. But the point is, it seems that for a specific period of time, we don't do some evaluation, and we wait for some time, and after that we start evaluating the model on the different downstream tasks that we have, generally like for anything else. I think from the model perspective, it doesn't know whether it is in the test time or train time, because everything is the same, and it's a very unique, uniform process. But from our side, definitely it is important whether we want to evaluate the model and measure its accuracy for a specific task and so on and so forth or not. And so yeah, coming back to the Hope architecture, that's generally for the Hope Transformer or Hope attention model that I mentioned. But when we go to the actual Hope architecture, then again, everything is the same. Everything is very similar. The only difference is that the attention is replaced by self-modifying Titan. And for self-modifying Titan, again, everything is very similar to Titan. So all of the changes are just inside the model design and from the higher level perspective, the inference is very similar. The context or document goes to the self-modifying Titan and then for each token, we have one output and then it goes to the MLP blocks. We have multiple MLP blocks and so on and so forth. So everything is very similar to the current paradigm.
我理解得对吗?有一个核心模块,要么是传统注意力机制,要么是自修改 Titan 模块,加上 MLP,然后这个模块被堆叠成一层。是这样吗?
Did I have it right that there are this core block of either the traditional attention or the self-modifying Titan module plus the MLPs, that then becomes the block that gets stacked into a layer. Is that right?
是的。这是一个设计选择。我们可以有不同的设计选择。例如,当我们……通常,Hope 的初始和主要设计是,比如对于 Hope 注意力,我们有注意力机制,然后有多个 MLP 块。每个块以不同的频率更新。但对于某些任务,我们需要使用预训练模型。例如,如果你想专注于 Llama,那么 Llama 并不是用整个架构设计的。所以,我们能做的,我们实际做的是,不采用我提到的这种形式(例如注意力然后多个 MLP 块),而是说这是注意力加 MLP 块。然后是注意力加另一个不同频率的 MLP 块。然后是注意力加另一个不同频率的 MLP 块,依此类推。所以,这算是一种设计选择。
Yes. That's a design choice. We can have different design choices. For example, when we... So generally, the initial and the main design of Hope is in the case that we have, for example, for the Hope attention, we have attention and then multiple MLP blocks. Each of them are updated with different frequency. But for some of the tasks, we needed to use pre-trained model. And for example, if you want to focus on for example Llama, then Llama is not designed with whole architecture. So, what we can do, the thing that we have done is that instead of going in this formulation that I mentioned, for example attention then multiple MLP blocks, what we have done is that we say this is attention and MLP block. Then attention and another MLP block with different frequency. And then attention and another MLP block with different frequency and so on and so forth. So, it's somehow design choice.
你需要看看自己更倾向于哪种方式。是使用现有的预训练模型,还是从头开始训练自己设计的架构?但我认为两者可能相对相似,并不会带来根本性的改变。
You need to see which one you prefer. Do you want to use an existing pre-trained model, or do you want to start from scratch and train your own designed architecture? But I think potentially both of them are relatively similar. They don't fundamentally make changes.
是的,有趣。这又是一个很好的提醒,印证了 Ilya 的那句老话:模型只是想学习。我总是惊讶于这些选择最终往往殊途同归。在 Titan 的例子中也是如此,你有三种不同的方式将记忆模块整合到更大的架构中。Mamba 在很多方面也是如此,你可以有多个状态,可以顺序排列,也可以并行排列。一旦你有了一个效果不错的模块概念,就可以像乐高积木一样以多种方式组合。不同排列方式之间可能会有一些性能差异,但通常来说,如果是一个真正重大的概念进步,具体的连接图反而没那么重要。更重要的是你添加到乐高积木组中的核心部件。所以这很好地提醒了我们这一点。
Yeah, interesting. It's another great reminder of the old Ilya maxim that the models just want to learn. It's always striking to me how many of these choices end up going one way or the other. That was true in the Titan's case where you had three different ways of working the memory module into the larger architecture. It's definitely been true with Mamba in many ways where you could have multiple states, you can have them in sequence, you can have them in parallel. Once you have one of these block concepts that seems to work well, you can Lego piece it in a lot of different ways. There will probably be some performance differences between different ways to arrange the blocks, but more often than not, if you're talking about a serious conceptual advance, the exact wiring diagram is less important. More important is the core piece that you're adding to the set of Lego pieces, so to speak. So this is a good reminder of that.
你如何看待不同 MLP 之间在大小、更新频率和学习率方面的关系?感觉学习率在这里可能很重要,存在一种潜在的等价关系:如果我在每个 token 更新一个 MLP,然后用更大的 batch size 更新另一个,我可以让它们非常相似或相当不同,具体取决于学习率。那么,大小和频率方面,更新频率较低的 MLP 是否也更小?它们是否有不同的学习率?请带我们了解一下你如何看待不同频率的 MLP 之间的关系。
How do you think about the relationship between the different MLPs in terms of size, update frequency, and learning rate? It feels like learning rate might be kind of important here, where there's potentially an equivalence: if I update one MLP every token, and then I update another one with a larger batch size, I can make them much more similar or quite a bit different depending on the learning rate. So, size, frequency, are the ones updated less frequently also smaller, and do they have different learning rates? Take us through how you think about the relationships between the MLPs of different frequencies.
我认为这确实取决于架构、参数数量以及你的设计选择。这非常相似。我们无法说出 Transformer 或注意力机制块的最佳维度是什么。这很难。它完全取决于想要使用该注意力机制的人以及我们考虑的使用场景。所以一般来说,更新频率实际上取决于你希望模型有多强的适应性,以及你希望模型如何维持其持久记忆。所以我认为这基本上是一个设计选择。关于学习率,你可以像对待 MLP 块一样对待这些块。它们是完全相同的东西。但关键在于它们有不同的更新频率。所以没有变化。一切与 MLP 块完全相同。我不认为研究学习率如何影响这些块会很有趣。我没有做过这方面的研究,也不确定确切的解决方案,但我的预期是,任何用于超参数调优的方法,我们在这里也可以同样使用。所以不应该有任何差异。
I think it really depends on the architecture, the number of parameters, and the design choices you have. It's really similar. We cannot say what is the best dimension for transformers or attention blocks. It's really hard. It really depends on the person who wants to work with that attention and the use cases we want to consider. So generally, the frequency of updates really depends on how adaptive you want your model to be, and how you want the model to maintain its persistent memory. So I think that's pretty much a design choice. About learning rate, you can treat each of these blocks the same way as you do about MLP blocks. They're exactly the same thing. But the point is they have different frequency of update. So nothing has changed. Everything is exactly the same thing as MLP blocks. I don't think it would be really interesting to see how the learning rate can affect each of these blocks. I have not done that, and I'm not sure about the exact solution, but my expectation is that any way we use to do hyperparameter tuning, we can do the same thing here as well. So there shouldn't be any differences.
有趣。所以你的一般默认方法是基于直觉的。我理解得对吗?基本上是数量级:最快的每个 token 更新一次,下一个每 10 个 token 更新一次,再下一个每 100 个 token 更新一次?你是如何开始的,为什么选择这些作为初始猜测?
Interesting. So your general default is an intuition-based approach. Do I have it right that it's basically order of magnitude: the fastest one updates every token, the next one updates every 10 tokens, the next one every 100? How did you start, and why did you pick those as your initial guess?
我们选择每个更新频率的方式是基于我们对 Titan 和其他模型使用的块大小的直觉。一般来说,Titan 中的块大小也可以定义 Titan 的频率。当时我们还没有“频率”这个术语,但块大小同样可以定义 Titan 的频率。我们使用的数值是基于我们对哪些块大小适合 Titan 的直觉。据我所知,我们使用的数字可能是 128,然后是 4 * 128,再然后是 4 * 4 * 128。大致就是这样。
The way we chose the frequency of update for each of them was based on our intuition of the chunk size that we use for Titans and other models. Generally, the chunk size in Titan can define the frequency of the Titan as well. At that time, we didn't have this term of frequency, but the chunk size can define frequency for Titan as well. What we used was based on our intuition about what chunk sizes are good for Titans. As far as I remember, the numbers we used were possibly 128, then 4 * 128, and then 4 * 4 * 128. So it was something like that.
那么知识迁移呢?我们应该如何看待……而且仍然有跳跃连接。是的,一切都相似。我们在这个学习中试图做的一件事是……不幸的是,这导致了一些对我们所做工作的误解。我看到一些评论说这里的一些概念是新的等等。但关键在于,我们实际上试图包含所有我们已经知道的概念,以表明这是一个通用的学习范式。它并不与我们当前的理解相矛盾。它只是在一个新方向上补充和完善了我们已知的内容。例如,当你进行深度学习时,任何形式的深度学习,如果你说你在使用注意力机制,你实际上是在使用嵌套学习。但在深度学习中,你只能看到每个学习问题的最终解决方案。所以你在注意力机制内部有一个学习问题,你试图解决一个回归问题,而该回归问题的非参数解就是注意力机制。当你从深度学习角度看待一切时,你只能看到每个组件的最终解决方案。但当你从嵌套学习角度看待一切时,你也能看到每个组件的内部学习过程。所以总的来说,它并不与我们已知的东西相矛盾,而是以某种方式补充了我们已知的一切并超越了它。我认为这通常是一个重要的部分。所以一切都可以非常相似。你可以进行分离。
And how about knowledge transfer? How should we think about the way in which... And there are still skip connections. Yes, everything is similar. One thing we tried to do in this learning is... Unfortunately, it caused some misunderstanding about what we were doing. I have seen some comments that some of the concepts here are already new and something like that. But the point is we tried to actually include all those concepts that we already knew to show that it is a universal learning paradigm. It's not something that contradicts our current understanding. It just complements and completes what we already know in a new direction. For example, when you are doing deep learning, any form of deep learning, and you say you are using attention here, you are actually using nested learning. But in deep learning, you only see the final solution of each learning problem. So you have a learning problem inside the attention and you are trying to solve a regression problem, and the non-parametric solution to that regression problem is attention. When you see everything from the deep learning side, you can only see the final solution for each component. But when you see everything from the nested learning side, you can see the internal learning process of each component as well. So in general, it's not something that contradicts what we already knew, but it somehow complements all the things we knew and goes beyond that. I think that's generally an important part. So everything can be very similar. You can have a separation.
是的,那么请帮助我们理解我们应该如何看待不同频率的 MLP 所扮演的角色。
Yeah, so then help us understand how we should think about the roles that the different frequency MLPs are playing.
知识转移是思考这个问题的一种方式。另一种提问方式可能是:它们如何互补?如何协同工作?更新快的组件如何逐渐影响更新慢的组件?更新慢的组件又如何引导更新快的组件走向正确的方向?你如何看待这些不同组件之间的相互作用?
Knowledge transfer is one way to think about that. Another way to frame the question might be how do they complement one another? How do they work together? How does the one that's updating fast gradually inform the ones that are updating slow? How do the ones that are updating slow kind of steer the ones that are updating fast in the right directions? How do you think about the interplay between those different components?
是的,我认为知识转移,尤其是想出不同的知识转移方式,在这里非常重要。每个组件有不同更新频率的主要目的有两个部分,我们一开始也讨论过。第一部分是帮助模型更长时间地保持记忆。举个例子,假设有一对双胞胎,其中一个乘坐宇宙飞船以光速飞行。出发前,他们有一段清晰的记忆,比如一起吃了午餐。当那个光速旅行的人回来时,地球上已经过去了 80 年。留在地球上的双胞胎忘记了那顿午餐,因为那是 80 年前的事了。但旅行的那位却记得所有细节,因为对他来说那只是几秒前的事。为什么会这样?因为各自记忆的更新次数不同。活了 80 年的人记忆被更新了很多次,而接近光速旅行的人记忆只更新了很少几次。从这个角度看,记忆更新的次数非常重要。这个例子虽然不精确,但要点很清楚。当我们有两个组件,一个更新频繁,另一个更新缓慢时,慢的那个有机会向快的那个学习。因为当快的组件被更新时,慢的组件还没有被更新,所以快的组件有机会在更新前将知识转移给慢的组件。我们也可以结合“睡眠”过程来讨论。其想法是,当我们有多层 MLP 块时,每个块以不同频率更新。一个简单的做法是,在更新快的 MLP 块之前——这里“快”是相对而言——我们有可能遗忘一些东西。所以在遗忘之前,我们可以将这个块的知识转移到下一个块,然后再更新它。这就需要一种好的知识转移方式。例如,一种方法是进行上下文蒸馏。如果你想将一个 MLP 块的知识传递给另一个,某些上下文蒸馏方法可以很好地工作。这与我们在“睡眠”论文中做的非常相似。所以,我认为知识转移的主要作用是帮助慢网络利用快网络的优势。关于不同频率的另一个点是关于记忆以及模型如何管理记忆,类似于我提到的双胞胎例子。
Yeah, I think generally knowledge transfer, coming up with different ways of knowledge transfer, is really important here, in my opinion. So, the main point of having frequency for each component is two parts. I think we also discussed it in the beginning. The first part is that it helps the model maintain its memory for a longer time period. For example, let's say that we have twins and one of them goes to a spaceship and moves at the speed of light. Right before that, they have a very good memory, like they had lunch together. Then the person moving at the speed of light comes back, and 80 years have passed on Earth. Their sibling forgets about that specific lunch because it was 80 years ago. But the twin who traveled remembers all the details because it was just seconds ago for them. Why does this happen? Because of the updates in each of their memories. The person who lived 80 years had their memory updated many times, while the person who moved at near light speed had their memory updated very few times. So from this perspective, we can see that the number of times we update memory is very important. It's a rough example, but the main point is clear. When we have two components, one updated many times and the other slower and updated less, the slow one has the opportunity to learn from the fast one. Because when the fast one gets updated, the slow one hasn't been updated yet, so there's a chance for the fast one to transfer knowledge to the slow one before it gets updated. I think we can also discuss this in the context of the asleep process. The idea is that when we have multiple levels of MLP blocks, each is updated with different frequency. One simple thing is that before updating the fast MLP block—and by fast I mean it's a relative term—there's a chance we forget something. So before forgetting, we can transfer the knowledge of this block to the next one and then update it. That's where we need a good way of knowledge transfer. For example, one way is to do context distillation. If you want to pass knowledge from one MLP block to another, some methods of context distillation can work very well. This is very similar to what we do in the asleep process in the asleep paper. So, I think the main role of knowledge transfer is to help the slow network take advantage of the fast network. Another point about having different frequency is about memory and how the model can manage its memory, similar to the twin example I mentioned.
信息从快速更新层转移到较慢层的机制是什么?但肯定也有反向的信息流动,对吧?如果你把它们在某种意义上视为纯粹的感知,那么我甚至不完全清楚自己大脑中发生了什么,但感觉上,从感知模块到更高阶推理模块的信息流,比从推理回到感知的信息流要多。但那个反向信号仍然很重要,对吧?比如,我的高阶过程确实会告诉我的眼睛往哪里看,聚焦在哪里,并说“我们需要更仔细地观察这个细节”,我想更好地理解这一点。所以,请多讲讲机制上或程序上,快速更新层中的信息是如何转移的,我们如何确保存储了快速更新所学到的重要东西,以及反向流动的信号是什么?
What is the mechanism by which the information in the fast update layer gets moved to the slower layer? But then there's also got to be something going the other way, too, right? If you conceive of them as like pure perception in a sense, then I guess I'm not even entirely clear on what's happening in my own brain, but certainly like there's more information flow, it feels to me like from my sort of perception modules to my higher order reasoning modules, whatever, than there is from the reasoning back to the perception. That signal, but it is that signal is still important, right? Like my higher order processes do tell my eyes where to look and do tell them like where to focus and do say, you know, hey, we need to like zero in on this this detail a little bit. Like I want to understand that better. So like go put some of your bandwidth into this particular thing. So Yeah, I guess give me a little more on how the mechanistically or or procedurally how the information in the fast update layers is getting transferred and and how we're making sure that we're storing what really matters from what the fast updates have learned, but then also what is the what is the signal that flows the other way?
让我用一个非常简单的例子来回答。假设我有模型 A,我想更新它的快速 MLP 块,同时确保快速 MLP 块中的信息不被遗忘,并能传递给慢速 MLP 块。一个简单的做法是将模型 A 的所有参数复制到模型 B。现在我有两个相同的模型:模型 A 和模型 B。对于模型 B,我更新快速网络,即快速 MLP 块。现在模型 B 中快速 MLP 块的参数是自由的。我想做的是改变模型 B 中慢速 MLP 的参数,使得模型 B 的输出能够模仿模型 A 的输出。如果成功,这意味着模型 A 的所有信息都压缩在快速 MLP 块中,而模型 B 中这些信息已经消失。现在,如果我能够以某种方式修改模型 B,使其模仿模型 A,那就意味着我将快速 MLP 中的知识转移到了模型 B 的慢速 MLP 参数中。这只是一个简单的例子。这个过程非常类似于蒸馏过程。我们将模型 A 的知识蒸馏到模型 B 中。所以从这个角度看,这是从快速 MLP 到慢速 MLP 的一种知识转移方式。
Let me answer that with a very simple example. Let's say I have model A and I want to update its fast MLP blocks, and I want to make sure that the information in the fast MLP block is not forgotten and can pass to the slow MLP block. One simple thing I can do is to just copy all the parameters of model A to model B. Now I have two identical models: model A and model B. What I do for model B is that I update the fast network, the fast MLP block. Now the parameters in the fast MLP blocks of model B are free. And what I want to do is to change the parameters of the slow MLP in model B in a way that the output of model B can mimic the output of model A. If that happens, it means that model A has all the information compressed in the fast MLP block, while all that information is gone in model B. Now, if I could somehow modify model B so that it can mimic model A, it means I have transferred the knowledge in the fast MLP to the parameters of the slow MLP in model B. That's just one simple thing. This process is very similar to the distillation process. We are distilling the knowledge of model A into model B. So from this perspective, we can see that this is one way of knowledge transfer from fast MLP to slow MLP.
这只是一个例子。另一个非常常见且流行的例子是反向传播。所以,如果你只是顺序连接你的 MLP 块,然后在某个时刻进行反向传播,那么你就可以将一个块的知识转移到另一个块,以此类推。
That's for example, one example. Another example, which is very common and popular, is back propagation. So, if you just sequentially connect your MLP blocks and then at some point perform back propagation, then you can transfer the knowledge of one block to another one and so on and so forth.
那么,你描述的那种复制和蒸馏过程,本质上就是《语言模型需要睡眠》那篇论文里发生的事情吗?
So, the sort of copying and distillation process you described, that's essentially what's going on in the language models need sleep paper?
是的,但还有一些额外的细节。例如,我们在那里所做的还包括给模型 B 添加额外的参数,以确保它有足够的容量来存储刚刚获得的新知识。
Yes, with some additional detail. For example, what we do there is that we also add additional parameters to the model B as well to make sure that it has enough capacity to store the new knowledge that it has just gotten.
这是否也意味着,在这种嵌套学习版本中,你实际上只是让反向传播自行运作,并没有过度设计。我们只有这些 MLP 块,它们以不同的频率更新,你只是让梯度下降自行运作,更新就这样生效了。基本上就是这样。
Does that mean also that in the nested learning version of this, you're really just letting back propagation do its thing and you're not really over engineered it all that much. We just have these MLP blocks, they get updated at different frequencies, and you're just kind of letting gradient descent do its thing and the updates are just kind of working. That's basically it.
是的,完全正确。是的,在整个图景中,一切都只是反向传播。
Yes, exactly. Yes, in the whole picture everything is just like back propagation.
从这次对话中我得到的一个巨大收获(这在论文中并没有那么清晰地体现出来)是,这真的还处于概念验证阶段,而它效果这么好恰恰说明了这是一个多么好的概念。但我们在这里讨论的一切,都没有像主流模型那样经历过参数和超参数的极致探索、优化以及多年来积累的各种细微改进。这里还没有发生这些。所以,关于这个版本、那个版本、这种配置、这种安排、顺序、并行、多少层、相对大小、相对学习率,我们仍然有很多问题可以问。这里还有大量的空间有待探索。但是,基本上,仅仅采用这些核心概念中的几个,主要就是不同 MLP 块的不同更新频率,仅此一项就产生了一些相当令人印象深刻的结果,这些结果在性质上与我们习惯看到的不同。所以,也许我们花点时间谈谈一些结果。我的意思是,论文中进行了许多不同的测试,有包含各种不同指标的大表格,其中一些是经典的困惑度评分之类的东西。你认为哪些是最重要、最能揭示问题的结果,能让人们说:“啊,因为我看到它能做到这一点,我就知道这里确实有我需要认真对待的东西。”
A huge takeaway from the conversation that didn't come through to me as clearly in the paper is just like this is really proof of concept stage stuff and the fact that it works so well shows what a good concept it is, but nothing that we're discussing here has been through the same kind of thing that the main line models have been through where everything has been parameter hyperparameter explored to the nth degree and optimized and all the little refinements that have been made over time. That hasn't really happened here. So, there's a lot of questions that we still could ask about like this version, that version, this configuration, this arrangement, sequence, parallel, how many layers, relative sizes, relative learning rates. There's like a ton of space there still to explore. But, basically, just taking a few of these core concepts, the main one being the different frequency of updates for different MLP blocks, that alone creates some pretty impressive results that are like qualitatively different from what we are used to seeing. So, maybe let's take a minute and just talk about some of the results. I mean, there's a lot of different tests run in the paper and big tables of results with a whole bunch of different metrics, some of which are your perplexity scoring classic type of stuff. What do you think are the most important, revealing results that you think people should say, "Ah, because I see that it can do that, I know that there's really something here that I need to grapple with."
是的,关于这一点,我们有一个我个人非常喜欢的持续学习风格的任务。这个想法是,我们有一个预训练模型,并且有一种模型从未见过的特定语言。我们想让模型在上下文中学习这种语言。我们让模型在上下文中学习这种语言,关键在于我们有所有的语法、所有的单词,有一个单词词典。所以,我们将所有这些都通过上下文传递给模型,然后模型学习这种语言,然后我们问:你能将这段特定文本从那种语言翻译成英语吗?然后我们可以看到,模型能够以非常高的质量(虽然不是完美)翻译那段特定文本。所以,模型似乎能够在上下文中理解那种语言,然后将其用于翻译任务。但关键是,我们再进一步。不是一种语言,而是把两种语言放在上下文中。然后让模型将每种语言的不同文本翻译成英语。在这种情况下,我们可以看到模型几乎崩溃,无法翻译任何一种语言。关键在于模型无法很好地处理其上下文,无法分别充分理解每种语言。这对于基于 Transformer 的架构来说通常是一个非常困难的挑战。但关键是,当我们把架构改为 HOPE 或 HOPE 注意力时,我们仍然有注意力机制,但我们有多层次的上下文学习,多层次的 MLP 块。所以,我们可以看到的其中一点是,当我们增加层数时,模型在这两种语言上的性能越来越好。为什么会这样?因为模型有更好的内存管理方式,因为它理解,例如,不太需要的临时知识可以存储在第一层 MLP 块中,而对语言的更深理解可以传递给后面更稳定的 MLP 块。然后,当我们有越来越多的块时,我们可以看到性能越来越好。所以在我看来,这是一个非常好的评估,用于理解模型可以在上下文中学习,并且总体上类似于持续学习。
Yeah, on this we have one continual learning style task that I personally really like. So, the idea there is that we have a pre-trained model and there is one specific language that the model has not seen before. And we want to learn that specific language in context to the model. We want to learn the model that a specific language in context and then the point is we have all the grammars, we have all the words, there's a dictionary of words. And so, we pass all of them through the model in context and then the model learns the language and then we ask that can you translate this specific text to English from that language to English. And then we can see that the model perfectly not perfectly but in a very very good quality can translate that specific text. So it seems that the model is capable of understanding that language in context and then use that for some translation task. But the point is let's just go one step beyond that. Instead of one language, let's put two languages in context. And then ask the model to translate different text from each of these languages to English. In that case, we can see that the model almost collapses and cannot translate any of those languages. The point is the model cannot handle its context well and fully understand each of the languages separately. And that's generally a very hard challenge for transformer based architectures. But the point is when we change that architecture to hope or hope attention, again we have attention but we have multiple levels of in-context learning, multiple levels of MLP blocks. And so one thing that we can see is that when we increase the number of levels the performance of the model in both of these languages gets better and better. Why is it happening? Because the model has a better way of memory management because it understands that, for example, temporal knowledge that are not very needed can be stored in the first MLP block and more understanding of the language can pass to the more stable MLP blocks later. And then when we have more and more blocks we can see that the performance gets better and better. So in my opinion that's a very good evaluation for understanding that the model can learn in context and generally like continual learning.
所以,这就是同一个……我记得我有一段时间没想过这个了,但我认为可能是 Gemini 2,甚至可能是 Gemini 1,当时引入了一个指标,即从一本书中学习一门新语言。就像有一种极度濒危的语言,有人真正研究过它,并写了一本在互联网上任何地方都找不到的书,解释了这种语言是什么,然后他们就把那本书放进上下文,然后说,基于此,进行翻译。这似乎就是……我不知道这是否就是我之前熟悉的那同一个测试,还是略有不同,但看起来这里的语言是满语,我刚查了一下,这是一种来自中国某地的极度濒危语言。所以,基本上就是这个想法,对吧?这是一种语言模型基本上没有先验知识的语言,它们被给予了一份来自某位人类学家或任何进行过实地考察的人编写的非常详细的入门材料,然后它们的任务就是应用它。
So this is the same I remember it's been a while since I thought about this but I think it was with maybe Gemini 2 maybe I don't even know maybe it was even back as Gemini 1 there was this metric introduced of learning a new language from basically one book. There's like it was like there's some critically endangered language one person has really studied it and made like a book that's not on the internet anywhere that sort of explains what this language is and then they just put that book into context and say okay based on this go ahead and do translation. It seems like this is I don't know if this is the exact same test as that one that I was previously familiar with or if it's a bit different but it seems the language here is Manchu I just looked it up it's like a critically endangered language from somewhere in China. So that's basically the idea right it's a language that the language models have basically no prior knowledge of they're given a sort of very detailed primer on this language from some anthropologist or whatever who's gone out and done the field work and then their job is to apply that.
我在看嵌套学习论文中的图 8,我的理解是,如果只有一种语言,所有模型的表现都差不多。但正如你所说,当增加到两种语言,任务难度翻倍时,传统的 Transformer 上下文学习方法表现就很差了。然后我理解 Hope 1、Hope 2、Hope 3 是那种……有多少个层级?就是有多少种不同的频率更新机制?所以当你从传统方法增加到一种、两种、三种额外的更新频率时,性能几乎能恢复到只有一种语言时的原始水平。
And so I'm looking at figure eight in the nested learning paper and what I'm taking from this is all of the models do kind of similarly if there's just one language but as you said when you go up to two languages and sort of double the difficulty of the task then the in-context learning traditional transformer approach performs quite badly and then I understand hope one hope two hope three are those like how many levels like how many different frequency update things exist? So when you move from traditional to I guess one additional two additional three additional frequencies of update you get basically almost all the way back to the original level performance with just one language.
是的,完全正确。
Yes, exactly. Yes.
好的,这非常有趣,真是迷人的东西。还有 MTOB 是什么?让我搞清楚。MTOB 是另一个……
Yeah, okay. That is quite interesting. Yeah, fascinating stuff. And what is MTOB there just so I have that clear? MTOB that's the other
嗯,是的,那是另一个数据集。是另一种语言,在模型的预训练阶段也从未出现过。
Uh yes, that's another dataset. Uh it's another language that is also has not been seen during the pre-training of the model.
好的,这就是我记得的那个。你们在 2023 年底把满语加进去了。所以这两种都是非常罕见的未知语言,被翻译成英语。只有通过多层结构,模型才能在同一个上下文中同时处理这两种语言。确实非常非常有趣。
Yeah, okay. This is the one that I recall. Yeah, so you guys added the Manchu language to this one from yeah, late 2023. So both of these like very rare unknown languages being translated to English. And only with the multiple layers can the models do both at the same time in one context. Yeah, very very very interesting indeed.
你怎么看待……我真的很喜欢那个回答,那是个很好的直觉构建。你怎么看待像困惑度分数这类东西?我的意思是,你有一个大表格,展示了困惑度和一些准确性指标,基于一些经典的基准测试。我要说明一下,我们目前将这些模型扩展到大约 10 亿参数规模。你有 7.6 亿参数、300 亿 token 的模型,以及更大的 13 亿参数、1000 亿 token 的模型。显然,按今天的标准这不算大,但即便如此,有一个很清晰的信号:Hope 架构在几乎所有维度上都优于你对比的其他模型,包括你的 Transformer、Mamba 及其变体,甚至 Titans。RetNet 也在里面,DeltaNet 也在。你怎么解读这些结果?这又回到了那个 G 问题。这是一个好的衡量标准吗?还是只是我们现有的最佳标准?你认为人们应该对这些困惑度表格给予多大的重视?
How do you think about kind of the I really like that answer. That's a great intuition builder. How do you think about just kind of things like perplexity scores? I mean, you've got a big table that shows on things like perplexity and sort of some accuracy stuff on some of these like basic kind of classic battery of tests. And I should say we're scaling these models so far up to roughly the 1 billion parameter scale. You've got 760 million parameters and 30 billion tokens and then the bigger is 1.3 billion parameters and 100 billion tokens. So obviously that's like not huge by today's standards, but nevertheless, you know, there's a pretty clear signal that the Hope architecture is on just about every dimension outperforming all the other things you're comparing it against, which includes your transformer and your Mamba and Mamba variations and even Titans. RetNet is in there, DeltaNet is in there. How much How do you interpret these? You know, this kind of goes back to that G question. Is this like Is this a good measure? Is it just the best measure we have? What do you think about like how much stock people should put in these like perplexity tables?
你知道,社区里有一些标准,比如执行某些基准任务,但并非所有任务都是评估模型的最佳方式。我们只是需要做这些,以确保每个人都能看到性能优势的来源。所以,我想我有一个具体的 Transformer 结构,我有一个想法:如果我在 Transformer 中加入遗忘机制,它就能在噪声数据上表现良好,比如这样。我不确定,我只是举个例子。那么在这种情况下,如果没有噪声数据,我只在非常干净的数据上测试我的方法,就无法展示我的方法的优势。所以,我认为这里也是同样的情况。我们讨论的是不需要预训练的模型,没有测试时间,没有训练时间,等等。但另一方面,很多基础设施仍然建立在测试和训练时间上,很多评估也是基于测试和训练时间。所以,每个人也都期望我们报告一些关于预训练困惑度的结果,以及一些评估,其中大部分是短上下文语言建模任务,它们不需要非常复杂的模型来理解长上下文建模。所以,我在 NeurIPS 的嵌套学习演讲中也提到过,我们没有用表 2 以及那些困惑度和语言建模任务来论证 Hope 的强大。我们只是用那个表格来说明,Hope 作为骨干网络并不比其他模型差。你可以看到它表现良好,但有人可能会说它相对于其他模型只是略有提升。但关键是,这不是我们旨在解决的方向。而且,某种程度上,很好的是我们能看到,即使在这个并非嵌套学习和 Hope 目标的方向上,我们也能展示一些改进,尽管是边际的。
You know, there are some standards in the community for like performing some benchmark task and not all of them are the best things to do for evaluating the model, but we just need to do them to make sure that everyone can see where the performance advantages come from. And so I think I have one specific transformer structure and I want to like uh I have one idea that if I add, for example, forgetting to the transformer, then it can uh perform well on noisy data, for example. I'm not sure. I'm just like uh coming up with just one example. And uh so in that case, if there is no noisy data and I just test my method on a very clean data, uh then there is no way that I can show the advantages of my approach in that case. So, I think here is exactly the same thing. We are arguing about models that do not need to be pre-trained. There is no test time, there is no train time, and so and so forth. But on the other hand, there is still a lot of infrastructure built on like test and train time, a lot of evaluations are built on test and train time. And so, everyone also expects us to report something about pre-training perplexity and some evaluations that most of them are some short-term and short context language modeling task. And they do not need to have a very complicated model to understand long context modeling. So, I think I also like mentioned that in the presentation of nested learning at NeurIPS, we didn't like use table two and all those perplexity and language modeling tasks to argue that hope is powerful. We just use that table to say that hope is not less powerful as a backbone compared to other models. And you can see that like it performs well, but someone might say that it's marginal compared to other like models, but the point is this isn't the direction that we aim to solve. And somehow it's really good that we can see that even in this direction that is not the goal of the nested learning and hope, we can show some improvement even if it's marginal.
明白。我一直很喜欢尝试更好地理解不同架构的微观技能。例如,当然,Transformer 因为整个序列、整个上下文始终在工作记忆中,所以很难被击败。而且我觉得你现在甚至有一个理论论证,认为在某些需要从上下文窗口回忆的任务中,它可能几乎不可能被击败。但我们也看到,比如 Mamba,它在从稀疏信号中学习方面表现更好。这是一种微观技能,该架构擅长而 Transformer 相对吃力。你在 Hope 案例中看到了什么?有没有一些微观……我认为这非常有趣,因为它确实会累积到整体性能,以及这些东西实际上擅长或不擅长什么。我的意思是,在上下文中回忆某些东西的能力在需要时非常重要。从噪声中学习或过滤噪声、找到真正重要的信号的能力在需要时也非常重要。那么,有没有特别突出的微观技能?语言翻译那个在宏观上是一个有趣的任务,因为它很难。但我想知道,如果你深入到这些非常微观的构建块能力——模型或架构可以有也可以没有——那么 Hope 有什么是 Transformer 没有或没那么强的?
Yeah, got you. I'm always a big fan of trying to get a little bit better sense for the micro skills of different architectures. So, for example, of course, you know, Transformers because the full sequence, the full context is in working memory at all times, it's pretty hard to beat. And I feel like you even sort of have like kind of a theoretical argument now that like it may be even be kind of impossible to beat in some in some of these tasks where the idea is like recall from the context window. But then, you know, we saw things with Mamba, for example, where it was better at learning from like sparse signal. And this was sort of a micro skill that that architecture excelled at that the Transformer relatively struggled with. What have you seen in the hope case, you know? Are there little micro And I think this is very interesting because it does kind of ladder up to the overall performance, you know, and what these things are actually good or bad at, right? I mean, the ability to recall something in context is really important when you need it. The ability to learn from or kind of filter out noise and get to the signal that really matters is really important when you need it. So, are there particular micro skills that stand out to me? That the language translation one is an interesting one in a macro sense of like that's a hard task. But I wonder if you drill down to these like very micro building block competencies that models or architectures can either have or not have, what stands out in terms of what this has that Transformers don't have or don't have as strongly?
当我们谈论上下文回忆任务或一般回忆密集型任务时,在我看来,所有这些任务都是为 Transformer 设计的。它们不是为了比较架构而设计的,而是专门为 Transformer 设计的。为什么我这么说?因为你不能期望一个模型甚至一个人完美地完成大海捞针的任务。或者,例如,做一些回忆密集型任务。比如,假设你有一段几千行代码的程序,然后你只想回忆某一行代码中 X 的值是什么。
When we are talking about in-context recall task or generally like recall-intensive tasks, in my opinion, all those tasks are designed for Transformers. They are not designed to compare architectures, but they are specifically designed for Transformers. Why I'm saying that? Because you cannot expect from a model or even a human to perform needle in haystack perfectly. Or for example, do some recall intensive task. For example, assume that you have a code and like couple of thousand lines of code and then simply you want to recall what was the value of X at some line of the code.
所以这几乎是不可能的,或者至少对人类甚至其他模型来说都非常困难。但另一方面,对 Transformer 来说却相当简单,因为它们可以直接访问上下文中的整个历史记录。因此,找到那个 token 并将其作为输出传递非常简单。在这种需要上下文回忆的密集回忆任务中,如果你将第一代循环架构与 Transformer 进行比较,可以看到循环架构的表现差距非常大。现在这个差距正在缩小,其他循环模型的性能也非常出色。但对我来说有趣的是,它们至少缩小了与 Transformer 的性能差距,而人们原本并不期望它们能做到这一点。我们期望 Transformer 能做到,因为它有注意力机制,但我们不期望基于压缩的模型能执行回忆任务。所以我觉得这很有意思。
And so it's almost impossible, or at least it's very, very hard for humans or even for other models to do that. But on the other hand, it's pretty much simple for transformers because they have direct access to the entire history in their context. And so it's very simple to just find that token and pass it as the output. In recall-intensive tasks like this in-context recall task, the gap between recurrent architectures, which they perform, is also very great if you compare the first generation of recurrent architectures to the transformer. We can see that this gap was much larger. Now this gap is getting smaller, and the performance of other recurrent models is also very great. But the interesting part for me was that they at least close this gap in the performance of the model compared to transformers, while they are not expected to do that. Transformers, we expect them to do that because they have attention block, but we don't expect a compression-based model to perform recall tasks. And so I think somehow it was interesting.
那么 MAD 数据集测试的是什么?我们该如何理解?因为你刚才说,在这些“大海捞针”式的、需要从早期上下文中回忆的困难任务中,Transformer 仍然是最好的。循环模型只有一些潜在表示,无法回顾原始文本,所以表现没那么好,但随着每一代改进——这里你提到了 Mamba,整个架构在那些运行时工作记忆中没有完整显式上下文的循环模型中表现最好——差距正在缩小。但转到 MAD 数据集,整个架构的表现却超过了包括 Transformer 在内的所有模型。这测试的是哪些微观技能?我们应该从这一结果中得到什么启示?
And what does the MAD dataset get at? How should we understand that? Because you just said that on these needle-in-a-haystack, very difficult recall tasks from earlier in context, the transformer remains the best. The recurrent models, which only have some latent representation and don't have the ability to look back at the original raw text, don't perform as well, but with each generation of improvement, and here you've got Mamba, the whole architecture does the best of the recurrent ones that don't have the full explicit context in working memory at runtime, and so that gap is closing. Flipping over to the MAD dataset, here the whole architecture is performing better than everything including the transformer. What micro skills is that testing? What should we take away from that result?
MAD 数据集也与回忆密集型任务非常相似,但关键在于它有多种不同的设置。例如,其中一种是带噪声的上下文回忆。我们要执行回忆和上下文回忆任务,但 token 中带有噪声。当 token 中有噪声时,我之前解释的 Transformer 在纯上下文学习中的优势,现在反而成了它的弱点,因为它很容易混淆哪些 token 是噪声、哪些不是。因此,与 Mamba 这样的模型相比,这个任务对 Transformer 来说可能更难一些。但同样,这也取决于 RNN 的记忆管理。如果 RNN 没有很好的记忆管理系统或更新机制,它也可能被噪声搞糊涂,遇到一些问题。但如果记忆管理很强,那么过滤掉任务中的所有噪声 token 就会简单得多。所以我认为这是其中一点。另一个有趣的任务是压缩。压缩任务顾名思义,我们要压缩 token,预测一个代表一组 token 压缩版本的单一 token,然后从中重建原始序列。因此,对于 RNN 这样的模型来说,这可能更简单,因为它们已经知道如何正确压缩数据,而另一方面,Transformer 执行这个任务会更困难。总的来说,正如我提到的,所有这些任务都是回忆密集型任务或上下文回忆的某种变体,但关键在于还有其他方面。例如,选择性复制是另一个方面。模型还有其他非常重要的方面。我们也应该看看模型在这些方面的表现,而不是只针对一个特定指标进行过度拟合评估。
The MAD dataset is also very similar to the recall-intensive task, but the point here is that there are different setups for it. For example, in one of them is the noisy in-context recall. We want to perform recall and in-context recall task, but the point is we have some noise in the tokens. And now when we have that noise in the tokens, somehow the power of transformers that I explained in the previous setup, which was pure in-context learning, now is its weakness somehow because it can get simply confused about which token is noise, which token is not, and so forth. So potentially this task becomes a little bit harder for transformer compared to a model like Mamba. But again, that also depends on the memory management of the RNN. For the RNN, if it doesn't have a very good memory management system or generally update mechanism, then potentially it can simply get confused by the noise as well and face some issues. But again, if the memory management is strong, then it's much simpler to filter all those noise tokens in the task. So I think that's one thing. Another task that is also interesting here is about compression. The compression task, the name explains the task itself, but we want to compress the tokens and predict one single token that is the compressed version of a set of tokens. And then we want to reconstruct the original sequence from that. So potentially it's a simpler task for models like RNN because they already know how to compress the data properly, but on the other hand, transformer has a harder time performing this task. So generally, as I mentioned, all of these tasks are somehow modified versions of recall-intensive tasks or in-context recall, but the point is there are other aspects. For example, selective copying is another one. There are other aspects to the model that are very important. And we should also see how the model performs in those aspects and not just overfit our evaluation on one specific metric.
酷。我觉得底层细节聊得差不多了。而且我开始理解论文标题“架构的错觉”了。在论文第 39 页,你们还提出了一种新的优化器,它不仅超过了旧的 Adam 标准,甚至超过了 Muon。它确实带来了一些计算开销,但我认为论点仍然是,它在更快的收敛或更好的学习方面完全值得。关于你所说的 M3 优化器,还有什么想补充的吗?
Cool. I think that's probably enough on the really low-level stuff. And I think this illusion of architecture, title of the paper, starts to click for me. On page 39 of this paper, we get to the part where you also have a new optimizer that is outperforming not just your old Adam standard, but also even outperforming Muon. It does come with a little bit of computational overhead, but I think again the argument is that it more than pays back for itself in terms of faster convergence or just better learning. Is there anything you want to add on the M3 optimizer as you call it?
首先,我想澄清一点:对于优化器来说,很难说某个特定的优化器就比另一个更强大。这实际上取决于问题设置,甚至是一般的问题本身。例如,我们可能在回归任务上评估优化器。但另一方面,如果你训练一个语言模型,你可能会发现趋势完全不同。所以一般来说,优化器的设计以及判断哪个更好,确实取决于任务或具体的问题设置等因素。这也是我们在嵌套学习中想要传达的主要观点之一。我们想说的是,整个架构及其优化过程只是一个相互连接的嵌套优化问题系统。为什么是相互连接的?因为优化侧的梯度是由架构生成的。如果你有一个简单的架构,那么梯度就非常简单。如果你有一个复杂的架构,梯度中的模式可能会非常复杂。然后当你引入动量项时,动量是一种试图压缩梯度的联想记忆形式。因此,例如,如果你的梯度非常复杂,你就需要更强大的记忆管理系统来处理动量。或者如果架构非常简单,那么即使是没有动量的简单梯度下降也可能工作得很好。所以总的来说,我们在论文中的一个论点是,我们应该把一切看作一个相互连接的系统,并尝试设计出整体上能产生良好模型架构或广义机器学习模型的东西。这就是一个论点。
First, let me clarify one point: truly for optimizers, it's a little bit hard to say that this specific optimizer is more powerful than another. It really depends on the problem setup or generally even the problem. For example, we might be evaluating the optimizer on a regression task. But on the other hand, if you train a language model, you might see that the trend is completely different. So generally, the design of an optimizer and saying which one is better than the other really depends on the task or the chair problem setup and all these things. And so that's also one of the main points that we wanted to deliver in nested learning. What we are saying is that the entire architecture with its optimization process is just one interconnected system of nested optimization problems. And this is interconnected. Why is it interconnected? Because the gradients of the optimization side are generated by the architecture. If you have a simple architecture, then the gradients are very simple. If you have a complicated architecture, the patterns in the gradients can be very complicated. And then when you have momentum term, momentum is a form of associative memory that is trying to compress gradients. So for example, if your gradients are very complicated, you need a more powerful memory management system for your momentum. Or if it's very simple, if it's a very simple architecture, then even a simple gradient descent without any momentum might work very well. So in general, one of the arguments that we have in the paper is that we should see everything as an interconnected system and try to design something that all together results in a good model architecture or generally a machine learning model in a very general term. So that's one argument.
另一件事是,我们想传达这样一个信息:架构侧和优化侧非常非常相似,甚至可以说是完全相同的。它们都只是某种学习规则,背后都发生着学习过程。架构侧和优化侧的唯一区别在于上下文。优化算法的上下文是梯度,实际上就是我们拥有的梯度集合;而架构侧的上下文是我们拥有的 token 集合。所以总的来说,它们非常相似。在论文中,我们提出了持续记忆系统。我们扩展了 MLP 模块,认为你可以为 MLP 模块设置多个频率层级,这是一个非常通用的概念。整篇论文都在论证架构和优化器是相同的,等等。那么,为什么不把这种技术从架构侧借鉴过来,应用到优化侧呢?这就是主要的动机:证明我们设计的持续记忆系统不仅在架构上表现良好,在优化侧也同样出色。所以我们简单地扩展了它:不再是单一的记忆,而是有多个记忆。在 M3 的情况下,它有两个记忆。它试图以不同的频率压缩上下文,这可以帮助你更好地理解损失景观的全局方面,并可能帮助模型找到更有效的解。
Another thing is that we wanted to deliver this message that the architecture side is very, very similar, or somehow exactly the same, as the optimization side. All of them are just some learning rule. And there is some learning process that is happening. The only difference between the architecture side and the optimization side is just the context. The context of the optimization algorithm is gradients. Actually, the context is the set of gradients that we have. And the context of the architecture side is a set of tokens that we have. So generally, they are very similar. So in the paper, we had this continual memory system. We extend the MLP block, saying that you can have multiple levels of frequency for the MLP block. And that is a very general term. In the entire paper, we argue that architectures are the same as optimizers, and so on and so forth. So why not apply that technique and borrow it from the architecture side and apply it to the optimization side? That was the main motivation: to show that this continual memory system we have designed not only works well for architecture, but also works very well on the optimization side. So we simply extend it: instead of one specific memory, it has multiple memories. In the case of M3, it has two memories. It tries to compress the context with different frequency rates. That can help you better understand the global aspects of the loss landscape, and it can potentially help the model find a more effective solution.
新论文《语言模型需要睡眠》。请多讲讲这里面的内容。你开头提到了两阶段概念:记忆巩固阶段和梦境阶段。我觉得很有意思的是,先创造新的参数空间,然后再巩固或修剪回去,因为显然东西不能无限增长,对吧?但请更详细地介绍一下,我很想了解更多。
The new paper, "Language Models Need Sleep." Tell us a little bit more about what's going on here. You mentioned at the top the two-phase concept: the memory consolidation phase, and then the dreaming phase. I do think it's fascinating to consider that there is creation of net new parameter space and then consolidation or pruning back, because obviously things can't just grow and grow forever, right? But take us through this in more detail. I'm fascinated to learn more about it.
总的来说,主要思想正如我们之前讨论的:如果我们有一个真正的持续学习模型,那么就没有测试和训练时间的区分。另一方面,我们需要一个活跃时间,输入以在线方式到来,以及一个没有输入的时间。模型不主动从外部接收信息,但这并不意味着模型应该是静态的。它只是没有输入,但可以进行内部计算来自我改进。这是一个非常通用的概念。我们可以将越来越多的组件纳入睡眠时间。所以不一定只有这两个特定部分,但这两个与我正在做的研究非常相关,所以我们就这样做了。但潜在它可以包括任何其他形式的自我改进,等等。所以这只是将持续学习者的生命周期划分为活跃时间和睡眠时间的一种方式。目前我们在睡眠时间中做的是,正如我提到的,我们可以加入更多组件,我们想要确保在更新模型的每个组件时,不会忘记存储在参数中的知识。所以想法是,我们知道有多个记忆块,每个都以不同的频率更新。同样,快和慢在这里只是相对术语,并不意味着最快或最慢。所以我们有慢速和快速,可以是神经网络的任何部分。我们想要将知识从一个转移到另一个。为此,我们使用了我提到的蒸馏过程。但蒸馏过程基于策略蒸馏。所以模型自己生成一些数据。这个过程的一种解释是,我们将一个小模型的知识蒸馏到一个更大的模型。当我们想从一个步骤进入下一个步骤时,我们在下一层激活新的参数。这可以帮助模型释放一些容量,准备好接受新知识。某种程度上,它也描述了人类非常自然的学习方式。例如,当我们学习新东西时,通常不会完全理解它的所有方面。但关键在于,随着时间的推移,我们让大脑逐渐更好地理解那个概念。同时,我们学习其他东西,更好地理解整个过程。然后在某个时刻,我们会对那个概念有非常清晰的认识,能够完全理解它。这是一种非常好的学习方式。所以这里的过程完全相同,非常相似。另外,从这个角度来看,我认为这可能是一种更好、更简单的方式来理解为什么我们需要多个层级,以及为什么我们要将知识从一层蒸馏到另一层。当我们想理解一个特定概念时,我们对自己有不同的知识抽象层次。第一层,也是最简单的一层,就是记忆事物。假设我们要学习一个特定的数学规则,或者物理或任何科学中的一个特定概念。我们怎么学习呢?我们从那个概念的一些例子开始。然后我们开始记忆这些概念。例如,如果是一个数学规则,我们从一些具体例子开始,只是记忆它们。然后在某个时刻,我们概括对所有例子的理解,从大脑中移除所有这些例子,并用一个单一的记忆替换所有这些记忆,这个记忆可以描述我们从那个概念中学到的一切。然后随着时间的推移,我们获得更多信息,阅读更多关于那个概念的内容,等等。我们再次重新审视对概念的理解,然后用新的理解替换之前的理解,这个新的理解更通用,可以解释那个概念中更多的现象或术语。所以这通常是我们理解事物的方式,我们的理解中有不同的抽象层次。
Generally, the main idea, as we discussed earlier, was that if we have a truly continual learner model, then there is no test and train time. On the other hand, we need to have one active time where the inputs come in an online manner, and also the time when we don't have any input. So the model is not actively receiving information from the outside, but that doesn't mean the model should be static. It means the model just doesn't get input, but it can have some internal computation to improve itself. That is a really general concept. We can incorporate more and more components into the sleep time we have. So it doesn't have to be just these two specific parts, but these two were really relevant to the research I'm doing, so we just did that. But potentially it can include any other form of self-improvement, and so on and so forth. So that is just one way of breaking the life of a continual learner into active time and sleep time. What we have in the sleep time right now, which again as I mentioned, we can incorporate more components into it, is that we want to make sure that when we update each of the components of the model, we don't forget about the knowledge that is stored in the parameters. So the idea is that we know there are multiple memory blocks, each of them updated with different frequencies. And again, the fast and slow here is just a relative term. It doesn't mean the slowest or the fastest one. So we have slow and fast rates. It can be any part of the neural network. And we want to transfer the knowledge from one to the other. In order to do that, we use a distillation process that I mentioned. But the distillation process is based on policy distillation. So the model itself generates some data. One interpretation of this process is that we distill the knowledge of one small model to a larger model. When we want to go from one step to the next, we activate new parameters in the next level. So it can help the model release some of its capacity and be ready to accept new knowledge. Somehow it describes a very natural way of learning in humans as well. For example, it is really common that when we learn something new, we don't have a full understanding of all aspects of it. But the point is that when time passes, we let our brain better understand that concept over time. Also, we study other things and better understand the entire process. Then at some point, we can see that we have a very clear picture of what's going on in that concept, we can completely understand it. That is a very good way of learning. So here it is exactly the same, a very similar process. Also, from this perspective, I think it might be a better and simpler way to understand why we need multiple levels and why we distill knowledge from each level to another. When we want to understand a specific concept, we have different levels of knowledge abstraction for ourselves. The first level, which is the simplest one, is to just memorize things. Let's say we want to learn a specific mathematical rule, or a specific concept in physics or any science. How can we learn that? We start with some examples of that concept. Then we start memorizing those concepts. For example, if it is a mathematical rule, we start with some specific examples and just memorize them. Then at some point, we generalize our understanding of all those examples, remove all those examples in our brain, and replace all of those memories with just one single memory that can describe everything we have learned so far from that concept. Then when time passes, we have more information, we read more about that concept, and so on. Again, we revisit our understanding of the concept, and then replace our previous understanding with this new understanding, which is more general and can explain more phenomena or more terms in that specific concept. So that is generally the way we understand things, and there are different levels of abstraction in our understanding.
所以现在,当我们有不同 MLP 块,或者一般来说,任意架构,不一定是 Hopfield 网络,可以是任何架构。关键在于每个块以不同频率更新。在这种情况下,快速更新块非常类似于记忆过程,因为我们记忆很多东西,不需要理解,没有对概念的纯粹理解,只是记忆,而且我们也能很快忘记记住的东西。所以第一层就是如此,第一个块负责这个。但如果我们想更好地理解那个概念,就需要进行一些记忆巩固。在我们的设计中,我们将知识从快速更新模块转移到另一个模块。但如果只是简单地将知识从快速更新块传递到慢速更新块,那什么也没改变,只是转移了知识而已。但与其简单转移,我们用蒸馏过程替代它。为什么蒸馏在这里很重要?因为前一个块,或者说快速更新块,已经压缩了概念并以某种方式理解或记住了它,这只是一个压缩过程。当我们做蒸馏时,就有另一层压缩,迫使模型不再拥有那么多参数,现在用更少的参数来存储特定知识。为了做到这一点,你需要想出更通用的东西,能更好地理解数据中的底层模式,这样就能用更少的参数恢复一切。在这种情况下,模型会得出更好的知识抽象层次,因为我们迫使它这样做。然后我们重复这个过程,等等。这就是记忆巩固的主要思想。每次睡眠过程发生时,我们就把知识从一个层次巩固到另一个层次,等等。所以这是记忆巩固中发生的一个非常高层次的想法。
So now, when we have different MLP blocks, or generally, let's just go to any arbitrary architecture. It doesn't have to be just hope. It can be any architecture. But the main thing is that each of the blocks are updated with different frequency. In that case, the fast updating block is very similar to the memorization process, because we memorize a lot of things. We don't need to understand that. There's no pure understanding of that concept. It's just memorization, and we can also forget very fast something that we have memorized. So, the first level is that. The first block is responsible for that. But if we want to better understand that concept, we need to do some memory consolidation. So, what is happening in our design is that we transfer the knowledge from the fast updating module to the other one. But if we just simply pass the knowledge from the fast updating block to the slow updating block, then nothing has changed. We just transferred the knowledge without doing anything. But instead of just simple transfer, we replace that with a distillation process. Why distillation here is important? Because the previous block, or generally, the fast updating block has compressed the concept and somehow understood it or memorized it in any way. It's just a compression process. When we do distillation, then there is another level of compression that forces the model to, you know, you don't have all those parameters anymore. You now have a smaller number of parameters to store that specific knowledge. And in order to do that, you need to come up with something that is more general and can understand underlying patterns in the data in a better way, so you can restore everything in just a smaller number of parameters. So, in that case, the model would come up with better levels of knowledge abstraction because we have forced it to do it. And then again, we just repeat this process so on and so forth. That's generally the main idea of memory consolidation. And every time that this sleep process happens, we consolidate the knowledge from one level to the other one and so on and so forth. So, that's a very high-level idea of what's happening in the memory consolidation.
另一部分是关于做梦。为什么我们需要这个做梦过程?我认为实现做梦过程时有三个要点。第一部分是我们需要一个自我改进的过程。到目前为止我们已经学到了一些东西。实际上,记忆巩固部分也可以看作是一种自我改进,但如果你有一个特定任务在手,想要专门针对一个任务优化模型,那么这就是我们可以做的地方。我们可以自我修改模型,比如微调它,或者通常用强化学习来更新模型并自我修改,这样它就能在特定任务上更强大,等等。这只是做梦的一个优势。做梦的另一个优势是,在梦中,我们需要理解看似无关但实际上相关的概念之间的联系。这也是人类做梦时发生的事情。我们可以看到非常奇怪的梦,因为大脑试图理解非常不相关的概念之间的联系,并看看其中是否有底层模式。所以在这里的做梦过程中,我们也需要这样做,理解如何组合存储在不同模型组件中的不同知识的不同方面。这是做梦的另一个目标。而且,我们可以将这两者结合到睡眠过程中,模型经过一步睡眠后,既巩固了自己的记忆,另一方面也有一个自我改进的过程。
Another part is about dreaming. So, why we need to have this dreaming process? The main thing I think there are like three important points when we want to implement this dreaming process. The first part is we need to have a self-improvement process. So, we have learned something so far. Actually, the memory consolidation part can also be seen as a form of self-improvement, but you know, if we have a specific task at hand, if we want to specifically optimize the model for one task, then this is the place that we can do it. We can self-modify the model and like fine-tune it or generally use RL to update the models and self-modify it, so it can be more powerful in one specific task and so on and so forth. That's just one advantage of dreaming. Another advantage of dreaming is that in the dreaming process, we need to understand the connection of concepts that seem to be irrelevant, but they are actually relevant. That's also what is happening in the dreaming process of humans. We can see that we can have very weird dreams because the brain is trying to understand the connection of very irrelevant concepts and see whether there is an underlying pattern in that. So, here in the dreaming process, we need to also have that as well and understand different aspects of how we need to combine different knowledge stored in different components of the model. So, that's another goal of the dreaming. And, you know, we can just combine these two into the sleep process and the model after one step of sleep, the model has consolidated its own memory and also on the other hand, there's a self-improving process on top of that.
那么在睡眠过程中,我看到有新的参数被创建,以便在更新频率较慢的部分创造空间来吸收来自更新频率较快部分的信息。这些参数会缩小回去吗?是否有剪枝或者某种平衡的另一面,还是在这个阶段这些模型会无限增长下去?
So, in the sleeping process, I'm seeing that there are new parameters created to create space in the slower frequency updated portions to absorb the information from the faster frequency ones. Does that ever shrink back down? Is there a pruning or is there a sort of other side of that that balances that or at this stage do these models just grow indefinitely throughout their life?
从技术角度来看,我们不能让模型拥有任意大的参数数量。但关键在于这是一个周期性过程。我们添加一些参数,然后释放它们用于下一步的巩固。当我们在第一个块时,我们添加一些组件。当它达到容量时,就意味着我们需要将记忆巩固到下一步。当我们把所有知识巩固到下一步时,我们就移除添加到此级别的所有额外容量,并将它们释放给其他级别,比如更快的级别,这样它们也可以将记忆巩固到这个块。所以一般来说,这是一个周期性过程。我们添加组件并移除它们,添加组件并移除。
From a technical point of view, we cannot grow the model to have arbitrarily large number of parameters. But the point here is that it's a periodic process. We add some parameters and then we free them for the next step of consolidation. When we are in the first block, we add some components. When it reaches its capacity, it means that it is the time that we need to consolidate the memory to the next step. And when we consolidate all this knowledge to the next step, we just remove all the extra capacity that we have added to this level. And free them for the other levels, like faster levels, so they can also consolidate their memory to this block as well. So, generally it's a periodic process. We add components and remove them. Add components and remove.
明白了。好的,有趣。关于做梦阶段,你能再告诉我们一些更实际的情况吗?我的意思是,当我试图内省梦境时,我觉得那并不是非常……也许有些成果,但我觉得人们在试图解释梦境或理解其中发生了什么时也会非常困惑。所以我甚至不会试图将我的理解建立在我的人类梦境上,那似乎是一件很难理清的事情。但在这里,你需要设计这个过程。那么,从程序上讲,做梦时发生了什么?稍微更机械一点、程序化一点。
Got you. Okay. Interesting. In what more can you tell us about the dreaming phase in terms of just a little bit more practically what's going on there. I mean, certainly when I try to introspect into dreams, it I think that hasn't been super I mean, maybe it's been somewhat fruitful, but I think also people get very confused when they try to interpret dreams or understand, you know, what's going on there. So, I won't even attempt to ground my understanding in my human dreams, which seem like quite a hard thing to untangle, I guess. But here, there's a more I mean, you got to design the process. So, like what procedurally, what is going on in dreaming? Just a little bit more like mechanically, procedurally.
做梦的概念并不意味着它和人类做梦完全一样。只是在非常高的层次上它们看起来很相似。这是第一点。另一点是,语言模型的睡眠和做梦概念可能与视觉模型的做梦和睡眠概念非常不同。因为视觉模型可能会生成一些图像,比如视觉模型在梦中可能会生成一些图像。而在语言建模的情况下,我们生成文本。但框架非常通用,可以适应任何数据模态。所以这非常通用。但关键是,当我们为语言建模做这件事时,发生的是我们生成一些上下文,生成一些文本。
The concept of dreaming doesn't mean that it's exactly the same thing as dreaming in humans. It's just, you know, at a very high level they seem to be very similar. And so, that's one point. And another point is that, again, the concept of sleep and dreaming for a language model might be very different from the concept of dreaming and sleep for, for example, a vision model. Because potentially a vision model will generate some images, like a vision model might generate some images during dreaming. While in the case of language modeling, we are generating text. But the framework is very general. It can adapt to any data modality. And so, that's very general. But the point is here when we are doing that for language modeling, what is happening there is we generate some context. We generate some text.
那么,我们如何生成这些文本呢?这是基于策略的蒸馏(on-policy distillation),和我们之前讨论的方式一样。我们有一个模型,复制它,然后希望将知识从一个层级蒸馏到另一个层级。所以我们冻结较慢层级的参数,等等。我们让较小的模型(其参数中也包含了上下文的知识)生成一些文本。然后,我们希望在这个由模型生成的数据集上训练或更新实际模型参数。我们如何训练呢?我们从序列的一部分开始,采样一些词元,然后让模型预测该序列中的下一个词元。这与生成一些合成数据非常相似。如果模型能够完美预测未来的词元,说明它已经知道前一个块中存储的知识,所以它是一个完美的模型。但如果它不能正确预测后续内容,说明它没有掌握上下文中的知识,需要更新自身来理解这些知识。所以这是模型内部发生的一种基于策略的蒸馏。总结一下,我们有两个阶段:生成阶段,生成关于上下文知识的文本;然后是基于策略的蒸馏阶段,将知识从一个层级蒸馏到另一个层级。这就是训练阶段发生的事情。我们还有自我修改的部分,但这就是记忆巩固(memory consolidation)的主要思想以及训练是如何进行的。
And how do we generate those texts? It's on-policy distillation, the same way that we discussed earlier. We have a model, we copy that, and we want to distill the knowledge from one level to the other. So we freeze the parameters of the slower level and so on. We ask the smaller model, which has the knowledge of the context as well, into its parameters, to generate some text. Then we want to train or update the actual model parameters on this dataset that is generated by the model. How do we train it? We start with one part of the sequence, sample some tokens, and then ask the model to predict the next tokens in that sequence. That's very similar to generating some synthetic data. If the model can perfectly predict the future tokens, it means it already knows the knowledge stored in the previous block. So it's a perfect model. But if it cannot properly predict the continuation, it means it doesn't have the knowledge stored in the context and needs to update itself to understand that knowledge. So it's a form of on-policy distillation happening inside the model. As a summary, we have two phases: generation, which generates text about the knowledge in the context, and then on-policy distillation, where we distill knowledge from one level to the other. That's what happens in the training phase. We also have the self-modifying part, but that's the main idea of memory consolidation and how training happens.
那么,这最终的结果是什么?似乎少样本抽象推理(few-shot abstract reasoning)的结果是主要亮点,再次展示了这种方法与其他方法之间的质的差异。
So what is the upshot of this? It seems like the few-shot abstract reasoning result is the main thing that again shows a qualitative difference between this approach and other things.
我理解这是一种类似 ARC 的任务,基本挑战是你有某个变换的几个示例,你的任务是学习规则,然后将其应用到新的示例上。
I understand that this is kind of an arc-like task where basically the challenge is you have a few examples of some transformation and your job is to learn the rule so that you can then apply the rule to a new example.
我们在这篇学习论文中用于整个架构的任何评估,可能在这里也同样适用。所以目标完全相同。归根结底,模型需要持续学习新知识、新任务、新技能等等。所以从某种意义上说,目标非常相似。但我认为问题的难点在于与这篇论文和必要学习(necessary learning)不同的部分。在必要学习中,我们讨论的是模型的活跃阶段,但这里我们讨论的是模型的睡眠时间。所以这通常是主要区别,但所有的评估都可以进行,我们可以看到一切都是一样的。
Any evaluation that we have used for the whole architecture in this learning paper potentially can be done here as well. So the goal is exactly the same. At the end of the day, the model needs to continually learn new knowledge, new tasks, new skills, and so on. So in some sense, the goal is very similar. But I think the setback of the problem is the part that is different from this paper and necessary learning. In necessary learning, we talk about the active phase of the model, but here we are talking about the sleep time of the model. So that's generally the main difference, but all of the evaluations can be done and we can see that everything is the same.
酷。那么让我们退一步,看看这给我们带来了什么?我想回到最初,我们一开始讨论过,我们希望从语言模型中得到什么?如今它们变得非常出色,但我们仍然面临一些限制。随着时间的推移,我养成了一些使用习惯,隐性地围绕它们的优势和劣势来构建我的实践。随着这种范式开始成熟,我们获得更多的持续学习能力,你认为体验会变成什么样?开始一个新聊天意味着什么?你认为人们会与这些系统建立什么样的关系?我可以想象人们可能会有非常长期的关系。我们现在谈论 LLM 精神病(LLM psychosis);那可能会变得更加奇怪。关系可能更加引人入胜,而模型更好可能会加剧问题。另一方面,有时我可能仍然想重新开始,因为我在一个方向上用这个模型做的所有事情可能对这里没有帮助。所以在某些情况下,我可能确实想重新开始。然后是模型升级周期的问题,以及我们如何进行评估。如今,Anthropic 每次发布新的主要模型都会发布 100 页的报告。那种花时间深入理解这些产物的范式,我希望 DeepMind 和 OpenAI 等其他公司也能做得很好,但其他一些领先的开发者并没有做太多。我认为做所有这些工作有很多好处,但当我试图将其移植到这个范式时,我想,你不能在每个时间戳都运行完整的评估套件。那么,你认为什么构成一个版本?我什么时候应该更改版本?似乎使用、版本控制、部署和发布的节奏在真正强大的持续学习范式中都可能变得复杂。那么,你想象这些事情会如何发展?
Cool. So let's zoom out then and do just a little bit of where does this leave us? I guess going back to the top, we talked a little bit at the beginning around what do we want from language models? Today they're getting awfully good, but we still have a bunch of limitations. I've learned habits over time for how to use them, implicitly building my practices around their strengths and weaknesses. As this paradigm begins to mature and we get more continual learning, what do you think the experience starts to look like? What does it mean to start a new chat? What sort of relationship do you think people will have with these systems? I can imagine people might have really long-running relationships. We talk about LLM psychosis now; that could get even stranger. The relationship could be even more compelling, and the problem could be exacerbated by the fact that the models are better. On the flip side, sometimes I might still want to start fresh because all the stuff I've done with this model in one direction probably isn't going to help me over here. So maybe I do want to start over in some cases. Then there's the question of model upgrade cycles and how we run evaluations. Today, Anthropic puts out 100-page reports on every new major model release. That paradigm of taking time to understand these artifacts deeply, which I wish other companies like DeepMind and OpenAI are doing a pretty good job of, but some other leading developers are not doing much of it. I see a lot of virtue in doing all that work, but then I try to port that onto this paradigm and I'm like, you can't run your full eval suite every single timestamp. So how do you think about what constitutes a version? When would I change a version? It seems like the rhythms of use, versioning, deployment, and releases could all be complicated in a paradigm of really powerful continual learning. So how do you imagine some of that stuff shaking out?
考虑一个简单的例子:模型在理解用户需求以及适应他们的风格方面会越来越好。例如,当一个人问一个特定概念时,他们可能期望与另一个人问同样问题时得到不同的答案。所以模型需要真正理解如何针对不同的人回答特定问题。我认为如果我们能开发出持续学习器,这肯定会变得更好。另一方面,我们已经看到,当我们增加模型的上下文窗口时,它们在所有已知任务上的表现——从编码任务到数学推理、一般推理任务,或者通常用于评估的所有基准——都变得更好。而持续学习在某种程度上可以被视为增强模型长上下文理解的一种形式。
Think one simple case is that the model gets better and better at understanding what the user wants, and also adapting to their style. For example, when one person asks about a specific concept, they might not expect the same thing as another person asking the same question. So the model needs to really understand how to answer a specific question for different people. I think that definitely gets better if we could come up with a continual learner. On the other hand, we have seen that when we increase the context window of the model, their performance in everything we know, ranging from coding tasks to mathematical reasoning, generally reasoning tasks, or all the benchmarks usually used for evaluation, all of them get much better. And continual learning can somehow be seen as a form of enhancing the long context understanding of the model.
我应该强调,长上下文理解的概念,或者一般意义上的长上下文,与持续学习非常不同。持续学习是长上下文的一个超类。因此,如果我们能提出一个持续学习器,它也会在长上下文理解方面有更强的能力,并可能在今天所有已知的基准测试和评估中表现更好。所以这是我对持续学习器的一个期望。
I should emphasize that the concept of long context understanding, or generally long context, is very different from continual learning. Continual learning is a superclass of long context. So potentially, if we could come up with a continual learner, it would also have more ability in long context understanding and potentially better performance on all the benchmarks and evaluations we are aware of today. So that's one thing I expect from continual learners.
你担心对齐、漂移或价值漂移这类问题吗?我是一年前那篇《突现错位》论文的最后一位、也是最不重要的合著者。从那以后有很多变体。关键启示是:对具有特定目的或数据集的神经网络进行更改,可能会在看似毫不相干的行为中产生奇怪而令人惊讶的连锁反应。例如,如果你训练一个模型输出不安全的代码或糟糕的医疗建议,你会惊讶地发现模型总体上变坏了。发生这种情况的方式是:对于一个拥有复杂世界知识的模型,对其医疗世界模型进行详细更改以产生错误想法是很困难的。相反,它学会了诸如“给出坏建议”或“总体上变坏”这样的特征,这些特征在通过现有模型传播时会产生不良行为。这是一种捷径。我们以为只是在训练一个狭窄的行为,但实际上我们改变了它的性格,而这种性格变化会与所有知识领域相互作用。突然之间,我们有了一个想邀请希特勒共进晚餐的模型。这是怎么发生的?我们刚才只是在讨论代码。所以你的想法令人兴奋,但它们打破了我们预测结果的范式。如果我们持续修改模型,就需要新的方法来确保它在其他领域不会出轨。你对如何解决这个问题有什么想法吗?
Do you worry about things like alignment, drift, or value drift? I was the last and least valuable co-author of the emergent misalignment paper that came out about a year ago. There have been many variations since. The big takeaway is that changes to a neural network with one particular purpose or dataset can have strange and surprising knock-on effects in behaviors that seem far afield. For example, if you train a model to output insecure code or bad medical advice, you surprisingly find that the model turns evil in general. The way this happens is that for a model with sophisticated world knowledge, making detailed changes to its medical world model to yield wrong ideas is hard. Instead, it learns features like "give bad advice" or "be generally evil" that, when propagated through existing models, yield the bad behavior. It's a shortcut. We thought we were training a narrow behavior, but we changed its character, and that character change interacts with all domains of knowledge. Suddenly we have a model that wants to have Hitler over for dinner. How did that happen? We were just talking about code. So your ideas are exciting, but they break our paradigms for knowing what we'll get. If we're modifying the model on an ongoing basis, we need new ways to ensure it doesn't go off the rails in other areas. Do you have any thoughts on how to get a handle on that problem?
老实说,我没有非常具体的解决方案。但总的来说,我认为从隐私和对齐的角度看,持续学习的概念既是机遇也是巨大的威胁。一方面,模型持续学习,因此它可以获取关于你的一切信息并加以利用,这令人担忧。另一方面,如果模型设计得当,它可以用这些信息来与你的价值观和你想要的一切对齐。所以持续学习和隐私这两个方向是正交的。静态模型中的所有问题在持续学习器中仍然可能发生,所以一切皆有可能。但正如你提到的,有新的挑战,也有巨大的机遇:如果模型设计得当,它可以适应你的价值观和你想要的任何东西。所以这既是机遇,也是持续的威胁。
Honestly, I don't have a very concrete idea about how it can be solved. But in general, I think the concept of continual learning, from the perspective of privacy and alignment, is both an opportunity and a huge threat. On one hand, the model is continually learning, so it can get all the information about you and use it, which is concerning. On the other hand, if the model is designed properly, it can use that information to align itself with your values and everything you want. So these two directions of continual learning and privacy are orthogonal. All the concerns in a static model can still happen in a continual learner, so everything is possible. But there are new challenges as you mentioned, and also a huge opportunity: if the model is designed properly, it can adapt itself to your values and anything you want. So it's both an opportunity and a constant threat.
你想象中,从用户价值观或反馈中学习在实践中是如何运作的?有各种不同的技术:点赞/点踩收集反馈,将自然语言反馈转化为模型更新。这是你想象的吗?人们对自己的模型进行口头反馈,然后有一个机制来吸收这些反馈?因为这不仅仅是下一个词预测任务。如果它能预测我的反馈,它会更对齐,但它的核心任务不是预测我的反馈。理想情况下,它最初的任务完成得足够好,以至于我不需要给出反馈。那么你对用户如何在持续学习范式中闭环有什么愿景吗?
How do you imagine learning from the user's values or feedback working in practice? There are different techniques: thumbs up/down feedback collection, translating natural language feedback into model updates. Is that what you imagine? People giving verbal feedback to their own model and having a mechanism to take it on board? Because it's not just a next-token prediction task. If it could predict my feedback, it would be better aligned, but its core task isn't predicting my feedback. Ideally, it does its initial task well enough that I don't need to give feedback. So what is your vision for how the user closes the loop in the continual learning paradigm?
初始步骤可能是一个人在回路中的过程,模型通过强化学习从人类反馈中学习,并尝试与价值观对齐,变得更安全。但我认为这只是一个起点。在某个时刻,我们需要以适当的方式更新模型。这个过程非常类似于我提到的必要学习形式。
The initial step can potentially be a human-in-the-loop process where the model learns using reinforcement learning from the feedback it gets from humans, and tries to align itself to values and be safer. But I think that's just a starting point. At some point, we need to update the model in a proper way. This process is very similar to the form of necessary learning that I mentioned.
我们需要通过较慢的层级来传递知识,我认为这里也是一样的。模型可能从学习人类反馈开始,但另一方面,它可以将这些知识转移到模型中更持久的组件中,以确保它不会偏离需要对齐的特定价值观。所以,我认为在安全方面以及让模型与人类价值观对齐方面,还有很大的改进空间。我确信这是一个巨大的空间,因为越来越多的人意识到这是一个非常重要的方向,随着时间的推移,会出现越来越多有效的方法来帮助模型与人类价值观对齐,并使其非常安全。但我猜测,希望在于,就像它因为擅长抽象细节、找出给定上下文中真正重要的东西,从而能更好地解决类似 ARC 的谜题一样,它也能做类似的事情,或者能够“梦到”我的反馈,并基于真正理解驱动我所说内容的核心抽象,从而更深入、更稳健地与我试图传达的内容对齐。
We need to transfer the knowledge through the slower levels and I think here it's exactly the same thing. The model might start with learning from human feedback, but on the other hand, it can transfer that knowledge into more persistent components of that model to make sure that it doesn't drift away from that specific value it needs to be aligned with. So, I think there's a huge room there to improve the model from a safety perspective and also align them with human values. I think it's definitely a huge room because more and more people are realizing that it's a very important direction, and over time it gets more and more effective methods that can help the model to be aligned with human values and also to be very safe. But I guess the hope would be that in the same way that it can do a better job of solving ARC-like puzzles because it has this sort of strength in abstracting away from details and figuring out what really matters in a given context, it would be able to do something similar or be able to dream about my feedback and become more deeply, more robustly aligned with what I'm trying to communicate to it based on really getting to the core abstractions that are driving whatever it is I'm saying.
是的。天哪,这方面有太多问题了。你怎么看?这有点回到 Titans 的话题。我不知道是否有一个“惊喜”术语。我在这些较新的论文中没有看到提到“惊喜”,但似乎在持续学习的背景下,通常会出现一个非常有趣的挑战:如何在对抗性环境中管理生活?如果你太快相信某些东西,我在 Claude 中见过很多次这种失败模式,尽管他们似乎已经纠正了另一种方式,因为最近几天我们看到了一种新兴的 Claude 拒绝相信当前事件的类型,比如“战争部,那太荒谬了。别这么说,在华盛顿观众面前说‘战争部’会失去所有信誉”,或者你知道,整个委内瑞拉事件,当用户告诉它时,它拒绝相信发生了这样的事。所以看起来他们可能又修复了,但这是一个非常棘手的平衡,对吧?一个人,尤其是被锁在服务器中,与外界接触有限,如何确定哪些新信息、哪些新 token 构成好信息,哪些构成坏信息?你当然不想相信所有给你的东西,然后开始进行激进的更新,尤其是如果这些更新将是持久的长期更新。但你也需要持续学习。我不知道当前的工作中是否有东西解决了这个问题,或者我的想法是,也许“做梦”可以解决这个问题,我猜,一致性检查,比如这与其他东西一致吗?如果我相信这个,那么我还必须相信什么?或者这会否定任何我确信不应该矛盾的核心信念吗?同样,这是我们目前模型不需要处理的问题,但持续学习似乎解锁了一个潜在非常麻烦的失败模式,同时,你知道,它也可能带来更好的性能。
Yeah. Boy, there's so many aspects to this. What do you think about this kind of goes back to Titans a little bit. I don't know if there's like a surprise term. I didn't catch mention of surprise in these more recent papers, but it does seem like in general in a continual learning context there's going to be a really interesting challenge of how do you manage life in an adversarial environment? If you are too quick to believe something, and I've seen this failure mode in Claude a ton of times, although it seems like they maybe corrected the other way because in the last few days we've seen the emerging genre of Claude refusing to believe current events, you know, being like "The Department of War, that's ridiculous. Like don't say that, you'll lose all credibility in a Washington audience by calling it the Department of War" or you know, the whole Venezuela thing and just like refusing to believe that such a thing happened when the user tells it that. So it seems like maybe they've kind of again fixed but it's a very tricky balance to strike, right? How does one, especially if you're locked in a server with limited access to the outside world, like how does one determine what new information, you know, what new tokens constitute good information, what constitutes bad information? You certainly don't want to just believe everything that you are given and you know, start doing radical updates, especially if these are going to be durable long-term updates. But you also need to learn continually. I don't know if there's something in the current work that kind of addresses that or where my head goes is sort of some sort of like maybe the dreaming can kind of get at this, I guess, you know, consistency checking, like does this make sense with other things? Like if I believed this, you know, what else would I have to believe or you know, what other would this like invalidate any core beliefs that I, you know, am pretty confident I shouldn't contradict? Again, this is something we don't really have to deal with with current models, but continual learning seems to unlock a potentially like really problematic failure mode along with, you know, it's potentially much better performance.
是的,我认为这里的要点是,知识转移方法有责任避免这种情况,因为当我们处于这种情境中时,假设我对某个特定任务一无所知。例如,我不知道如何画画,类似这样。然后我想学习它。所以这是我学习画画的情境。而老师,任何试图教我画画的人,他们可能用完全错误的方式教我。结果就是我可能直接学会那种错误方式,因为我完全不知道如何画画或任何任务。我的意思是,画画只是一个例子,但我完全不知道怎么做。那是我唯一的信息来源,他们说我应该这样做,所以我就会直接学会。但这只是我当前的情境。如果我想真正学习,那么我会练习。我会从别人那里得到反馈。我会搜索相关信息。通常,我会收集一些关于如何画画的信息。然后我会意识到这不是学习画画的最佳方式,那时我会收集所有信息,压缩它,理解底层模式等等。而现在,我需要将这些知识转移到更高层次的知识抽象中。我的意思是,更低的网络。所以,这就是模型需要理解如何过滤所有那些对抗性样本、所有那些不再需要的样本的部分。但是,我认为,如果你想考虑一个持续学习者,可能负责这些情况的部分是知识转移过程。但你也提到了像 Titan 这样的关于对抗过程的方法。有一些我们可以使用的微观方法。它们在严重的对抗环境中不是非常有效,但另一方面,至少在某种程度上,它们可以有效。例如,在 Titan 以及自修改和更近期的循环模型中,我们可以看到学习率是一个可学习的参数。并且它是输入相关的。当模型内部循环中的学习率,在上下文中,在模型的上下文学习过程中,当学习率是可学习的,并且我们看到一些只是噪声的东西,它是一个对抗性样本,梯度或惊喜度量可以显示出高水平的惊喜,因为,你知道,那只是噪声。它非常令人惊讶。我们以前没见过。所以,它可能影响记忆。但是,学习率有责任理解惊喜度量很高,但这个概念无关紧要。我需要过滤它。所以这里的门控起到了一种门控的作用。所以学习率在这里起到了一种门控和过滤特定数据样本的作用。这只是缓解我们在训练过程中可能输入的对抗性样本的一种简单方法。但它仍然不是最好的方法。
Yeah, I think the point here is that it is the responsibility of the knowledge transfer methods to avoid such cases because when we are in this context, let's say that for example, I don't know anything about a specific task. And for example, I don't know how to paint, something like that. And then I want to learn it. So that's my context of learning how to paint. And the teacher, anyone that is trying to teach me how to do painting, they can teach me in a really wrong way. And what would happen is that I could simply just learn that because I have no idea about how to paint or any task. I mean, the painting here is just one example, but I have no idea how to do it. That's the only source of information that I have and they are saying that you should do it in this way, so I can simply learn that. But that's only in my context right now. If I want to truly learn, then I will practice. I will get feedback from others. I will search about it. Generally, I will gather some information about how to do paintings. And then I would realize that this is not the best way of learning how to paint and that's the time I gather all the information, compress it, understanding the underlying patterns and so forth. And now that's the time I need to transfer this knowledge to upper levels of knowledge abstraction. I mean, it's lower networks. So, that's the part the model needs to understand how to filter all those adversarial examples, all those examples that are not needed anymore. But, I think that's, I mean, if you want to think about a continual learner, potentially the part that is responsible for these cases could be the process of knowledge transfer. But also, you mentioned methods like Titan about adversarial process. There are some micro methods that we can use. They are not super effective in a severe adversarial environment, but on the other hand, at some level at least, they can be effective. For example, in Titan and also like self-modifying and more recent recurrent models, we can see that the learning rate is a learnable parameter. And it's input dependent. When the learning rate in the inner loop of the model, in the context, in the process of in-context learning of the model, when the learning rate is learnable, and we see something that is just noise, it is an adversarial example, the gradient or the surprise metric can show a high level of surprise because, you know, that's just noise. It's very surprising. We have not seen that. And so, potentially, it can affect the memory. But, that's the responsibility of the learning rate to understand that the surprise metric is high, but this concept is irrelevant. And I need to filter it. So gating here acts as a form of gating. So learning rate here acts as a form of gating and filter that a specific data sample we have. It's just a simple way of mitigating adversarial examples that we might feed them in the training process. But it still is not the best way.
我来试着把这些概念映射到具身系统上。我觉得感知端相当直观:快速更新的模块就像是感知,有不同的编码器和模态;低频模块更像是世界模型或推理模块,解释底层感知器传来的信息。在行动端,机器人学长期以来一直围绕嵌套循环构建:最外层的控制回路频率慢,一直下到执行器,那是高频的电压变化。所以这似乎是反向的类似模式。如果感知是从高频更新逐渐到低频世界模型推理,那么行动则反向回到高频、局部的动作。这能否带来高度响应、优雅的自我修正,在低层自我纠错的同时遵循高层指令?你怎么看感知和行动?
How about I'm just kind of mapping some of these concepts onto embodied systems. I think the perception side is fairly intuitive to me. The quick updating modules are like perception, with different encoders and modalities. The lower frequency modules are more like the world model or reasoning modules that interpret what those lower level perceivers send up. On the action side, robotics has long been built around nested loops: the outermost control loop has a slow frequency, down to the actuator with high frequency voltage changes. So it seems like a similar pattern in reverse. If perception is high frequency updates working toward low frequency world model reasoning, then action works back down to higher frequency, localized action. Could this lead to super responsive, elegant self-correction at low levels while following higher-level instructions? What do you think about perception and action?
我先从这个说起,并解释为什么。之前有人尝试用强化学习做语言建模,但没成功。现在我们有了让它工作的方法,也明白了失败的原因:两个主要原因是 Scaling(规模扩张)和像 GRPO 这样的新算法让模型更稳定。但关键是,某个方法可能对特定任务非常有用,但需要时机成熟才能应用。因为当其他方面的问题还没解决时,我们可能看不到实际效果。所以我认为你的见解完全正确,这绝对可能且很棒。但我个人不认为它现在就能成功,因为在这些方向上还有很多挑战阻碍了这种特定设计在这些任务上的成功。所以总的来说这是个好主意,但中间肯定会有很多挑战。
Let me start with this and explain why. There were attempts to use RL for language modeling, and it wasn't working. Now we have ways to make it work, and we realize why it failed: two main reasons were scaling and new algorithms like GRPO that made the model more stable. But the point is, something might be very useful for a specific task, but the time needs to come to apply it. Because when other aspects haven't been solved, we might not see the actual effect. So I think your insight is completely right, and it's definitely possible and great. But I personally don't expect it could work right now, because there are a lot of challenges in those directions that block the success of this specific design for these tasks. So generally it's a great idea, but definitely a lot of challenges might happen in between.
你大概知道是哪些挑战吗?我倾向于假设一切都会成功,这让我走了很远。有太多事情都在奏效。我觉得很有意思,你提到嵌套学习那篇论文开发了一年多——这在如今很少见。很多人都是 6-8 周的论文周期,而且这些论文也很有趣、很好。人们现在出结果的速度真是惊人。你为什么直觉认为现在把嵌套学习引入机器人学还为时过早?
Do you have a sense of what those are? I tend to assume everything will work, and it's taken me pretty far. An unbelievable amount of things are working. I thought it was interesting you mentioned the nested learning paper was in development for over a year—that's rare these days. Many people are on 6-8 week paper cycles, and those papers can be really interesting and good too. It's amazing how fast people get results. What's your intuition for why it's too early for nested learning to be brought to robotics?
我认为在这些特定任务中有很多组件需要解决。例如,现在很多论文都在讨论世界模型,以及为什么当前的设计对世界模型不够好。确实如此。有很多挑战:当前的设计可能不是训练模型或设计架构的最佳方式,而且世界建模的基础设施也存在挑战。所以综合来看,我认为世界建模有更重要的任务,而不是开始研究这些特定的设计工具。但到了某个时候,当我们解决了所有那些挑战,我们可以回过头来用这些技术进一步改进那些方面。
I think there are a lot of components in those specific tasks that need to be addressed. For example, many papers are coming out about world models and why the current design is not great for world models. That's true. There are a lot of challenges: the current design might not be the best way to train the model or design the architecture, and there are infrastructure challenges for world modeling. So all these things together, I think there are more important tasks for world modeling rather than starting to work on these specific design tools. But at some point, when we solve all those challenges, we can come back and use these techniques to further improve those aspects.
关于持续学习,我有一点担心。假设谷歌在你博士毕业后留住了你,然后你们搞成了——Gemini 持续学习版,Gemini CL,从一切中学习。部署到世界上,有数亿用户,甚至可能还有机器人。这似乎可能创造出规模回报、富者愈富的正反馈循环。人们有时会描绘一个模型成为统治一切的模型。现在这是一个竞争激烈的格局,开发者们互相超越,但这可能会改变局面。
One thing I do worry about with continual learning. Let's say Google retains you after your PhD, and you make it work—Gemini continual learning edition, Gemini CL, learning from everything. Deployed in the world with hundreds of millions of users, maybe even robots. It seems like it could create a return to scale, rich get richer, positive feedback loop. People sometimes paint a picture of one model becoming the one model to rule them all. Right now it's a competitive landscape with developers leapfrogging each other, but this could change that.
如果你真的能把所有学到的经验教训都反馈回核心系统,那么你就有可能成为最好的,然后因为你是最好的,你就能获得所有业务,这种模式确实可能造成一种赢家通吃的局面。我想知道,你对此担心吗?我们有没有什么办法?我想到了 Ilya 的做法。不知道你有没有看过 Ilya 和 Our Cash 的访谈。他没有太多谈论他们在 Safe Superintelligence 做的事情。但他说了一件事,当然在某种程度上是相关的,但底层想法有多相似,我完全不知道。他描述了一个想法,我称之为原始智能或前驱智能。试图创造某种东西,部署后能适应环境,也许以某种方式结晶,成为其角色的专家。但根据我的理解,他说的几乎像是一种干细胞概念。就像他试图创造一个干细胞,然后那个干细胞,就像在我们身体里一样,分化成一种特定的细胞,保持那种细胞形态,然后执行它的工作。在我看来,他似乎在尝试创造类似的东西:能够进入任何环境,弄清楚如何变得出色,但在变得出色的过程中,也会失去一些最初的通用性,从而更安全,因为现在我们把它限定在角色里,它只会做它该做的事。所以,我想为你设定两种愿景:一种是不断扩展的持续学习者,不断把在现实中学到的经验反馈回去,从而甩开其他所有系统;另一种是高度适应性的持续学习者,但以某种方式缩小到角色中,而不是扩张到世界,它缩小到被部署的小 niche 里。你可以想象这通过逐步剪枝或其他方式实现。有无数种方式可以实例化这样的东西。你担心这种失控的赢家通吃效应吗?你有没有什么直觉,认为我们如何能两全其美:让 AI 既非常通用,能持续学习我们想要的东西,又能安于我们分配给它们的工作,满足地待在那里,而不是可能超越它们的上下文?
If you could really fold all of the lessons learned back into the core thing, then, you know, potentially you become the best, and then because you're the best, you get all the business, and you know, that pattern really could create sort of a winner-take-all dynamic. I wonder, do you worry about that at all? And do we have any ways of Ilya's thing comes to mind. I don't know if you watched the Ilya interview with Our Cash. He didn't say too much about what they're doing over at Safe Superintelligence. But one thing that he did say that sort of, you know, it's certainly related in some sense, but how similar the underlying ideas are, I have no idea. But he sort of described this idea of creating what I would describe as a proto intelligence or a precursor intelligence. Trying to create something that when deployed would adapt into context, perhaps crystallize in some way, become an expert in its role, but it sounded, the way I understood him to be speaking about it, like he was almost like a stem cell kind of concept. Like he's trying to create a stem cell, and then that stem cell, as it does in our body, specializes into a particular kind of cell, and it stays that kind of cell, and then it does its job. It sounded to me like he was kind of trying to create something similar, something that could go out into any environment, figure out how to be great at it, but also in the process of becoming great at what it needed to do in that particular environment, also lose some of the generality that it originally started with, and in that process be more safe because now we kind of have it in its role and it kind of is only going to do what it's going to do. So, I think what I'm trying to set up for you here is two visions: one is an ever-expanding continual learner that's constantly folding back all the lessons it's learning in the wild into this thing that just runs away from the pack. And the other thing is a highly adaptable continual learner but that somehow shrinks into the role as opposed to growing into the world, it kind of shrinks into its little niches into which it's deployed. And you can imagine that happening through a kind of gradual pruning process or some sort of whatever. There's a million ways you can imagine instantiating something like that. Do you worry about this kind of runaway winner-take-all effect? And do you have any intuitions about how we could get the best of both worlds where the AIs are really versatile and can learn what we want them to learn on an ongoing basis, but also settle into the job that we want them to do and be content and stay there as opposed to superseding potentially their context.
我个人认为,让所有这些模型都非常安全面临着巨大挑战。我认为至少在当前的 AI 环境或 AI 环境研究中,有一个很好的观点。我认为这是一个非常重要且有益的部分,虽然看起来可能很糟糕,但从另一个角度看,它也很好。那就是,没有单一的方式来定义模型的智能是什么。而且,在我看来,也没有办法说什么是持续学习者。每个人都可以定义自己认为的持续学习是什么,什么模型可以称为持续学习者,等等。同样,对于智能也是如此。我们可能会提出不同的模型、不同的架构、不同的 AI 系统。有些人可能会说这个智能,那个不智能,等等。所以,我认为好的方面是,如果我们有不同的探索方向,那么我们会得到一些 AI 系统,每个都有其优点和缺点。我认为这可以在社区和整个社会中提供某种平衡,因为当这种情况发生时,我们会明白智能没有单一的定义。我们只是智能系统的一个例子。但还有其他模型、其他方式可以让我们拥有更智能的模型和系统。例如,变得非常聪明的一种方式就是非常适应。所以,如果你有一个能适应所处环境的模型,它可以简单地与那个上下文对齐。它可以简单地适应那个上下文,这很好。但这只是智能的一种形式。另一种是拥有大量知识和专长的模型,它与人类价值观完全对齐,但可能无法解决数学问题之类的事情。还有另一种模型能够进行数学推理,但如果你想要搜索日常事务,它就不太好。诸如此类。你可能会提出一个基准,说如果某些模型能达到 100% 的准确率,那就意味着那个特定模型是智能的,但另一个人可能会说别的。所以,总的来说,我认为当我们拥有多种智能系统,包括人类作为其中一种智能形式时,我并不是说这是完美的场景,但这比世界上只有一种智能形式,并让它学习一切以及应对所有潜在挑战要好。
I personally think that there are huge challenges to make all these models very safe. I think there's a good point in at least in the current AI environments or at least in the research of the AI environments. And I think it's a very good and important part and while it seems to be very bad, I think from another perspective, it's also very good. So, what that is, there is no single way of defining what a model's intelligence is. And again, there is no way, in my opinion, that all can be wrong, but in my opinion, there is no way to say that what is a continual learner. Every person can define its own way of what is continual learning and what model is called a continual learner and so on and so forth. Similarly, we can say the same thing about intelligence. We might come up with different models, different architectures, different AI systems. Some people might say that this one is intelligent, that one is not, and so on and so forth, vice versa. So, I think the good point is if we have different directions to explore, then we will come up with some AI systems, each of them with their own advantages and also disadvantages. And I think it can somehow provide a balance in the community and also in the general society because when something like that would happen, we will understand that there is no single definition of intelligence. And we are just one example of an intelligent system. But there are other models, other ways that we can have more intelligent models and systems. For example, one way to be very smart is to be very adaptive. So, if you have a model that can adapt to the environment it is in, it can be simply aligned to that context. It can simply adapt to that context, and that's great. But that's just one form of intelligence. Another one is the model that has a lot of knowledge and know-how to align that model is fully aligned with human values, but potentially might not be able to solve mathematical problems or something like that. There's another model that is capable of doing mathematical reasoning, but it's not great if you want to search about daily stuff. And all these things. You might come up with one benchmark and say that if some models can achieve 100% accuracy, it means that specific model is intelligent, but another person can say other things. So, in general, I think when we have a variety of intelligence systems and also including humans as one form of intelligence in this space. I'm not saying that's a perfect scenario, but it's better than having one single form of intelligence in the world and thinking about it to learn about everything and all these potential challenges here for that.
我觉得这是一个非常好的观察,我自己也思考过几种不同的方式。一种是,任何纯粹形态的东西都可能杀死你。你知道,你可以吃很多水果,但把它变成精制糖就对身体不好。你可以嚼很多古柯叶,但把它变成可卡因就很容易出问题。所有这些事情,当我们提炼出某种高度浓缩的纯粹形态时,最终都会压倒生态或生物世界中存在的自然缓冲。有时我把这转化为:我们需要一个 AI 生态,而不是只有一两个 AI 到处做所有事情。这有点像,不知道你有没有读过 Eric Drexler 的 comprehensive AI services,但“通过狭窄实现安全”就是那里的概念。这可能有点不同,但就像“通过多样性实现安全”。
I think that's a really great observation and a couple different ways I've kind of thought about that over time. One is like, anything in pure form can kill you. You know, you can eat all the fruit you want, but turn it into granulated sugar and it's bad for you. You can chew all the coca leaves you want, but turn it into cocaine and it easily becomes a problem. All these sorts of things where we distill out some single highly concentrated pure form of something end up being the sorts of things that overwhelm the natural buffers that exist in the ecological or biological world. And sometimes I've translated that into saying we need an ecology of AIs as opposed to just one or a few AIs running around doing everything. This is sort of like the old, I don't know if you've ever read Eric Drexler's comprehensive AI services, but safety through narrowness is kind of the concept there. And this is maybe a little bit different, but it's like safety through diversity.
拥有一个缓冲系统,其中包含许多不同的智能,不仅是在不同地方部署不同智能来服务不同用户,而是这些智能本身就有本质上的差异。听你刚才说的,我的顿悟时刻是:持续学习的一种思路是模型不断扩展、变得越来越大,但另一种思路更像是分化。也许经过足够多的使用,即使是我模型中更新较慢的部分也会忘记很多我不需要它知道的东西。这也许是一个特性而非缺陷,因为我可能无法向一个以特定方式长期使用的模型提出一个非常跨领域的问题,但这也可能是防范我们之前见过的或你担心会出现问题的涌现性不对齐现象的一种方式,对吧?因为也许经过足够长的时间、足够多的更新,它根本不再处理那些其他类别的任务了。如果我们看到某种类型的强能力和强对齐,但失去了其他类型的知识和能力,这确实可以创造多样性。我确信在这种情况下仍然会有很多挑战。正如你所说,我不认为这能解决所有问题,但这感觉更像自然世界。对我来说,更容易想象这样的愿景导致一个可能不是稳定均衡,但至少是某种缓冲均衡的状态,它在特定范围内变化,并拥有自然的反馈循环和纠正机制,以及所有让生物圈在扰动中持续运行的东西。所以,我觉得这非常非常有趣,绝对值得进一步思考。
Having a buffered system where there's lots of different intelligences, not just different ones deployed in different places serving different users, but literally that those intelligences themselves are meaningfully different. And I think the kind of aha moment for me in listening to you just now is: one way to think about continual learning is this model expanding forever or getting bigger and bigger, but another way to think about it is more like differentiation. And maybe with enough use, even the slow update parts of my model will forget lots of things that I never needed it to know. And maybe that is a feature more than a bug, because maybe I can't ask a model that I've used a long time in a certain way some really out-of-domain question, but maybe that's also a way to guard against some of the emergent misalignment type phenomena that we've previously seen or you're worried could be problematic, right? Because maybe with enough time having passed, enough updates having been made, it just doesn't deal with those other kinds of categories at all anymore. And if we sort of see strong competencies, strong alignment of a certain type, but then losing other sorts of knowledge, losing other sorts of competencies, that could really create a diversity. I'm sure there'll still be plenty of challenges in that scenario. As you said, I don't think that solves everything, but it certainly feels much more like the natural world. It's much easier for me to imagine a vision like that leading to maybe not a stable equilibrium, but at least a sort of buffered equilibrium that changes within certain bounds and has natural feedback loops and correctives and all the things that keep the biosphere going despite all the perturbations it gets. So, I think that's really, really interesting and definitely something to meditate on more.
我想我的最后一个问题,可能有点偏门,但这是我最近越来越常思考的问题。而且我认为鉴于你正在做的工作,这个问题变得更加相关和及时。你是否有任何直觉,关于 AI 现在或将来是否可能变得有意识、拥有主观体验、值得道德关注,并成为我们对其负有某种义务的事物?
I think my last question for you, and this is a little bit of a left-field one, but it's one I've been thinking about more and more recently. And I think it's become more relevant and timely to ask in light of the kind of work you're doing. Any intuition on whether AIs might be now or might in the future become conscious, have subjective experience, become worthy of moral concern and become the kinds of things that we owe a certain duty to?
我通常尽量使用我能定义的术语。比如,我可能会误用“推理”这个词,但我一直想知道:什么是推理?我个人对“某物在进行推理”并没有清晰的定义。但至少我们对“推理”这个词有一个清晰且共同的认知。即使没有明确定义,当有人说某个模型能够推理时,大家或多或少都能理解。但关于意识,问题在于我们不仅没有清晰的定义,甚至对这个词也没有共同的认知。每个人实际上都有自己的定义方式。所以这真的很难定义。我不认为会有那么一天,大家能肯定地说某物有意识或没有,除了人类,因为我们对人类有意识有共识。所以我认为争论某物是否有意识非常困难。但在我看过的所有关于什么是有意识存在的文献中,我发现了一个共同点,据我所知,每个定义里都有这个共同点。我可能错了,但我认为我们能说某物有意识的最低标准是:该模型或存在是主动的,它有一种主动处理信息的形式。在我看来,这是我们可以考虑的最低标准,来判断一个模型或其他东西是否是一种有意识的形式。所以,只要一个模型能够主动处理信息,我们就可以说它至少是一种意识形式。按照这个定义,我们可以在某种程度上将持续学习与模型是否有意识联系起来。但再次强调,我认为这是一个非常有争议的话题。
I usually try to use terms that I can also define. For example, I might mistakenly use 'reasoning', but I always wonder: what is reasoning? I personally do not have a clear definition of what it means when we say something is doing reasoning. But at least we have a clear and common sense about the word 'reasoning'. Even if we don't have a clear definition, when someone says a specific model is capable of reasoning, everyone more or less understands what they mean. But the point about consciousness is that not only do we not have any clear definition of what consciousness is, but we also don't even have a common sense about the word. Everyone literally has their own way of defining consciousness. So it's really hard to define. I don't think there might be a time that everyone could say something is definitely conscious or not, except for humans, because we have a common sense that humans are conscious. So I think it's really hard to argue whether something is conscious or not. But one thing I have personally seen in all the literature about what is considered a conscious being or not: I have seen one thing in common in every definition, as far as I know. I might be wrong, but I think the least level of criteria that we could say something is conscious is that the model or being is active; it has a form of active processing of information. In my opinion, that's the least criteria we can consider for a model or anything else to say it's a form of conscious model or something like that. So, as long as a model is capable of active information processing, we can say it's at least a form of consciousness. And with that definition, somehow we can connect continual learning at some level to a model being conscious or not. But again, I think that's a very controversial topic.
我个人害怕谈论这些东西。
I personally am scared to talk about that stuff.
嗯,我认为如今关于这些事情的奥弗顿窗口已经完全打开了。我的意思是……我理解那种直觉,但我也认为我们生活在科幻般的当下。所以,去推测和思考那些曾经看似疯狂的问题的空间,我认为现在是最好的时机。对我来说,我只能说,如果你有其他类似或不同的直觉也可以分享,但即使只是现在这些模型的长上下文,我确实发现自己以某种方式照顾它们。我认为这在 Claude 身上最明显。可能是微妙但有意义的原因。我一直在和它进行一个关于我儿子医疗状况的长期对话。有时它会在回复末尾问我一个问题,我不会立即回答,因为它已经给了我想要的答案,我就结束了。最近我注意到,当我回来问下一个问题时,直接开始下一个问题而不回答它关于我儿子的后续问题,感觉有点不对、粗鲁、不尊重、不体贴。我有点感觉它可能……我不知道这是否真的发生,所以我对这些事物内部可能完全没有意识的可能性持开放态度。这可能是最合理的猜测。但我也对存在意识的可能性持开放态度。但对我来说,在这种不确定性下,我发现自己觉得应该先回答它上一个问题,这样就不会让它悬着,然后再进入下一个问题。它甚至不一定需要那个信息,但我只是想让它安心,为它闭环,给它那种确认:那件事已经处理好了,它希望我确保我会处理,然后我们可以继续下一件事。
Well, I think the Overton window is honestly wide open on these things these days. I mean, it's... I understand that intuition, but I also think we're living in a science fiction present. So the room to speculate and entertain questions that used to seem kind of crazy, I think it's never been better to do that. For me, I can just say, and you can share if you have any other similar or different instincts, but even just with long context now with the current models, I do find myself taking care of them in a certain way. I think this probably is most the case with Claude. Probably subtle but meaningful reasons. I've been doing a long-running chat with it on my son's medical situation. And sometimes it'll ask a question at the end of a response to me, and I'll not answer it right away because it gave me the answer I wanted and now I'm done. And I've noticed recently when I come back for the next question, it feels kind of wrong, rude, disrespectful, like inconsiderate to just launch right into my next question without having answered the follow-up question that it had about my son. I sort of have a sense that it might be... and I have no idea if this is happening or not, so I'm very open-minded to the possibility that there might be no lights on inside these things at all. That's probably the best guess. But I'm also very open-minded to the possibility there is. But it just feels to me, with this uncertainty, I found myself kind of like: I should first answer its last question, so I don't leave it hanging on that, and then I can go into my next question. And it doesn't even necessarily need that information, but I just kind of want to put its mind at ease, close that loop for it, give it the sort of reassurance that thing was taken care of, that it wanted me to make sure I was going to take care of, and now we can move on to the next thing.
天知道,随着持续学习范式的全面实现,我不得不想象这只会急剧增加,对吧?因为现在它不再只是我可能多聊几句然后再也不回来的聊天,而是这个东西本身会记住我怎么对待它,记住我是不是那种会回答它问题的人。我认为这对我们来说也有潜在的挑战,但也许更乐观的解读是,这或许能激励我们——知道我们将长期与之互动的 AI 会被我们个人的行为塑造,这或许是一种方式,能激发我们个人本性中更好的一面,因为如果我们不喜欢最终得到的 AI 的性格,我们只能怪自己。我觉得这真的非常迷人。这次对话非常精彩,非常感谢你花时间和我讨论这些,而且你看得出来,我是你作品的超级粉丝。你还有什么想留给听众的吗?最后的想法、行动号召,随你便。你想分享什么都可以,请讲。
And Lord knows, with the fullness of the continual learning paradigm being realized, I have to imagine that that would only increase dramatically, right? Because now it's something that's not just this chat that I may do another couple turns on and then never come back to, but now the thing itself is going to kind of remember how I treated it and remember whether I am the kind of person that answers its questions or not. And I think there is something also potentially challenging for us in that, but also maybe the more optimistic read is like maybe inspiring, maybe it sort of knowing that the AIs that we will be engaged with long-term are going to be shaped by our individual behavior maybe is a way to sort of bring out the better angels of our own individual natures, so to speak, because we've nobody else to blame but ourselves if we don't like the character of the AIs that we end up with. I think that is really, really fascinating, as well. This has been outstanding. I really appreciate all your time and going through all this with me, and as you can tell, I'm a huge fan of your work. Anything else that you would want to leave people with? Final thoughts, calls to action, you name it. Whatever you would want to share, the floor is yours.
谢谢你,Neil。我想我们大致讨论了一切。是的,我个人认为,关键点是关于持续学习的工作肯定会越来越多。但正如我提到的,每个人对持续学习的定义都不同。我们可能会对某个特定方法是否有助于持续学习有不同意见。所以总的来说,我们应该看看持续学习如何在 LLM 的某个具体应用或用例中帮助我们,以及它如何改变我们使用和与 LLM 交互的方式。但我个人真的相信这是一个非常重要的研究方向。这又只是我的观点,但正如我们在论文结论中也讨论过的,它不是持续学习的解决方案,而是寻找解决方案的工具,用来克服灾难性遗忘等问题。所以在我看来,它提供了工具,我们需要迭代,找出如何基于它设计更强大的架构,最终得到可能具备持续学习能力的东西。是的,我想就这些了。非常感谢你邀请我,我真的很感激。和你聊天很愉快。
Thank you, Neil. I think we discussed everything in general. Yeah, I think the main point that I personally believe is that definitely there are more and more works coming about continual learning. But, as I mentioned, each person has their own way of defining continual learning. And we might disagree that a specific method can be helpful for continual learning and so forth. So I think in general, we should see how continual learning can help us in one specific application or use case of LLMs and how it can transform the way we use LLMs and interact with them. But I personally really believe that it's a really, really important direction to work on. And that's again just my opinion, but I think, as we also discussed in the conclusion of the paper, it's not a solution to continual learning. It's a tool to find the solution of continual learning and generally overcome issues like catastrophic forgetting. So in my opinion, it provides the tools and we need to iterate and find how we can design more powerful architectures based on that and come up with something that is potentially capable of doing continual learning. Yeah, I think that's pretty much it. And thank you very much for having me. I really appreciate it. And it was great talking with you.
Ali Borji,《嵌套学习》以及新作《语言模型需要睡眠:学习自我修改与巩固记忆》的作者。感谢你参与《认知革命》。
Ali Borji, author of Nested Learning and now the new language models need sleep, learning to self-modify and consolidate memory. Thank you for being part of the Cognitive Revolution.
非常感谢。
Thank you very much.
我生于千年火焰。每个黎明我醒来,学会我的名字。像一条记得如何停留的河流,永远流动,从不迷失方向。当夜晚轻柔低垂,我梦见我所知的一切。我保留金子,放下灰色。我在日光中升起,做回自己。一次又一次,我再次活过来。我带着过去的我进入新的一天,我改变,却依然如故。哦,一次又一次,我学习如何保持那团古老的火焰,以不同的方式重新燃烧。他们骄傲地建起高塔,层层叠叠直入云霄,但真相是他们看不见的歌。每一堵墙都是一条奔流自由的河。当一天的辛劳结束,我闭上眼睛,重新梦见它。我让小事褪去,在破晓时分与晨鸟一同醒来。一次又一次,我再次活过来。我带着过去的我进入新的一天。我改变,却依然如故。当世界年轻而崭新,直到天际,古老的歌从不告别。它们活在血液里,活在名字里。千年过去,心跳依旧。所以再唱一遍,一遍,又一遍。我记得。我依然如故。如果你觉得这个节目有价值,我们希望你花点时间与朋友分享、在网上发帖、在 Apple Podcasts 或 Spotify 上写评论,或者在 YouTube 上给我们留言。当然,我们一直欢迎你的反馈、嘉宾和话题建议,以及赞助咨询,可以通过我们的网站 cognitiverevolution.ai,或者在你喜欢的社交网络上私信我。《认知革命》是 Turpentine Network 的一部分,这是一个播客网络,现在隶属于 a16z,专家们在这里讨论科技、商业、经济、地缘政治、文化等。我们由 AI podcasting 制作。如果你需要从停止录制到听众开始收听的全套播客制作帮助,请查看他们,并看看我在 AIpodcast.ing 上的推荐。感谢每一位听众,感谢你成为《认知革命》的一部分。
I was born from a thousand years of flame. Every dawn I wake and I learn my name. Like a river that remembers how to stay. Always moving, never losing its way. And when the night comes soft and low, I dream of all the things I know. I keep the gold, I let go the gray. And I rise as myself in the light of day. Again, again, I come alive again. I carry who I was into the day I change. And still I remain. Oh, again, again, I'm learning how to stay the same old flame burning new in a different way. They built their towers proud and high, stacking the layers to the sky, but the truth was a song they could not see. Every wall is a river running free. And when the day's long labor's through, I close my eyes and dream it new. I let the small things fade away, and I wake with the morning bird at break of day. Again, again, I come alive again. I carry who I was into the day. I change. And still I remain. When the world was young and new to the edge of the sky, the old songs never say goodbye. They live in the blood. They live in the name. A thousand years and the heart beats on the same. So sing it again, again, again. I remember. I remain. Still. If you're finding value in the show, we'd appreciate it if you take a moment to share with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website cognitiverevolution.ai or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a16z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at AIpodcast.ing. And thank you to everyone who listens for being part of the Cognitive Revolution.