Liquid AI:从有限算力中榨取最大智能

Liquid AI: Squeezing Maximum Intelligence from Limited Compute

拉明·哈萨尼 Ramin Hasani · The Cognitive Revolution · 2026-07-04 · 约 108 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

CEO Ramine Hassani 解释 Liquid AI 如何通过生物启发式神经网络和架构搜索优化边缘设备 AI,例如可在 iPhone 上运行的 10 亿参数模型。

CEO Ramine Hassani explains how Liquid AI's biologically inspired neural networks and architecture search optimize AI for edge devices, with proof points like a 1B parameter model running on iPhones.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 37)

全文 · Full transcript(中英对照)

0. 引言与嘉宾背景 Introduction and Guest Background

Host

大家好,欢迎回到《认知革命》,祝美国的朋友们独立日快乐。今天的嘉宾是 Liquid AI 的 CEO Ramin Hasani,这家公司由 MIT 的研究人员创立,正在开发设备原生基础模型。我先说明一下,在录制前我鼓励 Ramin 深入探讨 Liquid AI 的技术细节。正如你将听到的,他出色地展示了技术深度、差异化愿景和极具感染力的热情,在我看来,这期节目堪称经典。我们从团队在 MIT 开发的微型生物启发微分方程神经网络研究开始,这项研究也是他们创办公司的灵感来源。他们展示的一些能力,比如仅用 12 个液体神经元组成的控制模块就能停车,至今听起来仍像科幻小说。虽然这些系统尚未扩展到今天的能力前沿,但公司保留了液体哲学,即采用中立实证方法来设计和优化神经网络,使其在各种特殊约束下运行,最常见的是在内存和处理能力有限的边缘设备上运行。考虑到全球智能手机和笔记本电脑市场每年约 8000 亿美元,而全球 AI 数据中心建设才刚刚超过这个数字,推理需求可能会让世界上大部分地区无法负担前沿模型,而且许多企业和个人重视隐私和对自己信息的控制能力。这本身就是一个巨大的市场机会。Liquid 有扎实的证明,包括在 HuggingFace 美国下载排行榜上排名第五,以及与 Shopify 和梅赛德斯-奔驰等公司的显著合作。任何怀疑的人都可以快速下载并演示 Liquid 的 Apollo 应用,根据我的经验,即使是一个 10 亿参数的模型,结合少量注意力层和非常简单的门控学习卷积,虽然远非前沿,但在 iPhone 上运行速度足够快,可以成为基本用例(如私下搜索和分类本地文档)的真正选择。最有趣的是 Liquid 用于开发特定用例和运行时环境网络的网络架构搜索过程。他们发现代理指标常常误导过程,因此现在在客户实际使用的目标硬件上评估模型在真实下游任务上的表现。Ramin 分享了更多细节。简而言之,虽然基于注意力机制的架构在泛化方面仍然优于任何已知替代方案,因此继续主导前沿,但你的用例越具体,可用的算力资源越有限,他们的搜索过程就越可能找到一种非传统架构。目前,这正是 Mamba 和其他次二次创新架构大放异彩的地方。最后,Ramin 预告了 Liquid 即将推出的平台,允许客户以自助方式微调小模型用于自己的用例。出于多种原因,包括可能缓解对前沿模型的需求并改善全球 AI 的可及性,我个人非常期待看到它上线。话不多说,希望你喜欢这场高能量的对话,了解 Liquid AI 如何与联合创始人兼 CEO Ramin Hasani 一起从任何给定的计算资源中榨取尽可能多的智能。《认知革命》由 Mercury 赞助,这是一家金融科技公司,超过 30 万家雄心勃勃的公司和个人信赖它来管理财务。我已经把 AI 融入到生活的几乎每个角落:我的邮件、消息、日历。我甚至给了我的智能体 Mercury 虚拟卡,设置了低额度以及类别和商户限制供它们自主使用。但我的 AI 对财务数据的访问仍然有限。使用普通银行,我可能会导出一堆对账单让助手处理。但要获取实时最新信息,尤其是采取任何行动,试图让智能体通过浏览器使用银行太困难、太慢、太容易出错,不值得。这就是为什么 Mercury 的新对话界面 Command 如此重要。它直接内置于 Mercury,意味着你可以用自然语言访问财务信息,而无需将任何信息暴露在银行账户之外。无需导出、无需电子表格、无需将交易粘贴到第三方工具中。我真的认为很多人会喜欢这种方式。它已经可以帮助你采取行动,所有操作都受你账户中已设置的权限和审批策略约束。我对 2026 年银行业达到这种 AI 集成水平感到由衷钦佩。所以我邀请你和我一起进入未来。访问 mercury.com 了解更多信息,并在几分钟内在线申请。Mercury 是一家金融科技公司,不是 FDIC 保险银行。银行服务由 Choice Financial Group 和 Column NA(FDIC 成员)提供。感谢 Mercury 对《认知革命》的支持。现在开始节目。Liquid AI 联合创始人兼 CEO Ramin Hasani,欢迎来到《认知革命》。

Hello, welcome back to the Cognitive Revolution and happy 4th of July to everyone in the United States. Today my guest is Ramin Hasani, CEO of Liquid AI, a company founded by MIT researchers that's developing device native foundation models. I'll say upfront that just before recording, I encouraged Ramin to go deep into the weeds on the technical details of Liquid AI's work. And as you'll hear, he did a truly excellent job demonstrating a mix of technical sophistication, differentiated vision, and a contagious passion that, in my humble opinion, makes this episode an instant classic. We start with an overview of the team's research into tiny biologically inspired differential equation-based neural networks that Ramin and team developed at MIT and which inspired them to start the company. Some of the capabilities they demonstrated, such as parking a car with a control module that consisted of just 12 liquid neurons, still sound a bit like science fiction today. And while those systems haven't scaled up to today's capabilities frontier, the company has maintained the liquid philosophy, which today means taking a neutral empirical approach to designing and optimizing neural networks to perform under all sorts of exotic constraints, including most commonly the need to run on edge devices with limited memory and processing power. Considering the fact that the global smartphone and laptop market is roughly $800 billion per year, a number which the global AI data center buildout is only now surpassing, the demand for inference threatens to price much of the world out of the frontier model market, and that so many enterprises and individuals value privacy and the ability to control their own information. This is an absolutely massive market opportunity unto itself. And Liquid has serious proof points, including holding the number five spot on the HuggingFace United States Downloads leaderboard, plus notable partnerships with companies such as Shopify and Mercedes-Benz. Anyone who doubts can do a quick download and demo of Liquid's Apollo app, which shows in my experience that even a 1 billion parameter model, which combines a small number of attention layers with a very simple gated learned convolution, while admittedly far from the frontier, can run fast enough on an iPhone to be a real option for basic use cases such as privately searching through and classifying one's own local documents. Perhaps most interesting is the network architecture search process that Liquid uses to develop networks for particular use cases and runtime environments. Having found that proxy metrics too often lead the process astray, they now evaluate models on real downstream tasks on the actual target hardware that their customers intend to use. Ramin shares a lot more detail on their findings. But in short, while attention-based architectures continue to generalize better than any known alternative and therefore continue to dominate the frontier, the more specific your use case and the more limited the compute resources you have available, the more likely their search process is to land on an exotic architecture. This for now is where architectures like Mamba and other subquadratic innovations really shine. Toward the end, Ramin teases a platform that Liquid will soon be introducing to allow customers to fine-tune small models for their own use cases on a self-served basis. For multiple reasons, including the potential to ease demand for frontier models and improve access to AI globally, I for one will be very excited to see that come online. And so without further ado, I hope you enjoy this high energy look at how Liquid AI is squeezing as much intelligence as possible out of any given computational resource with co-founder and CEO Ramin Hasani. The Cognitive Revolution is brought to you by Mercury, the fintech that more than 300,000 ambitious companies and individuals trust to run their finances. I've wired AI into nearly every corner of my life. my email, my messages, my calendar. I even gave Mercury virtual cards to my agents with low limits and category and merchant restrictions for their autonomous use. But still, my AI's access to my financial data has remained limited. With a normal bank, I might export a bunch of statements and have my assistant process them for me. But for real-time up-to-date information and certainly for taking any action, trying to get your agent to use the bank via the browser is just too hard, too slow, and too error-prone to be worth it. And that's why Mercury's new conversational interface, command, is such a big deal. It's built directly into Mercury, which means you get natural language access to your finances without exposing anything outside of your bank account. No exports, no spreadsheets, no pasting your transactions into third party tools. I really think a lot of people are going to prefer it this way. And it can already help you take actions too with everything bound by the permissions and approval policies that you've already set up in your account. I am genuinely impressed to see this level of AI integration in banking in 2026. And so I invite you to join me in the future. Visit mercury.com to learn more and apply online in minutes. Mercury is a fintech company, not an FDIC insured bank. Banking services provided through Choice Financial Group and Column NA members FDIC. Thank you to Mercury for supporting the Cognitive Revolution. And now on with the show. Ramin Hasani, co-founder and CEO at Liquid AI. Welcome to the Cognitive Revolution.

Ramin Hasani

谢谢邀请。我很期待这次对话。多年来我一直从远处关注 Liquid AI,对你们在 MIT 时期共同开发的一些架构非常着迷。我也对公司过去几年的发展轨迹很感兴趣,它变得更以客户为中心,成为一个商业实体。那么首先,你能否讲讲 Liquid AI 从创立到今天的大致故事,以及公司现在的使命是什么?

Thanks for having me. I'm excited for this conversation. I've been following Liquid AI from afar for a number of years and fascinated by some of the architectures that you guys have developed going back to your time at MIT together. Also fascinated by the trajectory that I perceive the company to have taken over the last years as it's become more customer focused and commercial entity. So maybe for starters, how would you tell the kind of broad story of Liquid AI leading up to what you're doing today and what the company's mission is today?

Ramin Hasani

当然可以。Liquid AI 成立于三年半前,我们从 MIT CEL 孵化出来,基于一项我们实际上已经研究了大约 10 年的技术——基本上是在创办公司之前十年。我们在 MIT 的目标函数始终是:将尽可能多的智能压缩到最小规模的算法中。

Yeah, absolutely. So Liquid AI three and a half years ago, we spun out of MIT CEL building on a technology that we have actually have been working on it like 10 years before like basically like a decade before when we started the the whole company. We've been our objective function like at MIT has always been like maximizing the amount of intelligence that we can into smallest format of algorithms.

1. 效率与真实世界机器人 Efficiency and Real-World Robotics

Ramin Hasani

效率一直是我们研究的基石。我们一直专注于机器人技术和进入现实世界的系统。液态神经网络的概念是在我的博士论文中发现的。我和我的四位联合创始人一起研究了如何构建适用于机器人的机器学习解决方案,这些方案不需要数百万或数十亿参数,因此可以直接部署在 CPU、NPU 或安装在物理系统上的较小 GPU 上,同时提供与更大规模 AI 系统相当的可靠性。

Efficiency has been the cornerstone of our research. We have been working specifically on robotics and systems that come into the real world. The idea of liquid neural networks was discovered during my PhD thesis. Together with my co-founders, we researched how to build machine learning solutions for robots that don't have millions or billions of parameters, so we can host them directly on CPUs, NPUs, or smaller GPUs mounted on physical systems, while delivering the reliability of much larger AI systems.

Ramin Hasani

本质上,我们试图构建替代算法,在算法空间中进行创新,看看如何构建能够泛化到未见数据的机器学习系统。在现实世界中,分布偏移是一个真实存在的问题。想象一下在开放世界中部署机器人——自动驾驶汽车、飞行无人机、固定翼飞行器。这些系统会迅速遇到分布外的情况,因此你需要构建能够从容应对分布外情况的系统。人类非常擅长这一点,自然学习系统(如动物)也能很好地处理分布外情况。学到的概念可以泛化到未见过的数据。我们的想法是构建具有更好分布外泛化能力的系统。这不仅是我们的目标,也是整个 AI 领域的目标,尤其是在机器人技术和智能体在环境中行动的闭环环境中。

Essentially, we try to build alternative algorithms to get creative in the algorithmic space, to see how we can build machine learning systems that generalize beyond the data they have seen. In the real world, distribution shifts become a real issue. Imagine deploying a robot in an open world—an autonomous car, a flying drone, a fixed-wing vehicle. These systems will rapidly go out of distribution, so you need systems that are comfortable with being out of distribution. Humans are extremely good at this, and natural learning systems, like animals, follow a nice trajectory out of distribution. Learned concepts can generalize to unseen data. Our idea was to build systems with better out-of-distribution generalization. This has been a goal for the entire AI field, especially in robotics and closed-loop environments where an agent acts in an environment.

Ramin Hasani

自然地,我们开始研究大脑——动物的大脑。我想从第一性原理出发:神经元如何交换信息。我们开始研究蠕虫的大脑,特别是秀丽隐杆线虫。原因是在 2015 年,当我与联合创始人 Matias Lechner 开始这项博士研究时,这是唯一一个我们了解其整个神经系统的动物。这种动物仅用 300 个神经细胞就表现出大量的感觉反应行为。这令人着迷,因为它比当时任何用于控制的神经网络都要小得多,却能比世界上最好的机器人系统做出更灵巧的动作。所以我们想:让我们先了解蠕虫大脑中神经元如何交换信息,然后构建更复杂的神经回路,遵循神经系统的进化路径。

Naturally, we started looking into brains—animal brains. I wanted to take a first-principles approach: how neurons exchange information. We started looking into the brain of worms, specifically C. elegans. The reason is that in 2015, when I started this research as part of my PhD with my co-founder Matias Lechner, this was the only animal whose entire nervous system was known. This animal exhibits massive sensory-reactive behavior with only 300 cells in its nervous system. That was fascinating because it's much smaller than any neural network performing control at that time, yet it can do better dexterous movements than the best robotic systems. So we thought: let's understand how neurons exchange information in the worm's brain, then build more complex neural circuits, following the evolution of nervous systems.

Ramin Hasani

存在描述秀丽隐杆线虫两个神经元之间动力学过程的方程。由于蠕虫很小,其神经元不会产生尖峰;它们是分级神经元,以电子方式运作,类似于人工神经网络。它们非常可微,这使得它们很有吸引力,因为我们可以应用学习理论。一旦我们构建了相邻的两个、四个、八个或 100 个神经元并进行训练,我们就可以将反向传播作为微分编程应用于受自然启发的系统。这些微分方程表现良好。你可以让它们变得任意复杂,从而更接近生物学,尽管计算足迹会变得更复杂,难以扩展。但我们可以抽象掉那些描述两个细胞之间神经动力学的复杂微分方程,并用 sigmoid 函数、门控 sigmoid 函数来表示它们。现在我们有了矩阵乘法,通过基于 Transformer 的架构和注意力机制来捕捉输入的影响。这些都是为了扩展机器学习解决方案而进行的简化计算。

There are equations describing the neuronal dynamics of two neurons in C. elegans. Because the worm is small, its neurons do not spike; they are graded neurons that behave in an electronic way, similar to artificial neural networks. They are very differentiable, which made them attractive because we could apply learning theory. Once we construct two, four, eight, or 100 neurons next to each other and train them, we can apply backpropagation as differential programming on a system inspired by nature. The differential equations were well-behaved. You can make them as complicated as you want, and they mimic biology more, though the computational footprint becomes more complex to scale. But we can abstract away those complex differential equations describing neural dynamics between two cells and represent them with sigmoidal functions, gated sigmoidal functions. Now we have matrix multiplications capturing input impacts with Transformer-based architectures and attention mechanisms. These are simplified computations to scale machine learning solutions.

Ramin Hasani

从神经科学的角度来看,探索是否将那些基于微分方程的计算——一种更精细的计算形式——引入每个神经元的行为中,能够解锁比人工神经网络更强大的能力,这非常有趣。早期的结果令人着迷。用 12 个神经元,我们可以让汽车自动平行泊车。用 19 个神经元,我们可以驾驶汽车。用 30 个神经元,我们可以让无人机自主导航。我们可以用这些更复杂的神经动力学来处理感官信息,我们称之为液态神经网络——‘液态’代表适应性。

From a neuroscience perspective, it was interesting to explore whether bringing back those differential equation-based computations—a more elaborate form of computation into the behavior of every single neuron—could unlock something greater than what we have seen from artificial neural networks. Early results were fascinating. With 12 neurons, we could autonomously parallel park a car. With 19 neurons, we could drive a car. With 30 neurons, we could autonomously navigate a drone. We could process sensory information with these more complex neural dynamics, which we called liquid neural networks—'liquid' for adaptability.

2. 液态神经网络简介 Introduction to Liquid Neural Networks

Ramin Hasani

我们称之为液态时间常数神经网络,这就是我们发表的 LTC 论文,我们将其命名为“液态”,是因为这些系统的动力学在训练后仍保持灵活。因此,模型能够对接收的新类型输入做出反应,甚至在反向传播过程中学习如何适应所接收的输入。与人工神经网络以及其他你见过的系统相比,它实际上在神经网络学习动态的灵活性上编码了更多的自由度。

We called it Liquid Time Constant Neural Networks, like that was the LTC kind of paper that we got out, and we coined the name 'liquid' for the fact that the dynamics of these systems stay flexible even after training. So the models would be able to react to new types of inputs that they receive and even learn during backpropagation how to be adaptive towards the inputs that you're receiving. It actually encoded a little bit more degrees of freedom on flexibility of the learning dynamics of a neural network compared to artificial neural networks, compared to other systems that you have seen.

3. 应用与优势 Applications and Advantages

Ramin Hasani

这种动力学形式实际上使我们能够将这项技术扩展到不仅仅是机器人领域,还应用于金融服务、医疗领域的预测 AI,以及我们应用过的许多不同地方,比如音频建模、多模态视频理解。我们反复应用这项技术,并看到了它的前景,这项技术确实非常惊人。每个神经元都比普通的神经网络复杂得多,但同时,你不需要那么多的计算点就能得到你想要的结果。我们还看到这些模型的分布外性能非常好,它们非常适合机器人应用。

That form of dynamics actually allowed us to scale this technology into not just robotics, but also applying it to predictive AI in financial services, in medical domain, and in many different places that we applied, like audio modeling, multimodal video understanding. We applied this technology over and over and saw promise that this technology is actually very amazing. Every neuron is much more complex than a normal kind of neural network, but at the same time, you don't need that many points of computation in order to get to the results that you want. We also saw that the out-of-distribution performance of these models is extremely good, and they were suited for robotics applications.

Ramin Hasani

我们与美国空军有过合作。当时,波音公司实际上在 MIT 资助了我的博士后研究,我们展示了即使用这些神经网络和少量神经元也能驾驶喷气式飞机。看到你能用由这些液态神经网络驱动的非常小的系统(而非人工神经网络)将分布外泛化推到多远,这非常有趣。

We had interactions with the United States Air Force. At that time, Boeing was actually hosting my postdoc at MIT, and we showed that you can actually fly even jets with these type of neural networks with a handful of these neurons. It was very interesting to see how far you can push out-of-distribution generalization with a very small set of a system powered by these liquid neural networks, as opposed to artificial neural networks.

4. 复杂性与可扩展性挑战 Complexity and Scalability Challenges

Ramin Hasani

正如我所说,液态神经网络中的每个节点都非常复杂;它是一个微分方程。你必须求解它,而且神经元越多,网络的前向传播和反向传播就越复杂。因此,我们需要考虑效率。这些系统是高度非线性的。我们尽量不牺牲非线性,因为向学习系统添加非线性可以让你构建更具表达力的系统。这是我的博士论文中的一个证明:非线性实际上直接起作用,尤其是在较小规模的模型上。但对我们来说,问题是如何扩展这些系统?你有非线性系统,但扩展非线性系统极其困难。可扩展性背后的原因是模型的计算复杂度甚至达到了立方复杂度。它甚至不是二次的;你谈论的是立方复杂度,这太高了。

As I told you, each node in a liquid neural network is very complex; it's a differential equation. You have to solve that, and the more neurons you have, the more complex is a forward pass and a backward pass through a network. Therefore, we needed to think about efficiencies. These systems are highly nonlinear. We try not to sacrifice on nonlinearity because adding nonlinearity to learning systems allows you to build more expressive systems. This is a proof that I have in my PhD thesis: nonlinearity actually directly contributes, especially on the smaller size of models. But then the idea for us was how can we scale these systems? You have nonlinear systems, but it is extremely difficult to actually scale nonlinear systems. The reason behind scalability is the computational complexity of the models hitting even cubic complexity. It's not even quadratic; you're talking about cubic complexity, which is too much.

Ramin Hasani

你有一组微分方程;你可以用数值求解器展开它们。运行数值求解器,逐步计算所需的输出。用于计算的步数越多,结果就越精确。但问题又回到了可扩展性;你无法真正用这些数值求解器达到无限精度。我们想到的一个想法是:如果我们直接用闭式解求解整个系统呢?实际上,你有一个由这些液态神经网络组成的微分方程系统,每个网络代表两个细胞在某种抽象层次上交换信息的神经动力学。我们看到,让我们取这个系统并尝试用闭式解求解。结果发现,自 1907 年以来,这类方程就没有闭式解了。

You have a set of differential equations; you can roll them out with a numerical solver. You run the numerical solver and step by step compute the desired outputs. The more steps you use to compute, the more accurate the results become. But then the problem becomes scalability again; you cannot really get to infinite precision with these numerical solvers. One idea that came to us was: what if we just solve the whole system in closed form? Literally, you have a differential equation system with these liquid neural networks, each representing neural dynamics of two cells exchanging information with each other at a certain level of abstraction. We saw that let's take this system and try to solve it in closed form. Turns out the closed-form solution for this type of equation hasn't existed since 1907.

5. 历史背景与闭式解 Historical Context and Closed-Form Solution

Ramin Hasani

1907 年,一位名叫 Louis Lapicque 的科学家对膜电位进行了数学建模,这种方程形式成为通道建模的基础,比如信息如何通过细胞内的离子通道传播,以及神经递质如何传播到另一个细胞。所以 1907 年是 Louis Lapicque 的膜电位方程,它是一个开放微分方程。然后我们看到 Hodgkin 和 Huxley 这两位科学家;他们开始真正从生物学角度奠定这类微分方程的基础,并为模型增加了一点复杂性。他们开发了一个神经元的神经科学模型,即神经元如何对膜电位做出反应。他们于 1953 年开始研究,并在 1963 年因开发出更准确、更精确的神经动力学表示而获得诺贝尔奖。从那时起,在每一本教科书中,这类方程都没有已知的闭式解。液态神经网络也属于这类方程。我们首次真正解决了这个问题。大约在 2022 年,我们首次用闭式解解决了液态神经网络中神经元之间的相互作用。这成为一篇《自然·机器智能》论文,于 2022 年 11 月发表,称为闭式连续时间系统。闭式解具有重大意义,因为现在我无需使用任何数值求解器就能运行液态神经网络。

In 1907, a scientist called Louis Lapicque modeled the membrane potential mathematically, and that format of equation became fundamental for channel modeling, like how information propagates through ion channels inside a cell and how neurotransmitters are propagated to the other one. So 1907 was Louis Lapicque's membrane potential equation, which is an open differential equation. Then we've seen scientists called Hodgkin and Huxley; they started working on really biologically grounding this type of differential equation and adding a little bit more complexity into the model. They developed a neuroscience model of one neuron, how a neuron reacts to a membrane potential. They started in 1953 and won a Nobel Prize in 1963 for developing a better and more accurate representation for neuronal dynamics. From then on, in every textbook, this type of equation does not have a known closed-form solution. Liquid neural networks were also part of that type of equations. For the first time, we actually solved that. Around 2022, we solved the liquid neural network interaction of neurons with each other in closed form for the first time. This became a Nature Machine Intelligence paper published in November 2022, called the Closed-Form Continuous-Time Systems. The closed form has massive implications because now I don't need to use any numerical solvers to actually run a liquid neural network.

6. 从生物学到液态AI From Biology to Liquid AI

Ramin Hasani

我现在不仅能有数百个神经元,还能让数十亿个神经元彼此相邻,并且真正在保持非线性的前提下扩展这些计算。然后在 2023 年 1 月,实际上是 2 月初,一篇关于我和联合创始人 Matias 的文章出现在《量子杂志》上。这是一篇人物特写,探讨了最终在神经元动力学上找到闭式解的意义,以及这对机器学习和整个脑科学有多么重要。然后我的收件箱里全是风投;硅谷开始热议,每个人都想给我们递上投资意向书,让我们真正启动并扩展这项技术,因为它从根本上不同于基于 Transformer 和注意力机制的架构。这项技术植根于生物学。大约在同一时间,替代模型也出现了,比如状态空间模型、更快的状态空间模型迭代、卷积神经网络,它们都以线性形式变得可扩展。因为要扩展替代模型,你必须将其线性化。而我们这里的东西属于另一个类别;它加入了一些我们从生物学和物理学中学到的算子。随着研究的深入,我们现在已经变得非常有竞争力,能够构建解决越来越通用的任务的模型。当我说通用任务时,我指的是对信号进行建模,超越预测性能——以人类理解的方式对语言、音频和视觉信号进行建模。因此,Liquid AI 的使命是构建每个规模上高效的通用 AI 系统。我们在使命中强调了“高效”这个词,我们非常在意这一点,因为我认为从基础模型公司的角度来看,我们可能是地球上最高效的基础模型公司。我稍后会解释原因。我们将使命定为高效机器学习,是因为我们看到了构建闭式解需要多少考量,然后将其扩展到一个能够建模语言、音频、视频和视觉的系统。我们必须深入这些方面,才能以计算上可行的方式将我们的神经架构扩展到今天的状态。关于技术的核心,我稍后会讲——但今天,这项技术已经变成了 Liquid Foundation Models。Liquid Foundation Models,或者叫 Elephants,现在非常受欢迎。在美国,从下载量和受欢迎程度来看,我们排名第五。我们在 Hugging Face 上每周有超过 100 万次下载。我们构建的这些小型模型下载量很大。美国这边排名靠前的是 Google、Meta、Microsoft 和 Nvidia,第五名是 Liquid AI。我们只用了大约 1000 块 GPU 就达到了这个水平。所以从基础模型的角度来看,达到这样的知名度并发布 50 多个模型实例,供企业使用——这就是我们基于从自然和过去获得的灵感所打造的 Liquid。

I can now have not only hundreds of neurons but billions of neurons next to each other, and I can actually scale these computations while keeping the nonlinearity as part of this thing. Then in January, actually beginning of February 2023, an article came out in a quantum magazine about me and my co-founder Matias. It was a profile about the implications of having some closed-form solution finally on neuron dynamics, and how important this can be for both machine learning and brain science as a whole. Then my inbox was full of VCs; Silicon Valley started talking, everybody wanted to throw a term sheet at us to really get started and scale this technology because this was fundamentally different than a Transformer-based architecture and attention-based architecture. This was grounded in biology. Around the same time, alternative models were coming out, like state space models, faster iterations of state space models, convolutional neural networks that came out to be scalable, all in a linear form. Because to scale alternative models, you have to linearize them. And we have something here that is another category; it adds some operators that we learned from biology and from physics. As we grew the research, we got to the point where we can now become very competitive and build models for solving more and more general-purpose tasks. When I say general-purpose tasks, I'm talking about modeling signals beyond predictive performance—modeling signals for language, audio, and vision in a way that humans understand. So the mission of Liquid AI became building efficient general-purpose AI systems at every scale. We coined the word "efficient" in our mission, and we care about it so much because I think from a foundation model company's perspective, we are probably the most efficient foundation model company on the planet. I'll tell you why in a minute. The reason we set the mission to efficient machine learning is because we saw how many considerations we had to make to build a closed-form solution, then take it and scale it to a system that can model language, audio, video, and vision in general. We had to go into these things so that we could computationally tractably scale our neural architectures to where we are today. The base of our technology—I'll talk about the core technology later—but today, the technology became Liquid Foundation Models. Liquid Foundation Models, or Elephants, are pretty popular right now. We are ranked number five in the US from the number of downloads and popularity. We have over 1 million downloads per week on Hugging Face. That's a lot of downloads of these small models we are building. The top ones from the US side are Google, Meta, Microsoft, and Nvidia, and the fifth is Liquid AI. We got ourselves there by consuming about 1,000 GPUs. So from a foundation model perspective, getting to that level of popularity and releasing more than 50 model instantiations that people are using in enterprises—that's where we brought Liquid, building on inspirations from nature and the past that I portrayed for you.

7. 赞助商插播:Anthropic Sponsor Break: Anthropic

Host

嘿,我们稍作休息,马上回来继续采访。今天的节目由 Anthropic 赞助,他们是 Claude 和 Claude Code 的开发者。在过去的几个月里,Claude 帮助我构建并完善了一个个人深度上下文数据库,其中包含了我过去整整 5 年的所有电子邮件、Slack 消息、推文、跨平台私信、视频通话和播客转录。在此基础上,我们还加入了描述我与数百个联系人、组织和想法关系的摘要文章。现在有了这个数据库,几乎没有什么事是 Claude 帮不上忙的。报税季,我让 Claude 帮我整理。它翻遍了我的收件箱,找到了我所有 10 份兼职工作的 1099 表格,并为我生成了一份关于支出和捐款的全面报告。对于我的天使投资,Claude 现在可以根据我与创始人的通话和邮件往来,以我的风投基金要求的格式起草投资备忘录。当有人需要帮忙时,Claude 通常能做得和我一样好。最近,一位朋友问我是否认识适合他正在招聘的职位的人选。起初我没想到任何人,但后来我想到去问 Claude,果然,它找到了两个很棒的候选人。Claude 是为那些不满足于“足够好”的头脑而生的 AI。它是一个真正理解你整个工作流程并与你一同思考的协作者。无论你是在午夜调试代码,还是制定下一个商业策略,Claude 都能拓展你的思维,解决真正重要的问题。所以,对于值得解决的问题,请访问 claude.ai/tcr 开始使用 Claude。网址是 claude.ai/tcr。同时查看 Claude Pro,它包含了本期节目中提到的所有功能。再次强调,网址是 claude.ai/tcr。

Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Anthropic, makers of Claude and Claude Code. Over the last few months, Claude has helped me build and refine a personal deep context database that now contains all of my emails, Slack messages, tweets, DMs across platforms, video calls, and podcast transcripts going back a full 5 years. On top of that, we've now layered summary articles describing my relationship with hundreds of contacts, organizations, and ideas. And now that this exists, there's almost nothing that Claude can't help with. For tax season, I asked Claude to help me get organized. It went through my inbox, tracked down 1099s for all 10 of my part-time jobs, and built me a comprehensive report on my expenses and donations. For my angel investing, Claude can now draft investment memos in exactly the form that my venture fund requires based on the calls I've had and the emails I've exchanged with the founders. And when someone needs a favor, Claude can often do it as well as I can. Recently, a friend reached out to ask if I know anyone who might be a fit for a role that he is currently hiring for. Initially, nobody came to mind, but then I thought to ask Claude, and sure enough, it identified two great leads. Claude is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow and thinks with you. Whether you're debugging code at midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. So for problems worth solving, get started with Claude at claude.ai/tcr. That's claude.ai/tcr. And check out Claude Pro, which includes all of the features mentioned in today's episode. Once more, that's claude.ai/tcr.

8. 神经元vs参数:范式转变 Neurons vs Parameters: A Paradigm Shift

Host

太棒了。后续有很多方向我想深入探讨。首先,一个非常根本的问题让我印象深刻:当你描述液态神经网络时,你用神经元数量来描述,而我们通常听到的是参数数量。我隐约觉得这背后有一种范式差异,导致了描述方式的不同。所以,帮我建立一点直觉吧——我猜这可能与非线性有关。这非常根本,对吧?因为虽然已经取得了惊人的进展,但在真正鲁棒的域外泛化等方面,我们还有很多工作要做。在对抗鲁棒性等方面,主流范式下我们仍有很多工作要做。所以,请帮我拆解一下。为什么我们数的是神经元而不是参数?这告诉我们什么?

Brilliant. So many different directions I want to go in follow-up there. Maybe for starters, one thing that on a very fundamental level jumps out at me is when you describe the liquid neural networks, you describe them in terms of the number of neurons as we're used to hearing about the number of parameters. And something tells me there's a kind of paradigmatic difference there that underlies the difference in description. So yeah, help me develop my intuition a little bit more for I guess it probably connects to nonlinearity. And this is so fundamental, right? Because there's so much incredible progress that has been made, but on things like really robust out-of-domain generalization, we've still got a lot of work to do. On things like adversarial robustness, we still have a lot of work to do in the mainline paradigm. So yeah, tell me, unpack a little bit. Why are we counting neurons versus parameters and what does that tell us?

9. 神经元作为计算单元与参数数量 Neuron as Unit of Computation and Parameter Count

Ramin Hasani

所以我认为早期我们确实开始把神经元数量作为计算单元来讨论,因为那非常有趣。你可以把每个神经元关联到一个可以用数学方程建模的过程。你可以分配类似 sigmoid 的门控。所以函数越多,单个细胞增加的参数就越多。在液态神经网络中,就参数而言,你可以说乘以 7,那就是那个微分方程中一个神经元的参数数量。但那个细胞本身,因为它所拥有的方程,计算的不仅仅是前向传播。它有内部反馈,内部有三重反馈机制。所有这些参数实际上与任何你构建的人工神经网络系统在参数方面是一一对应的。所以一个经验法则:当我说神经元数量时,你乘以 7 就能得到系统的参数数量。但这适用于液态神经网络的早期版本。随着我们开始为想要探索的数学函数类型标准化并行化方案,以及我们获得的灵感,我们最终收敛到了参数数量。今天,当我们谈论一个拥有比如 10 亿参数的液态基础模型时,我们字面上的意思就是 10 亿参数,就像任何其他正在构建的 GPT 一样。这说得通吗?

So I think early on we actually started talking about the number of neurons as the unit of computation in our research, because that was very interesting. You can associate every neuron with a process that can be modeled by a mathematical equation. You can allocate basically sigmoidal kind of gates. So the more functions you have, the more parameters get added to that single cell. In liquid neural networks, in terms of parameters, you can say multiply by seven, and that would be the number of parameters of a neuron in that differential equation. But that cell itself, because of the equations it has, computes more than just a forward pass computation. It has internal feedbacks, three-degree feedback mechanisms inside. All those parameters are actually one-to-one analogous to any artificial neural network system you would build in terms of parameters. So a rule of thumb: when I say number of neurons, you multiply by seven to get the number of parameters of the system. But this is for the early version of liquid neural networks. As we started standardizing parallelization schemes for the type of mathematical functions we want to explore, and inspirations we are getting, we just converged to the number of parameters. Today, when we talk about a liquid foundation model with, say, 1 billion parameters, we literally mean 1 billion parameters in the sense of any other GPT being built out there. Does that make sense?

Host

是的。但让我们再深入一点看看那个神经元。我也有点好奇:当你尝试用我们今天拥有的 sigmoid 类型函数来扩展原始的液态方法时,根本限制或第一个瓶颈是什么?我认为这些函数肯定是被精心选择以便在 GPU 上轻松运行的。我猜测单个神经元内部的这些多自由度可能在现有硬件上执行该过程时带来挑战。我们如何从几十个神经元中得到这些惊人的结果?那个明显的“苦药丸”问题会是:如果我们采用完全相同的范式,扩展到数百万或数十亿神经元会怎样?但你一定会在过程中遇到一些瓶颈。它们是什么?

Yeah. But let's zoom in maybe a little bit more on that neuron. I'm wondering a little bit too: what is the fundamental limit or what are the first bottlenecks that you hit when you try to scale the original liquid approach with the sigmoid type functions that we have today? Those are, I think, certainly carefully selected to be easily run on GPUs. I'm guessing that these multiple degrees of freedom inside a single neuron maybe present challenges in terms of executing that process on available hardware. How can we get these amazing results from tens of neurons? The obvious bitter-pilled question would be: what if we take that exact paradigm and go to millions or billions? But you must be hitting some bottlenecks along the way. What are they?

10. 主要瓶颈:串行到并行计算 Main Bottleneck: Sequential to Parallel Computation

Ramin Hasani

所以主要挑战是将顺序计算转化为并行计算。当你从单个神经元动力学开始转向多个神经元动力学时,系统的参数不再是标量或向量,而是变成矩阵和张量。所以现在你谈论的是矩阵乘法。你想把标量计算转化为张量计算,对吧?你希望能够并行化顺序计算。问题在于,并非总是如此:如果方程中有非线性关系,你不能轻易地在向量化计算和张量化计算之间建立一一映射。所以液态神经网络和其他任何非线性循环神经网络的这些非线性属性——我的意思是循环本身就给整个方程的数学增加了复杂度。更不用说如果循环本身带有非线性,那么用我们已知的典型线性代数方法将向量解耦或写成张量就会变得复杂得多。你必须能够将这些计算转化为张量,才能并行计算,就像一次计算一个那样。你明白吗?所以这就是数学挑战。这不仅仅与液态网络有关。任何你想要张量化的数学运算,如果它是非线性的,你就会遇到我们讨论的问题。这就是为什么出现的状态空间模型都是关于线性动力学的。我们谈论的是线性动力系统。而这些线性动力系统之所以重要,纯粹是因为我们没有合适的方法来扩展非线性系统。你可以对线性动力系统应用非线性。例如,在一个门控系统中,你可以先以线性方式执行动力系统,然后添加一个 sigmoid 函数。你并行计算所有矩阵等,但之后可以在整个张量上逐点应用非线性操作。这就是你能做的。但如果参数之间的关系本身由某种非线性支配呢?那么你就不能真正使用典型的线性代数。你总是必须将那个非线性系统近似为线性系统,然后才能并行化系统,对吧?这说得通吗?所以这就是根本瓶颈。

So the main challenge is turning sequential computation into parallel computation. When you start going from single neuron dynamics to multiple neural dynamics, the parameters of your system, instead of being scalars or vectors, become matrices and tensors. So now you're talking about matrix multiplication. You want to turn the scalar computations into tensor computations, right? You want to be able to parallelize sequential computations. The problem is that not all the time, if you have nonlinear relationships in an equation, you cannot trivially create a one-to-one map between a vectorized computation and a tensorized computation. So those nonlinear attributes of liquid neural networks and any other nonlinear recurrent neural network—I mean recurrence itself adds some degrees of complexity into the math of the whole equation. Let alone if the recurrence has a nonlinearity on it, it becomes a lot more complex to disentangle vectors or write them with typical linear algebra methods into tensors. And you have to be able to turn these computations into tensors to be able to compute in parallel, you know, like you compute them once at a time. You see? So that's the mathematical challenge. It's not related only to liquid networks. Any mathematical operation that you want to tensorize, if it is nonlinear, you're going to have the troubles we're talking about. That's why the state space models that came out are all about linear dynamics. We're talking about linear dynamical systems. And why those linear dynamical systems are important is just the pure fact that we do not have proper ways to scale nonlinear systems. You can apply nonlinearity to a linear dynamical system. For example, in a gated system, you can add a sigmoidal function after you perform the dynamical system in a linear way. You compute all the matrices and everything in parallel, but then you can do a pointwise application of a nonlinear operation on top of the entire tensor. That's what you can do. But what if the relationship between the parameters themselves is governed by some nonlinearity? Then you cannot really use typical linear algebra. You always have to approximate that nonlinear system into a linear system, then you would be able to parallelize the systems, right? Does that make sense? So that's the fundamental bottleneck.

11. 闭式解与可扩展性限制 Closed-Form Solution and Scalability Limits

Host

那么闭式解在多大程度上解决了这个问题,并且利用现有的计算资源,原始的液态网络范式到目前为止能够扩展到什么程度?

So how much does the closed form solution address that, and how far, with available computing resources, has the original liquid network paradigm been able to scale so far to present?

Ramin Hasani

好问题。所以,液态神经网络的特性之一是它们具有多重反馈机制。它们不仅仅是单一形式的反馈。它们有三层反馈,比如在两个细胞之间,在突触动力学本身中,因为我们想模仿大脑中的方式,对吧?所以它具有多级反馈。它具有相互嵌套的非线性。这种嵌套非线性本身也会增加复杂度。即使是液态神经网络的闭式解版本,从动力学角度来看,你现在可以拥有不是一百个神经元,而是十万个神经元,也许一百万到一千万个神经元。但非线性意味着你仍然必须顺序计算模型,因为电路中存在这些嵌套的非线性关系。所以即使是闭式解在可扩展性上也会受到限制,因为你不能非常有效地并行化它们。但也有很多人正在研究这些主题,比如并行化,不仅仅是并行化,还有加速顺序计算。

Fantastic question. So, one of the properties of liquid neural networks was the fact that they have multiple feedback mechanisms. They're not just one format of feedback. They have three layers of feedback, like between two cells, in the synaptic dynamic itself, because we wanted to mimic how it is done in the brains, right? So it has multiple degrees of feedback. It has nested nonlinearities on top of each other. That nested nonlinearity itself would also add degrees of complexity. Even the closed form solution version of a liquid neural network, from a dynamics point of view, you can now instead of hundred neurons, you can have hundred thousands of neurons, maybe 1 million to 10 million neurons. But the nonlinearity means you still have to compute the models sequentially because of these nested nonlinear relationships that happen in the circuit. So even the closed form solution would be limited in scalability in the sense that you cannot parallelize them very effectively. But there are also people working on these topics a lot, like parallelizing, not just parallelizing but speeding up sequential computations.

12. 顺序扫描与复杂度降低 Sequential Scan and Complexity Reduction

Ramin Hasani

所以我们讨论的是顺序扫描。你知道,扫描是一种可以叠加在矩阵乘法等操作上的算子。不是用立方或平方时间,而是用次平方时间来完成,对吧?所以实际上你可以降低计算复杂度,即使是对非线性算子,也不需要并行化或线性化它们。目前有持续的研究,比如我们自己也在做这类研究,来真正找到顺序运行这些非线性系统的方法,因为很明显,在做函数逼近时,拥有这些形式的非线性是有优势的。

So we're talking about sequential scan. You know, scan is an operator that you can run on top of, let's say, doing matrix multiplication. Instead of doing it in cubic or quadratic time, you can do it in sub-quadratic time, right? So you can actually reduce the complexity of computations even on nonlinear operators without parallelizing them, without linearizing them. There is ongoing research, like we are doing some of that research ourselves, to really find ways to sequentially run those nonlinear systems, because clearly there is an advantage to have those forms of nonlinearities when you do function approximation.

13. 缩放定律与架构偏差 Scaling Laws and Architecture Bias

Ramin Hasani

我们也在讨论缩放定律,这在这里变得非常重要。我相信缩放定律定义了架构。它是如何做到的?当我们说 Transformer 是革命性的东西时,我们谈论的是最大规模。Transformer 架构和注意力机制之所以如此出色,是因为它是无结构的。没有结构。当我在谈论液态神经网络中的嵌套非线性和所有那些东西时,Transformer 中没有这些,对吧?基本上只有矩阵乘法作为核心功能表示,而且它是无结构的。你实际上可以将任何矩阵乘以任何大小的任何矩阵。所以整个想法是这样的:你制造的神经网络越大,你就越希望它们变得不那么结构化。我们已经看到了 Transformer 在万亿参数规模上的成功。现在我们谈论的是数十万亿参数。即将到来的下一代模型,我们谈论的是万亿级别,而你可以用 Transformer 架构做到这一点。一旦你开始在那种架构中在规模上添加一点点偏置,事情就会变得完全混乱。

We're also talking about scaling laws, which become very important here. I believe scaling laws define architecture. How does it do that? When we are talking about transformers being this revolutionary thing, we're talking about maximum scale. The reason why transformer architecture and attention mechanism is such a brilliant architecture is the fact that it is unstructured. There's no structure. When I'm talking about nested nonlinearities and all this stuff that we have in liquid neural networks, you don't have that in transformers, right? You have basically just matrix multiplication as the core functional thing that represents, and it is unstructured. You can literally multiply any matrix into any matrix of any size. So the whole idea here is this: the larger neural networks that you make, the more you want them to become less and less structured. And we've seen the success of transformers at the trillions of parameters. Now we are talking about tens of trillions of parameters. The next generation of models that are going to come, we're talking about trillions, and you can do that with a transformer architecture. As soon as you start adding a little bit of bias in that architecture at scale, things become completely messed up.

Ramin Hasani

所以我们谈论的是液态神经网络,这些替代架构。它们有一个规模上限。存在一个参数区间,你可以做得更好,比如高达 1000 亿参数,高达 1 万亿参数。这是我们目前运作的范围。在这个范围内,模型架构越小,你就越想让它们专门化以解决特定应用。而且它们在数学上偏向于比其他类型的架构更好地解决某些类型的任务。所以我会说算法上的偏置:你加入的偏置越多——所谓偏置是指添加更多的非线性、在神经网络上添加多个门控层、添加循环、多种不同类型的循环、添加卷积。卷积本身,如果你只保留基于卷积的神经网络,它们也相当无结构,因为你可以将卷积应用于任何大小的神经网络。所以问题变成了:你想要多少偏置,以及解决什么样的问题?所以神经架构中的偏置成为,我认为,神经网络规模以及你最想解决的用例的函数。谱系是:你越往大规模走,你使用的数学算子就越无结构——纯矩阵乘法、纯卷积、纯顺序扫描、纯算子。忘掉门控,忘掉添加像遗忘门这样的花哨控制。在规模上,这是我们见过的,因为我们已经扩展了神经网络,并且对不同规模下的架构变化有很好的理解。

So we are talking about liquid neural networks, these alternative architectures. They have a scale ceiling. There is a regime of parameters where you can just do better, let's say up to 100 billion parameters, up to a trillion parameters. That's the range we are operating in right now. In this range, the smaller the model architecture, the more you want to specialize them for a certain application to solve. And they are actually mathematically biased to solve a certain type of tasks better than other types of architectures. So I would say biases on algorithms: the more biases you put on—and by biases I mean adding a lot more nonlinearity, adding multiple gating levels on top of a neural network, adding recurrence, multiple different types of recurrence, adding convolutions. Convolution itself, if you just keep convolution-based neural networks, they're also pretty unstructured because you can apply convolutions on any size of neural networks. So that question becomes: how much bias and what kind of problems you want to solve? So the bias in neural architectures becomes a function of, I would say, the scale of the neural networks as well as the use cases you want to solve the most. The spectrum would be: the more you go to larger size, the more unstructured mathematical operators you use—pure matrix multiplication, pure convolutions, pure sequential scan, pure operators. Forget about gates, forget about adding fancy control like forget gates and all those things. At scale, this is something we have seen because we have scaled neural networks and we have a really good understanding of architectural variation at different scales.

14. 液态神经网络的硬件需求 Hardware Requirements for Liquid Neural Networks

Host

只是为了校准一下,当你谈到已经达到数十万甚至接近一百万个神经元的液态神经网络时,它运行在什么样的硬件上?是像 CPU 超级计算机那样的吗?

Just for calibration, when you talk about the liquid neural networks that have gone up to hundreds of thousands or pushing a million neurons, what kind of hardware does that run on? Is that like a CPU supercomputer?

Ramin Hasani

你可以在 CPU 上做到,甚至是在简单的 CPU 或简单的 GPU 上。因为它们没有那么大——可能其中一个液态神经网络只有,我不知道,1 到 25 兆字节的文件大小。基本上,你可以直接把它们放进去,然后一个 CPU 就足够了,一个树莓派就足以执行计算。我们已经在许多预测性的专业 AI 应用中证明,这些系统可以非常非常强大,并且能做事。它们可以像 Transformer 变体这样的系统一样强大,像语音合成中的开放微分方程版本的神经网络一样强大。这是其中一个应用。正如我所说,预测序列:想象一下,有复杂格式的序列输入,多变量序列输入,你想对它们进行某种预测。这些神经网络实际上相当不错。

You can do that on a CPU, even on simple CPUs or simple GPUs. Because they don't have that much—probably one of these liquid neural networks would fit on, I don't know, one to 25 megabytes of file size. Basically, you can literally put them there, and then a CPU would be enough, a Raspberry Pi would be enough to perform computations. We have shown that on a lot of predictive specialized applications of AI that these systems can be very, very powerful and do stuff. And they can be as powerful as systems like these transformer variants, as powerful as the open differential equation version of the neural networks in speech synthesis. That's one of the applications. As I said, predictive sequence: imagine you have complex formats of sequences coming in, multivariate sequences coming in, and you want to perform some sort of prediction on top of them. These neural networks are actually pretty good.

15. 分布外泛化与持续学习 Out-of-Distribution Generalization and Continual Learning

Ramin Hasani

然后如果你谈论更多的分布外泛化,在有限的时间范围内,你不能进行开放式的持续学习。这些液态节点不是持续学习系统。它们更像是自适应的计算形式,因为它们的数学公式中有许多门控、许多反馈和许多依赖于输入的参数。

And then if you're talking about more out-of-distribution generalization, to the extent that you're within a bounded period, you cannot have open-ended continual learning. These liquid nodes are not continual learning systems. They are more adaptive formats of computations because of the many gating, many feedbacks, and many input-dependent parameters that they have in their mathematical formulation.

Host

你在这里使用的持续学习的定义是什么,它们不满足这个定义?

What is the definition of continual learning that you're using there that they don't satisfy?

Ramin Hasani

所以持续学习将是一个不断接收新数据并自我调整的系统。所以它也会改变系统的参数。液态神经网络:系统的动力学是依赖于输入的。这意味着新数据进来,系统的动力学会对这些输入做出反应,但系统的参数,像任何其他神经网络一样,是固定的。只是因为液态神经网络中每个神经元或每个节点的形式就像一个微分方程。你会有一个不同的动力学——动力学是你添加到参数数量中的另一个维度。它允许你压缩更多信息,压缩更多知识。这就是为什么在较小的模型实例中,没有免费的午餐,或者这里没有魔法。就像我们也在用液态神经网络做的那样。

So continual learning would be a system that is continuously receiving new data and retunes itself. So it also changes the parameters of the system. Liquid neural networks: the dynamics of the systems are input-dependent. That means new data comes in, the dynamics of the system would behave to those inputs, but the parameters of the systems, like any other neural networks, are fixed. It's just because the format of every neuron or every node in a liquid neural network was like a differential equation. You would have a different dynamics when you're—dynamics is another dimension that you add to the number of parameters. It allows you to compress more information, compress more knowledge. That's why with the smaller instance of the models, there's no free lunch, or there is no magic here. Like we are doing also with liquid neural networks.

16. 液态神经网络与适应性 Liquid Neural Networks and Adaptability

Ramin Hasani

关键在于,动态轴是我们为神经网络添加的东西,目的是什么?是为了让它们更具适应性。我给你一个非常具体的例子。想象一下你正在开车,突然下起雨来。根据雨的形式,你的自动驾驶系统——实际上控制着车辆——能够应对这类驾驶场景。如果你有一个传统的神经网络,它没有见过那种环境,它可能会产生偏差,因为雨水打在摄像头上是噪声,一种特定类型的噪声,或者说输入的一种特定适应。但现实并没有改变。整个环境的混杂变量也没有改变。对吧?这就是为什么液态神经网络会完全不同地处理输入。它们吸收输入,对输入应用低通滤波,对这类场景的适应性要强得多。这并不意味着它们改变系统的参数来变得更具适应性,因为这里有两个轴。一个是改变系统的参数,另一个是改变系统的动态。我认为持续学习属于持续改变系统参数的情况。在液态网络中,我们不这样做。

It's just that the axis of dynamics has been something that we added to the neural networks for what? For being more adaptable. So I'll give you a very tangible example here. Imagine you're driving and all of a sudden it starts raining. So depending on the format of the rain, your autonomous driving system that is actually taking control of the car could react to those kind of driving scenarios. And if you have a traditional neural network that hasn't seen that kind of environment, it might actually get biased because that rain that hits the camera is noise, a certain type of noise, or a certain type of adaptation of the input that is coming. But the reality hasn't changed. The confounding variables of the whole environment haven't changed. Right? So that's why liquid neural networks, because they react to the input completely differently. They absorb the input, they apply low-pass filtering on top of the input, they are much more adaptable to those kind of scenarios. That doesn't mean that they change the parameters of the system to be more adaptable because there are two axes here. One is changing the parameters of the system, the other is changing the dynamics of the system. I would consider continual learning to be attributed to where you continuously also change the parameters of the system. In liquid networks, we don't do that.

17. 架构搜索与客户方法 Architecture Search and Customer Approach

Host

明白了。好的,很有帮助。那么,让我们快进到现在。你们已经在市场上与客户合作了。我听到一次值得注意的对话,是你们的一位客户——Shopify 的 CTO 在 Latent Space 播客上,他对你和你们的技术评价很高。让我印象深刻的是,你们在某种意义上变得中立了。你们的方法中让我眼前一亮的是,你们开发了一个架构搜索过程,对客户的承诺不是‘嘿,我们开发了这个范式,就像微积分老师常说的,当你只有一把锤子时,所有东西看起来都像钉子。’所以你们明确向人们承诺,尽管我们来自这种特定网络的谱系,但我们不会盲目地将其应用于你的问题。相反,我们创建了这种更高层次的抽象或更元的过程,用于搜索架构空间,找到最适合你的东西。其中两个值得注意的细节是:第一,你们发现代理指标效果不太好。仅仅测量困惑度之类的指标,你们发现实际上需要更进一步,在模型将要执行的实际下游任务上测试它们。第二,硬件在环,将实际目标硬件纳入循环,在机器人、传感器、手机等运行的真实物理约束下测试架构。也许我理解 LFM 模型就是由此产生的,但在我们讨论 LFM 之前,你能大致描述一下我们放入这个架构搜索的不同类型的问题吗?然后也许还涉及我们针对什么样的硬件,以及在不同约束下,哪些不同的架构在不同问题上胜出。

Gotcha. Okay. Helpful. Okay. So, let's fast forward to the present. You guys are now in the market working with customers. One notable conversation I heard with one of your customers was with the CTO of Shopify on the Latent Space podcast who had some very good things to say about you and your technology. I was struck by the fact that you've gone neutral in a sense. What sort of jumped out at me in terms of your approach is that you've developed an architecture search process where the promise to customers is not that 'hey we developed this one paradigm and old calculus teacher used to say when all you have is a hammer everything looks like a nail.' So you're explicitly promising to people that even though we came from this lineage of this particular kind of network we're not just going to blindly apply it to your problem. Instead, we've created this sort of higher abstraction or more meta process for searching through architecture space to find the thing that's going to work best for you. And the two notable details of that are one: proxy metrics you found don't work that well. Just measuring perplexity or whatever, you found you needed to actually go farther and test models on the actual downstream tasks that they're going to be asked to perform. And second: hardware in the loop, actual target hardware in the loop to test the architecture subject to the very real and physical constraints of the robot, the sensor, the phone, whatever it's going to be running on. Maybe I understand that the LFM model came out of that, but maybe before we even get to the LFM, could you sketch out the range of different types of problems that we're putting into this architecture search? And then maybe also subject that implies of course what kind of hardware are we targeting and then what kind of different architectures are winning for different kinds of problems under different kinds of constraints.

Ramin Hasani

当然。我们内部开发了一个系统,称为自动化基础模型设计(AFMD)。这是一个元学习系统,将硬件纳入循环,并通过进化策略尝试许多不同的算子。标准是进化策略优化几个方面:优化设备上的内存消耗、优化延迟、优化速度,同时不牺牲质量。当我们谈论质量时,困惑度不是衡量标准。实际上,我们关心的是应用,即下游应用。也不仅仅是公共基准。我们谈论的是 100 个不同的基准。所以从元学习的角度来看,问题空间变得非常复杂。现在我告诉你为什么我们采用这种方法来设计架构。为什么?因为我们希望在构建架构的早期就消除所有人类偏见。我们在这样的公司文化中意识到的一件事是,我可以告诉你:即使在美国目前最大的基础模型实验室,如 Anthropic 和 OpenAI,也有一群人,我称他们为架构的复仇者或后训练或预训练的复仇者。这些群体中,通常只有一小部分人在做决策,比如‘哦,你知道,你要调整这个架构的这部分,让它表现更好。’为什么?因为根据我的个人经验,这样开始效果更好。如果你真正思考,这是所有基础模型实验室都存在的问题,你不能说有人对此有解决方案。但现在递归自我改进的过程实际上正在解决这个问题,因为人们终于意识到必须把决策权交给算法。你必须接受苦涩的教训。你有苦涩教训的人。所以你必须用一种系统的方法来找出解决你问题的真正架构。你可以构建一个通用计算机。我以神经网络缩放定律的形式与你分享的见解,正是来自我们对架构空间的大规模探索。事实是,在较小类别的模型中,架构上的某些偏见有帮助,但在较大的模型实例中,你不需要对系统施加偏见,你可以直接使用纯卷积、纯 Transformer,以及非常简化的纯循环结构。你不需要添加任何形式的专门处理,比如门控、门控 Delta 网络,还有 Mamba、Jamba 以及许多其他变体。但你想做的是完全无偏见。即使是我们自己,在第一天创办 Liquid 时,我们实际上有 SSM 的发明者之一 Jimmy Smith 作为我们的创始科学家。他发明了 S5,基本上将并行扫描思想引入了 SSM 领域。然后我们还有 Stefano Massaroli 和 Max Poli。

Absolutely. So there's a system that we developed in house. We call it Automated Foundation Model Design, AFMD. It's a meta-learning system that puts hardware in the loop and tries out many different operators with an evolution strategy. The criteria is an evolution strategy optimizing for a couple of things: optimizing for memory consumption on that device, optimizing for latency, optimizing for speed, while no sacrifice on quality. When we talk about quality, perplexity is not the measure. It's actually the application, the downstream applications that we care about. It's not just also public benchmarks. We're talking about 100 different benchmarks. So the problem space from a meta-learning perspective becomes a very complex kind of problem. Now I'll tell you why we took this approach to design an architecture. Why? Because we wanted to remove all the human biases early on as we are building architectures. One of the things that we realize culturally at companies like this is what I can tell you: even at the largest foundation model labs in the US right now, Anthropic and OpenAI, there are a bunch of people, I call them the avengers of the architectures or avengers of post-training or pre-training. These groups of people, there's usually a very small set of people that are calling the shots on like 'oh you know what, you're going to tweak this portion of this architecture so that it performs better.' Why? Because in my personal experiences, it has started working better. If you're really truly think, and this is something that is broken in all the foundation model labs, you cannot say that somebody has a fix to this. But now the recursive self-improvement kind of process is actually fixing for that because now people are just finally realizing you got to give it to the algorithms. You have to be bitter lessons. You have bitter lesson people. So you got to give it to a systematic way to actually find out what is the true architecture for the problems that you want to solve. You can build like a general purpose computer. The insights that I shared with you in the format of the scaling laws of neural networks is coming out of our massive exploration of the space of architectures. The fact that in the smaller category of models, some biases on the architecture help, but in the larger instances of the models you don't need to bias the systems, you can actually go pure convolutions, pure transformers, and you can go pure recurrences that are very simplified. You don't need to add any form of specialized kind of treatment like gating, gated delta nets, and then there's the Mambas and the Jambas and a lot of different variations of these architectures that are coming out. What you want to do though, you want to be completely unbiased. Even us ourselves, day one, when we actually started Liquid, we had one of the inventors of SSM actually as our founding scientist, Jimmy Smith. He invented S5, he basically brought the parallel scan idea to the SSM world. Then we had Stefano Massaroli and Max Poli.

18. 算子的统一理论 Unified Theory of Operators

Ramin Hasani

他们实际上在设计 hyena 层级结构,其中有一个卷积路径在运行,所有这些都属于线性系统范畴。然后我们提出了基于非线性超控制理论的基础模型,也就是液体神经网络。房间里有很多大人物,你必须做出决策。所有实验室都面临同样的问题:你必须弄清楚。我们想,好吧,让我们从根本上解决这个问题,系统性地、从第一性原理出发,把所有感兴趣的算子——任何我们认为可能产生通用计算机的算子,比如一个系统,液体神经网络本身就是一个通用计算机,理论上如果你能做到,你实际上可以扩展它。所以你把所有那些方程、所有那些东西都放到一个统一的理论中,使用线性输入变化算子。算子空间实际上可以封装在这个线性输入变化或更一般的液体输入变化算子中,因为输入依赖性是我们过去 10 年一直在讨论的。这极其重要,我们现在看到在 Transformer 中,输入依赖性也非常重要,它自然地来自注意力架构,但方式非常非结构化。它不像在 RNN、液体神经网络、SSM 中那样结构化。在所有这些中,你不会在 Transformer 中看到这些门控变量非常结构化。现在我们实际上有各种卷积算子的变体、各种循环算子的变体,以及注意力本身——有分组查询注意力、原始 Transformer,很多不同的东西。然后在这些动力系统的所有变体之上,还有各种类型的卷积,我们引入了液体计算块,例如门控、双门控卷积或门控,对不同的动力系统或不同的算子施加某种类型的偏置。我们添加了它们,空间变成了大约 50 到 100 个不同的算子。然后你想要构建混合模型,以便减少——基本上,目标是什么?目标是在不损失准确性的情况下最大化计算效率。这就是我们设定的目标:搜索空间的目标函数也是这个。让我们运行,让我们把所有算力都投入到这个问题上,然后看看系统将如何设计它。我们开始对此进行缩放定律研究:我们从 1000 万参数的神经网络一直扩展到 720 亿参数模型的缩放定律,这些模型采用混合结构。所以我们在早期做了大量工作,Liquid 的早期阶段,比如 2023 年和 2024 年,一直在某种处理器上验证哪种架构最有效。

They were actually designing the hyena hierarchy, with a convolutional path going on, all of these are linear systems within the realm of linear systems. Then we came out with nonlinear hyper control theory enabled foundation models, which is liquid neural networks. A lot of big heads in the room, you gotta call the shots. Again the same problem in all the labs: you gotta figure out. We thought, okay, let's fundamentally solve this problem, let's systematically, first principle, put all the operators of interest—whatever we think could be an operator that can give rise to a general-purpose computer, like a system that can, say, a liquid neural network on its own is a general-purpose computer, you can actually scale it if you can theoretically get there. So you put all those kinds of equations, all those kinds of things in a unified theory, with linear input-varying operators. The space of operators can actually be encapsulated inside this linear input-varying, or in general liquid input-varying operators, because input dependence is something we have been talking about for the last 10 years. That's extremely important, and we see that right now in transformers it's also extremely important to have that input dependence, coming naturally with the attention architecture as well, but in a very unstructured way. It's not as structured as in RNNs, in liquid neural networks, in SSMs. In all of those things, you don't see these gated variables to be very structured in transformers. Now we actually have various variants of convolution operators, various variants of recurrent operators, and then the attention themselves—there's group query attention, the original transformer, many different things. Then on top of all these variations of dynamical systems, and also various types of convolution, we brought the liquid computational blocks, for example the gated, the double gated convolution or gating, a certain type of biases onto different dynamical systems or different operators. We added them, and the space becomes something around 50 to 100 different operators. Then you want to build hybrid models so that you can reduce—basically, what's the goal here? The goal is maximizing efficiency of computation without loss of accuracy. That's the goal we set: the objective function of the search space is also this. Let's run, let's have all our compute thrown at this problem, and let's try to see how the system is going to design this. We started doing scaling laws on this: we went as small as a neural network of 10 million parameters to running the scaling laws up to 72 billion parameter models in these hybrid structures. So we have done a massive early on, early days of Liquid, like 2023 and 2024, has been always proving out on a certain type of processor what's the most efficient type of architecture that can come out.

19. 消除人类偏见:双门控卷积 Removing Human Bias: The Double Gated Convolution

Ramin Hasani

结果发现,当我们把所有偏见都抛开——所有我们围绕算子添加的门控机制——它们被简化了。在 Mamba 风格的架构中,你有一堆门控机制;在门控 Delta 网络中,你有一堆算子;在线性注意力中,有门控变体。每个人都在手动调整网络中的某个参数,因为他们在自己的实验中观察到了某些东西。事实证明,如果你想得到最有效的架构形式,所有这些都必须去掉,最终得到的是双门控卷积,它实际上来自我们设计的这个大规模搜索空间,比如 AFMD,这个原始系统。这是出现的候选架构之一,结果证明它在被称为 CPU 的通用计算机上表现非常好,因为 CPU 没有特殊结构。在 CPU 上,在来自 AMD、高通、英特尔的所有 CPU 上,每个制造处理器的公司,ARM Pro,所有 ARM 处理器,我们试图测试不同的算子:什么样的通用结构和计算图能给你最简化的、没有任何手动调整特征的东西?它完全来自我们做的系统测试。这成为了事实上的架构,我们实际发布的 LFM2 结构。但在探索这些的过程中,我们观察到许多不同的候选方案也出现了。然后我们试图弄清楚,例如,对于 AMD、高通或英特尔驱动的 AIPC 内部的 NPU(神经处理单元),这些变体中哪一个会是更好的神经架构,能给那些硬件提供商和硅片厂商带来效率提升,或者计算速度、延迟、内存占用方面的提升,同时拥有不牺牲任何质量的计算图。所以这就是全部:用系统方法去除所有人类偏见,针对两种架构,甚至我们自己的偏见。在 CPU 计算上,Liquid 2 中唯一保留下来的就是这种双门控卷积。正如我告诉你的,我们最初在液体神经网络中有嵌套计算;这种嵌套计算格式变得非常有趣,成为我们构建的这些极其简化的神经网络的一部分。事实上,我们选择非结构化 1D 卷积作为层:我们网络的 70% 到 80% 由这些门控卷积(我们拥有的双门控卷积)构成。它们极其简化,并取代了注意力机制。它们大大降低了计算复杂度,大大减少了内存占用,大大提高了计算速度,并且在规模上也是如此。

It turned out that when we put all of our biases away—all the gating stuff that we are putting around operators—they got simplified. In Mamba style architectures, you have a bunch of gating mechanisms; in gated delta nets, you have a bunch of operators in there; in linear attention, there are gated variants. Everybody is tweaking a certain parameter in the network by hand because in their own experiments they observed something. Turns out all of this has to go away if you want to get to the most efficient form of architecture, and it became the double gated convolution that actually came out of this massive search space like AFMD, this original kind of system that we designed. This was one of the candidate architectures that came out, which turns out to be very good on general-purpose computers called CPUs, because CPUs don't have a special kind of structure. On CPU, on all the CPUs coming out of AMD, Qualcomm, Intel, everyone that is building processors, ARM pro, all the ARM processors, we tried to get to the place where we tested different operators: what's that generic kind of structure and computational graph that gives you the most simplified, no added hand-tuned features anywhere? It's just literally coming out of the systematic tests that we've done. This became the de facto architecture, the LFM2 structure that we actually announced. But while exploring these things, we observed many different candidates also popping up. Then we tried to figure out, for example, for an NPU (neural processing unit) that is inside an AIPC powered by AMD or Qualcomm or Intel, which of these variants would be a better neural architecture that gives those hardware providers and silicon portals a boost on the amount of efficiencies that they are unlocking, or the speed of computation, latency, memory footprint, while having a computational graph that doesn't sacrifice any format of quality. So this was the whole thing: removing all the human bias with a systematic approach on two architectures, even our own biases. The only thing that actually remained in Liquid 2 on CPU computation is this double gated convolution. As I told you, we have nested computation in liquid neural networks originally; this nested format of computation became very interesting to be part of these very simplified neural networks that we built. The fact that we have unstructured 1D convolutions as the layers of choice: 70 to 80% of our networks are structured by these gated convolutions, double gated convolutions that we have. They are extremely simplified and they replace attention. They reduce the computation complexity by a lot. They reduce the memory footprint by a lot. They maximize the speed of computation by a lot, and at scale as well.

20. 门控机制与输入依赖 Gating mechanisms and input dependence

Host

同时,在质量方面,我的意思是,你已经看到其中一些模型实际上与基于 Transformer 的替代方案极具竞争力。我想你还问了一堆其他后续问题,但我先在这里暂停一下,看看有没有后续问题。好,我们稍后进入用例。我对此很感兴趣,也想谈谈硬件的未来以及你的工作对硬件未来的影响。但我认为这对很多人,包括我自己,可能会有所帮助。尽管我已经对此进行了一定深度的研究,但我还是想更好地理解它。关于门控这个概念——我先给出我粗略的理解,然后你来改进和深化。我特别注意到,当我深入研究 Mamba 并对此感到非常兴奋时,似乎这些门控机制反复使用的技巧是:我们希望应用于数据的变换是依赖于输入的。所以仅仅学习一个变换是不够的。我们想要学习一个变换,但同时还要有一些——通常门控是相对低维的。有时它可能只是一个应用于该变换的标量。当然它可能更复杂,但这是一个相对简单的机制,它告诉我们:对于这个学到的变换,我们如何根据当前考虑的输入来修改它。这似乎很棒,似乎解锁了巨大的潜力。请帮我理解更多——你认为我遗漏了什么,或者帮我加深直觉,为什么这在所有这些不同的架构中是一个如此强大且反复出现的主题。

And at the same time on the quality, I mean you have seen like some of these models are actually extremely competitive to the transformer-based alternatives. I think you asked a bunch of other follow-up questions as well, but I would pause here for any follow-ups. Yeah, let's get into use cases in a minute. I'm definitely interested in that and I also want to talk about the future of hardware and what your work implies for the future of hardware. But I think it would probably be helpful for a lot of people, including myself. Although this is something I've studied in some depth, I still would like to grok it better than I do. This concept of gating — I'll give you my rough and ready understanding and then you can improve it and deepen it. I especially noticed this with Mamba when I went down that rabbit hole and became very excited about it. It seems that the kind of trick that is played over and over again with these gating mechanisms is we want the transformation that is done on the data to be input dependent. So it's not enough to learn a transformation. We want to learn a transformation but then also have some — relatively typically the gate is relatively low dimensional. Sometimes it could just be like a scaler that's applied to that transformation. And it could be obviously more complicated than that, but it's some relatively simple mechanism that says for this learned transformation, here's how we're going to modify it given the input currently under consideration. And that seems to be great. It seems to unlock a tremendous amount. Help me understand more — anything you think I'm missing there or help me deepen my intuition for why that is such a powerful and recurring theme in all these different architectures.

Ramin Hasani

那是一种液体结构,你知道。这就是我说的那种依赖于输入的东西实际上进入了 Mamba。所以在 Mamba 真正发布之前——不是之前,而是 Mamba 发布前一年半——我们发表了一篇名为 Liquid S4 的论文。因为如果你只读那篇论文的摘要,我们实际上是第一次引入这种依赖于输入的 SSM 的想法。基本上是依赖于输入的 SSM。那里的想法是,让我们把我们发现与表示学习、学习这种依赖于输入元素的能力有很大关系的基本构建块,真正把那种门控形式带到神经架构中,带到 SSM 中,然后也带到卷积中,以及我们今天正在设计的其他系统——液体基础模型也是如此。那种门控形式实际上增加了很多。你应该这样想:自然而言,这很有道理。如果神经网络不是关于前向传播是自适应的——因为一旦你有了输入依赖性,从学习理论的角度来看,所有的魔法都发生在反向传播中。当你反向计算梯度时,那个依赖于输入的算子本身会表示自己。所以当你从看到的数据中学习时,你也在学习某种动力学,依赖于输入的动力学。所以我之前谈到的第二个轴——不仅仅是神经网络的参数数量,还有神经网络实际学习的动力学。你在反向传播中学习到了那种动力学的某种表示。现在,那个门控的复杂性将带来巨大的变化。而且事实证明,那个门控在语言建模中也极其重要,尤其是当你拥有像 RNN 这样的序列模型时,比如经典控制中的连续时间 RNN,离散化的 RNN 如 LSTM,然后你还有液体神经网络——同样,这些线性 RNN 的非线性版本就是 SSM。当你添加这种类型的门控时,你也可以提高语言能力。似乎对于离散序列,这种门控也很有帮助。它不仅适用于连续时间序列,也适用于离散序列,因为我认为语言是一个离散的信息序列。所以你会拥有那种由这种门控添加的序列动力学。所以思考这种特征是一个非常正确的思维方式。对我们来说,我们认为这种输入依赖性是我们所做的基本发现之一,我们在生物学中也了解到它会发生,它来自物理学。如果你只是把数学放在一起,你会意识到:哦,这个动力系统的特别之处在于它具有这种非线性输入依赖性。这是一种结构。至于偏见,我的意思是门控的复杂性,甚至门控的存在,都是你添加到系统中的一种偏见。现在,在大规模下是否需要它?这是一个大问号。我们将看到当神经网络有 100 万亿个参数时,矩阵乘法是否足以真正像通用计算机一样运行,达到世界上的 AGI?这是一个大问号。但同样,我们需要把关于架构的讨论和对架构的痴迷放在一个视角中。另一个视角:我们必须把这些东西放在学习理论本身中。现在我们正在谈论不同的学习方案。例如,你可以用下一个词预测来训练模型,你可以在词建模的上下文中训练它们,你可以在序列化的长期视野和长期历史中训练它们。有很多方法可以构建学习算法本身的目标函数。这也会对学习过程有很大贡献。架构——我会告诉你,我们发现架构的目的和架构研究最重要的应用在于效率。真正实现高效的计算形式而不损失质量。这是一件极其重要的事情,因为随着模型变得越来越大,AI 的需求呈指数级增长,你正在谈论资源分配问题。

That's a liquid structure, you know. That's when I say this input-dependent kind of thing is actually coming into Mamba. So right before Mamba actually got out — not right before, but a year and a half before Mamba came out — we released a paper called Liquid S4. Because if you just read the abstract of that paper, we are actually for the first time introducing this idea of input-dependent SSMs. Basically input-dependent SSM. The idea there was that let's bring the fundamental building block that we found has a lot to do with representation learning, with capacity of learning this input-dependent element, to really bring that format of gating to neural architectures, to SSMs, and then bringing that to convolutions as well with the other kind of systems that we are designing today — liquid foundation models as well. That kind of format of gating actually adds a lot. You should just think about it: naturally it makes a lot of sense. If the neural network is not about the forward pass being adaptive — because once you have input dependence, from a learning theory perspective, all the magic happens in the backward pass. When you're computing the gradients backwards, that input-dependent operator itself is going to represent itself. So when you're learning from that data that you're seeing, you're also learning some sort of dynamics, input-dependent dynamics there. So that second axis that I was talking about — it's not just the number of parameters of the neural network, it's also the dynamics that the neural network actually learns. You have some representation of that dynamics being learned in the backward pass. Now the complexity of that gate is going to make a huge change. And also it turns out that gate is extremely important in language modeling, especially when you have sequence models like RNNs, classical control like continuous-time RNNs, discretized RNNs like LSTMs, and then you have liquid neural networks — again, not very nonlinear versions of these linear versions of RNNs is an SSM. And when you add this type of gating, you can improve the language capabilities as well. It seems that for discrete sequences, this gating also helps a lot. It's not just for continuous-time sequences but also discrete sequences, because I would consider language to be a discrete sequence of information. So you would have that sequential dynamics that is kind of added by this gating. So it is a very right frame of mind to think about that kind of feature. And for us, we think that this input dependence is one of the fundamental discoveries that we've done, that we learned also in biology it happens, it comes out of physics. If you just put the math together, you're going to realize: oh, what is special about this dynamical system is the fact that it has this nonlinear input dependence. That's kind of a structure. And by biases, I mean the complexity of that gating, even the existence of that gating, is a bias that you're adding to your system. Now, is it needed or not at scale? That's a big question mark. We're going to see when neural networks are 100 trillion parameters, is matrix multiplication enough to really perform like a general-purpose computer, to get to AGIs of the world or not? That's a big question mark. But also, we need to put this discussion about architecture and the obsession about architecture in perspective. Another perspective: we have to put these things in the learning theories themselves. Right now we are talking about different learning schemes that are coming. For example, you can train a model with next-token prediction, you can train them in a word modeling kind of context, you can train them in a sequential long-term horizon and long-term history. There are so many ways you can construct the objective function of a learning algorithm itself. That would also contribute a lot to the learning process. Architecture — I'll tell you that we found the purpose of the architecture and the most important application of architectural research has been in efficiency. Really getting to efficient formats of computation without loss of quality. That's an extremely important thing because you're talking about the resource allocation problem right now as the models are becoming bigger and the demand of AI is exponentially going higher.

21. 下一代智能的效率与架构 Efficiency and Architecture for Next-Gen Intelligence

Ramin Hasani

我们需要更高效的版本来真正大规模运行系统,否则如何让所有人都能使用 AI?我们看到,大型实验室正在耗尽全球的算力,因为它们既要训练模型,又要为所有人托管这些模型。因此,效率成为我们研究的这些模型架构的一个基本属性。如果你想构建能够实现下一代神经网络——抱歉,是下一代智能系统——的架构,使其像人脑一样以 20 瓦的功率进行运算,并且已经是 AGI,那将是一场超越架构本身的巨大探索。它涉及记忆研究、学习算法、数据特征、从统计学习视角进行的先验研究,以及学习理论本身。学习理论的局限性也施加在当今的学习系统上,因为学习理论的定义在大规模下实际上已经失效了。现在,一切都是独立同分布的,我们在构建平均机器。你已经看到了 AI 系统输出的写作质量和序列,就是因为这个原因。在某种程度上,多智能体已经成为解决方案,测试时 Scaling 也成为解决仅使用自回归建模作为预训练方法的学习系统所存在问题的方案。但我认为,不仅需要在架构上创新,还需要在整体上创新:数据、算法、模型和学习算法共同设计未来,这是该领域的终极圣杯。

You want to have more efficient versions of the systems to actually run the systems at scale, otherwise how can we provide access to AI to all? We are seeing instances where the larger labs are wiping out the compute off the planet because they have to train and then host these models for all of us. So efficiency becomes a fundamental property of these model architectures that we are doing research on. If you want to get into architectures that enable the next generation of neural networks—sorry, next generation of intelligence systems—in a way that they can become like the human brain, which performs computation with 20 watts of power and is already an AGI, that is the massive exploration that goes beyond architecture. It goes into memory research, learning algorithms, data features, prior research from a statistical learning perspective, and learning theory itself. The limitations of learning theory also get imposed on today's learning systems because the definition of learning theory itself is actually broken at scale. Right now, everything is i.i.d., we are building averaging machines. You've seen the quality of writing and the sequence that comes out of AI systems because of that. To some extent, multi-agent has been becoming the solution, and test-time scaling became a solution for the caveats we are seeing from just learning systems with autoregressive modeling as a pre-training method. But I think innovations are needed not just on architecture but on the whole thing: data, algorithms, models, and learning algorithms all together can design the future, the ultimate holy grail in this space.

Host

是的,这非常有趣。所以你的观点是,随着我们进入递归自我改进时代,真正的进步将更多来自新的学习范式、新的目标,然后这些进步将通过架构优化变得高效。但架构是在范式问题之后才出现的,即我们到底在学习什么,以及我们如何给出信号。

Yeah, that's very interesting. So basically your perspective is that as we get into the recursive self-improvement era, the real advances will come more from new learning paradigms, new objectives, and then those advances will be made efficient through architectural optimization. But the architecture comes after the paradigm question of what exactly are we learning and how are we giving a signal.

Ramin Hasani

这只是其中的一个组成部分。从数据问题、数据表示来看,一个轴是架构,另一个轴是学习算法本身。然后我们进入递归自我改进,这实际上是新系统的持续学习特性——持续学习研究已经进行了大概四十多年,现在有了一个花哨的名字。这就是我对整个领域的描述。你不能孤立地看待架构,认为它是改变一切的根本因素。

It is one component of it. Looking at it from a data problem, data representation, from one axis it would be architecture, and the other axis would be the learning algorithms themselves. Then we get into that recursive self-improvement, which is kind of the continual learning characteristics of the new system—the fancy name for the continual learning research that has been happening for more than probably four decades now. That's how I would characterize the whole space. You cannot just look at architecture in isolation as the fundamental thing that changes everything.

Host

是的,你让我有点想起 Ali Beirus 的《架构的幻觉》。如果有时间,我们或许可以谈谈嵌套学习,但好吧,我们还是先专注于你的工作。那么在 LFM 中,这是架构搜索的结果,一个令人惊讶的简单结果是:注意力层的数量减少了,但仍然至关重要,而其他层——我们在 SSM 注意力混合中也看到过类似情况——但我认为这个搜索过程的一个惊人发现是,非注意力层实际上可以非常简单,只要它们有一些门控机制。所以你有门控,然后就是一个非常简单的卷积,我认为它只考虑很短的 token 跨度,对吧?就像只向后看?

Yeah, you remind me a little bit of Ali Beirus's "The Illusion of Architecture." If we have time, maybe we can touch on nested learning, but okay, let's stay focused on your work for the moment. So in LFM, this is the result of this architecture search, and the kind of surprisingly simple thing that comes back is some reduced but still critical number of attention layers, and then the other layers are—we've seen this kind of with SSM attention hybrids as well—but I think the kind of surprising revelation from the result of this search process is that the non-attention layers can actually be extremely simple as long as they have some gating. So you've got the gate and then you've got just a real simple convolution that I think only considers a very short span of tokens, right? Is it like just back?

Ramin Hasani

所以,如果我们要更新多年前的标题《注意力就是一切》,我们可能会说,注意力确实仍然需要,至少在一定规模下,但你也需要在其他层上使用门控,但你实际上不需要任何超级疯狂、花哨、复杂的机制,比如状态空间模型之类的东西。事实证明,保留门控才是真正带来最大价值的部分,然后你可以在门控后面放一个非常简单的机制,然后你可以有 70% 的这种机制和 30% 的注意力,当然受限于资源约束,这最终成为了成功的公式。我有没有说错什么?

And so if we were going to update the headline "Attention is All You Need" from however many years ago now, we would maybe say attention is something that you really do still need, at least at certain scale, but you also need gating on your other layers, but you don't actually need anything super crazy, fancy, sophisticated, the state space model, all that kind of stuff. Turns out that keeping the gate is actually the part that really drives the most value, and then you can have a really simple mechanism behind the gate, and then you can have 70% of that and 30% of attention, and subject to obviously resource constraints, that ended up being the winning formula. Am I getting anything wrong there?

Ramin Hasani

不,你说到了点子上。这基本上取决于你所在的领域和系统的目标。你是想构建超级智能吗?你是想构建最强大的 AI 系统版本吗?你需要最无偏的算法版本。注意力机制是一种极其丰富、无偏的算法形式。即使注意力的计算复杂度是 n 的平方,也许我们真的需要 n 的平方才能达到那个水平,甚至可能需要更复杂的架构。我们一直在尝试降低架构的复杂度,仅仅是因为我们整个人类都受到资源限制。我们现在就受到资源限制。所以我认为我们目前的发现只是表明,随着模型规模的扩大,存在一个可以遵循的架构梯度。你遵循的梯度是:对于较小的模型和专用模型,你可以引入尽可能多的偏置,比如你带来的这些门控机制,你可以在计算图中使用任意多的感兴趣算子,这都会有效,并且如果你真的在最大化线性时间复杂度——比如你想实现线性注意力系统,最快的那种——它会给你带来某种提升。如果速度如此重要,你甚至愿意牺牲一点质量,你可以引入线性,整个系统可以是线性的。你甚至不需要那些混合模型。

No, I mean you're touching on the right things. This is basically it's like the regime you're operating in, the goal of your system. Are you trying to build superintelligence? Are you trying to build the most powerful version of the AI system? You need the most unbiased version of an algorithm. Now attention is an extremely rich, unbiased format of algorithms. Even if the computation complexity of attention is n squared, maybe we really need n squared to really get to that kind of level, and maybe even we need more complex architectures. We've always tried to reduce the complexity of architectures for the sheer purpose of the fact that we are resource-constrained as humanity as a whole. We are resource-constrained right now. So I would say the discoveries that we have right now just show that there is a gradient on architecture that you can follow as you scale models. The gradient that you're following is the fact that for smaller kinds of models and specialized models, you can put as many biases like these gating mechanisms that you're bringing in, and you can play around with as many operators of interest in your computational graph, and it is going to work and it is going to give you some sort of a boost if you're really maximizing for linear time complexity—like you want to implement linear attention systems, just the fastest kind. If speed is so important and you're actually willing to sacrifice a little bit quality, you can bring in linearity and the whole system could be linear. You don't even need some of those hybrids.

22. 缩放与架构复杂性 Scaling and Architecture Complexity

Ramin Hasani

它只是提升了准确率,因为正如我们所见,O2 复杂度,也就是在特定规模和层级上的计算复杂度,是我们达到所需性能所必需的。网络越大,你可以让它越无结构。这大致就是从我们开始设计神经架构的那整套算法方法中学到的东西。

It just boosts that accuracy because as we see, the O2 complexity, basically the computational complexity at a certain level, at a certain scale, is needed for us to really get to those performances that you want to do. The larger the network becomes, the more unstructured you can make it. That's kind of the learning from that whole algorithmic approach that we started designing neural architectures.

Host

是的,非常有趣。好的,那让我们看看光谱的另一端,当你与客户合作时。有哪些有趣的例子,在资源受限、领域狭窄的情况下,其他类型的偏置实际上在架构搜索过程中胜出?

Yeah, very interesting. Okay, let's look at the other end of the spectrum then as you work with customers. What are some interesting examples of when given resource constraints, given the narrowness of the domain of interest, other kinds of bias are actually winning in the architecture search process.

23. 领域特定架构:生物学与长上下文 Domain-Specific Architectures: Biology and Long Context

Ramin Hasani

好问题。例如,如果你去生物学领域,想要建模序列数据,在生物学中我们讨论的是 DNA 数据。从词汇量的角度来看,DNA 数据非常有限,对吧?它们不像语言那样有最大词汇量。在 DNA 语言中,它非常简单,但你需要处理的序列长度,比如对于人类或细菌,我听说大约在 1 到 1000 亿个序列元素之间。对于这么长的上下文,当你没有大词汇量时,你不需要注意力机制。所以你可以在这类数据上直接使用纯卷积、纯 SSM、纯液态神经网络(如循环网络及其并行版本)以及线性注意力。对于极长的上下文,你无法用其他方式做到。原因是上下文变得如此之大,注意力机制的二次成本就显现了。所以在生物数据上,你会希望有一些结构。

Great question. So for example, if you go to biology, and you want to model sequential data, in biology we're talking about DNA data. DNA data from a vocabulary perspective is very limited, right? They're not like language with a maximum amount of vocabulary. In DNA language, it's very simplified, but the lengths of those sequences you have to process are, say for a human being or maybe for a bacteria, I heard it's somewhere between one to 100 billion sequence elements. For that long context, when you do not have a large vocabulary, you don't need attention. So you can actually run on this kind of data with pure convolutions, pure SSMs, pure liquid neural networks like recurrent networks and parallelized versions of these recurrences, and linear attention. For extremely long context, you cannot do it any other way. The reason is that context has become so large that the quadratic cost of attention just kicks in. So on biological data, you would want to have some sort of structure there.

Ramin Hasani

然后还有像视频建模这样的领域。在这些地方,你可能需要不同的架构甚至不同的学习算法。你已经看到了扩散模型的成功,例如。扩散也是一种先验形式,你将其施加在某个架构上,但关于扩散和学习算法之间,或者扩散实际上是架构的一部分,仍存在争论。你也可以建立这种联系和区分。所以在视频和场景理解上,可能需要扩散的元素。有些人仍然相信通过自回归建模也能达到目标,但我们会看看这是否成立。

Then there are places like video modeling. In those places, you might want to have different architectures and even different learning algorithms. You have seen the success of diffusion, for example. Diffusion is also a format of a prior that you put on a certain architecture, but still there is debate between diffusion and learning algorithms, or diffusion is actually part of the architecture. You can make that connection and distinction as well. So on video and scene understanding, probably elements of diffusion would be needed. Some people still believe that with autoregressive modeling you could get there anyway, but we will see if that stays true.

Ramin Hasani

当我们谈论音频信号时,如果只有音频而没有语言,只是纯粹的声音,比如语音到语音、噪声到信号、信号到信号,这些地方循环神经网络仍然非常非常强大。它们极其强大,尤其是在低数据 regime 下。我认为在数据不多的小规模场景中,具有更多偏置的模型能带来很大价值。循环神经网络和架构中的偏置可以通过它们拥有的反馈机制帮助你在低数据 regime 下填补空白。架构越复杂,越接近你要解决的数据集的动态特性,你构建的学习系统就越好。

When we are talking about audio signal as well, if you have audio alone and language is not part of it, just pure voices, voice to voice, noise to signal, signal to signal, these are places where recurrent neural networks are still very, very powerful. They are extremely powerful, especially in the low data regime. I would say models that have a lot more biases in them, in these smaller regimes when you do not have that much data, this is where you can bring a lot of value. Recurrent neural networks and biases in architecture can help you in the low data regime to fill out that gap with the feedback mechanisms they have. The more complex architectures, the closer the architecture is to the dynamics of the datasets you are trying to solve, the better a learning system you are building.

Ramin Hasani

在物理建模中,有一类模型如物理信息神经网络。这又是另一种架构,非常适合物理模拟。所以这是光谱的另一端:数据的属性、上下文长度以及所有其他考虑因素会极大地改变架构。然后,你又想将其扩展到最大的 regime,使其再次无偏。Transformer 能够通过增加规模来击败这些,但在较小规模下,Transformer 无法击败我们描述的任何其他形式的动态系统。

In physics modeling, there is a class of models like physics-informed neural networks. These are again another architecture that would be really good for physical simulations. So this is the other side of the spectrum: the property of the data, the lengths of the context lengths, and all the other considerations would change the architecture by a lot. Then again, you want to scale this to the largest regime, make it again unbiased. A transformer would be able to add scale to beat this, but at a smaller scale, transformers would not be able to beat any of the other formats of dynamical systems that we described.

24. 资源受限部署与液态神经网络 Resource-Constrained Deployment and Liquid Neural Networks

Host

那么受微分方程启发的原始液态神经网络呢?很明显,蠕虫只消耗很少的瓦特。在实践中,我们是否遇到资源如此受限的情况,以至于今天必须使用这些极其特定或极其偏置的架构来实际部署?

And how about the differential equation inspired original liquid neural networks? It's clear that the worm is running on a very small amount of watts. Do we in practice encounter things that are so resource constrained that you have to go to these extremely specific or extremely biased architectures to actually deploy today?

Ramin Hasani

是的,100%。想想那些需要低延迟的地方,计算必须在微秒内完成。所以你无法承受更大的计算复杂度。我们想要最简单的系统来处理它。在那里你可以有自适应系统,多种不同形式的系统。另一件事是物理模拟,正如我提到的。你有物理数据进来,你想构建工厂中化学反应的数字孪生。在这些地方,你会走向极其偏置甚至只是基于微分方程的模型,比如液态神经网络。今天,它们被应用于许多不同的应用。我看到多年前的原始仓库仍然是开源的,人们仍在构建基于序列数据、物理数据或传感器数据的预测性机器学习模型。这是有道理的,因为系统是连续时间的,使用连续时间动态系统来建模这些行为是合理的。然后,借助像云这样更大的实例,你也可以引导你的智能体以自动研究的形式尝试各种方法。

Yeah, 100%. Think about places where you have to bring in latency, where computation has to happen in microseconds. So you cannot really afford to have larger computational complexity. We want to have the most simple type of system that actually handles that. There you can have adaptive systems, many different formats of systems. The other thing is simulation in physics, as I mentioned. You have physical data coming in, you want to build a digital twin of a chemical reaction that happens at a factory. In those places, you would go towards an extremely biased and maybe even just differential equation based models like liquid neural networks. Today, they are getting applied to many different applications. I see that the original repository from years ago is still open source, people are still building predictive machine learning models on sequential data or physical data or sensor data. It makes sense because the systems are continuous time and it makes sense to have a continuous time dynamical system to apply to modeling those behaviors. Then you can, with larger instances like clouds of the world, also direct your agents to try out a bunch in an auto research kind of format.

25. 开源与硬件协作 Open Source and Hardware Collaboration

Ramin Hasani

你可以直接告诉他们,‘嘿,去挑选最适合这类数据集的神经网络,然后根据你的偏好探索可能的空间。’机器学习人员现在大多在编排自动化的智能体流水线,就像我们正在构建的那样。这就是现状。就 Liquid 自身而言,我们正努力为开源社区做贡献。我们开源了一些来自搜索空间的实例,比如 LFM 2。LFM 3 将是下一代。目标是下一代必须在我们在意的标准上超越上一代,同时不牺牲质量。这对我们来说极其重要。然后,基于这些神经网络,我们决定要造多大的网络,这就成了我们模型的开源版本。我们还与 AMD、高通等芯片公司紧密合作。我们试图理解他们的芯片路线图、硬件路线图,并据此告知他们在下一代 ASIC 中需要做什么。当我们理解了智能的算法层面以及他们可以支持的各种事物,以大幅降低成本或满足他们想要启用的用例约束时,这些就是需要考虑的因素。我们甚至可以为他们构建特定的基础模型图。这就是我们与半导体公司合作的项目。

You can just direct them to, 'Hey, go pick the best neural networks that would be the best fit for this type of dataset, and here is the space of possibilities you want to explore based on the biases you have.' Machine learning people are now mostly orchestrating automated agent pipelines as we are building. That's what's happening. In terms of Liquid ourselves, we are trying to contribute to the open source community. We are open-sourcing some of the instances coming out of our search spaces, like LFM 2. LFM 3 is going to be the next generation. The objective is that the next generation should always beat the previous generation on the criteria we care about, without sacrificing quality. That's extremely important for us. Then, based on these neural networks, we decide how large a network we want to make, and this becomes the open-source version of our models. We also work very closely with silicon companies like AMD and Qualcomm. We try to understand their silicon roadmap, the hardware roadmap, and based on that, we inform what they have to do even in the next generation of their ASICs. When we have this understanding of the algorithmic aspects of intelligence and the variety of things they could support to dramatically reduce cost or satisfy the constraints of the use cases they want to enable, these are the considerations. We can even build specific foundation model graphs for them. That's a project we do with semiconductor companies.

26. 设备基础模型与商业应用 Device Foundation Models and Commercial Applications

Ramin Hasani

在商业方面,我们也把模型应用到许多不同的类别中。Liquid 在我们所谓的设备基础模型领域非常活跃——即数据中心之外的任何处理器。我们试图将技术应用到这些地方。在数据中心内部,我们应用于受限用例,比如超低延迟的 AI 应用、极长序列,以及希望 AI 系统占用极小内存的场景。你需要这些基础模型的高效实现。我们在内存、速度、延迟方面做出贡献,同时不牺牲质量和成本。这些就是定义 Liquid 可以介入的用例的要素。任何新一代架构,无论是来自架构搜索还是新模态,我们都会检查。例如,我们正在探索推荐、搜索、产品目录等多种应用,并且非常擅长理解多模态系统。这就是我们所做的。我们的模型目前已在 Shopify 投入生产,它们正在提升点击率等 Shopify 关心的内部指标的质量。我们在数据中心之外的模型被用于汽车内部。例如,我们为车载智能提供动力。最近,我们与梅赛德斯-奔驰签署了一份历史性合同,我们的模型将为车内的音频和视觉元素提供动力。无论何时你想与汽车对话,新的语音都将来自 Liquid 基础模型。我们用一个模型来控制它,这个模型能提供你迄今为止见过的最佳模型的质量,比如你见过的音频模型,但同时它只有 600 MB。它可以装进车内最小的处理器。这改变了游戏规则,因为你正在重要的地方——汽车、手机——大规模实现本地 AI。世界上有数十亿台设备。移动设备可以非常有用。顺便说一句,移动业务本身就是一个 5000 亿美元的市场,绝对惊人,而且与数据中心市场一样大。所以你可以想象这里有一个平行关系。高效市场和受限智能市场是 Liquid 正在追求的目标。笔记本电脑市场、可穿戴设备、机器人、制造、物联网系统——任何有处理器的系统,Liquid 都可以在上面引入智能。我的远大目标是真正在世界上各种硬件和处理器之上构建一个智能层。那将是我想要的。

On the commercial side, we also take models and apply them in many different categories. Liquid is very enabled in what we call device foundation models—anything with processors outside of data centers. We try to apply our technology to those places. Inside data centers, we apply to constrained use cases like ultra-low latency AI applications, extremely long sequences, and wanting a very small memory footprint for your AI system. You want cost-efficient implementation of these foundation models. We contribute to memory, speed, latency without sacrificing quality and cost. These are the elements that define use cases where Liquid can come in. Any new generation of our architecture from architecture search or new modalities, we check. For example, we are exploring many different applications across recommendations, search, product catalog, understanding multimodal systems very well. That's something we do. Our models are in production at Shopify right now, and they are improving the quality of click-through rates and other internal criteria that Shopify cares about. Our models outside data centers go inside cars. For example, we power in-car intelligence. Recently, we signed a historical contract with Mercedes-Benz where our models will power the audio and visual elements inside the car. Whenever you want to talk to your car, the new voice will come out of a Liquid foundation model. We control that with a model that gives you the quality of the best models you've seen so far, like the audio models you've seen, but at the same time it is 600 megabytes. It can fit inside the smallest processor inside that car. That changes the game because you're enabling local AI at scale on places that matter—on cars, on mobile phones. There are billions of devices in the world. Mobile devices can be extremely useful. By the way, the mobile business itself is a $500 billion market, absolutely insane, and as big as the data center market. So you can imagine there is a parallel here. The efficient markets and constrained intelligence market is something Liquid is going after. Laptop market, wearables, robotics, manufacturing, IoT systems—anywhere we have a processor inside a system, Liquid can bring intelligence on top of that. The aspirational goal I have is to really build an intelligence layer on top of the diverse formats of hardware and processors available in the world. That would be something I would want to have.

27. 边缘计算的巨大机遇 The Massive Opportunity in Edge Compute

Host

我很高兴你提到了那个数据。我记下了它,准备在某个时候强调一下。值得重复一遍。全球年度智能手机市场规模约为 5000 亿美元,而投入数据中心建设的资金也有数千亿美元。数据中心市场会变得比智能手机市场更大,但它现在才刚刚达到这个规模并超过年度智能手机市场。所以从智能角度来看,存在大量未充分利用的暗算力,这还不包括笔记本电脑市场。所以每年大约有 1 万亿美元的算力分布在世界各地的桌面和口袋中,这是一个巨大的可开发基础。

I'm glad you said that stat. I had that noted to make a point of at some point. It's worth repeating. The global annual smartphone market is about $500 billion for all of the hundreds of billions that are going into the data center buildout. And it will get bigger than the smartphone market, but it's only now getting to the scale and getting bigger than the annual smartphone market. So there is a lot of dark compute out there from an intelligence perspective that is not being anywhere close to maximized, and that does not include the laptop market. So there's a trillion dollars generally speaking worth of compute going out into the world on an annual basis that are sitting in people's desks and pockets, and there's a massive substrate there to take advantage of.

Ramin Hasani

我们必须这样做。我们需要利用它。我的意思是,因为我们没有足够的能源来承载它。

And we need to do it. We need to take advantage of it. I mean, because we don't have enough energy to host it.

28. 实现大规模智能供电的难度 Realizing the difficulty of powering intelligence at scale

Ramin Hasani

我觉得我们现在才意识到,要真正以规模驱动这种智能形式有多难。我们必须聪明地工作,不能只做最简单的格式或更受限的用例。比如,如果你有一个一次性预测任务或数据提取任务,你不需要调用最花哨的智能类型,比如那种“无肉”级别的模型来为你做数据提取。你可以有各种不同的智能系统做许多不同的事情。然后,对于世界上最复杂的问题,你可以使用云上最先进的 AI 系统。所以,数据中心之外的处理器世界必须被赋能,这一点正在发生。我认为实际上有一些讨论正在出现;硅谷投资者存在滞后,他们现在才意识到,“哦,现在说得通了。”这个差距现在正在被填补。我认为总的来说,我们正处于高效 AI 的激动人心的时代。

I feel like we are realizing right now how difficult it is to really power this format of intelligence that we have at scale. We really have to work smart; we cannot just do the simplest format or maybe more constrained use cases. For example, if you have a one-shot predictive task or data extraction task, you don't want to call the fanciest type of intelligence, like the meatless level kind of model, to perform data extraction for you. You can have a variety of different intelligence systems doing many different things. Then, for the most sophisticated problems in the world, you can go to the most sophisticated AI systems that exist in the cloud. So it's inevitable that the world, the processor world outside of data centers, should get enabled as we speak. I think there is actually some talk coming out; there's a lag in Silicon Valley investors that are realizing, 'Oh, now it makes sense.' That gap is getting answered right now. I think we are in exciting times for efficient AI in general.

29. 异构计算与架构搜索的挑战 Challenge of heterogeneous compute and architecture search

Host

那么,一个挑战是,显然随着价值数万亿美元的算力投入世界,它是超级异构的,对吧?有许多不同的设备、不同的芯片等等。那么,你的工作中有多少是试图更接近最优利用这些算力?显然,我们之前讨论过扫描、Transformer 和数据中心;这些已经进行了大量优化,以尽可能接近最大吞吐量。当你进行架构搜索,目标是某个随机手机或工厂里的传感器时,你需要在内核和扫描等基础上做多少工作,才能进行这种搜索?

So on the one challenge, obviously with this other trillion dollars worth of compute that's going out into the world, it's super heterogeneous, right? There are many different devices, different chips, etc. So how much of the work that you're doing is about trying to get closer to optimal use of that compute? Obviously, we talked about scans earlier and the transformers and the data centers; that has been optimized a ton to get as close to max throughput as possible. When you go then do an architecture search and you're targeting some random cell phone or sensor in a factory somewhere, how much do you have to work on kernels and scans as kind of foundation to even be able to do that search?

Ramin Hasani

是的,我的意思是,这是不同抽象层次之一。我们正在测试 AI,看看它们在内核设计方面能有多好,针对它们在开源中看到的东西,比如开源中可用的内核,比如它们看到的 GPU 结构,但不包括 NPU,那些最隐蔽的 NPU,其 IP 未公开。那些需要另一种架构搜索,因为我们没有那些知识。但我会说内核设计已经达到我们可以开始自动化的水平。所以你可以有循环来榨取最佳的事后优化。我们做的一些工作,比如我们到目前为止讨论的,是在设计基础模型之前进行的优化,在开始预训练基础模型之前的所有考虑。Liquid 在进入预训练模型之前采用的方法是,在硬件参与之前先进行大规模架构搜索。但设计完架构之后,还有各种训练量化、改变系统位数、进入内核级别以进一步榨取系统性能。所有这些事后工作也可以由内核工程师自动编排并执行这些优化。我们将其作为事后优化步骤。话虽如此,我们内部也有能力在内核级别定义算子。例如,矩阵乘法可能有 100 种方式来实现,带缓存或不带缓存,CPU 工作负载和 GPU 工作负载之间的分配,如何完成整个前向传播。当你执行所有这些优化时,它们原则上也可以成为那个巨大搜索空间的一部分。一些硅合作伙伴正在问我们这些问题:我们如何甚至在设计基础模型之前就先弄清楚竞争图,但又没有训练完成时的保证——因为那是一个昂贵的过程,数百万美元必须投入到神经网络的完整训练中——即使网络很小,你也希望在事后进行各种推理优化,比如量化感知优化。但原则上,我们绝对可以将其作为搜索的一部分推出。

Yeah, well, I mean that's one of the different layers of abstraction. We're testing out the AI to see how good they can get on the kernel design side of things, for the things they have seen in the open source, like the kernels available in the open source, let's say the GPU structures they have seen out there, but not the NPUs, the most hidden type of NPUs whose IPs are not disclosed. Those require another architecture search because we don't have that knowledge. But I would say kernel design is at a level that we can actually start automating. So you can definitely have loops where you can juice out the best kind of post-hoc optimization. Some of the things we work on, like whatever we talked about so far, has been optimization that we do before design of a foundation model, all sorts of considerations before we start pre-training a foundation model. The approach that Liquid takes before getting into pre-training the model is to run this massive architecture search before getting the hardware in the loop. But then right after you design the architecture, there are all sorts of quantization of training, changing the bits of the system, going into the kernel level to try to juice out even more off of the system. All those post-hoc stuff is also something that you can automatically orchestrate with a kernel engineer that performs those optimizations. We do that as a post-hoc optimization step. That being said, we have the capability in-house to define operators at the kernel level as well. For example, you could say matrix multiplication has probably 100 ways that you can actually structure it, with caching or not, distribution between CPU workload and GPU workload, how you do the entire forward pass. As you perform all of these optimizations, they could also in principle become part of that massive search space as well. These are the kind of questions that some of the silicon partners are asking us: how can we, even before designing a foundation model, figure out the competition graph first, but without having the guarantees that when the training is finished—because that's a costly process, multiple millions of dollars has to go into full training of a neural network—even if the network is small, you want to have all sorts of inference optimization before post-hoc, like quantization-aware optimizations. But yeah, in principle we can definitely launch something like that as part of our search.

30. 给硬件制造商的建议 Advice for hardware makers

Host

那么你认为硬件制造商应该做哪些不同的事情?你告诉他们优先考虑什么,以帮助我们实现这个更高效的分布式 AI 未来?

So what do you think hardware makers should be doing differently? What are you telling them to prioritize to help us realize this more efficient distributed AI future?

Ramin Hasani

我觉得随着技术的每一次迭代,硬件提供商应该构建的抽象层次正在提高。例如,我们谈论 CUDA 模式已经很长时间了。现在不再有 CUDA 模式了。为什么?因为如果你看看 AMD,现在内核级别有智能体;它们可以自动化那个栈。基础模型公司应该——抱歉,硬件公司应该开始工作的栈是智能层,就像 Nvidia 正在用他们的 Neatron 项目所做的那样。

I feel like with every iteration of a technology, the level of abstraction that the hardware providers should be building for is coming up. For example, we've been talking about CUDA mode for a long time now. There is no CUDA mode anymore. Why? Because if you look at the AMDs, there are agents at the kernel level now; they can automate that stack. The stack that foundation model companies should—sorry, hardware companies should start working on is the intelligence layer, like what Nvidia is doing with their Neatron project.

31. 硬件公司需要智能层 Hardware companies need an intelligence layer

Ramin Hasani

你知道,如果你看看这个项目实际上有多成功,可能几十亿美元投入到了这些 Nemotron 项目的设计中。如果你真的看看它对 Nvidia 的作用,基本上就是在 Nvidia 的算力之上构建一个智能层,让企业更容易进入并购买 Nvidia 解决方案。解决方案的销售总是围绕这些展开。所以从这些硬件提供商的企业业务角度来看,如果他们想卖出更多硬件,他们就需要提升自己的技术栈,从内核级优化(那些是后处理的事情)进入智能层。这是否意味着他们必须在某种程度上成为基础模型公司?是的,他们必须能够自己训练那个智能层。Nvidia 就是一个成功的例子,这个项目如何为他们带来回报,比如构建 Nemotron 项目。如果你看看其他基础模型公司,比如其他硬件公司,它们还没有做到这一点。它们严格地开始优化自己的模型以适应现有的开源架构。那也非常有价值,你必须这样做,这是理所当然的。但再次强调,如果你想创造差异化的价值并取得成功,我认为你必须能够将技术栈提升到智能层。而且那个智能层应该自然地适配在你的硬件之上。如果你的硬件有局限性,比如相对于竞争对手,软件总能帮你弥补。例如,如果你谈论 AIPC 上的最大 token 速度,在 Intel 电脑、AMD 电脑和 Qualcomm 电脑之间,它们可能有各自的权衡。而最终胜出的公司是那个真正拥有高效智能层在其硬件之上的公司,这样他们可以说,最终你是在这个系统上运行 token。我的 CPU 带宽是比你低一点还是高一点,或者我的内存容量是低还是高,这都不重要。有了智能层和软件优化技术,我实际上可以让自己成为赢家。例如,我作为硬件提供商构建的所有笔记本电脑或 PC 解决方案,都自然地带有智能层。Nvidia 现在也作为 CPU 领域的竞争对手大举进入这个游戏。正如你看到的,Google 通过 Android 生态系统进入,他们正在用 Android 设备(比如 Android 笔记本电脑)取代 Chromebook。这就是他们正在推出的东西,我想它叫 Aluminium OS 之类的,我不太确定具体是什么。然后 Meta 也会推出自己的设备。所以所有的处理器制造商,我感觉他们需要更接近那个智能层,并在规划时考虑软件和效率因素,比如提前规划,了解趋势以及为什么智能层很重要。因为现在 AI 是每个人都想要的东西。你想在硅片上运行的所有应用,都需要建立在已经提供的智能基础上。你想构建各种工具,基本上就是为解决特定问题而构建的应用,在传统的应用意义上。而那个市场可以在芯片之上被激活,如果你有一个非常好的技术栈,能在你拥有的东西之上启用智能。另一件事是,硅片在数据中心之外有巨大的多样性,但资源无论如何都是受限的。你没有那么多电力,所以你必须非常小心,哪些应用是你硬件介质和平台最重要的用途。然后对于那些用例,如果你已经有了智能层,你可以识别何时是淘汰一组设备的时机。例如,我坚信眼镜可能成为未来硬件的一种竞争形态。像眼镜这样的设备会是非常有趣的介质,我认为最终会有人做对。它实际上可能会取代笔记本电脑,对吧?但我们需要真正弄清楚它是什么。但如果你真的想从硬件角度定义市场并确定领域走向,我认为你必须大幅提升你的软件水平,真正动手并在基础模型方面的研发上投入更多精力和资本,因为那不可避免地是新的软件基础。

You know, if you see how this project is actually quite successful, it probably a couple of billion dollars actually went into design of these Nemotron projects. And if you really look at what it does for Nvidia, it's just like basically building an intelligence layer on top of Nvidia's compute so that you have an easier barrier of entry as an enterprise to buy Nvidia solutions. The sales of the solution are always around those things. So from the business of enterprise business of these hardware providers, if they want to sell more hardware, they're going to up their stack from that kernel level optimizations, those are kind of post things, and they get into the intelligence layer. Now does that mean they have to become a foundation model company to some extent? Yes, and they have to be able to train that intelligence layer themselves. Nvidia is a success example there, how this thing is actually paying dividend for them, like building the Nemotron project. If you look at the other foundation model companies, like the other hardware companies, they have not done this yet. They have started strictly like optimizing their models for the open source architectures that are out there. That's also extremely valuable. You have to do that. That's like a given. But again, if you want to create a differentiated value and you want to be successful, I would say you got to be able to bring the stack to the intelligence layer. And that intelligence layer should naturally fit on top of your hardware. If your hardware has limitations, let's say in terms of your competition, software can always help to give you that juice. For example, if you're talking about maximum token speed on an AIPC between an Intel computer, let's say an AMD computer, and let's say a Qualcomm computer, they might have their own tradeoffs. And the company that is going to have an edge is a company that actually owns that efficient intelligence layer on top of their hardware so that they can say at the end of the day, you're going to run tokens on top of this system. It doesn't matter what's the bandwidth of my CPU if it's a little bit lower than you or higher than you, or the amount of memory that I'm actually having in there is lower or higher. With the intelligence layer and software optimization kind of techniques, I can actually get myself to the place where I'm actually the winner here. And for example, all of my laptops or all of my PC solutions that I'm actually building as a hardware provider, they are naturally coming with an intelligence layer on top of that. Nvidia is entering this game also heavily as a competitor in the CPU space now. As you're seeing devices like Google actually entered with Android, with the Android ecosystem that they have, they're replacing Chromebook with Androids, right, like with Android laptops. This is kind of what they're putting out there. I think it's called Aluminium OS or something like that, I don't actually know what exactly it is. And then Meta is going to come with a sort of devices themselves as well. So all the processor builders, I have a feeling they need to get a little bit closer to that intelligence layer and try to leverage software and efficiency considerations into account when they are going for their planning, like planning ahead, where things are going and why that intelligence layer is important. Because right now AI is something that everybody else wants. All the applications you want to do on silicon, you want to do that on a base of intelligence that is already provided to you. You want to build harnesses, you want to build all sorts of harnesses, basically applications you can build for solving a certain problem in a traditional sense of an application, right? And that market can get enabled on top of chips if you have a very nice stack that enables intelligence on top of what you have. The other thing that I would say is that silicon, if you're going outside of data center, there is a massive diversity of silicon, but then the resources are constrained no matter what. You don't have that much power, so you got to be very careful about what sort of applications are the most important uses of your hardware medium and platform. And then those use cases, if you already have an intelligence layer, you could identify when is the time to sunset a set of devices. For example, I strongly believe glasses could be a format of competition for hardware later on. Devices like glasses would be a very interesting kind of medium, and I think somebody's going to get it right eventually. It might actually be a replacement for laptops, right? But we need to really figure out what it is. But if you really want to define a market and define where the field is going from a hardware perspective, I would say you got to up your software game by a lot and really get hands-on and invest a little bit more energy and capital into the R&D side of things on the foundation model side, because that's the new base of software, inevitably.

Host

那么这是否意味着未来会有大量的垂直整合和耦合?我购买的硬件会自带模型,未来可能不那么容易更换,因为模型针对硬件进行了大量优化,反之亦然,以至于这些东西在未来不像今天这样模块化?

So does that imply a future where we have just a lot of vertical integration and a lot of coupling? The model will come with the hardware that I buy and it may not be so swappable in the future because the model is heavily optimized for the hardware and vice versa such that these things are not so modular in the future as they are today?

Ramin Hasani

或者你为什么要换模型?这引出了那个选择问题。你基本上有一个默认选项。现在你想换掉它,当然可以,但为什么要换呢?如果这个模型,即其中的智能层,它不是固定的。它是一种可适应的系统。它是一个自我改进的系统,可以调用你的数据集等。你实际上可以有一个平台来对该系统进行完整的微调。它是启用的,而且不是只有一个模型可以加载到系统中,也不是只有一个云端模型或一个设备端模型可以使用。你必须能够协调这个模型的许多不同实例,才能构建应用。

Or why do you want to change the model? It brings you to that choice kind of question. You're given basically a default. Now you want to switch this thing, by all means, but why do you want to switch it? If this model, the intelligence layer that is in there, it's not fixed. It's kind of an adaptable kind of system. It is a self-improving system with a call to your datasets and stuff. You can actually have a platform that does full fine-tuning of that system. It is enabled, and there is not just one model that you can load into the system, and there's not just one cloud model or one on-device model that you're going to use. You got to be able to orchestrate between many different instantiations of this model to be able to build application.

32. 本地模型效率与生态系统优化 Local model efficiency and ecosystem optimization

Ramin Hasani

我们讨论的是一类模型,它们实际上可以运行在笔记本电脑的硬件之上。所以我认为你需要拥有这样的模型。但如果默认就能提供这种效率,并且开箱即用,那这就是你的选择。我认为这正是英伟达在企业级市场试图推广的,比如他们的 Nvidia NEM 项目。当他们去销售这些产品时,这非常有道理,因为你已经有一个项目,已经有一个多模态模型加载在你购买的 PC 上。我为什么要切换呢?因为他们已经为我做了各种优化,而且运行速度极快。我为什么需要更换呢?而且模型本身是可调的;我可以使用他们的 Megatron 框架来实际微调模型。如果你不想那样做,而只是想选择另一个模型来托管,正如我所说,这必须被提供。你拥有硬件,你必须为整个开源生态系统和可用的模型进行优化。但同时,这也会给你和你的客户带来优势。

We're talking about a model class that would actually sit on top of, let's say, hardware on a laptop, for example. So I would say you need to have that. But if the default is just giving you that efficiency and it's ready to go, that's the choice. That's something that I think Nvidia is trying to propose in enterprises when they go and sell, like the Nvidia NEM projects. When they go out and sell these things, it makes a lot of sense because you already have a project, you already have a multimodal model that is loaded on top of, let's say, the PC that you bought. Why should I switch? Because they have already done all sorts of optimizations for me, and it is running extremely fast. Why do I need to change it? And the model itself is tunable; I can use their Megatron kind of framework to actually tune the model. Now, if you don't want to do that and you want to just choose another model to host it, as I said, this has to be given. You have your hardware, you have to be optimizing for the entire open source ecosystem and models that are available. But at the same time, it would give you an advantage to yourself and to your customers.

Host

好的,也许在最后几分钟,我们来谈谈一个实际应用?你有一篇关于本地协作的博客文章:无云、无等待、在消费级硬件上使用 LFM2 24B A2B 进行工具调用的智能体。那是一个混合专家模型:总参数 240 亿,其中 20 亿活跃。假设我想让它成为我生活的一部分。我感兴趣的是,你会如何指导我今天就把它设置好。作为参考,顺便说一下,我试图把播客的转录稿变成可以喂给我的智能体的东西。所以你可以把这看作既是在指导我,也是在指导我的智能体。我有一个深度上下文数据库,包含我过去五年的数字输出。这个播客会被录制,显然会被转录。它会进入这个数据库。所以你所说的一切,我所说的一切都会在里面,并且可搜索。它还有我所有的电子邮件和 Slack 消息等等。好的,酷。所以现在我的桌面上有 Claude,它可以在本地调用工具来获取数据,但随后它会将所有结果发送到云端,以决定哪些结果才是真正需要关注的。到目前为止,我对这种方式还算满意。我觉得收益肯定值得我承担的任何风险,但我可能更希望先通过本地模型运行这些数据,这样我就不必每次都把所有数据发送到云端。值得注意的是,经过过滤的数据,我可能最终还是会通过一个基础模型来运行,所以它不会完全跳过云端。但我也想节省一些 token,因为我想在 Fable 回归后,把它用在最合适的地方。那么,我该如何在我的本地计算机上使用这个模型获得真正好的性能呢?我需要做微调吗?我需要蒸馏策略吗?我如何实际利用你拥有的基础模型,尽可能多地恢复,比如说,Opus 甚至 Fable 在搜索和理解我的数据方面的性能?那要花多少钱?我应该期望能把这个过程推进到什么程度?我能达到与前沿模型多接近的水平?所以,告诉我所有我需要知道的,然后我会让我的智能体去执行。

Okay, maybe in the last few minutes, how about a practical application of this? You have a blog post on local co-work: no cloud, no waiting, tool calling agents on consumer hardware with LFM2 24B A2B. So that's a mixture of experts: 2 billion active out of 24 billion total parameters. Let's say I want to make that a part of my life. I am interested in how you would coach me on setting this up today. I have, for reference, and by the way, I try to make the transcript of the podcast something that I can feed to my agent. So you can think of this as partly coaching me, partly coaching my agent. So I've got this deep context database that is the last five years of my digital output. This podcast will be recorded, obviously transcribed. It'll go into this database. So everything you said, everything I said will be in there and searchable. And it's got all my email and Slack messages and everything. Okay, cool. So now I've got Claude on my desktop that can call tools locally to get data back, but then it sends all the results to the cloud to decide which of those results are actually the right ones to be looking at. And so far, I've been okay with that. The benefits are certainly worth whatever risk I'm taking, I feel, but I would maybe love to run that data through a local model first so that I don't have to send all my data to the cloud every time. Now notably, what gets filtered through, I'm probably still going to end up running through a foundation model, so it's not going to entirely skip the cloud. But also, I'd like to save some tokens, too, because I'm going to want to use Fable whenever I get it back for whatever it's most appropriately used for. So, how do I get really good performance on my local computer with this model? Do I need to be doing fine-tuning? Do I need a distillation strategy? How do I actually take the base model that you have and recover as much of, let's say, Opus or even Fable performance in terms of searching through, understanding my data as I possibly can? And how much is that? What should I expect in terms of how far I'll be able to push that process? How close to par with frontier models can I get? So, tell me everything I need to know and I'll go have my agent do it.

Ramin Hasani

是的。嗯,这是个很好的问题。显然,就目前而言,本地协作只是为了开阔思路。嘿,你知道吗?你可以实现这类应用。这种本地智能体形式将实现的基本上就是一台本地计算机。你有一个编排器。如果某件事非常复杂,它应该能够将其发送到云端,为你获取答案。如果它不敏感,你可以使用,比如说,较小的 PII 模型。它应该能够使用 PII 模型来真正过滤掉所有个人身份信息,然后为你发送到云端。你甚至不需要看到这些模型。这些模型应该在后台运行。你需要微调的模型就是那个编排器,它可以在许多不同的服务之间路由,甚至包括一些正在执行任务的较小的专用模型,以及一些云端模型,对吧?那个路由器就是计算机。就像本地计算机一样。这是一个定义。当你打开笔记本电脑时,它应该就是那样。你实际上只有那个,然后你开始使用你需要的所有服务。这就像你与助手交流的方式一样。它应该以助手的格式出现,完成各种工作。你可能会对用户界面只是那样而感到奇怪,看不到所有的文件格式和你需要做的事情,但你必须习惯它。现在,Claude Code 甚至用于 IDE 了。现成的 24B 模型能否达到为你完成所有这些工作的质量?不,不能。目前没有任何本地模型能做到。你必须微调它们。你必须让它们专门化,为你想要做的事情提供适当的解释给你的云端智能体。你基本上就是给它这些东西。但是 Claude 能否真正出去为你构建这个模型,比如进行微调并达到那个水平?不能,因为今天,即使是 Fable 级别的模型也无法做到。首先,你无法访问它,因为 Anthropic 实际上拥有自动调优之类的平台,自动化的模型生成模型的平台。有一些公司,包括我们自己,正在构建平台,使你能够对整个系统进行微调,成本从几十美元到,比如说,几千美元不等,就能让你获得云端质量,并且经过所有检查,达到生产质量的模型。而且它不会花费你数万美元。

Yeah. Well, that's a great question. Obviously, the local co-work, as it stands today, is just to open minds. Hey, you know what? This type of applications you can enable. The class of what this format of local agents is going to enable is basically a local computer. You have an orchestrator. If something is so complicated, it should be able to send it to the cloud, fetch the answers for you. If it's not sensitive, you have, let's say, smaller models that are PII models. It should be able to use the PII models to really filter out all the personally identifiable information, send it to the cloud for you. You don't even need to see these models. These models should be in the background. The model that you have to tune is that orchestrator that can route between many different services or even smaller specialized models that are doing stuff, and also some of the cloud models that are out there, right? That router is the computer. That's like the local computer. That's a definition. When you open your laptop, it should be just that. You literally have just that, and then you start working with all the services you want. It's the same way that you communicate with your assistant. It should be in the format of an assistant that does all sorts of those jobs. You might be weirded out by the user interface being just that and not seeing all the file formats and what you have to do, but you got to get used to it. Now, Claude Code is even for the IDE now. Is the 24B off-the-shelf going to be there on the quality that does all of those stuff for you? No, it's not. None of the local models today are there. You have to fine-tune them. You have to get them to be specialized for the stuff that you want to do, with proper explanations for your cloud agent. You basically give it these things. But is Claude able to go out there and actually build this model for you, like do the fine-tuning and get it to that place? No, because today, even Fable-level models would not be able to. First of all, you wouldn't have access to that because Anthropic is actually having access to the auto-tune and stuff, automated kind of models generating models kind of platform. There are companies, and also ourselves, we are building platforms that enable you to do fine-tuning of this whole thing that would cost you between tens of dollars to, let's say, low thousands of dollars to actually get you to the cloud quality for your, and with all the checks, production quality kind of model. And it's not going to cost you like tens of thousands of dollars.

33. 微调成本与平台 Fine-tuning cost and platform

Ramin Hasani

成本会在几十美元到几千美元之间,这就是我们考虑的规模,因为效率非常重要。我们不希望微调发生在你已有的算力上,如果你没有,你可以把它放在一个安全且合适的数据中心,从提供商那里托管,然后就能获得那种质量的模型。未来几个月内,我们会宣布一些平台,你可以直接在终端中接入。你什么都不用做,只需说“嘿,去调用这个平台来微调这个东西”。它会给你一个生产级的基础模型,你可以自己部署,要么用你的数据微调,要么根据用例来操作。它甚至不需要看到那些数据,它可以自己合成数据,训练模型成为一个完美可靠的工具调用者,了解自己的不足,然后去其他地方交付。

It is going to be between tens of dollars to actually like those thousands of dollars, you know, and that's kind of the scale that we are thinking about because again efficiency here matters a lot because we don't want this fine-tuning to happen on the compute that you already have or if you don't have it you would actually put it in a secure and proper kind of data center that you would host from the providers basically and you would be able to get to that quality of the models. There are platforms that hopefully as we go forward in the next few months we're going to announce some platforms that you can hook it directly in your terminal. You don't do anything. You just say, 'Hey, go call this platform for fine-tuning this thing.' This would give you a production grade foundation model that you can deploy for yourself either fine-tune on that data or depending on the use case. It doesn't even have to see that data, it can even synthetically generate data on its own and actually train the model to be a perfect reliable tool caller and understands when its shortcomings are and it can actually go away and deliver it to some other places.

Host

是的,所以我们非常接近那个目标了。等你的平台准备好,我不需要自己动手,这就是关键信息。

Yeah. So I would say we're extremely close to get to that. Wait for your platform to be ready. I don't have to DIY it is the message.

34. 智能的微型化 Miniaturization of intelligence

Host

这太棒了。也许最后一个问题,然后我就把时间交给你来收尾。你认为智能的微型化能走多远?一种直觉是,我们看到的生物世界可能已经处于某种帕累托前沿,我们大脑消耗的瓦数可能已经接近该功率下的最大值。但也许你有不同的直觉。随着我们接近技术的物理极限,你预计手机或笔记本电脑上的智能上限会是什么?

This has been brilliant. Maybe one last question and then I'll just give you the floor to close however you'd like. How far do you think this goes in terms of miniaturization of intelligence if you will? One intuition would be like the biological world that we see is maybe on some sort of Pareto frontier already and so the watts that go into our brain is maybe getting us close to the maximum that we could get for that amount of power. But maybe you have a different intuition. What would you expect in terms of upper limits of intelligence that I could have on my phone or on my laptop as we really get to the physical limits of the technology?

Ramin Hasani

当你思考智能时,我现在的看法是,基于 Transformer 的网络和当前可用的架构,通过规模给了我们上下文学习能力。从下一个词预测中涌现出来的东西——“涌现”这个词很重要——对我来说,智能是一种涌现属性。如果你想微型化智能并将其带入物理世界,我不认为用当前的算法集能接近人脑的每瓦特智能。你做不到,因为我相信人脑经过进化——我们还得考虑设计人类所投入的能量。生物进化是一个非常漫长的过程,而很多人说的“预训练过程,比如你要看整个世界”——人类不需要看整个互联网就能推理。不,但人类经历了多年的进化。所以我很大程度上归因于进化的方面。但话说回来,人类智能有多种上下文学习机制。当当前的 AI 系统做上下文学习时,它们学习的是一个算法的模糊表示,即最小二乘法。基本上,它们以模糊的方式搞懂了梯度下降。对于某些用例,通过一些例子,你可以让系统理解并给出下一个例子。这些是美丽的属性——我称之为这些系统的涌现属性。你把算法设为下一个词预测,就得到了一个模糊版本的梯度下降上下文学习。对于人类,你不仅有涌现的梯度。你可以通过例子学习,但也可以做强化学习。你可以在头脑中运行模拟。你可以做各种算法。你可以在大脑中做贝叶斯统计。你明白我的意思吗?所以你有多种算法从人类被设计和获得智能的方式中涌现出来。我相信如果你开始优化这些,如果你试图强行让系统成为强化学习学习者,但从轨迹中学习,那不是获得智能涌现属性的正确方式。我相信将要被微型化的智能形式——比如在最小时间内获得最大智能,就像人脑一样——我们需要提出这些标识符,这些令人困惑的算法,下一个词预测是其中之一。我们还需要在系统设计之初做什么,以便强化学习、好奇心驱动的智能从最终系统中以有限的能量涌现出来?所以我认为涌现属性是一个需要大量研究的方向,多种学习方式的变体将推动下一代人工智能。

You know, when you think about intelligence, the way I look at it right now is that transformer-based networks and also our current architectures that are available gave us with scale the in-context learning capability. The thing that actually emerged from next-token prediction—the word 'emerged' is important—intelligence for me is an emergent property. If you want to miniaturize intelligence and bring it to the physical world, I don't believe with the current set of algorithms we would be able to get close to the intelligence per watt that the human brain provides. You're not going to get there because I believe the human brain over evolution—we also have to consider the amount of energy that went into the design of humans. Biological evolution is a very long process, and a lot of the pre-training process that people say, 'Oh, you're going to see the entire world'—humans do not need to see the entire internet to be able to reason about something. No, but humans have gone through years of evolution. So I would attribute it a lot to the evolutionary aspect of things. But again, human intelligence came with multiple mechanisms of in-context learning. When our current AI systems do in-context learning, they learn a vague representation of one algorithm, which is least squares. Basically, they figured out gradient descent in a mushy way. For some use cases with some examples, you can make the system understand and give you the next example. These are beautiful properties—this is what I would call an emergent property of these systems. You set the algorithm to be next-token prediction, you got a vague version of gradient descent in-context learning. For humans, you don't only have emergent gradient. You can learn by examples, but you can also do reinforcement learning. You can run simulations in your head. You can do all sorts of algorithms. You can do Bayesian statistics in your brain. You see what I'm saying? So you have a diverse set of algorithms emerged from the way humans got designed and got intelligence. I believe if you start optimizing for those, if you try to force your way into the system to become a reinforcement learning learner but learning from trajectories, that's not the right way to get to emergent property of intelligence. I believe the format of intelligence that is going to be miniaturized—like being in the smallest amount of time maximum amount of intelligence, like the human brain—we need to come up with these identifiers, these confounding algorithms, which next-token prediction is one of those. What else do we have to do at the beginning of the design of these systems so that reinforcement learning, curiosity-driven intelligence, emerges with limited energy from that final system? So I would say that emergent property is something where there's a lot of research that has to go in that direction, and multiple variations of ways of learning that would enable the next generation of artificial intelligence.

35. 结束语 Closing remarks

Host

这可以是一个很好的结尾,但让我再给你一个机会。还有什么我没提到的,或者你想在结束前留给听众的话吗?

That could be a great note to end on, but let me just give you one more opportunity. Anything else that I didn't touch on or anything else you would just want to leave people with before we break?

Ramin Hasani

不,我想我们涵盖了很多信息。

No, I think we covered a good lot of information.

36. 好奇心与AI的未来 Curiosity and the Future of AI

Ramin Hasani

我认为我们谈了一些我通常不会谈论的事情。作为 CEO,我现在更多地谈论商业机会和智能的市场方面,但核心上,我们是科学家,在推动边界。我们每天都在积极思考。即使现在,我后台还在训练一些东西,我真的不想脱离那种心态。我认为每个人现在都能构建,时机再好不过了。智能体非常棒,给了你进行前沿研究的机会。如果你对某事好奇,我们需要找到更好奇的人,提升人们的好奇心,而不是恐惧。我是一个技术乐观主义者;我通常以非常积极的眼光谈论技术。我常说的一句话是,我们需要让人们达到科学家最初那样的好奇心水平。是什么让我们进入科学?是满足好奇心和理解我们周围的世界。这就是科学的目的。借助 AI 系统赋予我们的超能力,我认为每个人都能为理解世界、创造价值和满足好奇心做出贡献。但这需要重新调整我们对工作、自动化以及为我工作的智能体大军的偏见,还需要个人和企业在文化和思维方式上的改变。我期待这种改变。速度非常快,但我们无论如何都必须做到。

And I think we went into some things that I don't usually talk about. In my current role as CEO, I get to talk more about business opportunities and market aspects of intelligence, but at the core, we are scientists pushing boundaries. We actively think about it day-to-day. Even right now, I'm training some stuff in the back end, and I really don't want to get away from that mentality. I think everyone can now build, and the times couldn't be better. Agents are amazing, giving you the opportunity to do frontier research. If you're curious about something, we need to find people who are more curious and improve curiosity in people, instead of fear. I'm a techno-optimist; I usually talk about technology in a very positive light. One thing I always say is that we need to get to a level where people are as curious as scientists from day one. What gets us into science is satisfying our curiosity and understanding the world around us. That's the purpose of science. With the current superpowers given to us by AI systems, I think everyone could contribute to understanding the world, generating more value, and satisfying our curiosity. But it requires rebasing our biases about what is work, what is automation, what is an army of agents working for me, and a change in culture and way of thinking for individuals and enterprises. I'm looking forward to that change. The pace is extremely fast, but we have to do it one way or another.

Host

我喜欢这一点。我认为自己非常幸运,如今能过上一种极其由好奇心驱动的生活。这是未来积极愿景的一个候选,应该能激励很多人,尤其是因为正如你所说,我们已经有机会借助现有的 AI 系统进入那个阶段。这太棒了。我真的很享受。非常感谢你的时间。Ramin Hasani,Liquid AI 的联合创始人兼 CEO,感谢你加入认知革命。

I love that. I consider myself extremely fortunate to live a remarkably curiosity-driven life these days. That's one candidate for a positive vision for the future that should inspire a lot of people, especially because, as you note, there's an opportunity to move into that phase already with the AI systems we have. This has been excellent. I really enjoyed it. Thank you so much for the time. Ramin Hasani, co-founder and CEO of Liquid AI, thank you for being part of the Cognitive Revolution.

Ramin Hasani

谢谢。降下来,感受它降下来。降下来。从 300 个小光点开始。一个微小的死脑筋。学会在夜里驾驶。再加 12 个神经元。精准转向。谁说需要一座山才能重新设计?砍掉它,砍掉它。感受它的干净明亮。最大的机器,只是一个新设计。感受它,移动它,感受它,但从不凝固在时间里。流过你的手。液体落在大地上。升上云端,进入炸弹。液体,液体,适应性强,有生命。保持运动。保持活力。把思维带到我们生活的地方。万亿个沉睡的巨人一直就在我们口袋里。唤醒银色。唤醒光芒。不是每件小事都需要最华丽的山。只要足以让整个地平线歌唱。砍掉它,砍掉它,直到它精干明亮。伟大的想法。我的使命只是一个新设计。感受它移动。感受它。从不凝固在时间里。液体流过你的头。液体从大地落下。升上云端,进入炸弹。液体,液体,液体。适应性强,有生命。保持运动。保持轻盈。把思维带到低处。秘密是什么?好奇心。它要去哪里?无处不在,我们会在车里、池塘里、露天里。认知革命。智能无处不在。液体,液体。感受它降下来。慢慢降下来。升上云端,落到地面。我们让它流动。让它流动。

Thank you. Coming down low. Feel it coming down. Coming down low. Started with 300 little lights. A tiny brain dead. Learn to drive into the night. 12 more neurons. Pock it on a dime. Who said you need a mountain just to redesign? M cut it down, cut it down. Feel it clean and bright. Biggest machine, just a new design. Feel it, move, feel it, but never frozen in time. Flowing through your hand. Liquid falling on the land. Up the cloud and into the bomb. Liquid liquid adaptable alive. Keep it moving. Keep it alive. Bring the mind down low where we live. A trillion sleeping giants in our pockets all this time. Waking up the silver. Waking up the shine. Not the fanciest mountain for every little thing. Just enough to make the whole horizon sing. Cut it down. Cut it down till it's lean and bright. Big ideas. My mission just a new design. Feel it move. Feel it. Never frozen in time. Liquid flowing through your head. Liquid falling from the land. Up the cloud and to the bomb. Liquid, liquid, liquid. Adaptable alive. Keep it moving. Keep it light. Bring the mind down low. What's the secret? Curiosity. Where's it going? Everywhere we'll be in the car in a pond in the open air. cognitive revolution. Intelligence everywhere. Liquid liquid. Feel it coming down. Coming down slow. Up the cloud and to the ground. We let it flow. Let it flow.

Host

如果你觉得这个节目有价值,我们希望你能花点时间与朋友分享、在网上发帖、在 Apple Podcasts 或 Spotify 上写评论,或者在 YouTube 上给我们留言。当然,我们一直欢迎你的反馈、嘉宾和话题建议以及赞助咨询,可以通过我们的网站 cognitive revolution.ai,或者在你喜欢的社交网络上私信我。Cognitive Revolution 是 Turpentine Network 的一部分,该播客网络现在是 A16Z 的一部分,专家们在这里讨论技术、商业、经济、地缘政治、文化等。我们由 AI Podcasting 制作。如果你需要从停止录制到听众开始收听的全套播客制作帮助,请查看他们,并在 aipodcast 上看到我的推荐。感谢每一位听众,感谢你们成为认知革命的一部分。

If you're finding value in the show, we'd appreciate it if you'd take a moment to share with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website, cognitive revolution.ai, or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of A16Z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast. And thank you to everyone who listens for being part of the Cognitive Revolution.

互动版:逐字朗读 + 针对本期提问 →