与约书亚·本吉奥教授探讨主动学习与 GFlowNets

Active Learning and GFlowNets with Professor Yoshua Bengio

约书亚·本吉奥 Yoshua Bengio · ML Street Talk · 2022-02-22 · 约 93 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

约书亚·本吉奥教授讨论 GFlowNets、主动学习,以及如何在复杂组合空间中高效查询训练数据。

Professor Yoshua Bengio discusses GFlowNets, active learning, and how to efficiently query oracles for training data in complex combinatorial spaces.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 25)

全文 · Full transcript(中英对照)

引言与事务 Introduction and Housekeeping

Host

非常荣幸,我必须说,你们的问题给我留下了深刻印象。这说明你们确实做了准备,读了论文并进行了思考,我非常感激。谢谢。今天是一个极其特殊的时刻:我们邀请到了约书亚·本吉奥教授。说实话,我到现在还难以置信。不过,首先是一些事务性通知。我们刚刚推出了一个新的 Discord 社区,欢迎加入,打个招呼,介绍一下自己。如果你想成为管理社区的一员,或者只是帮我们做些事情,我们很乐意与你交流。应大家的要求,我们还增加了两种支持我们的方式。我们现在有了 Patreon 和周边商店。如果你有兴趣支持 MLST 的某些节目,请联系我们,因为我们很乐意与你对话。今年我们做了很多酷炫的事情。我们已经录制了大约六集尚未发布的节目,并且还预约了一些了不起的人。所以,是的,这将会非常精彩。一如既往,如果你喜欢这里的内容,请考虑点赞、订阅,并在 iTunes 上给我们的播客评分,因为这真的能帮到我们。我叫它 iTunes……是 iTunes 吗?Apple Podcasts?我不知道,随便叫什么吧。Weights & Biases 是开发者优先的 MLOps 平台,我们非常自豪他们赞助了本期节目。现在,跟踪机器学习实验很困难。靠“即兴发挥”的方法只能走到一定程度。我们需要一个平台,能够比较模型,可视化它们相对于之前所有运行的性能特征,并找出最佳参数类型。最重要的是,这个过程需要可重复。听起来要求很高,对吧?这正是 Weights & Biases 平台为你所做的。现在你甚至可以实时跟踪长时间运行的实验指标。我认为,在 ML DevOps 生命周期中,深入理解科学与工程之间的复杂交互非常重要。数据科学家需要有价值的反馈,他们需要沟通为什么运行某个实验,并分享关于下一步的笔记。报告能让这项工作保持良好组织,并与实际运行的 Weights & Biases 实验相关联,而不是在 Slack 上随意分享截图。实验完成后,创建报告并与团队分享非常容易。你也可以为自己添加笔记,以便日后探索。你可以保留工作日志,甚至可以内部或外部共享你的发现。这绝对是一个游戏规则的改变者。我非常相信这种工程严谨性。我现在是一家名为 Merge 的代码审查初创公司的 CEO,我喜欢拉取请求流程和工具如何固化软件开发周期中做出的重要集体决策。类似地,Weights & Biases 固化了模型开发、实验和部署过程中做出的重要决策。记住,今天就去 wandb.com/forward/mlst 查看 Weights & Biases。如果你有兴趣赞助未来的节目,请联系我们。Weights & Biases 目前正在赞助我们的首播节目,但我们还有很多其他内容即将推出,也有赞助机会,所以请告诉我们。干杯!

My pleasure, and I must say I've been really impressed by all your questions. It showed that you did prepare and read papers and think about it, and that's very much appreciated. Thank you. Today is an incredibly special occasion: we have Professor Yoshua Bengio on the show. Honestly, I just can't get over it. But first, a little bit of housekeeping. We've just launched a new Discord community, so please jump in there, say hello, introduce yourself. If you want to be part of the moderating community or just help us do stuff over there, we would love to talk with you. By popular demand, we've also added a couple of ways in which you can support us. We now have a Patreon and a merch store. If you're interested in supporting some of the episodes of MLST, then get in touch with us because we'd love to have a conversation with you. We're just doing so much cool stuff this year. We've already recorded about six episodes that we haven't released, and we've got some amazing people booked as well. So yeah, it's going to be incredible. As always, if you like the content here, please consider hitting the like and subscribe button and rating our podcast on iTunes, because it really, really helps us out. I called it iTunes... is it iTunes? Apple Podcasts? I don't know, whatever it's called. Weights & Biases is the developer-first MLOps platform, and we're extremely proud today that they are sponsoring this episode. Now, tracking machine learning experiments is difficult. Using the winging-it methodology can only get you so far. What we need is a platform where we can compare models and visualize their performance characteristics against all of the previous runs and figure out the best type of parameters to use. Most importantly of all, this process needs to be reproducible. Sounds like a tall order, right? Well, this is exactly what the Weights & Biases platform does for you. Now you can even follow metrics from long-running experiments in real time. I think it's really important to lean into the complex interaction between science and engineering in the ML DevOps life cycle. Data scientists need valuable feedback, and they need to communicate why they're running given experiments, and they need to share their notes around the next steps. Reports keep this work well organized and connected to the Weights & Biases experiments which were run, as opposed to just sharing random screenshots in Slack. It's so easy to create a report and share it with your team after you've finished with your experimentation. You could just add notes for yourself as well to explore later on. You can keep a work log, and you can even share your findings internally or externally. This is an absolute game changer. I'm a big believer in this kind of engineering rigor. I'm the CEO of a code review startup called Merge these days, and I love how the pull request process and tooling immortalizes important collective decisions which were made during the software development life cycle. Similarly, Weights & Biases immortalizes important decisions that were made during model development, experimentation, and deployment. Remember, check out Weights & Biases today by going to wandb.com/forward/mlst. And if you're interested in sponsoring future episodes, get in touch with us. Weights & Biases are currently sponsoring our premiere shows, but we have lots of other content coming and opportunities for sponsorship, so let us know. Cheers!

GFlowNets 与主动学习 GFlowNets and Active Learning

Host

约书亚·本吉奥教授刚刚发布了一系列关于 GFlowNets 的论文。GFlowNets 完全属于主动学习领域,这是一种模型,能够经济地向一个预言机(很可能就是现实世界)询问最显著的训练样本以继续学习。学习器可以选择或影响它得到的样本,我们希望学习一个能高效逼近预言机的函数。我们应该如何选择查询?我们如何不仅考虑预测器的值,还要考虑学习系统对预测器的确定程度?不确定性或熵的区域有点像我们进一步探索的有趣候选。我们需要能够想象或发明查询给预言机。现在,机器学习模型之所以样本高效,原因之一是可能输入样本的组合空间。我们无法在所有数据上训练,因为空间太大——浩瀚无边。所以你可能听说过一个相关的主动学习概念叫机器教学,这是一种交互式版本,人类交互式地选择最显著的数据来训练机器学习模型,最大化关于训练样本的信息增益。实际上,我们在这里学习的函数空间是高度结构化的。我们只需要在函数空间中大部分丰富信息存在的地方采样训练数据。我的意思是,如果你仔细想想,机器学习模型只是信号和标签之间的联合概率分布,这个分布有众数或密度区域或信息区域,实际上大部分区域只是虚无区域,需要更少的训练样本来学习和表示。现在,如果你和一个贝叶斯学派的人(比如我工作中的朋友康纳·坦)谈论如何学习这个分布,他们会比屁股上绑了炸药的灵缇犬还快地提起马尔可夫链蒙特卡洛。马尔可夫链蒙特卡洛是一种越来越流行的采样方法,用于获取关于未归一化分布或能量函数的渐近信息,特别是在贝叶斯推断中估计后验分布,你可能之前听说过。现在,你可以在不知道分布所有数学性质的情况下描述一个分布——所以如果你没有它的解析表示——只需从分布中随机采样值。马尔可夫链蒙特卡洛的一个特别优势是,即使只知道如何计算不同样本的密度,它也可以用于从分布中抽取样本。马尔可夫链蒙特卡洛的马尔可夫性质是这样的想法:随机样本由一个特殊的顺序过程生成,每个随机样本被用作生成下一个随机样本的垫脚石。这听起来可能非常复杂,但实际实现相当简单。马尔可夫链蒙特卡洛从一个初始猜测开始——只是一个可能从分布中抽取的值——然后我们通过在该示例的邻域添加随机扰动,从这个初始猜测产生一个样本链,每个从随机扰动分布中抽取的新提议要么被拒绝,要么被接受。当然,有不同的变体。我的意思是,特别是调整邻域中随机提议的选择方式,或者提议是否……

Professor Yoshua Bengio has just released a bunch of papers around GFlowNets. Now, GFlowNets exists squarely in the domain of active learning, which is a model that can economically ask an oracle — which is probably the real world — for the most salient training examples to continue learning. The learner can choose or have an influence on the examples it gets, and we want to learn a function which approximates the oracle efficiently. How should we pick the queries? How should we take into account not just the value of the predictor but also how certain we are about the predictors from the learning system? Areas of uncertainty or entropy are kind of like interesting candidates for us to explore further. We need to be able to imagine or invent queries to give to the oracle. Now, one of the reasons that machine learning models are so sample-efficient is because of the combinatorial space of possible input examples. We can't train on everything because the space is just too large — it's vast. So you might have heard of a related concept of active learning called machine teaching, which is an interactive version where the human interactively selects the most salient data to train a machine learning model, maximizing the information gain with respect to the training samples. Now, the reality is the function space that we're learning here is highly structured. We only really need to sample training data where most of the rich information exists in that function space. I mean, if you think about it, a machine learning model is just a joint probability distribution between signals and labels, and this distribution has modes or areas of density or information, and actually most of it is just areas of nothingness which require fewer training examples to learn and to represent. Now, if you spoke to a Bayesian person like my friend Conor Tan at work, you know, how to learn this distribution, they would bring up Markov chain Monte Carlo quicker than a whippet with a bum full of dynamite. Now, Markov chain Monte Carlo is an increasingly popular sampling method for obtaining asymptotic information about unnormalized distributions or energy functions, especially for estimating the posterior distribution in Bayesian inference, which is where you've probably heard of it before. Now, you can characterize a distribution without knowing all of the distribution's mathematical properties — so if you don't have an analytical representation for it — just by randomly sampling values out of the distribution. Now, a particular strength of Markov chain Monte Carlo is that it can be used to draw samples from distributions even when all that is known about the distribution is how to calculate the density for different samples. Now, the Markov property of Markov chain Monte Carlo is this idea that random samples are generated by a special sequential process, and each random sample is used as a stepping stone to generate the next random sample. Now, this might sound very complex, but the practical implementation is pretty simple. Markov chain Monte Carlo just starts with an initial guess — just one value that might plausibly be drawn from the distribution — and then we produce a chain of samples from this initial guess by adding random perturbations in the neighborhood of that example, and each new proposal drawn from that random perturbation distribution is either rejected or accepted. There are different flavors of this, of course. I mean, in particular, like tweaking how the random proposals in the neighborhood are selected or whether the proposals are...

GFlowNets 简介 Introduction to GFlowNets

Host

最简单的启发式方法是看它是否在函数下方。现在,马尔可夫链蒙特卡洛方法的思想是用相对少量的随机样本捕获一个分布,但现实远非如此。在高维空间中,当分布有许多相距很远的模式时,它实际上是指数级昂贵的。有很多面向人类的技巧试图在特定情况下使其工作良好,但我们缺少一个更通用的、可机器学习的方法。这是为什么我们还没有在许多机器学习应用中看到它的主要原因。假设我们要学习的函数具有底层结构,那么我们可以通过机器学习逃离马尔可夫链蒙特卡洛的指数时间,这就是 Bengio 所说的系统性泛化,即我们如何以有意义的方式从数据中泛化到远处。

Selected the simplest heuristic being whether it's below the function or not. Now the idea is that Markov chain Monte Carlo methods capture a distribution with only a relatively small number of random samples, but the reality is anything but. In high dimensions and where the distribution has many modes spread far apart, it's actually exponentially expensive. There's a bunch of human-oriented hacks to try and make this work well in specific cases, but we're missing a much more general machine-learnable solution. This is the main reason why we haven't seen it used in many machine learning applications yet. Assuming that the function we want to learn has underlying structure, then we can escape the exponential time of Markov chain Monte Carlo with machine learning, and this is what Bengio calls systematic generalization, which is to say how do we generalize far from the data in a way which is meaningful.

Yoshua

现在,GFlowNets 是一个主动学习框架,其目标是生成显著且多样化的训练数据,以最样本高效的方式增强我们的模型。要使 GFlowNets 工作,我们需要一个奖励函数和一个确定性的情节环境。听起来熟悉吗?是的,就像强化学习一样。现在,流网络是一个有向图,有源和汇,边在它们之间通过中间节点携带一定量的流。所以我认为一个好的思考方式是把它想象成水管。现在,为了我们的目的,我们定义了一个具有单个源的流网络。根节点,或者你可以说是网络的汇,对应于终止状态。它被设计用来找到通过我们系统的可能轨迹。好的,就把 AlphaZero 看作这些轨迹的一个好类比。现在训练目标是使它们大致按给定奖励函数的比例进行采样。这与 AlphaZero 形成鲜明对比,在 AlphaZero 中我们采样是为了最大化期望奖励。

Now GFlowNets are an active learning framework where the name of the game is to generate salient and diverse training data to augment our model in the most sample-efficient way possible. For GFlowNets to work, we need a reward function and a deterministic episodic environment. Does that sound familiar? Yes, just like reinforcement learning. Now a flow network is a directed graph with sources and sinks, and edges carrying some amount of flow between them through intermediate nodes. So I think a good way to think about this is pipes of water. Now for our purposes, we define a flow network with a single source. The root nodes, or you might say the sinks of the network, correspond to the terminal states. Now it's designed to find the possible trajectories through our system. Okay, and just think of AlphaZero as being like a good analogy for these trajectories. Now the training objective is to make them approximately sample in proportion to the given reward function. This is in stark contrast to AlphaZero, where we were sampling to maximize the expected reward.

想象机器与高尔顿板类比 Imagination Machine and Galton Board Analogy

Yoshua

所以 Bengio 的大想法是,我们可以在生成模型和真实世界之间建立一个交互循环。真实世界是昂贵的,所以为什么不在我们的脑海中训练一个想象机器,直到我们准备好向真实世界提出好问题呢?我们可以用想象的实验来训练我们的生成器,然后向真实世界发出查询。我们正在思考如何可视化 GFlowNets 的工作方式,这时想到了高尔顿板。高尔顿板,也称为豆机,是统计课程、科学博物馆和趣味小工具店中的常见道具。板上有几排交错的钉子,下面有一排桶。珠子从顶部的漏斗倒入,然后洒在顶部中心的钉子上。珠子碰到钉子时向左或向右弹跳,最终收集到底部的桶中。如果钉子精确对称地排列,珠子会在底部聚集形成熟悉的二项式钟形曲线。现在想象一下,钉子变成了带有可调节阀门的流量门,可以将珠子更多地导向左侧或右侧,以偏置流动路径。有了这样的机器,你可以调整阀门或流速来创建任何分布。例如,要创建均匀分布,我们会打开远离板中心线的门,将更多的珠子流导向通向边缘和角落的较少路径。或者要创建多模态分布,我们会安排门将流分成两个或多个流,然后这些流会在下面堆积成多个驼峰或模式。这里有很大的灵活性。确实,给定一个分布,通常有多种流量门解决方案来产生它。如果我们有一种智能的、有原则的方法来训练这些门,那该多好啊。

So Bengio's big idea is that we could have an interacting loop between a generative model and the real world. The real world is expensive, so why not train an imagination machine in our mind until we're ready and waiting to produce good questions to the real world? We could use imagined experiments to train our generator, then produce queries to the real world. We were thinking about a way to visualize how GFlowNets work when the idea of a Galton board came to mind. A Galton board, also known as a bean machine, is a common prop in statistics courses, science museums, and fun gadget stores. The board has rows of interleaved pegs above a bottom row of buckets. Beads are filled into a funnel at the top of the board and then sprinkled on the top center peg. The beads bounce either to the left or to the right as they hit the pegs and eventually collect into buckets at the bottom. If the pegs are precisely and symmetrically arranged, the beads will aggregate at the bottom into a familiar binomial bell curve. Now imagine that the pegs were instead flow gates with adjustable valves that could direct the beads more to the left or more to the right to bias the flow paths. With such a machine, you could adjust the valves or flow rates to create any distribution. For example, to create a uniform distribution, we'd open up the gates flowing away from the center line of the board to drive more bead flow to the fewer number of paths leading to the edges and the corners. Or to create a multi-modal distribution, we'd arrange the gates to split the flows into two or more streams that would then pile up in multiple humps or modes below. There's a lot of flexibility here. Indeed, given a distribution, there are generally multiple flow gate solutions to produce it. It'd be nice, wouldn't it, if we had an intelligent principled way to train these gates.

Host

在 GFlowNets 中,我们在流量调整背后放了一个神经网络,一个大脑,它可以优化门以匹配我们想要的任何分布。这里我们感兴趣的是在强化学习背景下对奖励函数进行采样。在这个背景下,这是一个强大的模拟和采样范式。你看,一旦大脑调整了流量权重,这样一个修改后的高尔顿板,或者更一般地说,一个流网络,可以快速高效地采样多样化的路径,从而得到奖励分布。重要的是要指出,这样做路径采样更加多样化。与经典强化学习不同,GFlowNet 不会只专注于它偶然首先发现的一小部分高奖励路径。相反,它根据奖励的比例随机采样广泛的路径。当然,高奖励路径会被更重地采样,但数量大得多的低奖励路径也会获得一部分采样。我们为什么还要费心处理这些路径呢?答案是我们需要平衡利用(高奖励)和探索(学习),以更好地学习奖励函数。这在处理高不确定性的复杂现实场景时尤其重要。例如,想想分子药物发现和设计,或者在丛林地形中导航。在这两种场景中,我们真的对特定路径可能如何发展知之甚少。我们可能会偶然发现下一个奇迹疗法或陷入流沙陷阱。为了找到全局最优路径,保持开放的选择很重要。

In GFlowNets, we put a neural network, a brain, behind the flow adjustments, a brain which can optimize the gates to match any distribution we desire. Here we are interested in sampling a reward function in the context of reinforcement learning. In that context, this is a powerful simulation and sampling paradigm. You see, once the brain has tuned the flow weights, such a modified Galton board, or more generally a flow network, can sample diverse paths quickly and efficiently, leading to the reward distribution. It's important to point out that the path sampling is more diverse doing it this way. Unlike classic reinforcement learning, a GFlowNet doesn't just fixate on a small number of high-reward paths it happens to find first. Instead, it stochastically samples a broad spectrum of paths in proportion to their reward. Sure, high reward paths will be sampled with higher weight, but the far larger population of low reward paths will get a share of the sampling as well. Why should we even bother with such paths? The answer is we need to balance exploitation (high reward) with exploration (learning) to better learn the reward function. This is especially important when dealing with complex real-world scenarios of high uncertainty. For example, think of molecular drug discovery and design, or navigating jungle terrain. In both those scenarios, we really know very little about how a particular path may play out. We might stumble into the next miracle cure or a pitfall of quicksand. To find the globally best paths, it's important to keep our options open.

GFlowNets 优势:多样性与鲁棒性 Advantages of GFlowNets: Diversity and Robustness

Yoshua

除了这种采样多样性,GFlowNets 还带来了神经网络的全部力量来发现潜在结构并学习奖励函数。这与它们的多样化采样相结合,也使 GFlowNets 在处理多模态分布时更加鲁棒,而多模态分布是贪婪算法和马尔可夫链蒙特卡洛的常见陷阱。如果存在连接多个模式的结构,GFlowNets 可以学习它并外推到新的模式。一旦发现,它们会通过设计驱动采样覆盖这些模式,并学习更多整体结构。GFlowNets 似乎为智能采样范式提供了一条有趣的新路径(双关语)。所以你可能会问:GFlowNets 与 AlphaZero 有何不同?嗯,AlphaZero 中的策略网络根据状态给你一组动作。AlphaZero 训练策略网络以最大化奖励,因此所有轨迹最终都达到最高奖励。而 GFlowNets 的训练使得动作按奖励比例分布。因此,它不会修剪掉所有低奖励轨迹,而是只是不那么频繁地采样它们。现在,GFlowNets 在探索方面有一个明显的区别。我的意思是,你可能会争辩说蒙特卡洛树搜索在开始时仍然进行广泛的探索,但尽管它快速收敛并修剪低奖励轨迹,它仍然从经过 softmax 缩放的底层概率分布中采样。所以实际上一开始就没有太多可探索的。总之,GFlowNets 比 AlphaZero 蒙特卡洛更好。

Beyond this sampling diversity, GFlowNets also bring the full power of neural networks to discover latent structure and learn the reward function. This combined with their diverse sampling also makes GFlowNets more robust when dealing with multimodal distributions, which are a common trap for greedy algorithms and Markov chain Monte Carlo. If there is structure linking the multiple modes, GFlowNets can learn it and extrapolate to new modes. And once discovered, they will by design drive the sampling to cover those modes and learn more structure overall. GFlowNets seem to offer an intriguing new path (pun intended) for an intelligent sampling paradigm. So you might ask: how are GFlowNets different from AlphaZero? Well, the policy network in AlphaZero gives you a set of actions given a state. AlphaZero trains the policy network to maximize reward, so that the trajectories all end up at the highest reward. Now what GFlowNets do is they train so that the actions are distributed in proportion to the reward. So rather than pruning away all of the low reward trajectories, it will sample them just less often. Now there is a manifest difference between GFlowNets in respect of exploration. I mean, you might argue that the Monte Carlo tree search is still doing wide exploration at the beginning, but in spite of its rapid convergence and pruning of low reward trajectories, it's still sampling from the underlying probability distribution which has been scaled with a softmax. So there's actually not that much to explore in the first place. So in summary, GFlowNets are better than AlphaZero Monte Carlo.

GFlowNets 简介 Introduction to GFlowNets

Host

Bengio 教授,这太棒了。能跟我们讲讲这项激动人心的工作及其应用吗?

Professor Bengio, this is amazing. Can you tell us about this exciting work and some of its applications?

Yoshua

是的,我觉得至少在过去六七年里,没有哪个新课题能像 GFlowNets 这样让我兴奋。而且它实际上比你刚才提到的还要丰富得多。我对 GFlowNets 的理解是,它是一种用于概率机器学习的通用可学习推理框架。一种理解方式是,它是多链马尔可夫链蒙特卡洛采样的可学习替代品。但不仅如此,它还可以用于估计概率本身,不仅仅是采样,还能估计像配分函数和条件概率这样难以处理的数量,这些通常需要对难以计数的项求和。所以我认为这可能是——我们仍处于起步阶段——概率建模的瑞士军刀,利用机器学习来处理看似棘手的问题,并通过大型神经网络的泛化能力高效完成。

Yeah, I don't think I've been as excited about a new topic at least in the last six or seven years as I am now with GFlowNets. And it's actually even much more than what you've been talking about. The way I think about GFlowNets is as a kind of framework for generic learnable inference for probabilistic machine learning. So one way to think about this is it's a learnable replacement for multi-chain Markov chain Monte Carlo sampling. But actually, so there's that, and I'll explain if you want why this is important and to use machine learning there. But also it can be used to estimate probabilities themselves, not just sampling, but also estimate intractable quantities like partition functions and conditional probabilities that would otherwise require summing over an intractable number of terms. So I think of this as potentially—there's still, we're still at the beginnings of this—a Swiss Army knife of probabilistic modeling that uses machine learning to be able to do things that look intractable but do them efficiently thanks to the generalization power of large neural nets.

高尔顿板类比 Analogy with Galton Board

Host

我们一直在想办法帮助听众直观理解 GFlowNet 的工作原理。我想跟你探讨一个可能性。不知道你是否听说过高尔顿板,也叫豆子机。这是统计学教授在入门课程开始时常用的道具,用来提供直观感受。它是一块板子,底部有垂直的桶,桶上方有交错排列的钉子。珠子从顶部放入,碰到钉子后向左或向右弹跳,最终落入底部的桶中。如果钉子精确对称排列,珠子会在底部形成漂亮的二项分布曲线。似乎 GFlowNet 在优化路径时,就是微调钉子向左或向右,从而改变珠子流动的偏向。这样,GFlowNet 可以调整钉子,使珠子在底部形成我们想要的任何分布,与奖励函数匹配。所以,这是理解 GFlowNet 的好方式吗?

We've been trying to think of a way to help our listeners visualize what a GFlowNet does. I wanted to run by a possibility to you. I'm not sure if you've heard of Galton boards, also called bean machines. They are a prop often used by statistics professors at the start of an elementary introductory course to give a visual intuition. It's a board with vertical buckets at the bottom and interleaved rows of pegs above the buckets. Beads are fed into the top of the board, bounce left or right as they hit the pegs, and eventually collect at the bottom. If the pegs are precisely and symmetrically arranged, the beads form a nice binomial curve at the bottom. It seems like what GFlowNets are capable of doing when they optimize the pathways is tweaking the pegs a little to the left or right to bias the flow of beads one way or the other. In this way, a GFlowNet could arrange the pegs so that the beads could form any distribution at the bottom that we want, matching the reward function. So is this a good way to think about GFlowNets?

Yoshua

是的,这是一个好方式。但它缺少一个非常重要的方面,很难直观呈现:所有这些钉子的权重——比如球向左或向右的概率——并不是像表格机器学习那样独立学习的,而是有一个神经网络,它以这个大板上的位置作为输入,告诉你在这个位置向左或向右的相对权重。这之所以重要,是因为它允许泛化,因为这块板子非常大,是指数级大小的。所以不可能为每个选择学习一个独立的参数。因此,你有一个或多个神经网络,它们在所有可能的位置之间共享统计强度,从而能够从训练过程中看到的有限数量的训练轨迹泛化到从未见过的位置和路径。这一点至关重要,否则你无法扩展到大规模问题,而这正是我们想要做的。

Yes, it is. But it's missing a really important aspect of it, which would be difficult to present visually: all of these peg weights—like the probability of the ball going left or right—are not learned independently as a tabular machine learning, but there is one neural net that knows about the locations in this big board as input and tells you how much relative weight to go left or right at this position. The reason this is important is because it allows generalization because this board is huge, it's exponentially large. So there's no way you're going to learn a separate parameter for each of these choices. So you have this neural net, or potentially several neural nets, that share statistical strength across all the possible positions so that it can generalize to places and paths that it has never seen from a finite number of training trajectories that it sees while being trained. That's crucial, otherwise you couldn't scale to large problems, which is really what we want to do.

与 AlphaZero 对比及多样性 Comparison with AlphaZero and Diversity

Host

从某种意义上说,研究通过将马尔可夫链蒙特卡洛的预热时间和稳定时间转移到离线训练来实现相同目标。记住,整个过程可以离线训练,然后在推理模式下一次性完成,而蒙特卡洛研究则需要在推理模式下也进行。另一个方面是,我们省去了从马尔可夫链蒙特卡洛高效采样所需的大量人工工程。还有一点就是多样性,宝贝。想想 GFlowNet 和 AlphaZero 在采样奖励路径分布上的区别。如果你观察分布,会发现 AlphaZero 在众数周围有一个小框,而 GFlowNet 则是整个分布。我们很清楚,保持多样性对于在搜索问题中发现有趣的垫脚石至关重要。最后,Bengio 发表的结果显示,GFlowNet 在某些问题上比马尔可夫链蒙特卡洛和 PPO 收敛得更快,并且能更快地找到分布函数中的更多众数。享受节目吧,各位。

Research in some sense because they achieve the same goal by offloading the burn-in time and the stabilization time of Markov chain Monte Carlo. Remember, this whole thing can be trained offline and then when in inference mode we can do it in a single shot, whereas with Monte Carlo research we actually had to do it in inference mode as well. The other thing is we're kind of offloading all of the human engineering required to sample efficiently from Markov chain Monte Carlo. And the other thing is diversity, baby. I mean, consider the difference between how GFlowNets and AlphaZero sample the reward path distribution. If you looked at the distributions, you would see that AlphaZero has a little box around the mode; GFlowNets is the whole distribution. We know very well that diversity preservation is critical in order to discover interesting stepping stones in search problems. Now finally, Bengio has published results showing the GFlowNets converge exponentially faster than Markov chain Monte Carlo and PPO on some problems and finds more of the modes in the distribution function faster. Enjoy the show, folks.

Bengio 教授背景 Professor Bengio's Background

Host

Joshua Bengio 教授被公认为全球人工智能领域的顶尖专家之一,堪称深度学习之父。他在深度学习方面的开创性工作为他赢得了图灵奖,这是计算领域的诺贝尔奖。他是蒙特利尔大学的教授,也是 Mila 的创始人和科学主任,Mila 是一个拥有超过 900 名机器学习与 AI 研究人员的著名社区。他是全球被引用次数最多的计算机科学家之一。我无法用言语表达我们今天能进行这次对话是多么荣幸。Joshua 最近在 GFlowNet 上做了大量工作,这是一种强化学习配置下的主动学习框架,其核心是从现实世界中请求显著且多样化的训练数据,以最样本高效的方式增强我们的学习模型。现在我们试图最小化路径分布与奖励分布之间的散度,然后根据奖励分布采样路径。这与传统强化学习形成鲜明对比,传统方法试图最大化期望奖励。这种方法很可能发现多样化的策略,而不是贪婪地找到单一策略后迅速收敛。

Professor Joshua Bengio is recognized worldwide as one of the leading experts in artificial intelligence, indeed a godfather of deep learning. His pioneering work in deep learning earned him the Turing Award, which is the Nobel Prize of computing. He's a full professor at the University of Montreal and the founder and scientific director of Mila, which is a prestigious community of more than 900 researchers specializing in machine learning and AI. He's one of the most cited computer scientists on the planet. And I can't even begin to articulate how honored we are today to have this conversation. Joshua has done a lot of work recently on GFlowNets, which are an active learning framework in a reinforcement learning configuration where the name of the game is to request salient and diverse training data from the real world to augment our learned models in the most sample-efficient way possible. Now we're trying to minimize the divergence between the path distribution and the reward distribution and then sample paths according to the reward distribution. This is in stark contrast with traditional reinforcement learning where we're trying to maximize the expected reward. This approach is likely to find diverse strategies instead of being greedy and converging quickly after finding a single one.

GFlowNets 与熵估计 GFlowNets and entropy estimation

Host

学习,即最大化期望奖励,作为目标是误导性的,我们应该转而执行对未来路径的推理,平衡期望奖励和相对熵。这些想法之间有联系吗?我的意思是,GFlowNet 似乎在采样与奖励函数成比例的路径,从而保持与奖励函数本身一样多的熵。

Learning, which is to say maximizing expected reward, is the objective is misguided and we should instead perform inference over future paths balancing expected reward of relative entropy. Is there a connection between these ideas? I mean, it seems like GFlowNets are sampling paths proportional to the reward function that will maintain as much entropy as the reward function itself.

Yoshua

是的,完全正确。这是将奖励函数转化为能够采样等价对应分布的机制。所以我完全同意卡尔在这里说的。但正如我所说,有趣的是,原则上我们可以用 GFlowNet 做很多事情,比如我们已经做了数学推导和一些小规模实验,现在有几篇论文。我们可以做超出采样的事情,例如估计熵本身。熵是出了名的难以估计,我在关于 GFlowNet 的演讲中提到过,我们可以用 GFlowNet 机制来估计动作分布或贝叶斯参数分布的熵,如果你要在世界中采取行动并且你的世界模型存在不确定性,你会希望最小化这个熵。这与卡尔的兴趣很契合。你希望能够选择一个行动,最小化你对世界如何运作的不确定性。我们知道可能发生了什么最新的事情,其中重要的一部分是估计这些探索性行动的奖励,就像孩子们玩耍一样。通过那个行动,我对世界的知识熵会减少多少?所以你需要能够计算那个奖励,而那个奖励基本上是你关心的某个东西的熵。事实证明,你也可以用 GFlowNet 做到这一点。

Yes, yes exactly. It's a translation of the reward function into machinery that can sample the equivalent, the corresponding distribution. So yeah, I completely agree with what Karl was saying here. But as I said, what's interesting is we can do things with GFlowNets in principle, like we've done the math and some small-scale experiments that we have now a number of papers. We can do things that go beyond sampling, but for example estimate entropy itself. So entropy is notoriously difficult to estimate, and you know, I mentioned in my talks on GFlowNets that we can use the GFlowNet machinery to estimate entropy of say an action distribution or a distribution over Bayesian parameters, for example, which would be something you'd like to minimize if you're going to take an action in the world and you have a model of the world that has uncertainty. And that connects well with Karl's, for instance, interest. You'd like to be able to choose an action that minimizes your uncertainty about how the world works. We know what are the latest things that may have happened, and a good, you know, an important part of that is estimating the reward for these exploratory actions, like children playing around. How much reduction in entropy of my knowledge of the world am I going to get through that action? So you need to be able to compute that reward, and that reward is basically an entropy over something you care about. And it turns out you can also do that with GFlowNets.

向 Karl Friston 提问 Question for Karl Friston

Host

我们下周实际上还会和弗里斯顿再次对话。你有什么问题想让我们问他吗?

We're actually speaking with Friston again next week. Do you have a question that you would like us to put to him?

Yoshua

嗯,他比我更偏向生物学方面,我相信有极好的科学机会去探索 GFlowNet 提供的这种机制如何被大脑用来完成它的一些工作,即使用神经网络对世界的概率结构(包括不确定性)进行建模,这是他所关心的。但同时也要考虑高层认知、全局工作空间理论(我非常关心的东西)、注意力等,它们都适合 GFlowNet 的图景。所以我认为在计算和理论神经科学以及机器学习的协同作用中,利用 GFlowNet 提出的那种概率建模,有巨大的研究潜力,可以提出一些关于大脑概率性工作的解释性理论。而且,我认为他会是参与其中的绝佳人选。非常迷人。

Well, um, he, you know, he's on the biology side of things much more than I am, and I believe there are amazing scientific opportunities to explore how the kind of machinery that GFlowNets offer could be used by brains in order to do some of the things they do, using neural nets to model the probabilistic structure of the world including uncertainty, which is something he cares about. But also taking into consideration things like high-level cognition, the global workspace theory, which is something I care a lot about, attention, they all kind of fit in the picture of GFlowNets. So I think there's a huge potential of research at the synergy of computational and theoretical neuroscience and machine learning, probabilistic modeling of the kind that GFlowNets propose, to come up with some proposals for explanatory theories about what the brain does that's probabilistic. And you know, I think he would be a great person to be part of that. Fascinating.

GFlowNets 中的多样性与探索 Diversity and exploration in GFlowNets

Host

沿着这条线再深入一点,社区中有很多人是生物启发式机器智能方法的坚定倡导者。其中一个关键思想实际上是多样性的发现和保持,无论是在知识的获取还是表示方面。具体来说,进化算法的倡导者将自己与基于梯度的单智能体整体方法(如强化学习)区分开来,他们指出自己的方法克服了所谓的欺骗和搜索问题,也就是说它们不会陷入局部最小值。你的方法似乎在基于梯度的强化学习框架中实现了非常类似的效果。我的意思是,我不认为它们是互斥的,但你对此怎么看?

Going a little bit further down that line, there are folks in the community who are huge advocates of biologically inspired approaches to machine intelligence. And one of the key ideas actually is diversity discovery and preservation, both in how knowledge is acquired and represented. I mean specifically, evolutionary algorithm advocates they differentiate themselves from gradient-based single-agent monolithic approaches like reinforcement learning, and they point out that their approaches overcome so-called deception and search problems, which is to say they don't get stuck in local minima. Your approach seems to be achieving something very similar in the context of a gradient-based reinforcement learning package. I mean, I don't see it as being mutually exclusive, but what's your take on this?

Yoshua

是的,多样性在探索时很重要。人类,尤其是年轻人,是探索机器。他们试图理解世界如何运作,并通过行动获取信息。我同意,搜索过程需要对多样性有大的奖励,比如尝试不同的方式来实现好的结果,比如更好地理解世界如何运作。结果发现,在 GFlowNet 框架中,你有一个训练目标能产生这种多样性和探索,但它是基于端到端训练大型神经网络的。这与通常的端到端训练有点不同,因为我们没有一个目标;我们试图优化的目标实际上是不可解的。但我们可以采样这些轨迹,我将其视为采样思想。我们的思维过程经历了一些解释链;它不完整,也不代表所有解释。但我们发现,GFlowNet 的训练目标使得这些随机化的世界视图足以给执行实际工作的神经网络提供训练信号。

Yeah, diversity is important when you are exploring. And humans, especially young ones, are exploration machines. They're trying to understand how the world works and they're acting in the world in order to get that information. Yeah, I agree that that search process needs to have a big bonus on diversity, like on trying different ways of achieving something good, like better understanding how the world works. So it turns out that in the GFlowNet framework, you have a training objective that yields this kind of diversity and exploration, but is based on training large neural nets end to end. Now it's a bit different from the usual end-to-end training because we don't have an objective; the objective we're trying to optimize is not tractable actually. But we can sample these trajectories, which I think of like sampling thoughts. Our thought process is going through some chain of explanation; it's not complete and it doesn't represent all the explanations. But what we found with our training objectives for GFlowNets is that these sort of randomized views of the world are sufficient to give a training signal to the neural nets that do the real job.

与多臂老虎机的联系 Connection to multi-armed bandits

Host

我很好奇探索与利用之间的权衡。在我们的节目中,这个主题在很多语境下都出现过,尤其是当我们与多臂老虎机研究者交谈时。GFlowNet 似乎捕捉到了探索与利用之间的平衡,但多臂老虎机研究者深入研究了这种权衡,并有非常原则性和严谨的方法来分析它。你认为他们的研究在多大程度上可以应用于未来的 GFlowNet 变体?你认为它是否会为微调探索与利用之间的权衡提供更多选择?

I'm curious about this trade-off between exploration versus exploitation. It has come up in so many contexts throughout our show, and one in particular as we talked to multi-armed bandit folks. GFlowNets seem to capture this balance between exploration and exploitation, but the multi-armed bandit folks dive deep into this trade-off and have very principled and rigorous ways to analyze it. To what extent do you think that their research could be applied to future GFlowNet variations? Do you think it might open up more options to fine-tune the trade-off between exploration and exploitation?

Yoshua

是的,我的意思是,老虎机研究与 GFlowNet 这条线非常非常相关。我们一直在使用的 GFlowNet,例如在药物发现中,它们就是老虎机。只是动作空间不是 n 选一;它是组合的,因为你构建这些片段。所以动作空间无法枚举,因此你不能应用典型的老虎机算法。但很多数学是完全适用的。事实上,我们在药物发现设置中使用的是 UCB(上置信界)目标来学习一个好的探索策略。这源于老虎机研究。它所做的就是将风险和期望奖励量结合起来,从理论上保证你能高效地探索并找到奖励所在的所有可能位置。所以在 GFlowNet 论文中,你经常描述为我们不仅要采样最大奖励路径,还要有更多的多样性。

Yeah, I mean the bandit research is very, very closely related to the GFlowNet thread. The GFlowNets as we have been using them, for example for drug discovery, they are bandits. It's just that the action space is not one out of n things; it's combinatorial because you build these pieces. So the action space is not something you can enumerate, so you can't apply the typical bandit algorithms. But a lot of the math is totally applicable. In fact, what we use in the drug discovery setting is a UCB (upper confidence bound) objective to learn a good exploration policy. And that comes out of the bandit research. What it does is it combines the risk and expected reward quantities together in a way that in theory guarantees that you will do an efficient exploration and find where the reward is, in all of the possible places where you can get the reward. So in the GFlowNet papers, you often describe it as we want to sample not only the maximum reward path, but also have more diversity.

GFlowNets 中的探索与不确定性 Exploration and Uncertainty in GFlowNets

Host

也许去发现一些我们不知道的东西,如果我们只是追求最大奖励的话。这涉及到我们知道自己不知道的事情。我们可能知道这条轨迹看起来奖励较低,但最终可能奖励更高。然而,探索和强化学习也从根本上处理那些我们不知道自己不知道的事情,这就是随机探索之类的方法发挥作用的地方。你能谈谈你的看法吗?因为在我看来,如果我根据自己认为的奖励分布进行采样,我仍然会遇到一个问题:可能存在欺骗性奖励,我需要退一步,我可能不知道搜索空间的某些区域,这样我不是又遇到同样的问题了吗?

To maybe figure out something that we didn't know if we were just to go to the maximum reward, and that speaks a little bit to the things that we know that we don't know. We maybe know that this seems like a lower reward trajectory, it might turn out to be a higher reward trajectory. However, exploration and reinforcement learning is also fundamentally addressing the things about the things that I don't know that I don't know, which is where stuff like random exploration comes in. Could you maybe comment on how you see this? Because it seems to me that if I manage to sample according to what I think is the reward distribution, I still have this problem of maybe there is a deceptive reward, I need to take a step back, I may not know some area of the search space, and don't I just run into the same problems again?

Yoshua

这里的关键技巧是,你需要让你的奖励分布或奖励函数模型能够捕捉不确定性,也许是用贝叶斯方式。贝叶斯方法与 GFlowNet 框架很契合,因为我们可以把奖励函数的参数视为潜在变量。你实际上并不知道奖励函数,你正试图通过实验来弄清楚它。所以 GFlowNet 不仅可以采样你应该做什么来获取信息,还可以采样潜在的奖励函数。我们实际上并不了解世界中的奖励会是什么。经典强化学习正如你所说,取期望值并试图最大化它,而 GFlowNet 方法则试图尽可能多地获取关于底层奖励函数的知识,所以你在试图最小化不确定性。你的 GFlowNet 模型正在建模不确定性,然后你可以将其用作在现实世界中执行动作的策略的奖励。

The important trick here is you need your model of the reward distribution or the reward function to be one that captures uncertainty, maybe in a Bayesian way. The Bayesian way fits well with the GFlowNet framework because we can consider the parameters of the reward function as latent variables. You don't actually know the reward function; you're trying to figure it out from experiments. So the GFlowNet can sample not just what you should be doing in order to acquire information, but also potential reward functions. We don't actually have knowledge of what the rewards are going to be in the world. Classical RL goes as you said, taking the expected value and trying to maximize that, whereas the GFlowNet approach is trying to acquire as much knowledge as possible about the underlying reward function, so you're trying to minimize the uncertainty. Your model with the GFlowNet is modeling the uncertainty, and then you can use it as a reward for the policy that is going to do actions in the real world.

Host

好的,所以我们讨论的是不同的 GFlowNet。有一个 GFlowNet 建模从现实世界获得的奖励的不确定性,这就像一个贝叶斯模型。然后你有另一个 GFlowNet 控制搜索策略,它的奖励是通过这样做或那样做能减少多少不确定性。

Okay, so we're talking about different GFlowNets. There's a GFlowNet that models the uncertainty in the reward you're going to get from the real world, and that's like a Bayesian model. And then you have another GFlowNet that controls the policy that searches, and its reward is how much uncertainty reduction you're going to get by doing this or that.

Yoshua

是的,你需要让你的模型有一部分意识到世界上有整个区域是你不知道的,或者世界的某些方面是你不知道的,这样它才能驱动探索。

So yeah, you need to have a part of your model that is aware of the fact that there are whole areas in the world that you don't know about, or aspects of the world that you don't know about, so that it can drive the exploration.

GFlowNets 与结构学习的潜力 The Promise of GFlowNets and Structure Learning

Host

我很想知道其中的魔力来自哪里。GFlowNet 的承诺是,我们可以在路径分布中发现尽可能多的模式。传统上在马尔可夫链蒙特卡洛中,我们必须手动将先验知识注入算法,以高效地找到新模式或信息区域,尤其是当它们相距很远或不够尖锐时。GFlowNet 的假设是,这些模式的结构在许多问题上是可学习的,即使在高维空间中也是如此。这有点像说我们得到了一顿免费的午餐。我记得你用过这个确切的短语来描述我们在这里做的事情。许多研究方向试图开发通用方法来发现这些结构,但都失败了。你认为 GFlowNet 将如何克服这个看似棘手的难题?

I would love to know where some of the magic is coming from. The promise of GFlowNets is that we can discover as many modes as possible in the path distribution. Traditionally in Markov chain Monte Carlo, we had to hack priors into the algorithm by hand to find new modes or areas of information efficiently, especially when they are very far apart or not very sharp. The hypothesis of GFlowNets is that the structure of these modes is learnable on many problems, even in high dimensions. It's a little bit like saying we're getting a free lunch. I think you used that exact phrase to describe what we're doing here. Many research avenues have tried to develop general methods to discover these structures and have failed. How do you think GFlowNets will overcome this seemingly intractable curse?

Yoshua

不能保证它们一定能做到,因为如果你试图发现的底层函数——比如你关心的奖励函数或能量函数——没有结构,那么即使你访问了有限数量的模式,即奖励高的区域,它也不会告诉你其他好地方在哪里,其他模式在哪里。所以不能保证它会成功。但如果存在结构,那么就有免费的午餐。我们知道机器学习在这方面很擅长。过去十年深度学习的成功告诉我们什么?它告诉我们你可以泛化。这些网络并不完美,但它们可以泛化。所以你可以这样想:机器学习问题是,给定一些好事物的例子,即你获得奖励的地方,你可以泛化到其他地方。监督学习的思考方式是:给定一个候选位置,告诉我我认为会得到多少奖励。GFlowNet 采样器正在学习逆函数;它要采样,但这基本上是同一件事,只是方向相反:给我一些样本,一些奖励高的好地方。我们现在在设计强大的神经网络方面有很多经验,这些网络可以用于在我们通常使用 MCMC 的空间中进行泛化。如果存在允许泛化的规律性,那么所有这些都可以被利用。

There is no guarantee that they will, because if there is no structure in the underlying function you're trying to discover—say the reward function or the energy function that you care about—then having visited some finite number of modes, regions where your reward is high, it's not going to tell you anything about what the other good places are, the other modes. So there's no guarantee that it will work. But if there is structure, then there is a free lunch. And we know machine learning is good at that. The last 10 years of deep learning and its success—what is it telling us? It's telling us that you can generalize. These nets are not perfect, but they can generalize. So you can think of it like this: the machine learning problem is, given some examples of good things, places where you got reward, you can generalize to other places. The supervised learning way of thinking about it is: given a candidate place, tell me how much reward I think I would get. The GFlowNet sampler is learning the inverse function; it's going to sample, but it's kind of the same thing, just going in the other direction: give me some samples, some good places where the reward is high. We now have a lot of experience in designing powerful neural nets that can be leveraged to generalize in those spaces where we normally use MCMC. And if there are regularities that allow generalization, then all of that can be put to use.

游戏中 GFlowNets 与强化学习对比 GFlowNets vs. Reinforcement Learning in Games

Host

我们之前提到强化学习通常应用于有明确奖励函数的场景,比如下棋。我很好奇,如果我们将 GFlowNet 应用于国际象棋,假设会发生什么。我认为,鉴于像 AlphaZero 这样的强化学习专门训练来选择最佳走法而不是多样化走法,显然如果给 AlphaZero 和 FlowZero 相同的资源,AlphaZero 可能会击败 FlowZero。但是,我认为如果给 FlowZero 更多资源,并训练到相同的等级分,比如与 AlphaZero 相同的 Elo 等级分,那么 FlowZero 可能会下出更加多样化和有趣的棋局,风格更加丰富。我甚至认为,即使资源相同但足够高,一个假设的 FlowZero 可能会持续达到更高的等级分,因为它可能找到更有趣的垫脚石,有潜力避免欺骗,因为它可以探索看似奖励较低但最终发展成高奖励的路径。我很好奇你有什么想法。

We mentioned earlier reinforcement learning often being applied in a context where you have this kind of solid reward function, so let's say games, you know, playing chess. I'm really curious what would happen hypothetically if we applied GFlowNet to something like chess. I think given the fact that reinforcement learning like AlphaZero is trained specifically to choose the best move rather than diverse moves, it seems obvious that maybe if given equal resources to both AlphaZero and FlowZero, AlphaZero would probably beat FlowZero. However, I think if FlowZero were given more resources and trained to the same rating, say the same Elo rating as AlphaZero, it seems like FlowZero would play significantly more diverse and interesting games with a wider variety of styles. I think you could even imagine that it could be possible, even if given equal resources but sufficiently high enough resources, that a hypothetical FlowZero would consistently reach higher ratings because it might find more interesting stepping stones, have the potential to avoid deception because it can explore seemingly lower reward paths that ultimately develop into higher reward. More curious if you have any thoughts on that.

Yoshua

这是个好问题。我认为我们通过 GFlowNet 开创的方法真正发挥作用的地方在于,从学习者的角度考虑,它拥有有限的计算资源。原则上,如果你有无限的算力,并且知道奖励函数,比如国际象棋或围棋的规则,那么你可以直接计算出在所有可能情况下最优的策略。但现在你资源有限,有算力预算,你想高效地使用它。这就是探索与利用的权衡所在。GFlowNet 旨在探索多样化的高奖励模式,这可能有助于发现纯粹利用性方法可能错过的创新策略。因此,在有限资源的情况下,基于 GFlowNet 的方法确实可能找到更好的垫脚石,避免局部最优,从长远来看可能带来更高的性能。

It's a good question. I would say where the kind of approach we've been pioneering with GFlowNets might be really paying off is if you think about it from the perspective of the learner having a finite computational amount of resources. In principle, if you had infinite compute and you know the reward function like the rules of chess or Go, then you could just crank and find the policy that's best in every possible setting. Now if you have finite resources, you have a budget of compute, you'd like to use it efficiently. And that's where the exploration-exploitation trade-off comes in. GFlowNets are designed to explore diverse high-reward modes, which could be beneficial in discovering novel strategies that a purely exploitative method might miss. So in a finite resource setting, a GFlowNet-based approach might indeed find better stepping stones and avoid local optima, potentially leading to higher performance in the long run.

GFlowNets 用于主动学习与探索 GFlowNets for Active Learning and Exploration

Yoshua

权衡变得重要。如果你有一个当前策略,但不确定它是否正确,你试图说:‘我该怎么玩才能最大程度地改进我的策略,减少它选错东西的不确定性?’这正是使用 GFlowNet 有意义的地方。根据我们研究过的更简单的问题,我预计它会收敛得更快。如果你把游戏次数放在 x 轴,策略质量放在 y 轴,那就是你获得收益的地方。换句话说,学习曲线在渐近线上可能都有收益,但有趣的是学习速度。这里你需要主动学习:不只是试图获胜,而是收集信息以便未来赢得更多。这是一个不同的目标,需要多样性、探索、对自身不确定性的建模以及主动学习策略。

Trade-off becomes important. If you have a current policy that you're not completely sure is the right one, and you're trying to say, 'How should I play so that I'm going to improve my policy the most, reducing the uncertainty that it picks the right things?' That's where it makes sense to use GFlowNets. Based on the simpler problems we've looked at, I would expect it to converge faster. If you look at the number of games on the x-axis and how good your policy is on the y-axis, that's where you would gain. In other words, the learning curve might gain asymptotically, but the interesting part is how fast you learn. Here you want active learning: not just trying to win, but gathering information to win more in the future. That's a different objective, requiring diversity, exploration, a model of your own uncertainty, and an active learning policy.

Host

你认为这在多大程度上可以不仅用于奖励最大化,还用于信息收集?例如,在大脑中,可能有一个类似的过程:我还需要检索什么才能回答问题?或者像谷歌这样的搜索引擎,主动做多件事来回答查询:‘这够了吗?这够了吗?’你看到与这些事情的关联了吗,还是它们本质不同,因为它们可能不是即时学习?高效获取信息正是主动学习思维发挥作用的地方。我认为这在部署 AI 对话系统时是一个非常大的实际问题,这些系统不只是闲聊,而是帮助用户获取信息。商业世界和搜索引擎中有巨大的需求。我们目前没有能做到这一点的算法,而且人类必须驱动这一点很痛苦。如果我们有系统能够明确建模自己对用户需求或信息查找位置的不确定性,并且我们需要强大的模型——不仅仅是简单的高斯分布——这就是 GFlowNet 的优势所在。它们可以表示组合对象上的非常复杂的分布。那么我们就可以得到更高效的人机界面。同样的方法可以用于科学发现:科学家计划实验以减少对其理论的不确定性。这是同一个问题:你有一系列问题要问自然,你试图尽可能少地问,以尽快理解发生了什么。

How much do you think this could be part of not only reward maximization but also information collection? For example, in the brain, there might be a similar process: what do I still need to retrieve to answer questions? Or in a search engine like Google, actively doing multiple things to answer a query: 'Is this enough? Is this enough?' Do you see connections to these types of things, or are they inherently different because they might not be learning on the spot? Acquiring information efficiently is where active learning thinking comes in. I think it's a very big practical problem in deploying AI dialogue systems that are not just chit-chat but help users get information. There's a huge need in the business world and search engines. We don't have algorithms that do that right now, and it's painful that the human has to drive. If we had systems that could explicitly model their own uncertainty about what the user needs or where to find information, and we need powerful models for that—not just simple Gaussians—that's where GFlowNets' strength comes in. They can represent very complex distributions over compositional objects. Then we could get much more efficient human-machine interfaces. The same methodology could be used in scientific discovery: scientists plan experiments to reduce uncertainty on their theories. It's the same problem: you have a series of questions to ask nature, and you try to ask as few as possible to quickly understand what's going on.

Yoshua

这与因果关系有根本联系吗?我想到了你与深入研究因果关系的人合作的一些论文。智能体能否通过提出这样的问题来揭示世界的基本因果结构?这会是机器学习和因果关系的统一吗?你问的问题都很好。我从事 GFlowNet 研究项目的主要动机之一是,我认为它是实现我在演讲中所谓的‘系统 2 归纳偏置’的理想工具。我们从神经科学和认知科学中知道很多关于我们如何思考的东西,我们可以将其带入基于深度学习的概率机器学习中。其中一个归纳偏置是我们因果地思考:我们不断问‘为什么’的问题,试图找到解释。这与经典 AI 相关——规则、逻辑、推理——我们还没有将其整合到深度学习中。GFlowNet 为我们提供了一个极好的切入点,因为它们非常擅长表示分布并在图上采样。推理或一组可能的推理来解释某事,或规划——这些都是图。你的想法可以被视为图。例如,一个句子的语义和句法分析是一个图,通常不止是树,还有包括知识图谱在内的连接。隐式表示这些分布并采样其中的片段作为想法的能力是我们思考的基础。回到因果关系,GFlowNet 可以帮助解决的一个难题是因果发现:在给定观察的情况下,世界的潜在因果结构是什么,包括不确定性?许多因果关系研究假设我们观察随机变量并进行推断,但在大量变量中发现因果图要困难得多,当学习者看到低级像素并必须找出因果变量及其关系时,甚至更难。

Is there a connection fundamentally to causality? I'm thinking of papers you've collaborated on with people deep into causality research. Could an agent learn to uncover the fundamental causal structure of the world by asking such questions? Could this be the unification of machine learning and causality? You're asking all the right questions. One of my main motivations for the GFlowNet research program is that I think it's an ideal tool for implementing what I call 'System 2 inductive biases' in my talks. There are lots of things we know from neuroscience and cognitive science about how we think, and we can bring that into probabilistic machine learning based on deep learning. One of the inductive biases is that we think causally: we constantly ask 'why' questions, try to find explanations. That connects with classical AI—rules, logic, reasoning—which we haven't yet integrated into deep learning. GFlowNets give us an amazing handle on this because they are really good at representing distributions and sampling over graphs. Reasoning or a set of possible reasoning to explain something, or planning—these are graphs. Your thoughts can be seen as graphs. For example, a semantic and syntactic parse of a sentence is a graph, often more than a tree, with connections including knowledge graphs. The ability to implicitly represent those distributions and sample pieces of them as thoughts is fundamental to how we think. Going back to causality, one of the hard questions GFlowNets can help with is causal discovery: what is the underlying causal structure of the world, including uncertainty, given observations? Much causality research assumes we observe random variables and make inferences, but it's much harder to discover the causal graph in a large set of variables, and even harder when the learner sees low-level pixels and must figure out the causal variables and their relations.

因果世界模型与抽象 Causal World Models and Abstraction

Host

从因果的角度看,我认为 GFlowNets 可以帮助我们做到这一点。这打开了太多问题方向,可能本身就能做一期节目。但让我先问几个基本问题。你提到学习因果结构是一个更难的问题。第一个问题就是如何表示因果关系。你提到了图——图是一种方式,当然你可以发展出同构的方式来将逻辑的某些部分表示为图,取决于图结构的丰富程度。但还有一个问题:当你试图构建——我认为称之为世界模型是合适的,对吧?我们试图构建一个积极的——这是我用的词。好的。所以关于这一点我有一个快速问题。对一些人来说,世界模型只是判别函数——仅仅是给定 x 下 y 的概率。对我来说,它更一般;它还包括 x 的结构。这也是你的观点吗?

Causally, and I think GFlowNets can help us do that. This opens up so many avenues of questions. I think it would probably almost be a future episode in itself, but let me just ask you about some of the basic ones. As you mentioned, learning the causal structure is a much more difficult problem. The first question is just how to represent causality. You mentioned graphs—graphs is one way, and of course you can develop isomorphic ways of representing certain parts of logic as graphs, depending on how rich you make the graph structure. But there's also the issue of when you're trying to build—and I think it's probably correct to call this a world model, right? We're trying to build a positive—that's the word I use. Okay, great. So I have one quick question about that. To some people, world model is only the discriminative function—it's just the probability of y given x. To me, it's more general; it's also the structure of x. Is that also your view?

Yoshua

是的,没错。

Yes, okay.

Host

那么在构建这些世界模型时,一些偏向判别式方法的人对生成式技术的质疑是:'你看,你要构建这个生成模型,它会比判别模型更复杂,因为它还必须学习 x 的结构。'但我认为这里可能的免费午餐是你可以学习 x 上的抽象结构。如果你学习这些抽象世界模型,丢弃所有不重要的细节,你就有可能拥有非常强大的预测编码。你怎么看?

And so in constructing those world models, some of the pushback on generative techniques from folks more skewed towards the discriminative side is: 'Hey look, you're going to try and build this generative model; it's going to be even more complicated than the discriminative model because it also has to learn the structure on x.' But I think the possible free lunch here is that you can learn abstract structure on x. If you learn these abstract world models, throwing away all the nitty-gritty that doesn't really matter, you can potentially have very powerful predictive encoding. What are your thoughts on that?

Yoshua

哦,这正是我近 20 年来一直在思考的。这也是我对深度学习感兴趣的原因之一,因为它是一种发现抽象表示的方法。从深度学习的早期,比如 2005 年左右,到我和 Ian Goodfellow 合写的论文,以及 2010 年左右我和蒙特利尔大学的一些同事合写的关于深度学习的论文,都围绕着这个想法:我们希望这些学习过程能够发现我们所谓的抽象因子。但现在我认为,不仅仅是因子(变量),更重要的是它们之间的相互关系,在语言的情况下,我们称之为因果机制。这里有一个基本的思考方式:如果你不引入世界中存在的抽象结构,那么表示输入分布 p(x) 就非常困难。换句话说,你需要大量数据来学习它,而且它不会很好地泛化。抽象的全部意义在于它赋予你非常强大的能力,可以泛化到新的环境,包括分布外的情况,这是目前机器学习最热门的话题之一:我们如何扩展我们的方法,使其在新环境中也能很好地泛化?从因果的角度思考,这些抽象的因果依赖关系是在分布变化中保持不变的东西——比如如果我去了月球,物理定律是一样的,但分布非常不同。我如何跨越这种分布变化进行泛化?这是因为学习者,如果我们接受了正确的教育,已经弄清楚了潜在的因果机制,至少足够多,以至于我们可以被传送到一个不同的世界,那里有相同的物理定律,我们可以预测会发生什么,尽管它看起来与我们的训练环境完全不同。所以抽象的想法实际上是,如果你引入抽象,数据的描述长度会大大减小,这就是你获得泛化的原因。

Oh, that's what I've been thinking for almost 20 years. It's one of the reasons why I've been interested in deep learning as a way to discover abstract representations. From the early days of deep learning, like mid-2005 or so, and in the paper that Ian Goodfellow and I wrote, and other papers I wrote with some of my colleagues at the University of Montreal on deep learning around 2010, they are all about that notion: we would like these learning procedures to discover these abstract factors, as we call them. But now I think it's not just the factors—the variables—but also, more importantly, how they are related to each other, which in the case of language is what we call causal mechanisms. Here's a fundamental way of thinking about this: if you don't introduce the abstract kind of structure that exists in the world, then representing p(x), the input distribution, is very difficult. In other words, you'll need a lot of data to learn it, and it's not going to generalize very well. The whole point of abstraction is that it gives you very powerful abilities to generalize to new settings, including out-of-distribution, which is one of the hottest topics in machine learning right now: how do we extend what we do so that it generalizes well in new settings? Thinking causally about these abstract causal dependencies as the things that are preserved across changes in distribution—like if I go to the moon, it's the same laws of physics but the distribution is very different. How do I generalize across such changes in distribution? It's because the learner, if we had the right education, has figured out the underlying causal mechanisms, at least enough of them, so that we can be transported to a different world where it's the same laws of physics and we can predict what's going to happen even though it looks completely different from our training environment. So the idea of abstraction is really that if you introduce abstractions, the description length of the data becomes way smaller, and that's why you get generalization.

Host

绝对。我对这些抽象类别非常着迷。我认为这是人工智能中最令人兴奋的事情。Douglas Hofstadter 谈到过认知类别,比如用'酸葡萄'这个概念来代表某种事物。几乎神奇的是,我们的大脑似乎会安排这些认知类别,我不完全清楚它们是涌现现象还是其他过程。你在 GFlowNets 中发现的那些模式就是某种类别。这些认知类别是抽象,还有因果关系和几何深度学习等。但我一直有一种直觉,深度学习并不能自己学习类别;它需要人类将先验知识放入模型中,就像我们在几何深度学习中所做的那样。你认为这会一直如此,还是我们可以拥有那种元层次的学习?

Absolutely. I'm fascinated by these abstract categories. I think it's the most exciting thing in AI. Douglas Hofstadter spoke about cognitive categories, like the concept of 'sour grapes' to represent a certain thing. Almost magically, our brain seems to arrange these cognitive categories, and it's not entirely clear to me whether they're an emergent phenomenon or some other process. The modes you're discovering in GFlowNets are kind of categories. These cognitive categories are abstractions, also things like causality and geometric deep learning. But I've always had this intuition that deep learning doesn't learn the categories on its own; it needs humans to put priors into the model, as we do with geometric deep learning. Do you think that will always be the case, or can we have that meta-level of learning?

Yoshua

是的,我真正想做的是构建能够发现自己的语义类别的机器——那些真正帮助它们理解世界的抽象类别。当然,如果我们帮助它们,它们会学得更好更快,就像我们教孩子一样;我们不会让他们自己发现世界。但我们确实有能力发明新类别——科学家们一直在这样做,对吧?或者艺术家、作家、哲学家、学者,以及找到问题新解决方案的普通人。我们一直在这样做。我们的大脑是一个发现新抽象的机器。当然,通常这只是在我们从文化输入中获得的所有东西之上的一点点,但这是我们现在在机器学习中不具备的能力,而这将是一个巨大的优势。所以现在我们不是在强化学习,也不是在主动学习,我们在谈论无监督学习。我们在谈论机器如何发现这些通常离散的概念,这些概念在某种程度上帮助它理解。换句话说,构建对许多事物的紧凑理解,这些理解可以跨多种设置泛化。是的,随着我在 GFlowNets 上的推进,这条道路在我心中变得越来越坚定。作为一个线索,我们最近有一篇论文,我想是在 ICLR 上,关于'离散值神经通信',它与全局工作空间理论有关。这里有一个有趣的直觉:如果你限制不同模块之间的通信——比如在大脑或机器学习系统中——使用尽可能少的比特,而离散是实现极少比特的方式,你可以获得更好的泛化。这有很好的理由。

Yes, what I really want to do is build machines that can discover their own semantic categories—abstract ones that really help them understand the world. And of course, they're going to learn better and faster if we help them, just like we teach kids; we don't let them discover the world by themselves. But we do have an ability to invent new categories—that's what scientists do all the time, right? Or artists, writers, philosophers, scholars, and ordinary people who find new solutions to problems. We do that all the time. Our brain is a machine that discovers new abstractions. Of course, usually it's just a little bit on top of all the things we got from our cultural input, but that's the ability that we don't have right now in machine learning, and that is going to be a huge advantage. So now we're not in reinforcement learning, we're not in active learning, we're talking about unsupervised learning. We're talking about how can a machine discover these often discrete concepts that somehow help it understand. In other words, build a compact understanding of lots of things that generalize across many settings. And yes, that path to build that is becoming more and more firm in my mind as I move forward with GFlowNets. As a clue, there was a paper we had recently, I think in ICLR, on 'Discrete-Valued Neural Communication' that's connected to the global workspace theory. One interesting intuition here is that if you constrain the communication between different modules—say in the brain or in a machine learning system—to use as few bits as possible, and discrete is the way to get very few bits, you can get better generalization. There are good reasons for that.

意识与大型神经网络 Consciousness and Large Neural Networks

Host

关于意识这个话题,我们不得不问你这个问题。OpenAI 的 Ilya Sutskever 发推文说:‘今天的大型神经网络可能稍微有点意识。’这引起了不小的风波。你怎么看这个说法?

On the topic of consciousness, we would be remiss not to ask you this. Ilya Sutskever of OpenAI tweeted: 'It may be that today's large neural networks are slightly conscious.' This has caused quite a storm. What do you make of this statement?

Yoshua

这类说法的根本问题在于我们不知道意识到底是什么。我们需要保持谦逊。我无法判断 Ilya 是对是错。我认为意识远比这些大型语言模型所拥有的要复杂得多,差距很大。我们需要与神经科学家和哲学家合作。我论文《意识先验》的标题有点随意,之后我学到了很多。我们有很多不了解的地方,但从神经科学中我们已知足够多的东西,可以启发我们构建具有类似意识处理机制的机器学习系统。为了减少争议,我们称其为‘意识处理机制’。‘意识’这个词在科学界长期是禁忌,但现在神经科学在测量大脑在有意识和无意识过程中的活动方面取得了进展。我宁愿探索假说,也不愿对当前神经网络是否有意识做出大胆断言。

The fundamental problem with such statements is that we don't know what consciousness really is. We need humility here. I can't say whether Ilya is right or wrong. I think there's more to consciousness than what we have in these large language models, by a big gap. We need to work with neuroscientists and philosophers. I was a bit liberal with the title of my paper 'The Consciousness Prior' and have learned a lot since. There's much we don't understand, but we know enough from neuroscience to inspire building machine learning systems with conscious-like processing. Let's say 'conscious processing machinery' to be less controversial. The word 'consciousness' has been taboo in science, but neuroscience is now making progress measuring brain activity during conscious vs. unconscious processes. I'd rather explore hypotheses than make bold claims about current neural nets being conscious.

Host

下个月 David Chalmers 会来我们节目。你有什么问题想问他吗?

We have David Chalmers on the show next month. Do you have any questions for him?

Yoshua

我喜欢 Michael Graziano 关于意识的假说,它试图用科学方式解释感受质和主观体验。前提是:不要从哲学扶手椅上试图弄清楚意识是什么,而是把它当作大脑中发生的现象。当人们报告主观体验时,我们可以测量大脑活动。我们能提出理论解释为什么我们感觉有主观体验吗?这为科学探究打开了大门。Graziano 的理论基于我们有一个世界模型,注意力聚焦于其中一部分。我们需要一个微型世界模型来控制注意力,从而在真实知识和抽象控制机制之间产生分离。这可能给我们带来笛卡尔二元论的错觉,我认为这是一种错觉,但必须植根于生物学现实。我想听听 Chalmers 对这种研究计划的看法。

I like Michael Graziano's hypothesis about consciousness, which tries to explain qualia and subjective experience in a scientific way. The premise is: let's not try to figure out what consciousness is from a philosophical armchair, but treat it as a phenomenon happening in the brain. We can measure brain activity while people report subjective experiences. Can we come up with theories explaining why we feel we have subjective experience? This opens the door for scientific investigation. Graziano's theory is rooted in the idea that we have a world model, and attention focuses on parts of it. We need a mini world model to control attention, creating a separation between real knowledge and abstract control machinery. This could give us the illusion of Cartesian dualism, which I think is an illusion but must be grounded in biological reality. I'd like to hear what Chalmers thinks about such a research program.

Host

谢谢。我有一个具体问题,基于你最近在半监督学习中关于插值一致性训练的工作。它通过未标记样本和插值伪标签之间的 mixup 强制线性,取得了显著改进。此外,ReLU 是分段线性的,在神经网络中占主导地位。你能评论一下吗?

Thank you. I have a nitty-gritty question based on your recent work on interpolation consistency training in semi-supervised learning. It found significant improvements by forcing linearity via mixup between unlabeled samples and interpolated fake labels. Also, ReLUs are piecewise linear and dominant in neural networks. Can you comment?

分段线性近似与抽象 On Piecewise Linear Approximation and Abstraction

Host

在环境空间中的线性细胞,它们被输入样本激活或关闭。我的问题是:为什么线性,无论是分段还是其他形式,主导了近似方法的最新进展?在我看来,这有点像回到了未来,如果你愿意这么说的话,我们似乎放弃了更平滑的非线性方法,回到了更新但更复杂的线性近似形式。

Of linear cells in the ambient space and they're activated turned off or on by input examples. So my question is: why is linearity, whether it's piecewise or otherwise, dominating the state-of-the-art in approximation methods? It almost seems to me like we've kind of gone back to the future, if you will, sort of leaving behind attempts at more smooth non-linear methods and gone back to newer, albeit more complicated, forms of linear approximation.

Yoshua

没错。我会说,大致线性的事物更简单,对吧?所以有一个正则化器,要求你尽可能大致线性或局部线性,这是一种平滑先验,对吧?这有助于泛化,但如果太强也会有害。因此,采用这种分段线性的解决方案是一个很好的折中。它要求尽可能少的片段,并且理想情况下以组合方式组织,这样它就不只是 ReLU,更像是离散的抽象逻辑,推理的东西位于顶部控制这些片段,但每个片段本身相当简单,比如是线性的。所以看待这个问题的一种方式是,看看经典 AI 研究人员使用的规则,每条规则都相当简单,几乎是线性的或非常简单的逻辑,但正是所有这些规则的组合赋予了这些系统表达力。当然,问题在于他们不知道如何正确训练它们。但我认为我们学会了用这些离散的方式将事物分解成更简单的部分,事实上,如果你有耐心,它就会自然出现,而且这些假设非常非常弱。所以从某种意义上说,这几乎是分段抽象。所以我们真的有点回到了——是的,我倾向于这么说,而不是分段线性。但线性当然是其中的重要部分,它是获得简单性的一种简单方式。

Right. I would say something that's roughly linear is simpler, right? So having a regularizer that says you want to be roughly linear or locally linear, at least to as much extent as you can, is a smoothness prior, right? So that's going to help generalization, but it could also hurt if that is too strong. And so having this piecewise linear kind of more type of solution is a good compromise. It says as few pieces as possible, and ideally organized in a compositional way, so that it's not just like a ReLU, it's more like the discrete abstract logic, reasoning things sitting on top that's controlling the pieces, but otherwise fairly simple in each how each of the pieces are, you know, like linear for example. So one way to look at this is if you look at the kind of rules that classical AI researchers were using, each rule is fairly simple, it's almost linear or it's very simple logic, but it's the composition of all those rules that gives the power of expression of these systems. Of course the problem then is that they didn't know how to train them properly. But I think we learn to come up with these discrete ways of breaking up things into simpler pieces, and in fact I think if you're patient about it, it just comes out naturally and they're very very weak assumptions. So in a way, it's almost piece-wise abstraction. So we're really kind of back—yes, that's what I would lean to rather than piecewise linear. But linear of course is a broad part of it, it's an easy way to get simple.

个人研究历程与 Gary Marcus 共识 Personal Research Journey and Convergence with Gary Marcus

Host

太棒了,Bengio 教授。我对您的个人历程很感兴趣。我们一直在讨论不同的轨迹,我想知道您过去十年的研究轨迹。我的一位朋友,心理学家和符号学家 Gary Marcus 教授,顺便说一句,他大概是您最好的朋友之一,他在 2012 年《纽约客》的一篇文章中指出,MLP 缺乏表示因果关系的方式,比如疾病和症状之间的关系。我认为这近年来一直是您关注的重点,正如我们讨论过的。他当时认为您有点过于“系统一”了,他谈到了需要异构架构以及获取抽象概念、组合性和外推能力,我认为这也是您过去十年左右的重点。我们非常喜欢看您和 Marcus 的辩论,顺便说一句,我们很乐意主持第二场辩论,如果您感兴趣,请告诉我们,我们会安排。但他常被视为异端,而且,暂且不谈符号与神经网络之争,我是否正确地认为你们在某种程度上已经趋同了?您如何从自己的角度描述这一点?

Amazing, Professor Bengio. I'm interested in your personal journey. So we've been talking about diverse trajectories, and I wanted to know about your own trajectory of research over the last 10 years. Now, one of my mates, a psychologist and symbolist professor Gary Marcus, presumably one of your best friends, by the way, he pointed out in his 2012 New Yorker article that MLPs lacked ways of representing causal relationships, such as between diseases and their symptoms. And I think this has been a significant focus of yours in recent years, as we've discussed. And he thought at the time that you were a bit too, quote, system one all the way, and he spoke then about the need for heterogeneous architectures and the acquisition of abstract concepts, compositionality, and extrapolation, which I think has also been a huge focus of yours in the last decade or so. We've really enjoyed watching your debate with Marcus, and by the way, we would love to host v2 of that debate, so if you're interested, you just let us know, we'll do that. But he's often viewed as a heretic, and you know, just forgetting about symbols versus neural networks for a minute, am I right in thinking that you've converged in at least some ways in your thinking, and how would you characterize that from your perspective?

Yoshua

是的,我在 90 年代是一个纯粹的神经网络亚符号连接主义研究者。我读研究生时研究神经网络,当时主流思想是那些经典的基于规则的 AI 系统,完全没有学习。而且它是主导的,这意味着像 Jan、Jeff 和我以及其他持不同想法的人组成的小团体不得不捍卫我们的观点,这可能导致了某种“我们 vs 他们”的思维方式,我认为这是不健康的。当然,我成熟了。其中一个重要的——我认为这段旅程中有几个转折点。嗯,其中之一是在 2000 年代,我意识到了抽象的重要性。更具体地思考这个问题,因为抽象意味着什么?我在想,我们希望在无监督深度网络的顶层拥有什么样的表示?因为那个十年我们主要做无监督深度网络,比如深度玻尔兹曼机等。我在想,应该是像词一样的东西,对吧?就像我们在顶层操作的那种概念。嗯,是词或者可能经过消歧的等价物,但当时我们似乎没有合适的工具。然后它一直是一个目标。接着在 2014 年,我们发现了注意力机制的力量,现在它与抽象紧密相关,因为它专注于少数事物。当然,这非常符合我们思考的特征:一个思想包含很少的元素,这意味着我们选择了这些元素,这就是注意力的作用。所以它更接近于构建像人类一样思考的机器这一理想。然后当然在 2017 年,我写了那篇关于意识先验的论文,在那里我发现了所有关于全局工作空间理论的工作,势头就起来了。当然现在,人类思考并使用符号,理解它们之间非常抽象的关系,我们需要构建能够做到这一点的神经网络。所以我想我与 Gary 的分歧,也许他也改变了,在于这将是神经网络来完成,对吧?只是我们要以特殊的方式训练它们,这正是 GFlowNets 的目标。

So yeah, I used to be in the 90s a pure neural net sub-symbolic connectionist researcher. I did my grad studies at a time on neural nets, at a time when the dominant way of thinking was these classical AI rule-based systems with no learning at all. And it was dominant, meaning that the little group like Jan and Jeff and I and others who were thinking otherwise had to defend our views, and maybe that led to a kind of us versus them, I think unhealthy, way of thinking. And of course I matured. And one of the big—so I think there are several turning points on that journey. Well, one of them in the 2000s was the realization of the importance of abstraction. And the way to think about this maybe more concretely, because what does it mean to be abstract? I was thinking, well, what would be the right kind of representation we want to have at the top level of our unsupervised deep nets, because we were doing mostly unsupervised deep nets like deep Boltzmann machines and stuff in that decade. And I was thinking, well, it would be things like words, right? Things like the sort of concepts that we manipulate at the top level. Well, it's words or the equivalent maybe with disambiguated, but yeah, it didn't seem that we had the right tools for that. And then it remained like an objective. And then in 2014, we discovered the power of attention, and now it's closely connected to abstraction because what it does is it focuses on a few things. And of course that's very much a characteristic of how we think: a thought has very few elements in it, that means we have selected those elements, and that's where attention comes in. So it's getting closer to this ideal of building machines that think like humans. And then of course in 2017, I wrote this consciousness prior paper where I discovered all the work on global workspace theory, and the momentum built up. And of course now, humans think and they use symbols and they understand the very abstract relationships between them, and we need to build neural nets that can do that. So I guess where I've maybe departed from Gary, but maybe he's moved too, is it's gonna be neural nets that do it, right? It's just that we're gonna be training them in a special way, and that's what GFlowNets really aiming at.

致谢与结束语 Appreciation and Closing Remarks

Host

我能说一句吗,我们问过很多嘉宾关于他们演变的问题,有时他们的回答会比其他人更尖锐。但我必须说,从我的角度来看,您的回答是我们迄今为止听到的类似问题中最有信息量、最优雅、最高尚的。所以向您致敬,太棒了。

Can I just say, we asked many guests these questions about their evolution, and sometimes they tend to be spicier than others. But I have to say, from my perspective, your answer was the most informative, the most gracious, and the most noble of answers that we've heard so far to similar questions. So kudos to you, that was awesome.

Yoshua

谢谢,我简直不敢相信。

Thanks, I just cannot believe it.

Host

我们总是做大量的准备,但现在我们已经知道我们大概只能问六个问题左右。所以我们有点像指数级地,你知道,对我们的问题有一个指数先验。

And we always do a hell of a lot of preparation, but it's got to the point now where we know that we're not going to get more than about six questions in. So we kind of like exponentially, you know, have an exponential prior on our questions.

Host

不过他很棒,你知道,我们请他给出相对三分钟左右的回答,他坚持做到了,这真的很酷。我的意思是,这对进行有趣的对话非常有帮助。我简直不敢相信我有多自豪,你知道,他欣赏我们投入了准备时间,并且提出了不错的问题,希望这些问题对他和我们的观众来说都很有趣。

Well, he was awesome though, with like you know, we asked him to give relatively sort of three-minute answers, and he stuck to that, which was really cool. I mean, that's very helpful to have an interesting dialogue. And I can't believe how proud I am, you know, that he appreciates that we put the prep time into it, and you know, had decent questions that were hopefully interesting for him as well as for our audience.

Host

那么 Kilcher 博士,光速文化,你怎么看?很酷,他,我是说,他嗯,是的,我认为他的……

So Dr. Kilcher, lightspeed culture, what's your take? It's cool, he's, I mean, he's um, yeah, I think his is the...

关于意识主张的讨论 Discussion on consciousness claims

Host

他最后提到的那件事,比如他的谦逊,在他所做和所答的一切中都闪耀着光芒。他就像在说,‘你知道,这是我能给出的最好答案’,但他似乎非常开放,并不那么……嗯。有人注意到他不在推特上。就像……不,那是个绝妙的问题。我觉得我们应该把那个问题发到我们的推特上,因为你知道,对于一年后看这个视频的人来说,那可能已经被遗忘了。但是,是的,OpenAI 的那个 Ilya 说模型可能有点意识。我对此感到恼火,因为我看了他在 Lex 上的采访,我知道……但如果说他的坏话,他永远不会来我们的播客。不过我觉得他本来也不会来,所以没关系。但我觉得这很糟糕。你怎么看?为什么?是啊,为什么?

The thing he mentioned at the end, like his humility, it kind of shines through everything he does and he answers. He's like, 'You know, here's the best answer I can give,' but he seems to be very open and not very... yeah. One notices he's not on Twitter. It's like... no, that was a brilliant question. I think we should post that question on our Twitter, because you know, there's that bit for people watching this in a year's time—it's probably forgotten about. But yeah, that Ilya guy from OpenAI said that the models might be slightly conscious. I was exasperated by that, because I watched his interview on Lex, and I know... but by saying bad things about him, he will never come on our podcast. But I don't think he would have done anyway, so it doesn't matter. But yeah, I think that's pretty bad. What do you say? Why exactly? Yeah, why?

Host

到底为什么?是啊,为什么?

Why exactly? Yeah, why?

Host

就像,‘哦,你不觉得这很糟糕吗?’不,他说……我认为因为 OpenAI 的很多人,你知道,就像理性主义者社区的人,他们真的相信我们面临 AI 接管世界、我们变成回形针的迫在眉睫的威胁。而且我认为接下来……我听了他在 Lex 上的采访,他听起来像个推销员,在谈论 Codex 以及它将如何彻底改变一切。老实说,我认为他们说的和现实之间存在着巨大的分歧。

It's like, 'Oh, you don't think it's bad?' No, he says... I think because a lot of the folks at OpenAI, they are, you know, like in the rationalist community, and they seriously believe that we're an imminent threat of the AI taking over the world and us being paper clips. And I think it's next... I listened to his interview on Lex and he sounded like a salesman talking about Codex and how it was going to revolutionize everything. And I honestly think that there's just such a divergence between what they're saying and reality right now.

Host

好吧,不要偏离我们今天的嘉宾太远,但我认为那只是种随想,你知道吗?就像,‘你知道,今天的大型神经网络可能有点意识’,对吧?然后你就想,‘是啊,嗯……’这确实是个随想,就像在推特上,你随便发的一条推文,它引出了有趣的问题。就像他提出的有趣问题:你是一团神经元,你只是一堆拼凑起来的物质,对吧?你有意识。所以显然,学习系统与数据结合——或者甚至不结合数据——会产生意识。那么为什么另一个在硅基中拼凑起来的、摄入数据的系统就不能有点意识或具有一些类似的属性呢?这基本上就是……

Well, not to drift too far away from our guest today, but I thought it was just kind of a shower thought, you know? Like, 'You know, the large neural networks of today might be a little bit conscious,' right? And you're just like, 'Yeah, well...' And it is a shower thought, like it's on Twitter, it's just something you tweet out, and it brings up interesting questions. Like he brings up interesting questions: you're a ball of neurons, you're just a slapped-together piece of matter, right? You have consciousness. So clearly something about learning systems combined with data—or maybe not even combined with data—gives rise to consciousness. So why can't another in silico slapped-together system of neurons ingested with data be slightly conscious or have some properties like that? And that's essentially...

Host

是的,Benjamin 拒绝给出——他谦逊的性格——他拒绝给出那个强烈的观点。但那会……因为可能只是我——这是我的看法,显然不是他的——但读了《意识先验》论文,这并不太遥远。他将意识表述为具有这些元素:我有我的内部状态,那是我大脑中所有我能带到前台的东西;然后我从外部世界获得一些输入,通过输入我进行过滤,就像用注意力机制一样,我看看我脑海中现在能聚焦什么,对吧?这是通过类似注意力机制的东西实现的。然后我把那个东西放到这些抽象概念中;我把脑海中的概念表示为一个稀疏因子图,通过聚焦其中的部分,我可以在稀疏因子图中进行推理,等等。现在显然像 GPT-3 这样的东西并不完全具备这些,至少不是显式的,但其中一些是存在的,对吧?它有一个输入,有大量的权重,所以它用注意力机制来看到它能聚焦什么。

Yeah, Benjamin refused to give, like, the humble person he is, he refused to give the strong take on that. But that would have... because I might just—this is my opinion, not his obviously—but reading the Consciousness Prior paper, it is not too far off. He formulates consciousness as having these elements: I have my internal state, which is sort of everything in my brain that I could bring up into my forefront; then I get some input from the outside world, and through the input I then filter, like with an attention mechanism, I look at what in my mind I could now bring into focus, right? And that is by use of something like an attention mechanism. And then I take that thing and I put it into these abstract concepts; I represent the concepts in my head as a sparse factor graph, and by focusing on parts of that I can then make inferences in this sparse factor graph, and so on. Now obviously something like GPT-3 doesn't have all of that, at least not explicitly, but some of it is there, right? It's: I have a piece of input, I have a giant amount of weights, so I use an attention mechanism to sort of see what I can focus on.

Host

是啊,但我觉得那是对意识的一种非常陈述性的描述,而其根源是关于现象学体验的,对吧?我知道我们讨论过计算主义和泛心论——我们别钻那个牛角尖——但他们肯定不认为这个模型能感受吧?嗯,但意识不是关于感受,对吧?它是关于意识到存在……哦,我甚至不知道它是什么。我只是说这听起来和《意识先验》论文的内容不太远。是的,我知道它叫意识先验而不是意识,但你知道。

Yeah, but I think that's a very declarative description of consciousness, and at its roots it's about the phenomenological experience, right? And I know we discussed computationalism and panpsychism—let's not go down that rabbit hole—but surely they don't think that this model can feel? Well, but so this is consciousness is not about feeling, yeah? It's about being aware of existence... oh, I don't even know what it is. I'm just saying that it sounded not too far away from what the Consciousness Prior paper was about. And yes, I realize it's called the Consciousness Prior and not consciousness, but you know.

Host

是的,我的意思是,我认为他以科学家应有的方式回答了这个问题,我对他的回答非常满意。那就是:a) 意识必须是大脑中某些神经元活动和放电之类的东西,否则我们就是在谈论魔法,那不在科学领域内;b) 无论那是什么,它显然非常微妙和复杂,我们还不知道,所以我们需要保持谦逊,这意味着我们不应该危言耸听。所以我们不需要明天就去烧书,因为我们创造了一个 GPT 什么的,每当它的轮子转动时它实际上在受苦——你知道,如果你问它一个太难的问题,它在转动,那是因为你在伤害它,它在受苦,所以我们需要立刻关掉它。但等等,我们不能关掉它,因为那样我们就是在谋杀一个有知觉的存在之类的。我们对人类思维和意识这类复杂行为的理解还太幼稚,远未达到那个阶段。所以从我的角度看,他完全 100% 科学地回答了这个问题,而有很多所谓的科学家花了很多时间进行不科学的思考。

Yeah, I mean, I think he answered it the way a scientist should answer it, and I was really happy with his answer. Which is: a) consciousness has to be some activity of neurons and firings or whatever in the brain, or else we're talking about magic, and that's not in the field of science; and b) whatever that thing is, it's obviously quite nuanced and complicated, and we don't know yet, so we need to have some humility here, which means we shouldn't be alarmist. So we don't need to be going and burning books tomorrow because we created a GPT whatever that anytime its wheel is spinning and it's actually suffering—you know, if you ask it a question that's too hard and it's spinning, it's because you're hurting it and it's suffering, and so we need to turn it off like right away. But wait, we can't turn it off because then we'd be murdering a sentient being or something. We're way too in our infantile understanding of this type of complex behavior of the human mind and consciousness to be at that point. So from my perspective, he answered it completely 100% scientifically, and there's a lot of folks out there who are supposed to be scientists that spend a lot of time with unscientific thinking about it.

Host

酷。我们来谈谈 Bengio 的想法中我真正觉得有趣的一点,除了因果关系和系统二之外,就是多样性的概念。我们和 Kenneth Stanley 讨论过开放性和多样性保持。我们也和 Friston 讨论过平衡相对熵的重要性等等。在经验学习中我们都有这些诅咒,对吧?统计诅咒、近似诅咒、维度……是的,我们必须提到维度。甚至,你知道,我们在谈论诅咒和高维空间中的马尔可夫链蒙特卡洛,我们需要假设这些模式周围存在某种结构。所以所有这些方法都是同时能够探索但不受诅咒的方式。那么你对此怎么看?

Cool. Let's talk about one of the things that I really found interesting about Bengio's ideas, other than the causality stuff and the system two stuff, is this notion of diversity. We've had conversations with Kenneth Stanley all about open-endedness and diversity preservation. We've also had conversations with Friston about the importance of balancing relative entropy and so on. And we have all of these curses in empirical learning, right? The statistical curses, the approximation curses, dimensionality... yeah, we have to mention dimensionality. And even, you know, we're talking about curses and the Markov chain Monte Carlo in the sense of it being a high-dimensional space, and we need to assume that there's some structure around where these modes are. So all of these approaches are ways of simultaneously being able to explore but not being cursed. So what's your take on that?

Host

嗯,我的看法是,我喜欢他对抽象如此感兴趣,因为对我来说,这不仅是宇宙中更大的谜团之一——至少对我来说,我不知道——抽象、唯心主义、柏拉图式思维,等等。我的意思是,关键在于他认为抽象是通向实用有用路径的关键,这是一个难题,一个非常难的难题。

Well, my take was that I like that he's so interested in abstraction, because to me that's been not only one of the larger mysteries—at least for me, I mean, I don't know—of the universe: abstraction, idealism, Platonic thinking, whatever. I mean, the whole point is just that he views abstraction as a key to pragmatically useful paths forward, and it's a hard problem, a really hard problem.

基于图的方法与抽象 Graph-based methods and abstractions

Yoshua

他目前专注于基于图的结构。我得承认,对我来说它们也很有吸引力。我不知道这是否是正确的前进方向,但看到很多研究在探索基于图或超图的方法,这确实很酷。它们似乎是一条有前途的道路。如果我们能继续以合理的速度进步,我认为未来几十年会很有趣。我还想提出,我们今天的大多数基准测试可能并不适合衡量抽象带来的进步。他的论点是,创建抽象实际上可能有助于学习能力,这是我们凭直觉都能理解的。如果我有好的抽象,我就能将知识从这里迁移到那里。在像 ImageNet 分类这样的任务或今天的大多数基准测试中,可能并不需要抽象;问题的难度并不要求引入抽象。因此,限制因素可能不仅是模型本身,还有我们衡量抽象所能带来进步的能力。我认为这在不久的将来会改变,因为人们正在进入多模态研究领域。在那里,概念——也许不是抽象,但至少是类似概念的东西——要重要得多。

His focus right now is on graph-based structures. I have to admit, for me too, they're quite seductive and appealing. I don't know if they're the right path forward, but it's definitely cool to see a lot of research looking into graph-based or hypergraph-based methods. They seem to be a promising path forward. I think we're in for some interesting decades ahead if we can continue to progress at a reasonable rate. I would also postulate that most of our benchmarks today aren't necessarily suited to measure progress with abstractions. His argument was that creating abstractions might actually help your ability to learn something, which we all intuitively understand. If I have good abstractions, I can transfer my knowledge from here to there. In something like ImageNet classification or most benchmarks today, the necessity of abstractions is probably not there; the hardness of the problem doesn't require abstractions. So the limiting factor might not only be the models themselves but also our ability to measure the progress one could make with abstractions. I think that's going to change in the near future because people are moving into multimodality research. There, the concept of concepts—maybe not abstractions, but at least something like concepts—is way more important.

Host

让我快速补充一下,Tim。对 Yannick 来说,有一个非常有趣的点:缺乏处理多模态数据的良好工具。很多时候我们只是丢弃了有价值的数据源,因为我们没有好的工具集。想想自动驾驶汽车:视觉与激光雷达的争论。为什么不能两者兼用?你可以装上一些便宜的激光雷达传感器。为什么我们不利用这些数据呢?这确实是因为我们没有好的工具来处理多模态数据。

Let me follow up quickly, Tim. There's something very interesting to Yannick: the lack of good tools to deal with multimodal data. A lot of times we're just throwing out valuable sources of data because we don't have good tool sets. Think about self-driving cars: the debate of vision versus lidar. Why isn't it both? You could throw on some cheap lidar sensors. Why wouldn't we take advantage of that data? It's really because we don't have good tools to deal with multimodal data.

Yoshua

我们在讨论中触及了一个关于寻找抽象本质的好问题。我想知道神经网络在哪里能找到抽象。愤世嫉俗的观点是,人类创造了这些归纳偏置,它们代表了抽象。在几何深度学习里,情况正是如此:我们加入偏置来缩小近似空间。Keith 和我昨天有一个有趣的想法:几何深度学习和因果表征学习之间存在类比。Keith 上网找到了一个因果模型的定义:它对对抗样本具有免疫力。现在的模型所做的是学习每个像素与某些事件之间的关系,这就是它脆弱的原因。

We got to a good point in the discussion about the nature of finding abstractions. I wonder where neural networks can find abstractions. The cynical view is that humans create these inductive priors and they represent the abstraction. In geometric deep learning, that's what's happening: we put the priors in to reduce the size of the approximation space. Keith and I had an interesting idea yesterday: there's an analogy between geometric deep learning and causal representation learning. Keith went online and found a definition of a causal model: it's immune to adversarial examples. What a model does now is learn a relationship between every single pixel and something happening, which is why it's vulnerable.

Host

这来自一篇论文。我可以找出来,但现在不引用具体出处。它说:因果模型的预测与非因果模型的预测有什么区别?关键在于,几乎根据定义,如果你有一个因果模型,那么如果你扰动输入,预测仍然是有效的输出,因为它反映了世界的因果结构。而如果是非因果的,它就有可能学习虚假的结构,这就是为什么会出现对抗样本——你在某个地方放一个彩虹像素,它就会搞乱分类,就是因为那个虚假的连接。

That was from a paper. I could dig it up but I won't get the reference right now. It said: what is the difference between a prediction from a causal model versus a non-causal model? The point was that almost by definition, if you have a causal model, then if you perturb the inputs, the prediction remains a valid output because it's reflective of the causal structure of the world. Whereas if it's non-causal, it has the potential to learn spurious structures, which is why you get adversarial examples—put a little rainbow pixel somewhere and it messes up the class because of that spurious connection.

Yoshua

同样,对抗样本的存在是因为不精确,因为我们没有完美的判别函数。我也可以说:如果我只有正确的判别函数,它不需要是因果的。如果我对输入空间进行了正确的划分,我就不会受到对抗攻击。真正的问题是:这在技术上是否等同于一个因果模型?如果我对输入空间进行了完美的类别划分,这在数学上是否等价于一个因果模型?谁知道呢。

In the same vein, adversarial examples exist because of inaccuracies, because we don't have the perfect discriminative function. I could also say: if I just had the correct discriminative function, it doesn't need to be causal. If I had the right partitioning of my input space, I'm not vulnerable to adversarial attacks. The real question would be: would that technically amount to a causal model? If I had the perfect partitioning of the input space into my classes, is there a mathematical equivalence to a causal model? Who knows.

Host

我认为如果你有完美的判别函数,它很可能就是你能从因果模型中推导出的判别函数。我不完全确定能否反向推导,因为从因果模型到判别函数可能存在信息损失。例如,在带有重写规则的生产系统中,因果系统的定义是:所有能到达特定输出的潜在转移图都是同构的。即使你有完美的判别函数,也可能存在多个同构的图。所以我不太确定这如何转化,但我认为对于判别来说,它同样有效。

I think if you have the perfect discriminative function, it's probably the discriminative function you would derive from a causal model. I'm not 100% sure you can go in reverse because there might be information loss going from a causal model to a discriminative function. For example, in production systems with rewriting rules, the definition of a causal system is one where all potential transition graphs to a particular output are isomorphic. Even if you have the perfect discriminative function, there may be multiple possible graphs that are isomorphic. So I'm not quite sure how that translates, but I think it would be just as good for discrimination.

Yoshua

我认为这与 NLP 中的语义讨论有关。像 Walled Suburbs 这样的人说神经网络没有语义。就像蓝色像素一样,在现实世界中,男性睾酮水平与车祸发生率有因果关系。这意味着你可以把模型带到另一个国家,因为它是一个因果因素,它会以同样的方式外推。但神经网络模型——人类所做的是提出正确的表征抽象,构建一个非常简约的模型。神经网络将一切与一切建模,语义都混为一谈。嗯,我不太热衷于与 NLP 人士讨论语义。为什么不呢?我不知道。这常常会偏离主题,陷入其他东西。

I think it's related to the semantics discussion in NLP. People like Walled Suburbs say that neural networks don't have semantics. In the same way as blue pixels, in the real world, male testosterone levels are causally linked to incidence of car crashes. That means you can take the model to a different country, and because it's a causal factor, it will extrapolate in the same way. But neural network models—what a human does is come up with the right representational abstraction, building a very reductionist model. A neural network models everything to everything, and the semantics are all one thing. Well, I'm not too keen on discussing semantics with NLP people. Why not? I don't know. It often veers away and fears into something else.

神经网络中的抽象 Abstraction in Neural Networks

Yoshua

这触及了语义的本质。我喜欢康纳·莱希很久以前说过的话,我同意他的观点:你必须从这些模型的视角来看待一切。如果我是 GPT-3,我的整个世界就是文本输入。人们不能因为 GPT-3 没有与现实世界的连接,植物生病了不会去看医生,就评判它。它们并不生活在现实世界,而是生活在互联网的文本世界里。在那个世界里,我不确定这些模型是否没有形成某种抽象层次。我不想声称这些东西不会形成抽象事物。它们形成的抽象类别可能与我们不同,但它们确实形成了某种抽象层次。当然,它们无法迁移,因为我们只给了它们一种模态,但它们可能能够在文本的不同领域之间迁移,有时也确实如此。所以,我不会对这些事情下绝对的结论。

Bit tears into semantics. I like what Connor Leahy said when we talked to him a long time ago, and I happen to agree with him: you have to see everything from the perspective of these models. If I'm a GPT-3, my entire world is text input. People can't judge GPT-3 by saying, 'You don't have a connection to the real world, you don't go see a doctor if your plant is sick.' They don't live in the real world; they live in the text world of the internet. In that world, I'm not sure if there isn't a level of abstraction happening in these models. I don't want to claim that these things do not form abstract things. It might not be the same abstract classes that we form, but they definitely form some level of abstraction. Of course, they can't transfer it because we only give them one modality, but they may be able to transfer it between different areas of text, which they sometimes do. So I just wouldn't be so conclusive with respect to these things.

Host

我觉得我们绕了一圈又回来了。在与兰德尔·贝利斯特里奥讨论样条之后,那几乎导致了对 MLP 的一种愤世嫉俗的解读,认为它们只是哈希表。但我们用的不是 MLP,而是 Transformer 和 CNN。实际上,如果你把抽象仅仅视为外推,我认为它们基本上是同义词。关键在于能够在训练集之外进行外推。那么那些归纳先验确实在产生抽象。但问题在于这些归纳先验是人类设计的。我们想要的是学习抽象,而这正是我认为没有发生的事情。

I think we've gone full circle now. After speaking with Randall Bellistrio about the splines, that almost results in such a cynical reading of MLPs that they're just hash tables. But we're not using MLPs; we're using Transformers and CNNs. Actually, if you think of abstraction just as being extrapolation, I think they are basically synonymous. It's about being able to extrapolate outside of your training set. Then those inductive priors are indeed producing abstractions. But the problem is humans design those inductive priors. What we want is to learn abstractions, and that's the thing that I don't think is happening.

Yoshua

我在某种程度上与扬尼克和康纳站在同一立场:如果抽象只是输入空间的压缩或编码,那么它们当然在学习抽象。它们丢弃了一些信息,保留了一些抽象的东西。但我认为这有点曲解了人们说抽象时的本意。传统上我们认为的抽象类型是对更一般、更长程结构的简化。而我们都知道,神经网络学到的很多所谓抽象是那种捷径,是低层次的、近乎虚假的抽象。这就是为什么它们那么容易失效,那么脆弱。这里存在模糊性:什么样的抽象才是好的抽象?我不知道。但我认为这有点偏离了重点。我们讨论的,也是本吉奥所说的,目标是弄清楚如何让机器学习通过简化、通过简单的抽象来学习那些在分布外更具泛化能力的结构。这才是真正的目标。其余的都只是语义问题,双关语。

I'm kind of on the same page as Yannick and Connor on the one hand, which is: if an abstraction is just a compression, an encoding of the input space, then of course they're learning abstractions. They are throwing away information and retaining some abstract thing. But I think that just kind of devolves into a bastardization of what people mean when they say abstractions. The types of abstractions that traditionally we think about are simplifications of more general, longer-range structures. Whereas we know, and I think we all know this for sure, that a lot of the so-called abstractions that a neural network learns are these kind of shortcuts, low-level borderline spurious kinds of abstractions. That's why they break so easily, that's why they're so brittle. There is this vagueness here: when is an abstraction a good abstraction? I don't know. But I think it all kind of misses the point. What we're talking about here, and this is a lot of what Bengio said, is that the goal is to figure out how to get machine learning to learn structures that, by virtue of their simplification, their simple abstractions, are more generalizable out of distribution. That's really the goal here. The rest of it is just semantics, pun intended.

Host

如果你放眼世界,许多文化和人类必定拥有相同的抽象。所以这一定意味着它不仅仅是你一生中学习的东西。

If you look across the world, a lot of cultures and humans must have the same abstractions. So it must mean that it's not just something you learn during your lifetime.

Yoshua

绝对正确。它是通过进化、通过物种、通过生命本身学习到的。

Absolutely, so I'm correct. It's learned by evolution, by the species, by life itself.

Host

但把我们构建正确抽象类比为进化通过随机搜索来完成的捷径,这可能是不同的性质。我们希望机器去学习,因为通常我们说希望机器学东西时,是希望它们像人类一生中那样摄取数据。但这些抽象以及形成抽象的能力似乎是在更基础、更共享的层面上发生的。

But is the analogy to us building in the correct ones as a shortcut for evolution doing it using essentially random search? That might be a different quality. We want machines to learn something because usually when we say we want machines to learn something, we want them to ingest data akin to what a human does during its lifetime. But these sorts of abstractions and the ability to form abstractions seem to be happening on a more fundamental, shared level.

Yoshua

是的,你正中靶心。我认为完全正确。有很多理由说明它是副现象,而且很多智能是具身的。我同意有很多事情在我们不知情的情况下发生。这显然是大多数人不理解的东西。也许这就是为什么在群体层面上,我经常与人沟通不畅,因为我从不认为学习是关于人类一生中学到的东西。对我来说,它一直是进化范式。你的神经元里编码了什么,你的 DNA 里编码了什么,细菌很久以前学到了什么,这些如何转化为人类的行为?我不知道为什么这么多人关注人类一生中学到的东西。为什么那是目标?我不确定。但我们有陷入过度简化论的风险。康纳·莱希说过,人类是否智能还是一个开放问题。如果你沿着那条思路走,很快就会说人类只是像 GPT-3 一样的哈希表。显然人类在某种程度上是智能的。你可以把它当作定义问题,但它不是二元的。为什么我们总是陷入这种非黑即白的观念,认为某物要么智能要么不智能?我不是这样看问题的。我认为智能是一个从零到无穷大或某个远超人类的巨大数字的频谱。所以它是一个连续体。这就是为什么我喜欢乔莱的智能度量,尽管它还没有给出定量的测量方法,但它至少朝着正确的方向思考:如何测量智能并将其定义为一类活动?然后我们就能摆脱这种非黑即白的思维。

Yes, you just put the pin in the center of the bullseye. I think that's exactly right. There's a lot to be said for it being an epiphenomenon, and a lot of intelligence is embodied. I agree that there's an awful lot of stuff going on unbeknownst to us. This is clearly something that most people don't have a grasp on. Maybe this is why at the population level, I'm frequently miscommunicating with people, because I never assumed that learning was about what a human being learns in a human being's lifetime. To me, it's always been the evolution paradigm. What's encoded in your neurons, what's encoded in your DNA, what was learned by bacteria a long time ago, and how did that translate into what human beings are doing? I don't know why so many people are focused on what a human being learns in their lifetime. Why is that the goal? I'm not sure. But we run the risk of being very reductionist. Connor Leahy said that it's an open question whether humans are even intelligent. If you go down that line, very quickly you start saying human beings are just hash tables like GPT-3. Clearly humans are intelligent in some way. You can take it as a matter of definition, but it's not a binary thing. Why are we always into this black and white concept of something is or is not intelligent? That's not how I view things. I think there's a spectrum of intelligence from zero to maybe infinity, or some really large number beyond what human beings are. So it's a continuum. That's why I like Chollet's measure of intelligence, because even though it doesn't give us a quantitative way yet to measure intelligence, it at least thinks along the right direction: how do you measure intelligence and define it as a category of activity? Then we can get past this black and white thinking.

Host

先生们,总是很愉快。当然,当然。

Well, gentlemen, always a pleasure. Absolutely, absolutely.

互动版:逐字朗读 + 针对本期提问 →