基于权重的学习:将模型权重视为数据

Weight-Based Learning: Treating Model Weights as Data

达米安·博思 Damian Borth · TWIML AI 播客 · 2026-07-27 · 约 46 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Damian Wirth 探讨将训练好的模型权重作为新神经网络的输入数据,从而更快地生成和分析模型。

Damian Wirth discusses treating trained model weights as input data for new neural networks, enabling faster generation and analysis of models.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 15)

全文 · Full transcript(中英对照)

引言 Introduction

Host

当今 AI 面临的最大问题之一是:随着高质量训练数据越来越难找,基础模型如何持续改进?一些研究者押注合成数据,另一些则押注推理时推理。今天的嘉宾押注的是一种截然不同的方向。每个训练好的模型都代表着数千甚至数百万 GPU 小时,用来探索什么有效。与其把这些权重仅仅视为训练过程的终点,如果它们也是下一个训练过程的起点呢?圣加仑大学 AI 与机器学习教授 Damian Wirth 将训练好的模型本身视为数据——可以被学习、被分析,甚至用来生成全新模型的数据。当我请他解释基于权重的学习背后的想法时,他是这样开始的。

One of the biggest questions facing AI today is how foundation models keep improving as high-quality training data becomes harder to find. Some researchers are betting on synthetic data, others on inference time reasoning. Today's guest has his chips on something very different. Every trained model represents thousands or even millions of GPU hours spent discovering what works. Instead of treating those weights just as the end of the training process, what if they're also the beginning of the next one? Damian Wirth, professor of AI and machine learning at the University of St. Gallen, sees trained models themselves as data. Data that can be learned from, analyzed, and even used to generate entirely new models. When I asked him to explain the idea behind weight-based learning, here's where he started.

Damian

所以我们基本上想到了一个非常简单的想法。如果我们把训练好的神经网络的权重作为输入,来训练一个神经网络,让它更好地理解我们已有的这些权重,会发生什么?想到你可以把权重作为一种输入模态,突然就给了你这样的机会:我们能不能更快地为特定任务创建新权重?或者当有人给我一个我不了解、从未见过的神经网络时,我们能不能更精确地分析它的权重?

So, we basically thought about this very simple idea. What happens actually if we take the weights of trained neural networks as the input to train a neural network to understand these weights that we have out there much, much better. Thinking about that, that you can treat the weights as an input modality gives you suddenly this opportunity of... can we be much, much faster in creating new weights for a particular task, or can we be much more precise in analyzing weights when somebody gives me a neural network that I'm not knowledgeable about and I never saw before.

Host

我是 Sam Charrington,这里是 Twilio AI 播客。十多年来,我一直在通过这样的对话探索塑造 AI 未来的想法和创新,帮助你理解什么是真实的、什么是下一步、什么才是重要的。让我们开始吧。

I'm Sam Charrington, and this is the Twilio AI podcast. For over a decade, I've been exploring the ideas and innovations shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters. Let's jump in.

基于权重的学习起源 Origin of Weight-Based Learning

Damian

我们在 2021 年左右开始了我们称之为基于权重的学习的工作。基于权重的学习是一种看待机器学习的有趣方式,目前这是我们的主要课题。我们也在遥感领域做一点工作,还有表格数据的表示学习。我们现在正把一切整合起来,更专注于权重空间学习。我认为这是一个非常有趣的未来方向,它解决了社区目前遇到的一些问题。从一个非常深奥的想法,发展到了出奇好用的东西。

We started like in 2021 the work on what we call weight-based learning. And weight-based learning is a quite interesting way of looking at machine learning in general. That's currently the major topic. We also do a little bit of work in remote sensing and then representation learning on tabular data. We're now walking and combining everything together to focus more on weight space learning. Which I think is a really interesting way forward. It solves a couple of problems that the community currently encounters. And started from a very, very esoteric idea to something that works surprisingly well.

Host

你知道,我们把权重视为训练模型的产物。我们从中获得一些效用,可能会用它们来做可解释性之类的事情,或者在量化时对它们进行操作。但这个想法似乎是,我们可以从这些权重中学到更多东西。

You know, we think about weights as the product of training a model. And we get some utility out of them. We may use them for things like explainability or manipulate them when we're quantizing or something like that. But the idea seems to be that there's so much more that we can learn from these weights.

Damian

没错。所以如果你想想机器学习,机器学习有一个概念:你有数据和某种输出。在经典的监督式机器学习中,是数据和预测。你在中间训练一个神经网络来模仿数据集,模仿数据集的分布。在这个非常昂贵的训练过程结束时,得到一组权重,一组定义神经网络的参数配置,就像神经网络的 DNA。这就是经典的机器学习——监督式、非监督式、自监督式——在过去 10 年推动了很多创新,并且随着生成式 AI 进入了新阶段。如果你看看过去几年发生了什么,越来越多的模型被发布,公开地在 Hugging Face 或 GitHub 这样的仓库中在线可访问。所以我们基本上想到了一个非常简单的想法。如果我们把训练好的神经网络的权重作为输入,来训练一个神经网络,让它更好地理解我们已有的这些权重,会发生什么?所以,再打个比方,语言模型,你拿一个大模型,在互联网上的每一个句子上训练它,最后你得到一个能够分析语言和生成语言的语言模型。你可以对像素采用同样的想法。你拿一个大模型,在互联网上所有像素上训练它,你就能分析像素,也能生成像素。我们对所有训练好的神经网络的权重做同样的事情,所以我们能分析神经网络的权重,也能生成神经网络的权重。虽然说起来简单,显然细节上还有更多内容,但想到你可以把权重作为一种输入模态,突然就给了你这样的机会:想想看,对于神经网络模型来说,语言翻译会是什么?生成词和词元,以及生成权重词元,会是什么?我们能不能更快地为特定任务创建新权重?或者当有人给我一个我不了解、从未见过的神经网络时,我们能不能更精确地分析它的权重?然后你就有了一个全新的世界,一个你可以用权重做各种事情的空旷空间。你把它带入社区,希望有人倾听并继续下去,建立社区,这在过去两三年里发生了,非常令人兴奋,因为有更多人关注这个,是的,权重令人兴奋,不仅作为学习的输出,也作为学习的输入。

Exactly. So if you think about machine learning, machine learning has this idea of you have data and some output. In classical supervised machine learning, data and some predictions. And you train a neural network in between to mimic the data set, mimic the distribution of the data set. And the outcome during this very expensive training procedure is a set of weights, a configuration of parameters that define the neural network like the DNA of the neural network. And this is classical machine learning supervised, unsupervised, self-supervised that fuels a lot of innovation over the last 10 years and with GenAI, moved to next stage. If you look at what happened over the last couple of years, more and more of those models have been published. Publicly are online accessible at repositories like Hugging Face or GitHub. So we basically thought about this very simple idea. What happens actually if we take the weights of trained neural networks as the input to train a neural network to understand these weights that we have out there much, much better. So, to take another analogy, you have language models, you take a big model, you train this on every single sentence on the internet, at the end you have a language model able to analyze language and to generate language. You can take the same idea for pixels. You take a big model, you train on all the pixels on the internet, and you can analyze pixels, and you can generate pixels. We do the same idea on all the weights of trained neural networks, so we can analyze weights of neural networks, and we can generate weights of neural networks. As straightforward as it is, obviously there's a little bit more into the details, but thinking about that you can treat the weights as an input modality gives you suddenly this opportunity of thinking about, okay, what would be language translation in with more like neural network models, right? What would be generation of words and tokens that are words, and generation of tokens that are weights, and can we be much, much faster in creating new weights for a particular task, or can we be much more precise in analyzing weights when somebody gives me a neural network that I'm not knowledgeable about and I never saw before. And then you have this new entire world, this empty space of things you can do with weights. That you kind of carry into the community and hope that there's somebody listening and continuing and building up a community which happened over the last 2 or 3 years, which is very exciting because there are more people about that, and yeah, weights are exciting, not only as the output of learning, but as the input for learning.

早期工作与灵感 Early Work and Inspiration

Host

你提到你在 2021 年开始了这项工作。你从哪里开始的?然后我们再一路聊到我们现在在权重空间学习方面的进展。

You mentioned you started this effort in 2021. Where did you start from, and then we'll work our way towards like where we are now with weight space learning?

Damian

最初这个想法来自 2020 年,我们在 21 年发表了第一篇论文,想法非常简单。我们能不能像对软件那样,对神经网络进行指纹识别或版本管理?在软件中,你可以做差异比较,对吧?你有一百万行代码,有人改动了某些东西,然后你做差异比较,你就能确切知道差异在哪里。那么,我们能不能对神经网络做同样的事情?可能对于神经网络,如果你在训练中做一次权重更新,每个权重都会有一点不同。所以从中提取不到太多信息。

Originally this idea came 2020 and we got the first paper published 21 and the idea was very simple. Can we fingerprint a neural network or version a neural network like we can do with software, right? In software you can do a diff, right? You have a 1 million lines of code, somebody changes something and then you do a diff, you know exactly where the difference is. So can we do this with neural networks? Probably with neural networks is if you do one update of weights during training, every weight is a little bit different. So there's not much you can extract from this.

Host

对,它们在局部相当不稳定。

Right, they're fairly unstable locally.

Damian

如果一切都不同,那就没有什么是不同的,对吧?所以我们想,我们能不能找到一个空间,让这些神经网络、权重、这些神经网络的模型被稍微压缩,更容易理解。我们开始思考这个问题。与此同时,有一项非常出色的工作,我最早知道的是 Thomas Unterthiner 和 Daniel Keysers 以及他们在 Google 苏黎世的同事,他们发表了一篇论文,使用权重作为输入,提取了一些统计特征和手工特征,然后预测这些权重的准确率。还有另一篇论文,我听说预测了泛化差距。

If everything is different, nothing is different, right? So we were thinking about can we find a space where these neural networks, the weights, the models of those neural networks are a little bit compressed and more understandable. And we started to think about that. And in parallel there was this really amazing work, the first work I know from Thomas Unterthiner and Daniel Keysers and colleagues from Google Zurich and they developed a paper that used weights as input, did some statistical features and crafted features and then predicted the accuracy of those weights. And in another paper where I listened that predicted the generalization gap.

初步构想与首篇论文 Initial Idea and First Paper

Damian

所以人们开始用权重来提取信息,而且都是手工设计的特征。我就在想,手工特征已经主导了传统机器学习,那为什么不试试端到端学习呢?于是我们发展出了自己的想法:把权重序列自动编码到一个低维空间,然后再重建出来。如果我们能对一群神经网络做到这一点,也许就能学到一个新的低维流形,这个流形实际上就是新网络所在的流形。也许这个流形就编码了关于训练数据、训练比例、学习率等等这些潜在生成因素的信息。我们开始研究这个,第一篇论文真的就是“好吧,这能行,我们可以压缩神经网络”。就像非常小的玩具神经网络例子,小得有点尴尬,只有几千到一万个参数。但它确实有效,我们能预测准确率,看到这个结果真的很高兴。然后我们就发表了。

So people started to use weights to extract information. And they were all handcrafted features. So I was thinking about handcrafted features that sold traditional machine learning. So why not end-to-end learning? So then we developed our idea of autoencoding sequences of weights into a lower dimensional space and then reconstructing it. And if we are able to do this from a population of neural networks, then we can maybe learn a lower dimensional manifold that actually populates where the new networks' populated manifold. And maybe that manifold encodes information about the currency, what training data was used, what training fraction, learning rate, and all these latent generating factors. And we started working on this, and the first paper was really like, "Well, it works. We can compress neural networks." Like very small tiny neural network toy examples, right? Embarrassingly small, like thousands or 10,000 parameters. But it was working, and we could predict the accuracies, and it was very nice to see that. And then we got published.

Host

你们具体能预测什么?

What exactly were you able to predict?

Damian

我们用一个自编码器,有编码器和解码器。我们用重建损失和中间的对比损失来训练自编码器。然后我们只取编码器,把未知的神经网络编码到潜在空间。这些嵌入我们放入一个简单的回归,比如线性回归头,来预测准确率。所以你给我一个神经网络,我观察这个神经网络,想法就是:我能预测这个神经网络的准确率吗?我能不能不用测试数据就测试这个神经网络?显然,这在小神经网络上有效,而且只在同质群体中,我们称之为“模型动物园”,即一群模型。它们都在相同的数据和架构上训练。这是朝着目标迈出的小步。但我们能提取出关于准确率、性能、泛化差距、用了什么激活函数的信息。如果你画出这个潜在空间,你会看到不同的初始化和它们如何演化,因为我们有 50 个 epoch 的模型训练和 1,000 个模型。所以能看到小的轨迹。所以是有结构的,它们都围绕一个潜在空间组织起来。

So we took an autoencoder. We have an encoder and decoder. We learned the autoencoder with reconstruction loss, a contrastive loss in the middle. And then we took only the encoder and unknown neural networks that we encoded into the latent space. And these embeddings we put into a simple regression, like a linear regression head, to predict the accuracy. So you give me a neural network, and I observe this neural network, and the idea was, can I predict the accuracy of this neural network? Can I test the neural network without the use of test data, right? Obviously, this worked on really small neural networks, only in a homogeneous population, we call this a model zoo, a population of models. So they were all trained on the same data, same architecture. First, small steps towards this goal. But we could extract this information about the accuracy, the performance, the generalization gap, what kind of activation function was used. And if you would plot this latent space, you would see different initializations and how they evolve, because we had 50 epochs for model training and 1,000 models. So little trajectories were visible. So there was structure, and they were all organized around a latent space.

Host

关于性能和准确率的元数据,这些是你们从基础模型那里得到的,所以在这个意义上是监督的。

The metadata around performance and accuracy, these are things that you had from the base models, and so supervised in that sense.

Damian

没错。所以我们需要能够在一群新网络上训练这个自编码器,也就是模型动物园,而且当时是在已知条件下实验室训练的,完全透明。我们知道哪个模型、哪个 epoch 有哪个准确率。我们显然有训练、测试和验证集。所以自编码器在 600 个神经网络上训练,在另外 300 个上测试。然后我们可以比较,我们在 Fashion MNIST 上预测 90% 的准确率,而模型有 92,然后我们用了 R 平方,我们超过了 2013 年和 Danny Kaser 的工作,所以我们很高兴,超过了原始权重空间。那个故事成功了。我们拿到了论文,我非常感谢审稿人。他们告诉我们这很小,但很有趣,我们得以发表,因为那是过去 4 年这段奇妙旅程的起点。

Exactly. So we needed to be able to train this autoencoder on some population of new networks, a model zoo, and at that time, laboratory-trained under known conditions, fully transparent. We knew which model and which epoch had which accuracy. We had obviously train, test, and validation splits. So the autoencoder was trained on 600 neural networks and tested on 300 others. And then we could compare, we predicted 90% accuracy on Fashion MNIST, and the model had 92, and then we were R squared, and we outperformed the work from 2013 and Danny Kaser, so we were happy, outperformed the original weight space. That story worked, though. We got the paper, and I'm very thankful to the reviewers. They were kind of telling us that it's small, but it's interesting, and we were able to publish because that was kind of the ignition for this amazing journey that we had over the last 4 years.

Host

所以我可以想象有很多不同的方向,包括扩大模型规模,试图从这个空间中获得更多见解。比如,下一步是什么?

So I can imagine lots of different directions, including scaling up the models, trying to get more insights out of the space. Like, what was next?

Damian

问题是,对于这类研究,一切都很明显。就像摆在那里,你只需要去做。我以前从未有过这种感觉,对吧?所以你有一个自编码器,你取编码器,就可以预测判别性的下游任务,比如准确率、泛化差距。但我们还有另一个东西叫解码器,那么我们能从这个空间采样来生成新网络吗?这很明显,对吧?第一篇论文里我们没有空间,所以我们还需要 1 年时间,在 22 年发表了一篇关于生成新网络的论文。希望那些新网络比标准初始化更好。它们不如最终或完全训练好的新网络。所以我们遇到了一些麻烦,这真的很有趣。我们训练了这个自编码器。均方误差非常低。我们取新网络,进行前向传播,重建它,我说损失很低,我们把权重插回新网络。结果完全搞砸了整个新网络。我们想:“我们怎么搞砸的?”是的,就像“嘿,均方误差很低”。现在显然了,对吧?均方误差是一个平均值。所以我们非常擅长重建平均权重,但那些决定这个函数是否工作的微小差异,它们太重要了。这和像素与图像有类比。就像当你有了图像生成模型,图像总是模糊的。所以人们努力驯服 Transformer 来处理高分辨率图像,把均方误差改成感知损失,还有一些关于量化和 GAN 的额外东西。所以我们知道重建给出了错误的损失,我们尝试归一化并调整损失,考虑了行为损失,然后我们能让那些模型好一点。但仍有缺失的 delta。我们当时生成的是模糊的权重,对吧?低频信息,高频信息缺失。我们很高兴,因为再次,人们对我们很好,说这是玩具例子,但很有趣。我们生成的数字很好。但然后我说:“好吧,我们不能三次都幸运,所以我们必须非常努力地扩大规模,对吧?”我的意思是,你幸运两次,但你知道,三次,你的业力在接下来几年就没了。所以我们真的,我非常感谢 Konstantin Shcherhold,他参与了初始阶段,他真的努力工作,我们有了这个想法:不是重建整个模型,而是把模型参数看作一个序列,你加窗口然后重建窗口。因此,我们会在某种程度上把原始模型的序列与自编码器的序列分离开。

The thing is, for this type of research, everything was very obvious. It was like lying out, and you just needed to do it. I never had this before, right? So you have an autoencoder, you take the encoder, so you can predict discriminative downstream tasks, like what's the accuracy, what's the generalization gap. But we have the other thing called the decoder, so can we sample from this space to generate new networks? And it's obvious, right? We didn't have space in the first paper, so we needed 1 more year to have a 22 paper published on generating new networks. Hopefully those new networks were then better than standard initializations. They were not as good as final or fully trained new networks. So there was some trouble that we had, which was really interesting. We trained this autoencoder. The mean squared error was super low. We took the new networks, we moved them in a forward pass, we reconstructed it, I said the loss is very low, we plugged the weights back to the new network. It totally screwed up the entire new network. We're like, "How did we screw up?" Yeah, it was like, "Hey, then the mean squared error is low." And obviously now, right? A mean squared error is an average. So we're very good at reconstructing the average weights, but the little things that make the difference of having this function working or not, they were so important. And there is an analogy to pixels and images. Like when you had generative models for images, the images were always blurry. So people tried hard and tamed transformers for high-resolution images, changed the mean squared error to perception loss, and there some additional things on quantizing and the GAN. So we knew that reconstruction gave the wrong loss, and we tried to normalize and play around with the losses, and thought about the behavior loss, and then we were able to make those models a little bit better. But there's still a little bit of delta that is missing. We generated at that time blurry weights, right? Low-frequency information, high-frequency information missing. And we were happy because again, people were kind to us and said like it's toy examples, but it's interesting. We generated the numbers are good. But then I said, "Okay, we cannot be three times lucky, so we have to work really hard to scale that up, right?" I mean, you're lucky twice, but you know, three times, your karma is gone for the next years. So we then really, and I'm very thankful to Konstantin Shcherhold who was part of that initial phase, and he really worked hard, and we had this idea of instead of reconstructing the entire model, think about the model parameters as a sequence that you window and then reconstruct the windows. And therefore, we would kind of detach the sequence left of the original model to the autoencoder one.

扩展与合作 Scaling up and collaboration

Damian

这很有意思,因为突然间我们可以扩展到 ResNet 及更深的网络,这促成了 2024 年与加州大学伯克利分校的 Michael Mahoney 的合作。他曾在一次讨论中说:‘Damien,你做的事情很有趣,但没用。’

And this was interesting because suddenly we could go to ResNets and beyond, and this led then to the work in 2024 in collaboration with Michael Mahoney from UC Berkeley. He actually said in one of the discussions, like, 'Damien, what you're doing is really interesting, but useless.'

Host

所以那就像是,‘当然,你知道,很多研究都是这样开始的。’

And so that would be, like, 'Sure, you know, a lot of research starts like that.'

Damian

所以,因为他那么说,我就请他帮忙扩大规模,对吧?所以我抓住了他,他后来成了论文的合著者之一。之前的工作也有大量合作,比如 Shalini 和 Boris Knyazev。他们参与其中,因为一开始,这个想法有点,如我所说,晦涩难懂。所以我们想知道,别人会怎么看?所以我们很早就让社区里的很多人来帮忙确认,我们是不是疯了,还是这个想法至少在一定程度上是有意义的。所以到了 2024 年,我们扩展到更大的网络,其他人也开始感兴趣,我们在会议上遇到了‘哦,还有其他人也在做类似的事情’。很多来自 Technion 的工作,比如 Hagai、Marwan、Gil Chechik、Yair,还有 Eliahu 等人。其实有件有趣的事。有个学生来到我们的海报前,他戴着那种徽章,徽章底部通常写着大学名字。但代替大学名字,他写着‘权重是新模态’,Eliahu Horwitz。我当时想,‘哦,这正是我在想的。’然后合作就开始了,形成了一个社区,我们后来意识到还有其他人,我们就说,‘为什么不办个研讨会呢?’然后一件事引发另一件事,然后,你知道,我们……

So, and then, because he said that, I asked him that you have to help to scale it up, right? So, I caught him, and he was then one of the co-authors on the paper. Also, the previous work was with a lot of collaboration, you know, Shalini and Boris Knyazev. They were part of that because in the beginning, the idea was a little bit, as I mentioned, esoteric. So, we were wondering, like, what do other people think about? So, we very early involved a lot of people from the community to double-check if we are the crazy ones or if this idea is, you know, it's to at least a particular limit, meaningful. So, we then in 2024, we scaled up to larger networks, and other people got interested, and we were at the conference, and we met, 'Oh, there exist other people that are doing similar things.' A lot of work from Technion, you know, Hagai, Marwan, Gil Chechik, Yair, and then people like Eliahu. It was actually a funny thing. There was one student that came to our poster and he had this kind of badge, and at the badge at the bottom, you have always the university written. And instead of the university, he had like 'weights are the new modality', Eliahu Horwitz. And I was like, 'Oh, that's exactly what I'm thinking.' And then a collaboration started, like, a community, and we then recognized that other people, and we said, like, 'Why not doing a workshop?' And then, one led to another and then, you know, we

Host

那是去年的事。我明白了。

That was last year. I see.

Damian

2024 年,我们有了第一个提案,2025 年,我们举办了第一次研讨会。然后,其他人意识到还有其他非常重要的领域。权重空间的一个棘手之处在于,当你有一个神经网络,有两层,为了简单起见,假设是全连接的,你可以改变神经元的位置。这实际上改变了权重的顺序,但函数是一样的。所以,权重空间中有一些排列对称性和其他不改变底层函数的东西。所以,我们在第一篇论文中已经有了一些增强,但有很多人对此非常专业,更有经验,更理论化。

2024, we then had the first proposal, in 2025, we had then the first workshop. And then, the other people recognized there are other areas that are very important. One of the tricky things with weight spaces is when you have a neural network and you have two layers, let's say for simplicity, fully connected, you can change the position of the neurons. And it's actually changes the sequence of weights because the order of weights, but the function is the same. So, there's little permutation symmetries and other things in the weight space that do not change the underlying function. So, we had it already in the first paper some augmentation, but there are a lot of people that are very specialized on that and much more experienced, much more theoretical on that.

Host

增强是指将这些恒等变换应用于你的训练、你的输入模型,并用它们来提高泛化能力,以及……

Augmentation in the sense of like applying these identity transformations to your training, your input models, and using them to increase generalization and

Damian

没错。是的,很简单,我们在 2021 年工作时,我们使用了对比损失,而要构建对比,你需要增强。所以,翻转图像或旋转图像很简单,但权重空间中的对应物是什么,对吧?你可以做排列,其他人后来发明了尺度增强等。这个领域很令人兴奋,因为你看到 NLP 中发生的事情,看到计算机视觉中发生的事情,你必须将其转化到权重空间。而且它奏效了,对吧?比如,增强发生了,我们怎么转化?有感知损失,我们怎么把它转化为行为损失,对吧?所有这些。所以,你可以从其他领域借用想法,而且这是一个空白领域,可以填充内容。

Exactly. Yeah, very simply, we were working in 2021, we had the contrastive loss and to build a contrast, you need to augment. So, it's simple to flip an image or rotate an image, but what's the counterpart in weight spaces, right? You can do the permutations and other people then invented scale augmentations and other things. So, the field was exciting because you saw things happening in NLP, you saw things happening in computer vision, and you had to translate it into weight spaces. And it worked, right? Like, augmentations happen, like how can we translate it? There's this perception loss, how can we translate it into behavior loss, right? And all those things. So, you can borrow ideas from other fields and it was an empty field to fill with content.

探索权重空间变换 Exploring weight space transformations

Host

沿着这些思路,听你谈到这些恒等变换,让我想到权重空间中的其他几何变换,比如笛卡尔到极坐标的变换之类的。我们看到谷歌的一篇量化论文,我忘了那篇量化论文的名字,它刚出来,或者其实是一年前出来的,但大约一周前被重新提起,他们做了笛卡尔到极坐标的变换。就像你在权重空间里可以做的各种事情,我很好奇这方面探索了多少。

Along those lines, hearing you talk about these identity transformations makes me think about like other kinds of geometric transformations in the weight space or like Cartesian to polar transformations or things like that. We saw a Google quantization, I forget the name of the quantization paper that just came out or actually came out a year ago, but it was revisited a week or so ago and they did some Cartesian to polar transformation. Like all kinds of stuff that you can do in the weight space that you might, I'm curious how much of that is being explored.

Damian

所以,正是如此。我们来自这个,这是动机的一个领域。我们遇到了其他人,如我所说,他们都在研究群的对称性,你可以做的操作。还有另一个世界,有这种模式连通性,可以进行重基,你可以将模型对齐到某些参考模型。如果你只考虑这种排列对称性,在损失景观中有一些是连通的。有很多点是相同的,但它们在权重空间上是不同的。只是理解或试图理解损失表面是什么样子,模型如何沿着轨道演化,这是 Bull 和加州大学圣地亚哥分校的 Rose 的伟大工作,所有这些工作帮助我更好地理解学习过程中实际发生的事情,然后希望我们能将其应用到我们的骨干学习中,因为一些分词、一些位置编码在自动编码器中是 Transformer 自动编码器,我们借鉴了一些想法,它帮助了我们。然后,人们也在发现 grokking 和相变,以及模型如何突然收敛或突然不工作然后突然工作。所以,这也有关联,我们很想进一步探索这个方向,以理解如何使训练更快、模型更好,或者给出保证,保证是个强词,但有点像模型在性能、准确性等方面可操作的区间。

So, that's exactly. So, we came from this, this was one area of motivation. We met the other people as I mentioned, that are all on the symmetries of the group, operations you can do. There's another world of, there's this mode connectivity that gets rebasing where you can align models along some reference models. And if you just think about this permutation symmetries, there is some in a loss landscape that is connected. There's so many points that are the same but they are different with respect to the weight space. And just understanding or trying to understand how loss surfaces look like, how models can evolve along orbits, this is great work from Bull and then the Rose from UC San Diego on all this work kind of helped me to understand better what actually happens during learning and then hopefully we could move this in our backbone learning, because some tokenization, some positioning coding in the autoencoder is a transformer autoencoder, we took some of the idea, it helped us. And then, the people are also discovering grokking and phase transitions and how models suddenly kind of converge or suddenly don't work and then suddenly work. So, this also is connected and we'd love to explore this direction more to kind of understand how can we make training much much faster, better models, or give guarantees, guarantees is a strong word, but kind of bands of where models are operational with respect to their performance, accuracies, etc.

与先前工作的联系 Connection to prior work

Host

是的,我想我们上次讨论权重空间时,我不认为我们把它称为权重空间学习,但这种内省权重的想法是和 Charles Martin 一起的,他当时在研究我们的 weight watcher、grokking 和 motor labs 等等。

Yeah, I think the last time we covered kind of weight space, I don't think we talked about it as weight space learning, but kind of this idea of like introspecting weights was with Charles Martin who with what he had to work on with our weight watcher and grokking and motor labs and all that.

Damian

所以,Charles 和 Charles 与 Michael Mahoney 合著,而 Michael 后来也合著了我们的论文。所以,是的,是的,这真的很有趣,因为 Michael 和 Charles 在做分析权重的工作,观察它们的形状如何变化,这样你就可以判断它们是否收敛,他们显然扩展了非常了不起的工作。所以,这也是并行发生的。我把这看作是一项分析性的工作,而我们的则是一项学习性的工作。

So, Charles and Charles co-authored with Michael Mahoney who was then co-authored in our paper. So, yeah, yeah, and it was really funny because Michael was doing with Charles the work of analyzing weights and looking how their shapes are changing so you can make a statement about if they converge or not and they extended obviously really amazing work. So, it did this also happen in parallel. I'm looking at this like an analytical piece of work and where ours is like a learning piece of work.

在Hugging Face模型上的扩展与训练 Scaling and Training on Hugging Face Models

Damian

嗯,所以我们希望,你知道,在某个时候我们能扩展我们的主干网络,处理更多样化的模型,使用不同类型的架构。我们现在有工作能训练了——这就是原因,对吧?你能从 Hugging Face 的开源权重仓库训练不同规模、不同架构、不同任务和模态的模型吗?我们能从 Hugging Face 下载所有东西,然后训练一个神经网络的基础模型吗?

Uh, so, you know, we hope that at some point we can scale up our backbone and process more diverse models, use different types of architectures. We have work where we are now able to train—and that's the reason, right? Can you train different-sized models with different architectures, tasks, and modalities from open-weight repositories at Hugging Face? Could we download everything from Hugging Face and train a foundation model of neural networks?

Host

而且你已经开始了这条路。我想……

And you've started down that path. I think the...

Damian

这个兔子洞,是的。

The rabbit hole, yes.

Host

那个兔子洞的海报——我最初在 GGC 看到的那张海报,让我找到你的那张,是关于在 Hugging Face 模型动物园上训练的,对吧?

The poster that that rabbit hole—the poster that I originally saw at GGC that led me to you was something about training on the Hugging Face model zoo, right?

Damian

没错。所以,在我们能够扩展之后,我们就在想,能不能——因为我们仍然局限于模型动物园,对吧?所以,比如,我可以告诉你我能生成一个新的神经网络,但我需要先训练这 1000 个神经网络才能生成那一个。所以你说这很棒,但现在我们有 1001 个神经网络了。那我们得到了什么,对吧?最终我们赢在哪里?所以下一步是,我们能不能在已有的模型上训练?而且有很棒的工作在分析 Hugging Face 如何像 Ilya Sutskever 和 Yarin 的模型图谱,而且有很多模型。所以我们能不能下载这些模型,然后不管它们是什么架构、在什么数据集上训练的,都用我们的机制来处理?这有点棘手,因为它们有不同的序列长度、不同类型的神经网络层。所以我们必须注入一些信息。但我们成功训练了第一个权重空间学习模型,可以对这类神经网络的权重进行生成和分析,而且是在 Hugging Face 模型上做的。这是 Daniel Fagg 的工作,真的很了不起。而且我们再次觉得做这个要难得多。但你知道,你必须扩展,你需要机制,你需要那些小技巧来处理那些不同的——你知道,分词器需要适应任意架构。就是这样,是的。

Exactly. So, after we were able to scale up, we then thought, can we—because we're still limited to the model zoo, right? So, like, I can tell you I can generate a new neural network, but I need these 1,000 neural networks to have trained before I can generate that one. So you tell me that's great, but now we have 1,001 neural networks. So what are we gaining, right? At the end, what are we winning? So the next step was, can we train on models that are out there? And there's amazing work on analyzing how Hugging Face looks like the model atlas from Ilya Sutskever and Yarin, and there are a lot of models. So can we download those models and then, independently of what kind of architecture they have or what dataset they trained on, use our machinery? And it's a little bit tricky because they have different sequence lengths, different types of neural network layers. So we have to put some information into it. But we were able to train the first weight-space learning model that can do generation and analysis of weights of this kind of neural network, and it does so on Hugging Face models. So this was work by Daniel Fagg, which is really amazing. And again, we thought it was much more challenging to do this. But you know, you have to scale, you need the machinery, and you need those little tricks how to handle those different—you know, the tokenizer needs to be adapted to arbitrary architectures. That's the thing, yeah.

Host

那么需要对 Hugging Face 庞大的模型库应用什么过滤器,才能将它们标准化成你可以处理的东西?

And what's the filter that needs to be applied on the Hugging Face vast library of models that normalizes them to something you can deal with?

Damian

首先,这也是其他人发现的,比如在 Yahoo。Hugging Face 上有很多内容没有文档。大约 30% 的模型没有任何有意义的元数据。所以你不知道——我的意思是,Hugging Face 很棒,对吧?

First of all, and this is work that others figured out too, like in the Yahoo. There's a lot of content on Hugging Face that's not documented. Around 30% of the models don't have any meaningful metadata. So you don't know—I mean, Hugging Face is amazing, right?

Host

就把那些都去掉?

Just get rid of all of those?

Damian

要知道哪些模型是有用的。所以我们做了一些实验——如果我们扩展,仅靠扩展本身是不够的。你需要增加模型动物园的多样性。所以我们想要多样化的模型。所以我们想要不同的数据集。主要集中在计算机视觉。语言是下一步。我们开发了一个评分函数,看模型有多受欢迎,是父模型还是衍生作品?里面有很多树。为了下载一组——我们有 20,000 个模型,其中 2,000 个通过了质量检查。从这些有数十亿参数的模型中,我们在开源权重模型上训练了主干网络。然后它可以采样不同的架构。我们可以采样 ITs、ResNets。这整个事情都集中在计算机视觉上。我们实际上能够采样一个 GPT-2 模型。它仍然是一个小模型,但存在领域变化,我们采样的这个模型被用作初始化,所以它比在常规语言数据集上训练更快。所以从计算机视觉模型到语言模型有一些知识迁移。这也很有趣,因为我们仍在试图弄清楚那些封装和编码的方式是什么。所以这有点像——下一步将是扩展它,并在不同的模态、任务和架构上训练。

Knowing which models are helpful. So we did a little bit of experiments—if we scale, scaling alone doesn't help. You need to increase the diversity of the model zoo. So we want to have diverse models. So we want to have different datasets. Mostly focus on computer vision. Language is the next. And we kind of developed a scoring function on how popular is the model, is it a parent or is it some derived work? There's a lot of trees in there. To download a set of two—we have 20,000 models and from them 2,000 models that pass some quality checks. And from these models that have billions of parameters, we trained the backbone on open-weight models. That then can sample also different architectures. We can sample the ITs, ResNets. This entire thing was focused on computer vision. And we're able actually to sample a GPT-2 model. It's still a small model, but there is a domain change, and this model that we sampled we used as an initialization so it can train faster as compared to training it on regular datasets that are language. So there is some knowledge transfer happening from computer vision models to language models. That's also interesting because we're still trying to figure out what are those ways of encapsulating, right, and then encoding. So that's kind of—and the next step would be then to scale it up and to train on different modalities, tasks, and architectures.

Host

你提到必须用一些技巧才能使用不同类型的模型,你特别提到了分词器。再深入讲一下,也谈谈你为此必须采用的其他技巧。

You mentioned that there were some tricks that you had to employ to be able to use different types of models, and you mentioned specifically tokenizer. Dig into that a little bit more and also talk about some of the other tricks that you had to employ to do this.

Damian

所以好的模型很重要。训练中使用的模型权重的多样性很重要。分词和 token 处理发生了很大变化,受到新加坡 Kai Wang 工作的启发。这帮助我们用一种不可知的方式识别或处理权重,这样我们就不受限于,你知道,这是层开始,这是层结束。

So good models are important. Diversity is important in the model weights you use for training. The tokenization and the processing of tokens changed strongly, inspired by work from Kai Wang in Singapore. So that helped us to identify or to process weights in an agnostic way so that we are not bounded by, you know, this is a layer starting, this is a layer ending.

Host

那么明确一下,我们说的是你的东西是一个模型,对吧?所以它有自己的分词器,还是说我们在说你摄入的模型的分词器的标准化,还是两者都有?

And so to be clear, are we talking about your thing is a model, right? And so it has its own tokenizer, or are we talking about normalizing the tokenizer of the models that you're ingesting, or both?

Damian

所以,好的,这是个好问题。我们从 Hugging Face 下载模型。我们按读取顺序剥离权重,非常笨拙。可能有更好的方法。所以我们有点破坏了度量结构。然后我们只是序列化或将其平铺。然后它是一个很长的序列,你知道,数百万或数十亿的参数。

So, okay, that's a good question. We take the models that we download from Hugging Face. We strip away the weights in reading order, very stupidly. There are probably better ways of doing it. So we kind of destroy the metric structure. And then we just sequentialize or get to rest that. And then it's a long sequence of, you know, millions of parameters, or billions.

Host

明白了。所以它只是数字,然后你的摄入过程中有一个分词器。

Got it. So it's just numbers, and then you've got a tokenizer as part of your ingestion process.

Damian

还要归一化。所以归一化也起着重要作用。在权重上归一化的不同位置——作为预处理、在分词期间、或在损失函数处。我们尝试了不同的方法,目前这有点取决于 Hugging Face 数据。我们在分词期间在批次内进行归一化。所以就是这样——在其他设置中,我不知道,但这很重要。分词很重要。我们仍在尝试损失函数,因为我们想获得一些高保真、高频信息。所以我们在这方面仍然有点吃力。举个简单的例子,当我们生成一个模型时,模型有点受损,所以我们需要一些微调步骤来恢复它。

Also normalize. So normalization also plays an important role. The different positions where you can normalize on the weights—as a preprocessing, during the tokenization, or at the loss function. We tried different things, and currently it depends a little bit on the Hugging Face data. We normalize inside the batches during tokenization. So that's—and in other setups, I don't know, but this is important. The tokenization is important. We're still playing around with the losses because we want to get some of the high-fidelity, high-frequency information. So we're still kind of suffering a little bit with that. So just as a simple example, when we generate a model, the model is a little bit damaged, so we need some fine-tuning steps to recover that.

恢复微调与神经网络基础模型的愿景 Recovering Fine-Tuning and the Vision of a Foundation Model of Neural Networks

Damian

很多时候这些微调步骤能很快恢复,但鉴于我们的解码器存在模糊权重的问题,我们无法得到完美的权重,对吧?所以这是我们正在努力的方向,因为如果你设想一个世界,我们能在所有这些数据上训练一个神经网络的基础模型,并按需采样你想要的任何模型,那我们实际上就取代了预训练,对吧?那为什么还要有预训练模型呢?你只需采样你需要的模型。这就是神经网络基础模型的愿景。

Very often these fine-tuning steps are very quickly able to recover, but given that our decoder has this problem of the blurry weights where we're not getting perfect weights, right? So that's something that we are working towards because if you think about a world where we could train on all this data a foundation model of neural networks and sample on demand your favorite model, whatever you need, then we would be actually replacing pre-training, right? So why have a pre-trained model? You just sample the model that you need. So that's the vision of this foundation model of neural networks.

Damian

我们还有一个额外的小问题,也带来一些麻烦。我们需要采样一个模型,并确定采样的锚点。而这个锚点必须是一个我们通过编码器处理的模型。所以这个模型越好,采样的模型就越好。这导致我们需要一个训练良好的模型来生成另一个训练良好的模型,如果是在同一领域,这说不通,因为既然我有了一个模型,为什么还要生成一个呢?

We have one additional little thing that also causes a bit of trouble. We need to sample a model and anchor where we sample. And this anchor needs to be a model that we put through the encoder. So the better this model, the better the sample models. This leads to a situation where we need a well-trained model to generate another well-trained model, which doesn't make sense if it's in the same domain because well, I have a model, why should I generate one?

Host

鸡生蛋、蛋生鸡的问题,对吧?

Chicken and egg problem, right?

Damian

没错。所以,我们有一篇即将在 CVPR 上发表的关于遥感领域的论文,我们利用我们的机制,从我们拥有的图像中生成遥感模型或遥感基础模型。然后我们面临领域变化,我们的编码器-解码器提供知识迁移,生成比 ImageNet 微调能达到的更好的模型。这里我们实现了真正的知识迁移,非常棒,我们能够超越或与当前模型(如 ICLR 上发表的 Terra FM)性能相当。我记得作者声称他们训练了 12,000 GPU 小时,而我们能在 350 GPU 小时内完成。这大约是 20、25、30 倍的差距,取决于你怎么算。这突然变得有趣,因为你是从模型而非数据训练。所以如果你想想,我们正在耗尽数据。这就是为什么缩放定律被以不同方式看待,大家都在转向测试时适应、测试时训练。我们用于训练大型模型的数据快用完了,但我们没有利用旧模型的权重。那为什么不利用这些权重、所有知识、人们投入的所有算力呢,对吧?

Exactly. So therefore, we had this paper that we're going to publish soon in CVPR about remote sensing where we take an image that we have, use our machinery to generate remote sensing models or remote sensing foundation models. Then we have the domain change where our encoder-decoder is providing knowledge transfer to generate models that are better than the ones that ImageNet fine-tuning would be able to reach. So here we have a true knowledge transfer which is really great, where we are able to outperform or be equal in performance with current models like Terra FM published at ICLR. I think the authors claim they trained for 12,000 GPU hours and we are able to do this on 350 GPU hours. That's a factor of, I don't know, 20, 25, 30 depending on how you count. And this is suddenly interesting because you train from models and not from data. So if you think about it, we're running out of data. That's the reason why the scaling laws are a little bit considered differently and everybody is moving into test-time adaptation, test-time training. We're running out of data to train the large models, but we are not using the weights of older models. So why not use the weights, all the knowledge, all the compute that people invested, right?

Host

这也是有趣的背景。如果所有数据,如果我们确实在耗尽数据,至少在视觉文本等特定领域,那么所有这些数据已经存在于一堆模型中了。为什么要复制它,为什么不找到不同的方法从现有模型中提取它呢?这基本上就是前提,对吧?

That's also interesting context. Like if all the data, if we're in fact running out of data, at least in specific domains like visual text, etc., then all that data is already in a bunch of models. Why replicate that and why not just find different ways to slurp it out of the existing models? It's essentially the premise, right?

Damian

这一切都说得通。在遥感社区,根据一些调查,我们有 70 个基础模型。而且还有人在训练第 71 个、第 72 个,对吧?那为什么不利用这些知识,将其全部压缩到权重空间学习表示中,然后按需采样模型呢?因为如果你有一个 ViT 基础模型,有 8 亿、9 亿参数,有人将其微调到某个任务,这个人,即实践者,需要使用全部 9 亿参数,这需要大量算力,而且性能可能只比 5000 万、4000 万参数的 ResNet 好一点。所以你在基础模型世界中突然被限制在大模型上。而用我们的机制,用我们的权重空间学习方法,你可以采样一个大模型,可以采样 ResNet,可以采样 EfficientNet,取决于你需要什么。对吧?你可以把架构作为期望输出,然后我们为任何架构采样参数,这对边缘设备、基础模型或其他用例都有帮助。

The whole thing makes sense. In the remote sensing community, we have 70 foundation models according to some surveys. And there are still people training the 71st, 72nd one, right? So why not take this knowledge, compress it all into a weight-space learning representation, and then sample models on demand? Because if you have a foundation model that is a ViT with 800, 900 million parameters, and somebody fine-tunes it to a task, this person, the practitioner, needs to use all the 900 million parameters, and that's demanding compute, and maybe the performance is a little bit better than a ResNet with 50 million, 40 million. So you're suddenly bounded in this foundation model world to large models. While with our machinery, with our weight-space learning approach, you could sample a big model, you could sample a ResNet, you could sample an EfficientNet, depending on what you need. Right? You can give the architecture as a desired output, and then we sample the parameters for whatever architecture, which is helping for edge devices, helping for foundation models, or helping for other use cases.

Host

所以你也可以在某种意义上把它看作一种压缩技术。

So you can also think of it as kind of a compression technique, in a sense.

Damian

没错。有趣的问题是我们压缩的是什么,对吧?然后有多少冗余,我们能否用剪枝和蒸馏在训练中做到这一点?到目前为止,我们用原始数据训练,但我们可以生成更小的模型,这些更小的模型比拿原始模型做师生蒸馏得到的更好。这也是我们在遥感场景中展示的。所以目标真的是拥有一个大型神经网络基础模型,然后能按需生成模型。不过还缺一块拼图,我们有一篇正在审稿的论文可能解决这个问题。所以如果你想想,我们需要一个模型作为提示来获得锚点以采样其他模型。可以是领域变化。所以真正理想的是你不需要这个模型作为提示,而是可以用你的数据集来提示。

Exactly. The interesting question is what are we compressing, right? And then how much redundancy is there, and can we do it with pruning and also distillation in training with that? Until now, we train with the raw data, but we can generate smaller models, and these smaller models are better than if you would take the original model and do distillation with a teacher-student. So that's also what we show in this remote sensing scenario. So the goal is really like having one big foundation model of neural networks, and then able to generate on-demand models. There's one missing piece, though, and we have a paper currently in review that might solve that. So if you think about that, we need a model as a prompt to get an anchor to sample other models. Can be a domain change. So what would be really nice is if you would not need this model as a prompt, but you could prompt with your dataset.

Host

所以给我一个在这数据上表现好的模型。

So give me a model that works well on this data.

Damian

没错。而且如果这能以保护隐私的方式完成,那就更好了,对吧?因为那样人们就能在不泄露数据集中个体成员的情况下使用它。开放。已经有类似的工作了,Kai、Andreas 和其他人正在研究,我们在这方面推进,因为我们有模型动物园,对吧?我们有数据集,也有模型。所以我们有图像,也有模型。为什么不训练一个像 CLIP 对齐文本和图像那样的对齐空间呢?然后我们可以用一个数据集编码器作为模型提示,进入这个空间并指向这个空间。我们不用模型提示,而是用数据集提示。所以,比如说你是一家银行、金融机构、医疗保健提供商等等。你不泄露你的数据。你有一个数据集,比如 100 个样本、1000 个样本。你创建一个数据集嵌入。这样你无法推断出个体成员或样本。你把这个嵌入给我们。我们提供权重。我们给你权重。你在继续训练时会快得多。所以这会非常有趣,对吧?因为那将打开所有不在 Hugging Face 上的潜在数据集。

Exactly. And if this could be done in a privacy-preserving way, it would be even better, right? Because then suddenly people could use this without revealing individual members of this dataset. Open. So there's work that does something similar already, and Kai, Andreas, and others are working on that, and we kind of move this forward because we have model zoos, right? We have datasets, and we have models. So we have kind of images, and we have models. And why not train an aligned space like CLIP did with text and images? Then we could use actually a kind of dataset encoder as a model prompt that could move into this space and point into this space. Instead of a model prompt, we use a dataset prompt. So, you know, you're a bank, you're a financial institution, health care provider, whatever. You don't reveal your data. You have your dataset of, I don't know, 100 samples, 1,000 samples. You create one dataset embedding. So you cannot infer the individual members or samples. You give this embedding to us. We provide you the weights. We give you the weights. You're much faster in continuing training. So this would be really interesting, right? Because that opens up to all the potential datasets that are not on Hugging Face.

Host

那么隐私保护的角度在那里吗?因为你的过程反正会从那个嵌入开始,所以谁产生它并不重要?还是说这是一种妥协,你可以用这个过程做到,但如果你有实际数据,你可能会得到更多?

And is there the privacy-preserving angle there because your process would start with that embedding anyway, so it doesn't matter who produces it? Or is that a compromise that you could do it with the process, but you could probably get more out of it if you had the actual data?

Damian

好问题。这取决于我们使用的方法,因为想象你有一个包含 1 万、100 万张图像的数据集。

Good question. So it comes by the method that we used because imagine you have a dataset with 10,000, 1 million images.

数据集提示 Data Set Prompting

Damian

你需要某种聚合。你不能用所有单独的图像来提示。你需要一个带有某个东西的提示。

You need some kind of aggregate. You cannot prompt with all individual images. You need a prompt with a thing.

Host

所以,你需要一个带有某个东西的提示。一个东西或一个向量,对吧?

So, you need a prompt with a thing. One thing or one vector, right?

Damian

而且你需要聚合这些知识。所以,就像你知道的,我们做一个句子,然后你有一个文本嵌入来提示你生成的图像。所以,它随之而来。显然,你必须检查,你知道,信任成员攻击,你必须添加一些噪声,等等。但想法真的是,一旦我们有了这个嵌入,我们能否从这个数据集嵌入生成现在用于权重的令牌或生成令牌的嵌入?

And you need to aggregate this knowledge. So, like, you know, we do a sentence and you have a text embedding to prompt your image that you generate. So, it comes with that. Obviously, you have to check for, you know, trust membership attack, you have to add some noise, etc. But the idea is really like once we have this one embedding, can we generate from this one data set embedding now the tokens or the embeddings that generate the tokens for the weights?

Host

这会打开局面,然后问题是,如果你能开放到人们不愿意分享或因法规不允许分享的数据集。

And this would open up and then the question is if you could open up to data sets that people are not willing to share or not allowed to share because of regulation.

Damian

一个金融机构可能很乐意分享,但他们不被允许。我有一个与德国联邦银行(德国国家银行)合作的外部博士生。他们不能分享,但你知道,他们可能出于那个原因使用那些嵌入。

A financial institution that would maybe love to share, but they're not allowed. I have an external PhD student with the Deutsche Bundesbank, the National Bank of Germany. They cannot share, but, you know, they might use those embeddings for that reason.

Damian

而且它也会让我们有点理解我们正在学习的空间,这个潜在空间,因为最有趣的事情是我们可以插值我们在 Hugging Face 上看到的模型。对吧?所以,一个数据集,按定义,从未出现在 Hugging Face 上,也没有模型在其上训练,这个数据集和数据集提示能否生成一些有意义的神经网络,这些网络存在于已知神经网络的嵌入空间之间?

And it would also allow us to kind of understand how is the space that we're learning, this latent space, because the most interesting thing would be that we could interpolate the models that we see on Hugging Face. Right? So, can a data set that, by definition, is never on Hugging Face and there's no model that's trained on that, can this data set and the data set prompt generate some meaningful neural networks that live between the space of known neural networks that are embedded?

Damian

所以,我们是否有一个表现良好的潜在空间?因为有一天,如果我们能证明我们有一个,这个机制可以用于生成神经网络,只需最少的预训练或完全不需要训练。所以,想象一个世界,你可以按需获得神经网络,在前向传播中按需超个性化。

So, do we have a well-behaved latent space? Because someday, if we could show that we have one, this machinery could be used for generating neural networks with minimum pre-training or training at all. So, think about the world where you could have on-demand neural networks, on-demand hyper-personalization in a forward pass.

Host

我的意思是,这听起来比我们今天思考神经架构搜索的方式简单得多,那是一个非常复杂的机制。

I mean, it sounds a lot simpler than the way we think about like neural architecture search today, which is a lot of very complex machinery.

Damian

完全正确。到最后,我的意思是,你想要处理的外部信号和什么类型的架构,对吧?我们只是生成权重。所以,我认为这项工作是对神经架构搜索的补充。

Exactly. To the end, I mean, the external signal you want to wrestle and what kind of architecture, right? And we just generate the weights. So, I consider this work complementary to the neural architecture search.

Host

啊,所以那可能定义结构,你可能提供适合该结构的权重。

Ah, so that might define the structure and you might provide the weights that fit into that structure.

Damian

完全正确,是的。或者,你知道,我想要这种类型的架构,对吧?给我一个 Transformer,给我任何东西,对吧?因为这是,按定义,按设计已经固定的。我们提供最好的权重。

Exactly, yeah. Or, you know, I want to have this type of architecture, right? Give me a transformer, give me whatever, right? Because this is, per definition, per design already fixed. And we provide the best weights.

Host

它在多大程度上产生权重时带有关于这些权重将用于的架构的知识,还是就像你有一个参数,你知道,它吐出的权重数量,然后你必须作为后处理以给定方式将其应用于架构?

To what degree does it produce weights kind of with knowledge of the architecture in which those weights will be used, or is it just like you've got a parameter, you know, the number of weights that it spits out, and then you have to as a post process apply that to an architecture in a given way?

Damian

完全正确。所以,到目前为止,使用模型提示,你给出一个架构;使用数据集提示,你不给。所以,你需要给出一个外部信号,比如给我一个 ResNet 18 或 ViT。但使用模型提示,你给出一个我们用权重填充的架构。肯定有更聪明的方法来构建一个标记器,它能理解处理了哪些类型的架构元素。所以,这将是下一步,以更了解这一点,并允许解码器生成条件架构。我们直到现在都不这样做。我们生成权重。我们希望在学习训练过程中,骨干网络看到了足够多的相同架构实例。所以,我们可能只能重建——不,我们肯定只能重建我们在 Hugging Face 集合中看到的架构。但如果我们能在此基础上采用这种条件信号,可能会好得多。有些人在做这个,比如新加坡、台湾和 KAIST 的人,他们在其上做扩散,然后你有一个条件数据集信号和条件架构信号。所以这也将是下一步。但总是在承诺下,我们可以在多样化的开放权重模型上做到这一点,所有混乱,对吧?因为那是从这个权重空间学习想法中能产生最多用途的地方。

Exactly. So, until now, with the model prompt, you give an architecture; with the dataset prompt, you don't. So, you need to give an external signal like give me a ResNet 18 or ViT. But with the model prompt, you give an architecture that we fill with weights. There are definitely more clever ways of building a tokenizer that kind of understands what type of architecture elements are processed. So, this will be the next step to be more knowledgeable about that and to also allow the decoder to generate conditioned architectures. We don't do this until now. We generate the weights. We hope that during learning and training, the backbone saw enough instances of the same architecture. So, we probably only can reconstruct—no, we definitely can only reconstruct architectures that we saw in our Hugging Face collection. But it would be probably much better if we could take this conditional signal on top. And some people are doing that, like people in Singapore, Taiwan, and also KAIST, they do diffusion on top of that, and then you have a conditional data set signal and conditional architecture signal. So this would be also the next step to go. But always under the promise we can do it on open weights models that are diverse, all the chaos, right? Because that's where the most use can be generated from this idea of weight space learning.

Host

模型是产生一个权重序列,然后你必须映射到架构中的位置,还是模型也产生一个权重和架构中的位置?

Does the model produce a sequence of weights that you then have to map to a position in an architecture, or does the model also produce like a weight and a position in an architecture?

Damian

所以,它生成一个权重序列,然后你必须适应架构,每个随后转换为权重的令牌都有一个位置编码。层内的哪个位置,哪一层或哪个块,然后有绝对位置。

So, it generates a sequence of weights that you have to then fit to the architecture, and every token that is then translated into a weight has a position encoding. Which position within the layer, which layer or which block, and then with absolute position.

Host

但那是输入侧的。

But that's on the input side.

Damian

在输入和输出上。因为它是一个自编码器,是的。

On the input and on the output. Because it's an autoencoder, yeah.

Host

啊,好的。明白了。明白了。明白了。

Ah okay. Got it. Got it. Got it.

Damian

所以,然后你得到一个权重序列,你必须将其适应正确的架构。你可以将其适应不同的架构,然后切割或切片切块。那将是一个有趣的实验,看看会发生什么。你需要多少微调步骤来修复它?这也是为什么我们仍然需要几个微调步骤来让模型快速达到性能的原因之一。因此,我认为条件化会帮助我们更快地生成更好的权重,这些权重足够多样化。所以这将是潜在的下一步,对吧?移动并条件化解码器,并查看解码器。

So, then you get a sequence of weights and you have to fit it to the right architecture. You could fit it to a different architecture and then cut it or slice and dice it. That would be an interesting experiment what happens then. How many fine-tuning steps do you need to repair that? And that's for me also one of the reasons why we still need a couple of fine-tuning steps to get the model very quickly up to performance. And therefore, I think that conditioning would kind of help us to generate better weights quicker, that are diverse enough. And so this would be potential next steps, right? To move and condition decoder and look into the decoder.

Host

当你产生这些权重时,你是像覆盖一个初始化模型,比如随机初始化或零初始化之类的,这样模型总是有效的,还是你必须考虑,嗯,你知道,模型实际上没有幻觉,它为这个位置生成了两个权重,而为那个位置没有生成权重,那种事情?

When you produce these weights, are you like overriding an initialized model, like random initialization or zero initialization or something, so that the model is always valid, or do you have to think about, well, you know, the model didn't actually hallucinate it and it generated two weights for this position and no weights for that position, that kind of thing?

Damian

我们完全覆盖。所以,我们替换模型中随机初始化的任何内容,然后我们完全覆盖并将检查点加载到该架构中。由于我们有序列并且知道位置编码,这在技术上是直接的。无论模型做什么,它可能会有重复的打嗝。没有可见的模式,就像我们可以观察到的伪影,那种一遍又一遍重复的东西。

We entirely overwrite. So, we replace whatever is in the model randomly initialized to have that, and then we entirely overwrite and load the checkpoint into that architecture. Since we have the sequence and we know the position coding, this is technically straightforward. Whatever the model does, it could have hiccups of repetition. There's no visible pattern that would be like an artifact that we could observe that is kind of repeating again and again and again and again.

生成多模型与任务向量 Generating Multiple Models and Task Vectors

Damian

但你也可以考虑不只生成一个模型,因为这是前向传播,成本很低,你可以生成一大堆类似的模型,对吧?然后拟合它们,你就突然有了这个新的自由度可以生成。但没错,我们相信我们的方法是我们可以完全覆盖已有的任何东西。目前有一些关于任务向量的有趣工作正在进行。

But you could also think about not only generating one model, but because it's a forward pass, it's very cheap, you can generate a whole bunch of similar models, right? And then fit them, and you suddenly have this new degree of freedom that you can generate. But yeah, we believe our way is that we can entirely overwrite whatever is there. There's some interesting work that currently is happening on task vectors.

Host

什么是任务向量?

And what is a task vector?

Damian

比如,你有一个基础模型,然后你针对一个任务微调,再针对另一个任务,第三个任务。你可以取微调后的模型与基础模型的差值,这就是一个任务向量。你可以用它做任务算术。在权重空间中,人们主要在语言模型上使用,有时也在计算机视觉中使用。当你处理这些任务向量时,我们在某些情况下观察到,它们比其他情况更容易学习和生成。在其他情况下,完整权重更容易,所以我们凭经验观察到这一点,但我们无法解释为什么在某些设置下一种更容易,另一种则不然。

So, you have a base model, for example, and you fine-tune to one task, another task, and the third task. You could take the difference of the fine-tuned model to the base model, and this would be a task vector. And you could do a task arithmetic with that. And then in weight space, people are using that in particular on language models, sometimes in computer vision. And when you work on those task vectors, we observe in some cases that it's easier to learn and easier to generate than in other cases. In other cases, full weights are easier, so empirically we observe that, but we cannot explain why for some setups the one is easier or the other.

频域与权重处理 Frequency Domain and Weight Processing

Host

你几次提到高频噪声、低频噪声和权重之间的区别,这让我想到在频域中做事情,比如对权重应用 FFT 之类的。这是人们在研究的方向吗?

You mentioned a couple of times the distinction between high frequency noise, low frequency noise, and the weights, which calls to mind like doing things in the frequency domain, applying an FFT to weights or something like that. Is that something that people are working on?

Damian

据我所知,没有。但我认为有趣的是,因为我们生成的是某种模糊的基础,然后你可以在上面添加一些高频信息。我认为在图像领域,这是通过变分自编码器重建,然后生成对抗网络强制高频信息来实现的。然后你结合两者的损失。所以,你可以在我们的权重领域做类似的事情。我不知道是否有人真的做过。但将权重移到其他领域然后处理是一个有趣的想法。不确定频率是否是正确的方式。但你也可以考虑很多预处理步骤,直到你进行权重空间学习、压缩或学习潜在表示、低维流形。也有人朝这个方向走,但让我们看看。我们目前不走那条路。

To the best of my knowledge, not. But I think what would be interesting is because what we generate is kind of this blurry base, and then you could add some high frequency information on top. And I think in the image domain, this is happening through we have a kind of variation autoencoder reconstructing, and then you have a generative adversarial network enforcing high frequency information. And then you combine loss of both of that. So, you could do things like that in our domain on weights. I would not know if somebody really did this. But moving the weights into some other domain and then processing is an interesting idea. Not sure if frequency would be the right way of doing it. But you could think also a lot about a lot of preprocessing steps until you would go and do the weight space learning, compression, or learning the latent representation, the lower dimensional manifold on that. There are people who are also moving into that directions, but yeah, let's see. We're not going along that path currently.

工作坊延续与社区 Workshop Continuation and Community

Host

今年研讨会还会继续吗?

And will the workshop be continuing this year?

Damian

还有另一个很棒的研讨会,由我们的一些同事在 ICML 组织。这是一个关于权重对称性的研讨会。我们考虑在下一个机会继续举办,那就是 NeurIPS,但让我们看看情况如何。目前,我们有一些博士生即将毕业,新的博士生也来了,所以社区有点断层。随着大量优秀人才完成博士学业,进入博士后岗位,肯定会有延续。最重要的经验之一是认识到有一个社区,一旦社区存在,就要努力发展共同语言、共同基准测试,以及作为社区一致的想法,并沿着这个方向前进。我认为现在是时候朝这个方向努力,把这些想法汇集起来,同时也有自由去探索其他想法和方向。

There's another great workshop that is organized by some of our colleagues at ICML. It's a workshop on weight symmetries. And we're thinking in continuing it for the next opportunity, which would be then NeurIPS, but let's see how this looks like. Currently, we have a couple of people that were PhDs are finishing, new PhDs are coming, so there's a little bit of a gap in the community. With the inflow a lot of great people are currently finished with the PhDs on the market already in postdoc positions. So, there will be definitely a continuation on that. And like one of the most important experiences was recognizing that there is a community and once the community is there, trying to develop a common language, trying to develop common benchmarking, trying to develop some ideas that are coherent as a community and moving along that way. I think this is now the time to try to move in that direction and bring those ideas together with the liberty of also going to the one or the other idea and direction.

互补研究方向 Complementary Research Directions

Damian

比如,Hagai Maron 有一些有趣的工作,他把权重空间学习的思想应用到梯度或激活空间。我在做表示工程。你也可以考虑不是权重,而是训练期间的梯度或激活的神经产物。Yedid 也有关于探测神经网络的有趣工作。所以,不是看权重,而是有受控的输入输出关系,从而在给定未知神经网络时,当我控制某件事时观察会发生什么。

There's interesting work from Hagai Maron for example, and he's doing taking the ideas of weight space learning, applying it to gradients or applying it to activation spaces. I'm doing representation engineering. Which you could also think about the neural artifact that is not the weights but the gradients during training or the activations. There's interesting work from Yedid coming around probing neural networks. So, not looking at the weights but having controlled input output relationships and therefore looking what happens given an unknown neural network when I control the one thing.

Host

Zabic 也在这些方面做了很多工作。

And Zabic's been doing a lot of work along those lines as well.

Damian

没错。所以,有相当有趣的工作可以视为对此的补充。而且总是这个问题,我们是否有这些产物的集合?我们能否学习某种表示,然后用它来预测神经网络的属性,或操纵和编辑行为、改变和修改?

Exactly. So, there's quite interesting work that could be considered complementary to that. And it's always this one thing, do we have these artifacts as a collection? Can we learn some representation and then use it to predict properties of the neural network or to manipulate and edit the behavior, change and modify?

结束语 Closing Remarks

Host

非常有趣的工作。非常感谢你参加节目,并与我们的听众分享一些相关内容。

Very interesting work. Thank you so much for jumping on and sharing a bit about it with our audience.

Damian

谢谢邀请,让我们看看三年后我们会怎样。

Thanks for having me and yeah, let's see where we will be in 3 years.

Host

当然。谢谢。

Absolutely. Thank you.

互动版:逐字朗读 + 针对本期提问 →