TabPFN:取代数小时模型调优的表格数据 AI

TabPFN: The AI That Replaces Hours of Model Tuning for Tabular Data

弗兰克·胡特 Frank Hutter · 机器学习街谈 · 2026-09-23 · 约 113 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Frank Hutter 解释 TabPFN 如何将上下文学习引入表格数据,使深度学习最终超越 CatBoost 和 XGBoost。

Frank Hutter explains how TabPFN brings in-context learning to tabular data, finally making deep learning outperform CatBoost and XGBoost.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 58)

全文 · Full transcript(中英对照)

引言与背景 Introduction and Background

Host

Frank,非常高兴你能来。欢迎来到 MLST。

Frank, it's amazing to have you here. Welcome to MLST.

Frank

谢谢邀请。很高兴能上这个节目。其实两年前就有人问我,嘿,你有可能上 MLST 吗?我说,啊,也许以后吧。现在能来真是太激动了。

Thank you for having me. It's great to be on the show. Actually, two years ago, people asked me, hey, is there a possibility that you might be on MLST at some point? And I was like, ah, maybe at some point. And yeah, super exciting to be here now.

Frank

嘿,我叫 Frank。在创办 Prior Labs 并担任 CEO(现在是联合 CEO)之前,我做了 12 年的机器学习教授。我对此非常兴奋。我主要专注于 AutoML(自动化机器学习),创办了 AutoML 研讨会系列并运营了八年。我们把它转变成了 AutoML 会议。我合著了第一本关于 AutoML 的书,做了第一个关于 AutoML 的慕课。所以我在那个领域做了很多,但后来也转向了深度学习,并在深度学习优化方面做了很多工作。所以 AdamW 和余弦学习率来自 Ilya Loshchilov 和我。所以这被用于训练世界上任何 Transformer。然后我们把两者结合起来,做表格数据的深度学习,这是以前行不通的,而通过 TabPFN 我们真正让它成功了,并且正在扩大规模,构建基础模型,非常兴奋能彻底改变这个世界。

Hey, my name is Frank. I have been a machine learning professor for the last 12 years before starting Prior Labs as a CEO now co-CEO. I'm very excited about that. I focused a lot on AutoML, automated machine learning, where I started the workshop series on AutoML and ran that for eight years. We transitioned that to a conference on AutoML. I co-wrote the first book on AutoML, did the first MOOC on AutoML. So I really did a lot in that space, but then also moved over to deep learning and worked a lot on optimization of deep learning. So AdamW and cosine learning rate is from Ilya Loshchilov and myself. So that's being used for training any transformer in the world. And then we put the two together and did deep learning for tabular data, and that is something that didn't used to work, and with TabPFN we actually made it work and are scaling this up, building foundation models, and are really excited to revolutionize this world.

表格数据的挑战 Challenges of Tabular Data

Host

我认为人们可能没有意识到的一件事是,表格数据无处不在,而且根据痛苦的经验,处理表格数据是一场噩梦。所以任何用表格数据构建过机器学习模型的人,你显然有 pandas 和 scikit-learn 里面有一些东西,但你必须处理缺失值,你知道怎么处理分类特征,你怎么做特征工程,你怎么做反映问题语义的变换。你真的需要知道你在做什么,而且感觉有点笨拙。你总是即使你能做出一个机器学习模型,你事后总是感觉有点脏,因为你觉得你做了所有这些捷径,并以某种方式破坏了它。跟我说说这个。

I think one thing that people might not appreciate is that tabular data is absolutely everywhere and speaking from bitter experience it's a nightmare to work with tabular data. So anyone who's built a machine learning model with tabular data, you've obviously like pandas and scikit-learn has got some stuff in there but you have to deal with missing values, you know what do you do with the categorical features, how do you do some feature engineering, how do you do transformations that reflect the semantics of the problem. You really need to know what you're doing and it feels kind of cludgy. You always even though you can make an ML model you always feel kind of dirty afterwards because you feel like you've done all of these shortcuts and kind of corrupted it in some way. Tell me about that.

Frank

是的,完全同意。表格数据非常脏,我认为这是深度学习花了这么长时间才真正做好的原因之一。有无数尝试将深度学习用于表格数据,比如 2019 年谷歌的尝试被大肆宣传,数千次引用。是的,表格数据的新事物,但它就是行不通。它不能泛化到新的数据集。它只是,你知道,你有异构数据,你有异常值,你有缺失值,你有分类值,你有二元值,你有有序值,你有缺失数据,随机缺失,非随机缺失,所有这些数据复杂性。是的,你需要真正考虑到这一点。而标准的深度学习方法有点像是把所有东西都当作一样的,就像任何图像都一样,你知道有像素,而且你通常有非常非常相似的空间关系等等。你可以拿图像,比如你可以有一个 ImageNet,基本上覆盖世界上各种不同的图像,并在那一个数据上学习。但对于表格数据,你实际上需要很多不同的表格,因为如果你有一个来自医学的表格,然后你有一个来自保险的表格,就单个行而言,你无法从一个表格中学到任何关于另一个表格的东西。但你可以做的是,在这些模式的层面上进行迁移,检测模式,不同特征如何交互等,潜在地检测因果关系,在那个层面上使用上下文学习,你实际上可以跨不同的数据集进行迁移,这就是为什么在上下文学习中我们终于看到了表格数据的突破。

Yeah, fully agreed. Tabular data is very dirty and I think that's one of the reasons that deep learning took so long to actually do well for it. There's been countless attempts for deep learning for tabular data like 2019 by Google was really hyped, thousands of citations. Yeah, the new thing for tabular data and it just doesn't work. It doesn't generalize to new data sets. It just, you know, you have heterogeneous data, you have outliers, you have missing values, you have categorical values, you have binary values, you have ordinal values, you have missing data, missing at random, missing not at random, all kinds of these data complexities. And yeah, you need to actually take that into account. And just the standard deep learning approach sort of like treated everything as the same, like any image as the same in terms of you know like there's pixels and you typically have very very similar spatial relationships between the pixels and so on. And you can take images like you can have an ImageNet that is just all basically cover all kinds of different images in the world and learn on that one data. But what you need for tabular data is actually yeah a whole lot of different tables because if you have one table from medicine and then you have another table from insurance there's just nothing you can learn in terms of the individual rows from one can't tell you anything about the other one. But what you can do and where you can do the transfers on the levels of these patterns, detecting the patterns how the different features interact etc., detecting causality potentially and on that level using in-context learning you can actually transfer across different data sets and that's why within context learning we finally saw this breakthrough for tabular data.

表格数据方法演进 Evolution of Tabular Data Methods

Host

是的,我的意思是,回想做这类事情我就做噩梦。我的意思是,例如,我想传统机器学习算法它们有点希望数据在某个范围内作为欧几里得空间,所以你可能必须归一化它、标准化它,一些算法,你知道,如果你使用 SVM,如果不同字段的范围不同,它可能实际上不能很好地工作,是的,这绝对是一场噩梦。但是,你能再解释一下吗?所以之前数据科学家的工作,你可能会做一大堆,你知道,交叉散点图,你会做一些像特征分析之类的,然后你把它转换成这个漂亮的凸欧几里得空间,而现在我们有这个使用 Transformer 的系统,我猜架构有点类似于深度集合,所以它实际上可以内在地理解离散数据,这是新的,这是非常非常新的,那么为什么我们现在处于一个与以前不同的世界?

Yeah I mean I'm just having nightmares thinking back to doing stuff like this. I mean for example I suppose traditional machine learning algorithms they kind of wanted the data as a Euclidean space within a certain range so you might have to normalize it standardize it some of the algorithms that you know like if you're using SVM it might actually not work very well if the ranges of the you know different fields were different and yeah it was an absolute nightmare but but can you just explain this a bit more so so before the the job of the data scientist you might do a whole bunch of you know cross scatter plots and you would do some like you know feature analysis and whatnot and you and you would convert it into this beautiful kind of convex Euclidean space and and now we have this system using a transformer and and I guess the the architecture is vaguely resembling deep deep sets so it can actually intrinsically understand discrete data and this is new this is very very new so why are we in a different world now to where we were before

Frank

它可以理解离散数据,并且通过实际上在单个列上工作的架构,或者网络中不同类型的预处理,实际上查看列的统计信息,然后确保你为该列有正确的嵌入。现在所有这些都嵌入其中,你实际上不再需要自己做了。处理缺失值、处理异常值等。预处理中有很多步骤。必须处理这些当然非常好,它让你快得多。我的意思是,这就是我们想要通过表格基础模型实现的事情之一,你知道,你真正帮助数据科学家更快地完成工作。

It can understand discrete data and also with architectures that actually work on individual columns or different types of pre-processings that are in the network that actually look at the statistics of a column and then actually make sure that you have the right embedding for that column. That's all embedded now and you don't actually really need to do this yourself anymore. Dealing with missing values, dealing with outliers, etc. There was all like a lot of steps in your pre-processing. Having to deal with that is of course very nice and it just makes you so much faster. I mean that's one of the things we wanted to get to with tabular foundation models is that you know you really help data scientists do their work so much faster.

数据科学家角色转变 The Changing Role of Data Scientists

Frank

我认为在当今这个时代,新的数据科学家真的不再需要了解 XGBoost 的一切了——某些超参数、改变它们会有什么变化等等。工作真的变了。关键在于真正思考这个问题的正确框架:数据从哪里来?会不会有数据偏移?有没有数据漂移?有没有公平性问题?有没有因果性问题?有没有隐私问题?当我部署这个模型时,会发生什么?这会如何影响数据的收集方式?那完全是另一回事,而且有一些非常令人兴奋的问题,现在人们根本没时间去看。能够在一个问题上投入更多有人文质量的时间,而不是摆弄超参数——我认为未来作为数据科学家,那会是一份更有成就感的工作。

I think in this day and age, new data scientists really don't need to understand everything about XGBoost anymore—certain hyperparameters, what changes if you change them, etc. The job really changes. It's all about actually thinking about the right framing of this problem: where does the data come from? Is there going to be data shift? Is there data drift? Are there problems with fairness? Are there problems with causality? Are there problems with privacy? When I field this model, what's going to happen? How is this going to affect how the data is being collected? That's a whole different ballgame, and there are very exciting questions there that people right now just have no time to look at. Being able to spend more human-quality time on a problem rather than fiddling with your hyperparameters—I think that's going to be a much more fulfilling job as a data scientist in the future.

表格数据为何不性感 Why Tabular Data Isn't Sexy

Host

我想,表格数据之所以不那么性感,还有另一个原因。我的意思是,很多数据科学家热爱深度学习革命的地方在于规模这个概念——你可以到互联网上去。即使,你知道,我在银行工作,可能在做文档识别,我的思维模式就是:哦,让我们收集所有文档,看看互联网上有没有其他文档。让我们构建一个数据集。让我们真正——因为有一种观念是,我可以直接扩展,我可以添加更多数据,效果就会更好。而表格数据的心态则非常——几乎像是缺乏想象力。人们只考虑我手头的这个特定数据集,而不是它如何能——你知道,不是我如何用其他数据集来丰富它,也不是我如何使用表格基础模型。

I suppose another reason why tabular data hasn't been very sexy. I mean, a lot of data scientists, what they love about the deep learning revolution is this notion of scale—that you can go out on the internet. I mean, even, you know, I'm working in a bank or something and I might be doing document recognition, and my headspace is very much, oh, let's go and gather up all of the documents and let's see if there's other documents on the internet. Let's build a data set. Let's actually—because there's this notion that I can just scale and I can add more data and it's going to work better. Whereas the mindset with tabular data is very much—it's almost like a lack of imagination. People are thinking only in terms of the particular data set I have here, not how it can be—you know, not how I can enrich it with other data sets or not how I can use a tabular foundation model.

合成数据与表格基础模型 Synthetic Data and Tabular Foundation Models

Frank

是的,绝对如此。人类喜欢把文本放到网上。他们喜欢把图像、视频放到网上。用生成式视频模型等制作一条病毒式传播的推文非常容易。但人们不会把表格数据集放到网上。这就不那么光鲜了。而且有时候思考表格数据有点抽象。但当涉及到我们的健康、我们的科学以及我们的金融系统等时,它就变得非常具体,而且非常清楚,这是一个真正重要、需要解决的问题。拥有这些在合成数据上预训练的表格基础模型,正是让我们能够捕捉到这一点的杠杆,而这是必要的。我们无法直接使用由我们、由某人为我们准备的表格数据语料库。我们只需要生成我们的合成数据,因为并没有表格——有一些表格数据,但主要就是大约 5 万个数据集。而且,是的,网上有数百万个数据集或数百万张表格,但有些表格就像维基百科上的——这个篮球运动员背上有这个号码。这不是一个你能合理地从中学到统计学习算法的数据集。这类数据集一直都不存在,我们需要生成它们,然后才能在上面学习。

Yeah, absolutely. Humans like to put their text online. They like to put their images online, their videos online. It's very easy to make a viral tweet with a generative video model, etc. But people do not put their tabular data sets online. This is just not very glamorous. And also it's kind of abstract sometimes to think about tabular data. But when it comes to our health or to our science and our financial system, etc., then it becomes just very concrete, and it's just very clear that this is a really important problem that needs to be solved. Having these tabular foundation models that are pre-trained on synthetic data was the lever that actually allowed us to capture this, and this was necessary. We just couldn't take the corpus of tabular data that has been prepared by us, by someone, for us. But we just needed to generate our synthetic data because there is no tabular—like, there's some tabular data, but it's mainly like 50,000 data sets that are there. And well, yeah, on the web there are millions of data sets or millions of tables, but there are tables like on Wikipedia—this basketball player has this number on their back. It's not a data set that you can actually reasonably learn a statistical learning algorithm from. And those types of data sets just haven't been there, and we needed to generate them in order to then learn on them.

表格数据的ImageNet时刻 The ImageNet Moment for Tabular Data

Host

为了强调这一点,表格数据已经迎来了它的 ImageNet 时刻,对吧?所以直到前天,打个比方,深度学习对表格数据都不起作用。而现在它比 CatBoost 和 XGBoost 好得多。TabPFN 是第一个真正从数据中学习、从而在其 supposed 任务上表现更好的算法。所以我们在数亿个数据集上预训练它,让它在测试集上尽可能好。这并不是偷看测试集。这不是很多人在过拟合时做的事。在测试时你不能看测试集,但在元训练时,在数亿个数据集上,你已经学会了真正解决这个问题。我们把整个数据集作为上下文输入 Transformer,Transformer 可以关注数据集的正确部分,以找出这个数据集中的模式,以及我应该应用到测试数据上以进行良好预测的模式。

And just to hammer that home, there has been an ImageNet moment for tabular data, right? So until the day before yesterday, figuratively speaking, deep learning did not work for tabular data. And now it works dramatically better than CatBoost and XGBoost. TabPFN is the first algorithm that's actually been learned from data to be better at what it's supposed to do. So we've pre-trained that on hundreds of millions of data sets to actually just be as good as it can possibly be on the test set. And that is not peeking at the test set. That's not what a lot of people do when you're overfitting or something. You don't get to look at the test set at test time, but at meta-train time, over hundreds of millions of data sets, you have learned to really solve this problem. And we're feeding the entire data set in context to the Transformer, and the Transformer can attend to the right parts of the data set in order to figure out what are the patterns in this data set and what are the patterns that I should then also apply to the test data in order to predict well.

合成数据与过拟合 Synthetic Data and Overfitting

Host

一个重要细节是,这些人合成生成了大量训练数据,这在某种程度上克服了过拟合问题——也许不是完全克服,你可以解释一下——但你知道,所以你合成数据,然后构建这个基础模型。

And an important detail is that these guys have synthetically generated a whole bunch of the training data, and this overcomes the overfitting problem to some extent—maybe not entirely, and you can explain that—but you know, so you synthesize data and then you build this foundation model.

Frank

是的,绝对如此。这里有很多强大的理论基础,我们是在近似贝叶斯后验,针对任何我们可以采样的先验。我们选择的先验包含大量因果性。所以我们基本上是在构建所有可能解释数据的结构因果模型的贝叶斯后验。与 LLM 和视觉模型等相比,我们是在合成数据上训练的。我们不得不在合成数据上训练,因为互联网上没有大量表格数据。但相反,我们是第一个完全在合成数据上训练却达到最先进水平的基础模型。这很棒,因为我们没有任何泄漏,没有任何记忆问题等。我们可以直接控制数据中确切的内容。我们不必担心偏差等,因为我们有一个从头生成合成数据的编码化流程。

Yeah, absolutely. So there's a lot of strong theoretical foundations where we're approximating the Bayesian posterior over any type of prior that we can sample from. And that prior that we chose is a prior that has a lot of causality in there. So we're basically building the Bayesian posterior of all kinds of structural causal models that could explain the data. And in contrast to LLMs and vision models, etc., we have trained on synthetic data. We had to train on synthetic data because there isn't a whole lot of tabular data on the internet. But rather, we're the first foundation model that's actually state-of-the-art yet entirely trained on synthetic data. And this is great because we don't have any leakage, we don't have any memorization issues, etc. We can directly control exactly what's in the data. We don't have to worry about biases, etc., because we just have a codified pipeline for generating the synthetic data from scratch.

TabArena与模型评估 TabArena and Model Evaluation

Host

你们运营这个网站,很多在家的人都知道,你们见过 LM Arena。它使用的算法本质上与象棋中的 Elo 算法相关,你知道,在象棋中你让人们对弈,然后根据结果的信息增益,每个玩家都会得到一个排名,并且根据参数随时间收敛。所以 Elo Marina 用语言生成做了这件事,人类可以评价一个是否比另一个更好。而你们有一个叫 TabArena 的东西。实际上是你自己运营这个网站。

You guys run this website, know many many folks at home, you've seen LM Arena. So that is using an algorithm related to the Elo algorithm in chess, essentially, where you know in chess you play people against each other and then depending on the information gain of the result, you get a rank for every single player, and it kind of converges over time depending on the parameters. So Elo Marina did that with language generation, and humans could rate whether one was better than another. And you've got this thing called TabArena. Actually run the site yourself.

Frank

是的。所以我的意思是,TabArena 最初是 Tabula 中很多不同人之间的合作。第一作者是 Nick Ericson,他当时在 AWS。他在那里构建了 AutoLue,这是迄今为止最好的表格数据 AutoML 系统,直到现在,当你把它与表格基础模型结合时,它会变得更好。

Yeah. So I mean TabArena started as a collaboration between a whole lot of different people in Tabula. The first author is Nick Ericson, who at the time was at AWS. He built AutoLue on there, which is the best AutoML system for tabular data there was until now, where you combine it with tabular foundation models and it gets much better.

TabArena与开放基准测试 TabArena and open benchmarking

Frank

嗯,还有另外一批人,我想来自五六个不同的机构。到现在,Nick 已经加入了我们的团队,还有其他几位作者也加入了我们的团队,因为他们都想真正把表格预测推到极致。而 TabArena 非常开放——当它是一个活的基准时,我们会纳入任何类型的新模型。所以当我们发现任何数据集有问题,或者有人说“嘿,为什么这个数据集在这里而不在那里?”或者“嘿,这是一个新数据集,让我们更新一下”,为了减少社区对这个数据集的过拟合,我们就会更新它。所以,如果你想参与,请务必加入,我们正在不断构建新的基准。我们有一个基准现在实际上已经超越了 Arena,超越了 TabArena。我们正在为关系数据、为特征选择构建基准,所有这些努力都是完全开源的,我们邀请任何人合作。

And yeah, also a bunch of other folks with, I think, five or six different affiliations. By now, Nick has joined our team and several other authors have also joined our team because they all want to actually push tabular predictions to the max. And TabArena is very open—we include any type of new model when it's a living benchmark. So when we see issues in any of the data sets, or people say, 'Hey, why is this data set here not in there?' or 'Hey, this is a new data set, let's update this,' in order to get less overfitting of the community to this data set, then we update it. So yeah, if you would like to get involved, by all means, we're continuously building new benchmarks. We have a benchmark that's actually now beyond Arena that goes beyond TabArena. We are building a benchmark for relational data, for feature selection, and all of these efforts are entirely open source and we invite collaborations with anyone.

Host

那它在概念上类似于 LM Arena 吗?就是有一大堆不同的类别,你抽样一个比较,然后由多样化的人群来评定哪一个更好。是类似的东西吗?

And is it conceptually similar to LM Arena? So there are a whole bunch of different categories and you sample a comparison and then a diverse population of humans rate one as being better. Is it a similar kind of thing?

Frank

是的。所以在 TabArena 中,你不需要人类评分那一步,你只需要一个训练集和一个测试集。相似之处在于你计算 ELO 分数,所以你基本上有这些锦标赛,你可以说,看,如果 ELO 分数比另一个高 100 分,那么这个算法在任何数据集样本上战胜另一个算法的概率就是这个。好的。所以这是与 LM Arena 的一个相似之处,但它基本上是一个完全开放的平台,发布所有东西,包括所有工件、所有被比较算法的所有预测,并且已经被开源社区的人复现过。所以这让任何人都更容易正确地对算法进行基准测试,而且运行你自己的基准测试在 GPU 上大约花费 2 美元。所以这是对任何人都完全包容的。

Yeah. So in TabArena you don't need to have that step where the humans rate, but you just have a training set and a test set. What is similar is that you compute ELO scores, and so you have basically these tournaments and you can say, look, if the ELO score is 100 points more than the other, then the probability that this algorithm wins against this algorithm on any of the data set samples is this. Okay. So that is one similarity to LM Arena, but it's basically an entirely open platform that publishes everything in terms of all the artifacts, all of the predictions of all the algorithms that are being compared, and has been reproduced by people in the open source community. And yeah, so that makes it much easier for anyone to benchmark the algorithms properly, and running your own benchmark costs something like $2 on a GPU. So this is something that's totally inclusive to anyone.

客观基准与数据泄露 Objective benchmarks and data leakage

Host

是的,这很有道理,因为我想对于语言模型来说,它并不客观。所以人类需要比较,因为我们基本上无法设计一个目标函数,而人类也有问题,因为人类喜欢提到《星际迷航》时的样子,他们喜欢 GBT 垃圾,各种奇怪的事情都在发生。所以在这种情况下,你有一个客观标准,你可以运行它。不过最后一件事:TAP PFN 的伟大之处在于你可以合成数据,而我们在基准测试中不想要的是——我的意思是,在普通基准测试中它发生得很离谱——你知道,因为有时你只是在基准测试上训练。但这里是否存在一种无意的泄漏形式,即你不断生成更多数据,而你可以看到基准测试上发生了什么?是否存在无意中泄漏一些数据的倾向?

Yeah, that makes a lot of sense because I guess with language models, it's not objective. So humans need to compare because we can't basically design an objective function, and with the humans that's a problem as well because humans love it when Star Trek is mentioned and they love GBT slop and all sorts of weird things are going on there. So in this case you have an objective criteria and you can run that thing. Just final thing though: so the great thing about TAP PFN is that you can synthesize data, and what we don't want in benchmarks—I mean, it happens grotesquely with normal benchmarks—you know, because sometimes you're just training on the benchmark. But is there an inadvertent form of leakage here where you're continuing to generate more data and you can kind of see what's going on on the benchmark? Is there a tendency to kind of inadvertently leak some data?

Frank

并不是我们合成数据然后这些数据进入基准测试,对吧?基准测试——那些都是来自社区的真实数据集,我们从数千个可能的数据集中精心挑选,然后因为各种原因丢弃它们,每一步都有详细记录。例如,我不知道,就像 beyond Arena,我们丢弃了所有少于 100 个样本的数据集。我们也可以包括那些,但你知道,少于 100 个样本可能没有那么多信号了,等等。所以我们丢弃重复项,我们丢弃——例如,我不知道,你可以把 MNIST 看作一个表格数据集,只是把每个像素当作一个数字,但那些数据集我们丢弃,因为我们只想要那些你实际上会使用表格机器学习算法的数据集,并且这样做是有意义的,以免被各种看起来像表格但实际上你应该只使用视觉分类器的数据集分散注意力,因为 MNIST 有空间模式等等,相邻像素之间的空间相似性是你想要利用的,如果你把它当作表格数据集,那么你有点假装表格数据中有那种类型的数据,而通常你不希望在表格数据中有这些模式。所以是的,我们因为各种原因丢弃数据集,这些原因都有很好的描述,而且再次,所有这些论文都是开源合作。所以我们很高兴任何愿意投入精力来整理这些数据集的人。这是一项糟糕的工作。现在有了智能体容易多了,但仍然很难正确定义包含特征应该是什么。然后是的,我们只是运行这些实验,保持网站更新,运行——查看各种帕累托曲线,比如预测性能有多好,训练方面有多好或多快,推理延迟等等,然后查看竞争方法的各种统计数据。

It is not the case that we synthesize data and then that data goes onto the benchmarks, right? It's the benchmarks—those are all real data sets that come from the community, and we curate them from thousands of different possible data sets and then we drop them for various reasons, and every step is exactly documented. For example, I don't know, like for beyond Arena we dropped everything that's less than 100 samples. We could have also included those, but you know, with less than 100 samples maybe you don't have as much signal anymore, etc., etc. So we drop duplicates and we drop—so for example, I don't know, you could see MNIST as a tabular data set just taking every pixel as a number, but those data sets we drop because we only want to have data sets where you would actually use a tabular machine learning algorithm and where that would make sense in order to not just get distracted with all kinds of data sets that look tabular but really you should just use a vision classifier for MNIST and because there are spatial patterns and so on and spatial similarity between neighboring pixels that you want to exploit, and if you cast this as a tabular data set then you sort of pretend that there is that type of data in tabular data and that you don't want to have these patterns in tabular data typically. And so yeah, we drop data sets for various reasons that are very well described, and again, all of these papers are open source collaborations. So we're happy about really anyone who wants to put in the effort to curate these data sets. This is terrible work. Much easier now with agents, but still it's hard to define correctly what the inclusion characteristics should be. And then yeah, we just run these experiments and keep the website up to date and run—look at all kinds of Pareto curves like how good is the predictive performance, how good or how fast is it in terms of training, in terms of inference latency, etc., and yeah, look at all kinds of statistics of the competing methods.

博士导师与AutoML历史 PhD advisors and AutoML history

Host

显然你在传奇人物 Kevin Murphy 指导下完成了博士学位。

And apparently you did your PhD under the legendary Kevin Murphy.

Frank

是的,确实。实际上有三位导师:Holger Hoos、Kevin Leyton-Brown 和 Kevin Murphy。Kevin Murphy 可能是最不相信自动算法配置的人,但他仍然对能做的其他一些事情感到兴奋。实际上最后当他去谷歌时,他对神经架构搜索变得兴奋起来。所以那时我们实际上合作得比以前更多。

Yeah, indeed. Actually three supervisors: Holger Hoos, Kevin Leyton-Brown, and Kevin Murphy. Kevin Murphy was maybe the least believer in automated algorithm configuration, but he was still excited about some other things he could do. And actually in the end when he went to Google, then he became excited about neural architecture search. And so then we actually collaborated more than before.

Host

你的背景真的很有趣,因为我记得当我还是数据科学家时,大约 10 年前,有很多关于 AutoML 和神经架构搜索等话题的讨论,事实上你是那篇 NAS 论文背后的人之一,我记得很多年前读过它。是的,请简要描述一下那段历史。

You've got a really interesting background because I remember when I was a data scientist, about 10 years ago, there was lots of discussions about things like AutoML and neural architecture search, and in fact you were one of the guys behind that NAS paper and I remember reading about that at the time years and years and years ago. And yeah, just sketch out that history.

Frank

是的。所以 AutoML 真的可以追溯到我的博士研究。我的博士研究是自动算法配置,就是让设计算法中的所有细节决策更加自动化。它来自 SAT 求解器,你需要做出数百个不同的分类空间决策,这真的很烦人。

Yeah. So AutoML goes really way back to my PhD. My PhD was automated algorithm configuration, which was all about making all of the nitty-gritty decisions in designing algorithms more automated. It came from SAT solvers where you needed to make like hundreds of different decisions at categorical spaces and this was really annoying.

局部搜索与AutoML早期工作 Early Work in Local Search and AutoML

Frank

在我的硕士阶段,我针对一个特定问题写了一个局部搜索算法,并研究了不同类型的问题,总是需要做所有这些决策。于是我开始用局部搜索来自动化这个过程,以找到我自己局部搜索算法的更好参数。然后我更多地接触到机器学习,并用机器学习来找到我设定算法的更好参数。接着我更多地转向使用机器学习来找到机器学习算法的更好超参数。这导致了模型选择,即从不同类型的算法中选择一个,但每个算法都有很多自己的超参数。然后你有了这些层次化的、非常复杂的空间,你想以高效的方式进行优化,但你也想泛化到分布的不同部分。所以你不想被卡住并过度调参到某一类问题分布,而是真正泛化到你之后会看到的各种问题。然后还有另一个设计空间,即神经架构,同样的方法可以直接应用。它非常分类化、非常结构化,有各种层次化的决策等等。所以我们开创了贝叶斯优化用于此。而且确实看到了神经架构搜索中的很多问题,比如如何对不同算法进行基准测试等等。运行这些算法非常复杂、成本很高,这使得很难做好实证科学。

And so in my master's I wrote a local search algorithm for one particular problem and worked on different types of problems, and always you needed to make all these decisions. So I started automating that with local search to find better parameters of my own local search algorithm. Then I got exposed much more to machine learning, and I used machine learning to find better parameters of my set algorithms. Then I moved more towards using machine learning to find better hyperparameters of machine learning algorithms. That led to model selection, choosing one of different types of algorithms, but each algorithm would have a lot of hyperparameters of its own. Then you have these hierarchical, really complex spaces, and you want to do optimization in an efficient way, but you also want to generalize to different parts of the distribution. So you don't want to be stuck and overtuned to one particular type of problem distribution, but really generalized to all kinds of problems you're going to see afterwards. And then there was this other design space of neural architectures for which the same types of approaches really directly applied. It's very categorical, very structured, all kinds of hierarchical decisions, etc. So we pioneered Bayesian optimization for that. And also definitely saw a lot of issues in neural architecture search in terms of how to benchmark different algorithms and so on. It's very complex to run these algorithms, very costly, and that makes it very hard to do good empirical science.

AutoML需求与民主化 The Need for AutoML and Democratization

Host

确实,我想。我记得多年前我读博士时,我用的是一个叫 Weka 的东西,你可以尝试,哦,让我们用支持向量机,让我们用贝叶斯网络,让我们做核岭回归。这是一个很棒的工具箱,你说我有一个预测问题,你有一堆数据,你有一些信号和标签,你可以直接原型化所有这些不同的方法,你可以做交叉验证。我猜在过去,机器学习主要是手动的,有一些超参数优化等等,但我认为这个 AutoML 的东西有点暗示了这个结构性组件。所以如果我们能针对特定问题搜索机器学习模型结构的空间,岂不是更好?

Exactly, I suppose. I mean, I remember when I was doing my PhD all those years ago, I was using something called Weka, and you could try, oh, let's use a support vector machine, let's use a Bayesian network, let's do kernel ridge regression. It was a wonderful toolbox where you say I've got a prediction problem and you've got a bunch of data and you got some signals and labels, and you could just prototype all these different approaches and you could do cross validation. I guess in the olden days machine learning was mostly manual, and there was some hyperparameter optimization and so on, but I think this AutoML thing was kind of hinting to this structural component. So wouldn't it be better if we could search the space of machine learning model structures in respect of the particular problem?

Frank

是的,绝对。我很喜欢你提到 Weka,因为那是我们合作的第一批基础算法或基础分类器库。我们做了 AutoWeka。记得 Weka 吗?它的形象是这只 weta 鸟,所以我们有一个小机器人骑着一只 weta 鸟,就是 AutoWeka。是的,你描述了它。各种各样的人使用它,通常来自科学、生命科学和生物学等领域,他们不是机器学习专家,而 Weka 里有几十种不同的分类器、几十种预处理器,以及这些的各种超参数。你知道人们会用什么吗?我不知道,比如用默认超参数的 SVM,因为他们的朋友告诉他们会很好,但这并不是获得最佳性能的方式。所以你真的很想自动化这个,以便为大众提供更好的方法,而这正是 AutoML 背后的主要原则:将最先进的机器学习民主化给每个人,包括那些没有机器学习博士学位的人。

Yeah, absolutely. I love that you mentioned Weka because that was the first base algorithm or base library of classifiers that we were working with. We did this AutoWeka. Remember Weka? The image for this was this weta bird, and so we had a little robot riding a weta bird to be AutoWeka. Yeah, you described it. All kinds of people use this, often people from the sciences, life sciences and biology, etc., who are not machine learning experts, and there's all this stuff in Weka like dozens of different classifiers, dozens of pre-processors, all kinds of different hyperparameters of these. And you know what people would use? I don't know, like an SVM with default hyperparameters because their friend told them that would be good, and that's just not how you get the best performance. So you really want to automate this in order to have better approaches for the masses, and that was really the leading principle behind AutoML: to democratize state-of-the-art machine learning to everyone, also those without a PhD in machine learning.

神经架构搜索与多目标优化 Neural Architecture Search and Multi-Objective Optimization

Host

我想 Weka 的问题在于,我总是对神经网络很兴奋,但在 Weka 中,神经网络模型非常慢,而且总是比使用更简单的模型更差。所以我想然后我们有了像 Keras 这样的框架,然后我们开始尝试不同类型的激活函数,也许我们可以在这里有一个 CNN 层,在那里有一个 MLP,它非常可组合。你几乎可以构建一个神经网络架构,并将你之前做过的冻结模型组合在一起或微调它们,但再次我们回到了这个极其手动的过程。那么自动机器学习是如何触及神经网络空间的?

I suppose the problem with Weka was I was always so excited about neural networks, but in Weka the neural network model was incredibly slow and it was always worse than using just simpler models. So I suppose then we had frameworks like Keras for example, and then we started experimenting with different types of activation functions and maybe we could have a CNN layer here and an MLP there, and it's very composable. You could almost just construct a neural network architecture and compose together frozen ones that you'd done previously or fine-tune them, but again we're back to this incredibly manual process. So how did the automated ML touch the neural network space?

Frank

是的,我的意思是,基本上如你所述,人们厌倦了手动做这件事,并思考,嘿,我们怎么能真正自动化这个空间?设计空间应该是什么样子?空间中通常正确的元素是什么?我应该让我的网络多宽、多深?那是早期的设计空间。但后来卷积神经网络出现了,注意力机制出现了,然后你有各种不同的混合架构,然后在此基础上,对于不同的芯片,你有不同的延迟等等。所以你有硬件意识在里面。然后你有算法的性能,比如准确率,但你也有延迟。你有内存消耗。你有所有这些不同的目标。所以它变成了多目标的,并且真的成为了一个非常令人兴奋的模型开发或方法开发的大 playground,以便搜索这些空间。

Yeah, I mean, basically as you describe, people were annoyed of having to do this manually and were thinking, hey, how can we actually automate the space? What should the design space look like? What are typically the right elements in the space? And how wide should I make my network and how deep should I make my network? That was sort of the early design spaces. But then convolutional neural networks came around, attention came around, and then you have all kinds of different hybrid architectures, and then on top of that, for different chips you have different latency and so on. So you have this hardware awareness in there. Then you had the performance of the algorithm, like accuracy, but you also had latency. You had memory consumption. You had all these different objectives. So it became multi-objective and was just really a big playground for very exciting model development or method development in order to search through these spaces.

向基础模型的转变 Transition to Foundation Models

Host

太棒了。所以我想有一个老故事。正如你所说,在你的博士期间,你专攻 SAT 求解器,这有点像是如何搜索可能性空间,然后发生了一些非常大的事情,我猜是在 GPT 模型时刻左右,我们有了基础模型的概念。所以数据科学家不再试图构建低级模型并搜索可能性空间,而是越来越多地使用这些在大量数据上训练过的基础模型,然后以此作为起点。那么这种转变是什么样的?

Amazing. So I suppose there's that old story. So as you were saying on your PhD you were specializing in SAT solvers, which is just kind of how can I search the space of possibilities, and something really big happened I guess it was around the maybe the GPT model moment something like that where we had this concept of a foundation model. So rather than data scientists kind of trying to build low-level models and searching the space of possibilities, what they increasingly were doing was using these foundation models that have been trained on loads and loads of data and then working from that as a starting point. So what was that transition like?

Frank

是的,有一个非常令人兴奋的转变。ChatGPT 在 2022 年 11 月问世,而在此之前一周,我们实际上发表了一篇 TabPFN 论文,这是第一个用于表格数据的基础模型,它彻底颠覆了表格数据。你不再需要做模型选择,而是有一个预训练模型,它实际上会在上下文中使用整个数据集,并跨越数百万个不同的数据集学习如何在一个前向传递中为测试数据进行预测。AutoML 在 15 年中的一个特别令人兴奋的部分是元学习的故事:跨越不同类型的数据集学习神经网络的参数,以便在未见过的数据集上表现良好。

Yeah, there was a really exciting transition. ChatGPT came out in November 2022, and a week before that we actually published a TabPFN paper, which was the first foundation model for tabular data, which completely turned tabular data on its head. You didn't need to actually do this model selection anymore, but you would have one pre-trained model that would actually use the entire data set in context and learn across millions of different data sets how to make predictions for the test data in one forward pass. One particularly exciting part of AutoML over 15 years was the story of meta-learning: learning the parameters of neural networks across different types of data sets in order to actually work well on unseen data sets.

TabPFN与表格数据学习算法 TabPFN and Learned Algorithms for Tabular Data

Frank

你其实可以把这件事想得更大,去学习、去思考学习整套算法,让它能泛化到新类型的数据集。而 TabPFN 正是 AutoML 的自然演进:我们学习的是整套算法,它在前向传播中执行,并且根据你喂进去的数据集,学出一个不同的分类器。所以它不只是一个网络、一个分类器,而是一个能根据输入学出不同分类器的网络。你可以把这个网络导出成 ONNX,放到传感器上,于是这个传感器跑的不只是一个分类器,而是一个机器学习算法,它会根据输入实际输出不同的分类器。

And you can actually think this bigger and learn and think about learning entire algorithms that generalize to new types of data sets. And TabPFN is really this natural progression of AutoML where we learn this entire algorithm that is executed in a forward pass and depending on the data set that you feed in, it learns a different classifier. So it's not just one network that is one classifier, but it's a network that can learn different classifiers depending on the input. And you could take this network and write it out as ONNX, put it on a sensor, and then you have a sensor that runs not just a classifier but a machine learning algorithm that depending on the inputs will actually give you different classifiers as output.

Frank

所以你可以这样理解:像 XGBoost 这类传统算法,全都是手工推导、手工编码出来的。而 TabPFN 是第一个针对表格数据、真正端到端学出来的算法,它是在大量不同的数据集上完整学出来的,目标是能在这些数据集上表现良好。当遇到一个新数据集时,我们试图在每个数据集未见过的测试部分上优化交叉熵损失。你当然看不到测试集,不能偷看。但在训练时,这正是你真正想优化的损失指标:你想在这个测试部分上表现好。我们会看数亿个数据集,确保算法在这些数据上表现良好,然后它也能泛化到新的数据集。

And so you can see this as, for example, traditional algorithms like XGBoost etc. — all of these algorithms are hand-derived, hand-coded etc. And TabPFN is the first algorithm for tabular data that's actually fully learned end to end over a large set of different data sets that it's supposed to work well on. And we're trying to optimize a cross entropy loss on the unseen test portion of each of these data sets when you have a new data set. You of course don't see the test set. You can't peek at that. But at training time, that is the loss metric you actually want to optimize: you want to do well for this test portion. And we look at hundreds of millions of data sets and make sure that the algorithm works well on those, and then it will also generalize to new data sets.

Frank

这就是 AutoML 在算法开发上的美妙之处。它让你可以非常声明式地表达:对于这类数据集,你应该表现良好。在 TabPFN1 时,我们用的数据集相当简单,之后几年我们把它做得越来越复杂。比如在 TabPFN2 里,我们加入了缺失值和异常值,以及各种类似的数据复杂性。无信息特征以前是处理不好的,类别特征以前也表现不佳。我们把这些复杂性越来越多地放进训练数据集的构造过程里。目标就是在带有这些数据复杂性的数据集上表现良好,而神经网络和深度学习只需要发挥它的魔力,真正去优化这个目标,靠纯监督学习就能做得很好。所以这里没有什么魔法。然后出来的就是一个在前向传播中执行的算法。

And that is the beauty of these AutoML for algorithm development. It just lets you be really declarative. You can say: for these types of data sets, you should work well. And for TabPFN1 we had a fairly simple set of data sets, and then over the years we made this more and more complex. For example, with TabPFN2 we put in missing values and outliers and all kinds of other data complexities like that. Uninformative features was something that was broken before. Categorical features was something that didn't work well before. And we put more and more of these complexities into the creation process for the data sets that we would train on. And then the objective was to do well on data sets with these types of data complexities, and well, neural networks and deep learning just needs to do its magic and actually optimize for that objective, and it can do that well by pure supervised learning. So there's no magic there. And then out comes an algorithm that executes in a forward pass.

Frank

我们不再需要那种漫长的训练循环,测试时也不再需要搜索超参数等等,我们只需要做一次前向传播。AutoML 的其他部分依然存在,但它存在于推导这个算法的过程中。所以我们仍然可以搜索神经网络架构的超参数,可以搜索预训练用的学习率等等。但所有这些在测试时都消失了,数据科学家得到了他们一直想要的东西:一个开箱即用、非常快的方法。

And we don't have this long training loop anymore. We don't have this search over hyperparameters etc. at test time, but we can just do a forward pass. All of the other part of AutoML still is there, but it's there in coming up with this algorithm. So we can still do a search over hyperparameters of the architecture of our neural network. We can search over our learning rates etc. for the pre-training. But all of that at test time falls away, and data scientists get what they always wanted, namely a really fast method that works out of the box.

Frank

现在我有了这个模型,下一个组成部分就是你刚才说的自适应推理。它和普通 Transformer 里的思维链适配非常相似。你输入一个提示,实际上你的提示在对 Transformer 做条件化,它在做某种针对你所给输入的计算。所以在你的上下文里,不需要从头训练,你可以把一些表格数据放进去,然后可以说,在一次前向传播中,它本质上就创建了一个专门适用于你这个场景的模型。

Now I've just got this model, and the next component is what you were talking about, this adaptive inference. So it's very similar to chain of thought adaptation in a normal transformer. You put a prompt in, and what it's actually doing is your prompt is conditioning the transformer, and it's doing some kind of computation which is specific for the input that you give it. So in your context, without actually training the thing from scratch, you can put some tabular data in there, and you're saying in a single forward pass it is essentially creating a model that works specifically for your case.

Host

是的,完全正确。而且这确实就是一次前向传播。

Yeah, absolutely. And this is literally a single forward pass.

Frank

在思维链和 LLM 里,你是自回归地做这件事,一次生成一个词。但在表格基础模型里,它确实就是你需要预测的那一个词,仅此而已。因为你有训练数据 X train 和 Y train(标签),然后你有 X test,也就是你想预测的那个数据点,而你只想预测 Y test。如果你有很多很多 X test,那么 X test 2 不应该依赖于你为 X test 1 预测出的结果。所以存在这种独立性,它确实就是一次一个词,你不做自回归展开。这让我们能够直接优化最终真正重要的目标函数。

So with chain of thought and LLMs, you're auto-regressively doing this. So you produce one token at a time. But with tabular foundation models, it's literally the one token you need to predict, and that is it. Because you have your training data X train and the Y train, the label, and then you have X test, which is the data point you want to predict for, and you only want to predict the Y test. And if you have many many X tests, then X test two shouldn't depend on what you predicted for X test one. So there's this independence, and it's really just one token at a time, and you don't roll out auto-regressively. And that lets us actually just directly optimize for the objective function that matters in the end.

Host

你谈到过这在概念上和高斯过程相似,我想我们也应该把它放到语境里:在大多数机器学习模型里,甚至当前的 Transformer 里,你可以说它们在近似某种贝叶斯模型,但它们近似的是最大似然的点估计,而你的模型实际上是在近似纯粹的高斯过程。它实际上是在近似所有可能取值上的不确定性。为什么你的模型能做到这一点,而其他模型做不到?这背后发生了什么?

So you've spoken about how this is conceptually similar to Gaussian processes, and I suppose we should also just contextualize that in most machine learning models, even in current transformers. You can say that they're approximating some kind of Bayesian model, but they're approximating a maximum likelihood point estimation, whereas your model is actually approximating the pure Gaussian process. It's actually approximating the uncertainty across all of the possible values. Why is it possible for your model to do that and other models don't do that? Like what's going on there?

Frank

是的。我们直接近似的是贝叶斯后验预测分布。其他类型的贝叶斯推断方法通常首先做的是得到函数上的后验分布。也许稍微退一步讲。你有一个函数上的先验,或者等价地说,参数上的先验。然后你观察到一些数据,接着你想推断函数上的后验,或者描述函数的参数上的后验。为了计算这个后验,你可以用 MCMC,也可以用变分推断,而两者都有各自的问题。MCMC 就是非常慢,而变分推断——嗯,数学很复杂,近似有时候并不完美,而且有时也有点慢。

Yeah. So what we approximate is directly the Bayesian posterior predictive distribution. What other types of Bayesian inference methods first do typically is to get a posterior distribution over the functions. So maybe to back up a little bit. So you have a prior over functions or equivalently a prior over parameters. And then you observe some data, and then you want to reason about the posterior over functions or the posterior over parameters that describe the function. And for computing that posterior you can use MCMC or you can use variational inference, and both of them have their issues. MCMC is just really slow, and variational inference is — well, there's complex math and the approximations sometimes don't work out perfectly, and it's also sometimes a bit slow.

Frank

但一旦你有了函数或参数上的后验,那么通常要得到后验预测分布,也就是给定 x 和数据时 y 的概率,这其实通常很容易。如果你有函数后验的基于样本的近似,那基本上就是对这个基于样本的近似求和。但棘手的部分是潜变量上的后验,也就是参数或函数上的后验,而这一步我们完全跳过。我们直接去到贝叶斯后验预测分布,它是一个一维分布。它只是 y 的概率,回归里是一个标量,分类里就是 k 个类别上的概率,而我们可以通过一次前向传播得到它。

But then once you have this posterior over functions or parameters, then typically in order to get the posterior predictive distribution, the p of y given x and the data, that is actually typically easy. If you have a sample-based approximation of the posterior over functions, then that's basically just a sum over that sample-based approximation. But the tricky part is the posterior over the latents, over the parameters or the functions, and that step we just entirely skip. We just directly go to the Bayesian posterior predictive distribution, which is a one-dimensional distribution. It's just the p of y, the scalar for regression, or for classification just the probability over the k classes, that we can do in a forward pass.

Frank

我们怎么在一次前向传播里做到这一点?嗯,我们从先验中采样。我们从函数先验中采样函数,然后从每个函数中采样数据点。

How do we do that in a forward pass? Well, we sample from the prior. We sample functions from this prior over functions, and then we sample data points from each of these functions.

高斯过程先验与元学习 Gaussian Process Priors and Meta-Learning

Frank

所以如果你考虑高斯过程,我们有一个高斯过程先验。我们可以从中采样不同的函数。你有某个核等来指定高斯过程。所以这告诉你一些关于它有多颠簸等信息。然后你从中采样函数。然后你从每个函数中采样数据点,称其中一些为训练点,一些为测试点。你基本上学习从训练点预测测试点或测试点的 y 值。如果你能对这个函数的数百万个样本做到这一点,那么你实际上已经学会了在你的网络中捕捉先验的结构。网络已经学会了近似这个后验分布,关于在这些缺失的测试数据点上 y 值应该是什么。它已经从你的先验的数百万个样本中学会了这一点,然后给定一个新的数据集。这是你第一次真正有一个数据集。在此之前,它真的只是来自先验的样本。所以然后你有一个真实的数据集,你可以在一次前向传播中计算贝叶斯程序预测分布,就像你为你的数百万个训练数据点中的每一个计算它一样。所以你基本上元学习了这个后验预测分布,或者元学习了近似任意数据集输入的后验预测分布,纯粹通过两件事。你需要能够从你的先验中采样,你需要能够以监督学习的方式拟合强大的神经网络。而这两件事实际上都相当容易,如果你有一个可以从中采样的机制先验。这就是我们的主力,用来改进我们的先验,然后深度学习机制。我们只是搭便车于人们在 LMS 中所做的,架构变得更好,优化器变得更好,我们的方法也直接变得更好,当然我们也适应表格数据的架构。是的。所以我们可以计算这个后验分布,所以在贝叶斯模型中传统上我们整合所有可能的假设,在这个空间中也是可能的。

So if you think of Gaussian processes, we have a Gaussian process prior. We can sample different functions from that. You have some kernel etc. that's specifying the Gaussian process. So that's telling you something about how bumpy it is etc. And so you sample functions from that. And then you sample data points from each of these functions and call some of them your training points and some of them your test points. And you basically learn to predict the test points or the y-value of the test points from the training points. And if you can do that for millions of samples of this function, then you have actually learned to sort of capture the structure of the prior in your network. And the network has learned to actually approximate this posterior distribution over what the y value should be at these missing test data points. It has learned that over these millions of samples from your prior and then given a new data set. This is the first time you actually have a data set. Before that it was really just samples from the prior. So then you have a real data set and you can compute the Bayesian procedure predictive distribution for that in a forward pass just like you computed it for each of your millions of training data points. So you basically meta-learned this posterior predictive distribution or meta-learned to approximate the posterior predictive distribution for arbitrary data set inputs purely by two things. You need to be able to sample from your prior and you need to be able to fit strong neural networks in a supervised learning fashion. And both of these are actually quite easy if you have a mechanistic prior that you can sample from. And that's sort of been our workhorse to improve our priors and then sort of the deep learning machinery. We just piggyback on what folks do in LMS and architectures get better, optimizers get better and our methods also directly get better and of course we adapt the architectures to tabular data as well. Yeah. So we can compute this posterior distribution and so traditionally in Bayesian models we kind of integrate over all of the possible hypothesis and it's possible in this space.

Host

我想试着区分这与 Transformer 之类的东西有何不同。所以你在这里所做的是使用合成数据训练这些模型。所以你提出了大量关于不同类型结构化数据中因果关系的先验。你不能用语言来做到这一点,例如。我的意思是,是的,你可以有一个上下文无关文法,你可以生成一堆语言,但它只是我们刚刚发明的一些奇怪的不可描述的语言。这里真正酷的事情是你可以说,好吧,有一些原则是表格数据共有的,我可以将这些表示为先验。我可以从中采样,在你的预测架构中,输出空间可能是,嗯,它将是离散的。它可能是分类的,这意味着有相对较少的例子,或者如果是一个回归问题,你可以将其离散化为可处理的数量。所以本质上你可以端到端地做这件事,这在许多其他类型的模型中是不可能的。

I guess I want to try and distinguish how this is different from something like transformers. So what you've done here is you've trained these models using synthetic data. So you've come up with a whole bunch of prior about causal relationships in different types of structured data. And you couldn't do this with language for example. I mean yeah you could have like a context free grammar and you could generate a bunch of language but it would just be some weird inscriptable language that we've just invented. The really cool thing here is you can say okay there are these principles that are common to tabular data and I can represent those as prior. I can sample from them and in your predictive architecture as well the output space it might be well it's going to be discrete. It might be categorical which means there are relatively few examples or if it's a regression problem you can discretise it to a tractable number of things. So essentially you can actually do this thing end to end in a way that wouldn't be possible with many other types of models.

Frank

是的。这绝对正确。所以如果你的先验,例如,只是一个,你说它真的很简单,只有线性曲线,你可以从中采样,那么出来的贝叶斯后验预测分布实际上是贝叶斯线性回归。如果你从高斯过程先验中采样,出来的是一个贝叶斯后验预测分布,实际上就是 GP 后验或其近似。如果你有一个贝叶斯网络或神经网络,你从中采样,那么结果是一个贝叶斯网络预测。所以再次不是关于神经网络参数本身的后验,而是一个后验预测分布,取所有可能的神经网络,从你的贝叶斯神经网络中,并对它们进行积分。如果我们的先验是结构因果模型的空间,那么我们的后验就是对所有可能的结构因果模型的积分,这些模型可能导致数据。对于每个模型,说这个 SEM 导致这个特定数据的可能性有多大,以及 SEM 实际上会为测试点预测什么,然后你对所有这些进行积分,当然有大量的它们。它可能是可数无限的。但我们只看到有限的数量,比如数亿。但我们的网络太小了,不可能记住这些。所以它实际上学会了在整个空间中泛化,即使只是从看到有限的数量。

Yeah. That's absolutely true. So if your prior is for example just a you say it's really simple there's only linear curves you can sample from that what will come out as a Bayesian posterior predictive distribution is actually Bayesian linear regression if you sample from Gaussian process prior what comes out is a Bayesian posterior predictive distribution which is actually just the GP posterior or an approximation thereof if you have a Bayesian network or neural network that you sample from then outcomes a Bayesian network prediction. So again not a posterior over the neural network parameters itself but a posterior predictive distribution take all the possible neural networks out there from your Bayesian neural network and integrate over them. And if our prior is the space of structural causal models, then our posterior is well the integral over all possible structural causal models that could cause a data. Say for each of them, how likely is this SEM to cause this particular data and what would the SEM actually predict for the test points and then you integrate over all of them and of course there is a huge number of them. It's probably countably infinite. But we only see a limited number of them like hundreds of millions. But our network is so small it can't possibly memorize that. So it actually learns to generalize across the space even just from seeing a finite number.

Host

所以我想有趣的是,你可以扩展这个,在某个规模上,它会渐近收敛到高斯过程会是什么,根据你展示的图表,它已经相当接近了,这真是令人难以置信地兴奋,因为我想我们还没有说的是,你可以把这看作是表格数据的 ImageNet 时刻。所以当你看这个 tab arena 时,是在 25 或 26 左右,有一个阶跃变化,所以你知道 Kaggle 上的人们在使用 XG boost 和 cat boost,还有,你知道,另一个是什么,light GBM,对吗?是的,有一个绝对巨大的阶跃变化,有几件事在发生,所以现在你可以生成基本上任意多的训练数据,你可以把它烘焙到模型中,我认为重要的概念是它正在做一种摊销推理。所以以前当你做训练和推理时,你必须做很多工作,现在你有一个单一的前向传播,这意味着本质上你有一个固定的计算量。

So I suppose the interesting thing is that you can scale this up and at some scale it will asymptotically converge on what the Gaussian process would have been and it's already reasonably close to that based on the graph that you've shown and this is just incredibly exciting because I suppose what we haven't said yet is that you can think of this like the ImageNet moment for tabular data. So when you look on this tab arena was it around 25 26 there was a step change so you know folks on Kaggle they're using XG boost and cat boost and uh you know is it what was the other one light GBM is that right yeah and there was this absolutely massive step change and there's a couple of things going on so now you can just generate essentially as much training data as you want you can bake it into the model and I think the important concept is that it's doing a kind of amortized inference So whereas before when you're doing you know training and inference you had to do a lot of work now you have a single forward pass which means essentially you have a fixed amount of computation.

Frank

是的,没错。所以我的意思是我们基本上学习这个网络,训练它以便在我们作为输入馈送的数据集类型上尽可能好地表现。我们历史上所做的是生成相对简单的数据集,用 TabPFN v1 就像超级简单,TabPFN 2 稍微复杂一点,有异常值等。但即使使用 TabPFN 3,它们仍然是相对简单的独立同分布数据集,而且与一些数据集中存在的数十亿个数据点相比,数据点的数量相对较少。所以用 TabPFN v1 就像 10,什么 1,000 个数据点,2 是 10,000 个数据点,2.5 是 100,000 个数据点,3 是百万个数据点。所以我们正在接近。我们正在以每年两个数量级的速度接近,但仍然没有达到十亿,而且例如 Kaggle 上的许多数据集如此之大,以至于当你使用表格基础模型进行前向传播时,它们还没有赶上。

Yeah exactly. So I mean we basically learn this network that's trained in order to do as well as possible over the types of data sets that we feed in as input. And what we've done historically is to you know generate relatively simple data sets with TabPFN v1 like super simple TabPFN 2 a bit more complex with outliers etc. But even with TabPFN 3 they are still relatively simple iid data sets and there is you know relatively low number of data points compared to you know the billions of data points that are out in some data sets. So with TabPFN v1 like 10 what 1,000 data points two 10,000 data points 2.5 100,000 data points three or million data points. So we're getting there. We're getting there like two orders of magnitude a year but still we're not at a billion and a lot of the data sets on Kaggle for example are so large that you know the tabular foundation models haven't caught up there when you use them in a forward pass.

TabPFN中的扩展与数据复杂度 Scaling and data complexity in TabPFN

Frank

你可以围绕它搭建一个框架,我们有 Scaling(规模扩张)模式和思考模式等,基本上你可以用包装器来扩展测试时计算。然后它也可以处理 Kaggle 的数据复杂性,但目前为止它还没有被用于那么多 Kaggle 竞赛,因为并没有很多 Kaggle 竞赛处于这种非常小的数据规模,比如 10,000 个数据点,我们可以用 TabPFN 2 处理,或者 100,000 个数据点,我们从去年十一月开始可以应对。一百万我们从大概两个月前开始可以应对。只是并没有那么多那种规模的 Kaggle 竞赛。但总的来说,其实表格类竞赛本来就不多,因为语言和视觉等受到了更多关注。但表格数据在世界上如此重要,所以我肯定看到现在再次有更多关注这类模态,这也将导致未来有更多 Kaggle 竞赛。我们正在我们的先验生成中专注于越来越复杂的数据集。例如,非独立同分布的分组数据、时间数据、表格中的文本等。所有这些都是你在 Kaggle 上随处可见的数据复杂性,而这些在我们的基准测试中完全没有,在我们的数据生成过程中也完全没有,我们正在改变这一点。现在它已经在我们的基准测试中了。所以我们有这个 beyond arena,包含了所有这些数据复杂性,然后这表明 TabPFN 2.6 对一些数据复杂性并不特别好。TabPFN 3 在其中一些上有所帮助,但仍然不够好,所以这是我们目前正在优化的信号。

You can have a harness around it where we have the scaling mode and thinking mode, etc., where you can basically have wrappers around this and use test-time compute in order to scale up. And then it could also work for the data complexities of Kaggle, but for now it hasn't been used for that many Kaggle competitions because there aren't a whole lot of Kaggle competitions that are in this really small data regime of like 10,000 data points that we could deal with TabPFN 2, or well, 100,000 we can tackle since last November. A million we can tackle since like two months ago. There just haven't been a whole lot of Kaggle competitions of that size. But in general, just actually not that many tabular competitions anyways, because language and vision etc. has gotten a lot more attention. But tabular data is so important in the world that I definitely see a lot more focus on these types of modality now again, and that will also lead to more Kaggle competitions in the future. And we are focusing on ever more complex data sets in our prior generation. So for example, non-IID group data, temporal data, text in our tables, etc. All of these are data complexities that you see everywhere on Kaggle and that were just nowhere in our benchmarks, that were nowhere in our data generation processes, and we're changing that. It's in our benchmarks now. So we have this beyond arena that has all these data complexities in there, and then that shows that TabPFN 2.6 was not particularly good for some data complexities. TabPFN 3 helped on some of them but is still not quite there, and then so that's a signal we're optimizing for right now.

多模态与上下文窗口 Multimodality and context window

Host

是的,我想我们稍后再回到多模态的内容,因为那真的非常有趣。但我想一个很好的思维框架是,你知道,当你使用一个普通的语言模型时,有一个上下文,你有这种二次复杂度,你知道如果你使用 Claude,例如,当你达到一百万时,它真的在挣扎,并且有点变慢。我认为这里也是类似的情况,即上下文现在基本上就是你放入的表格,嗯,所以现在是大约 100,000 行吗?

Yes, I think let's come back to the multimodality stuff later because that's really really interesting. But I suppose a good mental frame for this is, you know, like when you use a normal language model there's a context and you have this quadratic complexity and you know things just slow down if you're using Claude for example when you go up to a million it's really struggling and it's slowing down a little bit and I think it's a similar thing here which is that the context now is essentially the table that you put in and um is so is it around 100,000 rows now?

Frank

是的,所以它最多可达一百万行,比如一百万行乘以一千列。所以如果你把它扔进 LLM,你会有十亿个元素,你需要对每个数字进行分词。所以每个数字可能需要三到四个 token。因此你的上下文中会有三到四十亿个 token,LLM 不会很乐意。所以它们开箱即用并不适用于这类数据。它们不是为表格数据而生的。它们只需要以某种序列方式读取表格,因为它们是序列模型。所以如果它们一次读取一行,那么它们不理解如果你交换两行或交换两列,它还是同一个数据集。所以这就是为什么它们没有利用这些不变性,并且它们有大量参数对于捕捉世界的语义、捕捉世界知识非常重要,但实际上你并不需要这些来捕捉那种统计推理,你需要这些来服务于 XGBoost 等的标准接口,你只有 X_train、Y_train、X_test 并想预测 Y_test,然后你甚至看不到列名,所以你真的不需要知道任何关于语义的东西。如果你确实知道语义,那么你可以做更多。所以如果你有比如五行左右,你知道,我不知道,这是一个 turn 数据集,你确切知道列名是什么等等,那么世界知识就派上用场了,LLM 会很棒。但如果你有一百万行,那么你实际上想了解这些数字的统计特性,而这就是 LLM 真正一败涂地的地方。

Yeah, so it's up to a million rows and like a million rows times a thousand columns. So if you were to throw this into an LLM, you would have a billion elements and you need to tokenize each of these numbers. So that would maybe be three four tokens each. So you have three four billion tokens in your context and LLMs wouldn't be very happy there. So they just don't work out of the box for this type of data. They're not made for tabular data. They would just need to read the table in some sort of sequence because they're sequence models. So if they read like one row at a time then they don't understand that if you switch two rows or you switch two columns it's the same data set. And so that's why they don't exploit these invariances and they have a whole lot of parameters that are super important in order to capture the semantics of the world, in order to capture world knowledge, but that you actually don't need in order to capture sort of these statistical reasoning that you need in order to serve this standard interface of XGBoost etc. where you just have the X_train, Y_train, X_test and want to predict Y_test, then you don't even see column names so you really don't need to know anything about semantics for this. If you do know semantics then you can do more. So if you have like five rows or something like that and you know, I don't know, this is a turn data set and you know exactly what the column names are etc., then in comes the world knowledge and LLMs would be great. But if you have a million rows then you actually want to learn about the statistics of these numbers and that's where the LLMs just really fall flat on their nose.

权衡及与XGBoost的比较 Trade-offs and comparison with XGBoost

Host

是的,我想我们应该对比一下这里的权衡。所以你知道我可以训练一个 XGBoost 模型,它会非常昂贵,但我想好处是我可以然后进行流式并行化的逐行预测,我可以遍历数据集,而你的系统的好处是它实际上将整个表格、整个数据集作为输入,这意味着它可能在学习全局关系、一阶和二阶关系,你知道表格级的不变性。所以它建模的复杂程度远高于 XGBoost,而美妙之处在于因为它是一个基础模型,它自动完成这些。

Yes, I suppose we should contrast what the trade is here. So you know I could train an XGBoost model and it would be very expensive and but the I suppose the good thing is I can then do streaming parallelizable row-wise prediction and I can go through the data set whereas the good thing about your system is that it's actually taking the entire table the entire data set as an input which means it's learning potentially global relationships first and second order relationships um you know table-wise invariances. So it's modeling at a level of sophistication which is far away from XGBoost and the beauty of it is because it's a foundation model it just does it automatically.

Frank

是的。它已经学会在前向路径中做到这一点。既然你提到 XGBoost 可以并行流式处理大量预测。所以历史上 TabPFN 在推理时非常慢,因为是的,你总是有整个数据集传入前向传播,所以训练阶段和预测阶段之间没有真正的区别,而现在就像 LLM 一样,我们实际上有一个 KV 缓存,所以训练和测试之间有区别,在测试时你只需要关注这个 KV 缓存,你不需要关注每个训练数据点。你也不需要保留你的数据,因为你知道出于隐私原因等,你不想在做出任何预测时输入所有数据,但你只需要保留存储在 KV 缓存中的权重。所以现在快多了,在 GPU 上它也越来越接近 XGBoost 的预测时间。

Yeah. It has learned to do this in a forward path. Since you mentioned that XGBoost you can stream a lot of predictions in parallel. So historically TabPFN was very slow at inference time because yeah you always had this entire data set that you pass in into the forward pass and so there wasn't really a difference between the training stage and the prediction stage and now just like in LLMs you actually we have a KV cache and so there is a difference between training and test and at test time all you need to attend to is to this KV cache and you don't need to attend to each of the training data points. You also don't need to keep your data around because you know also for privacy reasons etc. You don't want to feed all your data when you make any prediction but you just need to keep the weights that are stored in your KV cache. So it's much faster now and on GPU it's getting close to XGBoost prediction times as well.

与AutoML及代码生成的比较 Comparison with AutoML and code generation

Host

我们之前将这与 AutoML 之类的东西进行了对比,但你的幻灯片实际上相当不错。我来这里之前读了你的演示文稿,你正确地观察到,如果你将比如一千行数据输入到一个语言模型中,仅仅是一个裸语言模型,它不会做任何有趣的事情,因为它不理解数据中的结构,它没有建模那些不变性。但我不确定这是一个完全公平的比较,因为现在我们生活在 Claude Code 和代码生成的世界中,我可以对智能体说,你知道这里有一堆数据,它会生成 Python 代码,将其加载到数据表中,并且它会理解很多语义。它会理解这个或那个的含义,你知道这是一个风险相关的东西,保险相关的东西,会计相关的东西,它会构建所有这些结构化模型,我可以告诉它结晶出一个机器学习模型。所以我可以说不妨这里有一些文本数据,让我们使用 mini LM 编码器,让我们在上面加一个 XGBoost 头之类的。美妙之处在于它是自适应的。

We were contrasting this to something like AutoML before but you had quite a good slide actually. I was reading your deck before I came here and you made the observation correctly that if you fed let's say a thousand rows into a language model just a bare language model it's not going to do anything interesting is it because it doesn't understand the structure in the data it's not modeling those invariances but I'm just not sure that's an entirely fair comparison because now we live in the world of Claude Code and codecs and I could say to the agent you know here's a load of data and it would be generating Python code it would load it into a data table and it would understand a lot of the semantics. It would understand what the meaning of this was or that you know this is a risk thing, an insurance thing, an accountancy thing and it would be building all of these structured models and I could tell it to crystallize a machine learning model. So I can say okay well there's some text data here and let's use the mini LM encoder and let's have like a XGBoost, you know, head on it or whatever. And the beauty of it is that it's adaptive.

LLM与智能体用于数据科学 LLMs and agents for data science

Host

那么下周当出现预测错误时,Claude 会理解错误发生的原因,它会生成一些合成数据,可能会加入一些符号规则,它可能会说,如果召回率低于这个值,我就把它升级到 Haiku 模型之类的。所以你看,我们现在处于这个——这几乎就像梦想成真,你知道,就像几年前:如果我们能有这个,那该多棒啊?那么它和那种自适应架构相比如何呢?

So next week when there's a prediction error, Claude will understand why the error happened and it will generate some synthetic data and it might put some symbolic rules in and it might say, well, if the recall is less than this, then I'm going to escalate it to a Haiku model or something like that. So you see now we're in this—it's almost like the dream, you know, like a few years ago: wouldn't it be amazing if we could have this? So how does it compare to some kind of adaptive architecture like that?

Frank

是的。所以我完全同意 LLM 在编码方面非常出色。智能体在特征工程方面很棒。它们在探索性数据分析方面很棒。它们作为用户界面也很棒。实际上,让你再次检查所有数据是否正确输入。就像和你聊天,数据实际上来自哪里。所以对于特征工程,它们非常出色。实际上,我提到了 TabPFN——我们在 ChatGPT 发布前一周发表了它,然后当 ChatGPT 出现时,我们非常兴奋,我们做的第一件事就是放弃 TabPFN,实际上用 ChatGPT 做智能体式数据科学,在 TabPFN 之上进行特征工程。所以你实际上做的是需要世界知识的特征工程。我不知道,这是一个非常简单的例子:如果你知道患者的身高和体重,你想预测一些事情,知道他们是否肥胖会很有帮助。那么当然你可以计算身体质量指数,这只是一个简单的计算,你可以做这个计算,或者你可以让网络在第一层早期学习做这些类型的计算等等。但如果你不需要——如果你可以在外面用一个编码智能体来做——那就容易多了。所以这就是我们做的第一件事,写这篇 TabPFN 论文,用 LLM 做自动化数据科学。所以这绝对是我看到的关系类型:对于数据准备、清洗等,LLM 非常棒,我们到处都在使用它们。但最终,它们需要调用一个模型,是的,它们可以调用 XGBoost,但它就是不如那些每天都在变得更好的表格基础模型。

Yeah. So I'm absolutely in agreement that LLMs are amazing for coding. Agents are great for feature engineering. They're great for exploratory data analysis. They're great as a user interface. Actually, let you double check that all of the data is entered correctly. Like chat with you where the data actually comes from. And so for this feature engineering, they're amazing. Actually, I mentioned TabPFN—we published sort of a week before ChatGPT and then when ChatGPT came out we were like so excited and the first thing we did is like drop TabPFN and actually do agentic data science with ChatGPT to actually do feature engineering on top of TabPFN. So you actually do the feature engineering where you do need the world knowledge. I don't know, this is a super simple example: there is sort of if you know the height of a patient and the weight of a patient and you want to predict something where it would be helpful to know if they're obese or not. Then of course you can compute body mass index and it's just you know a simple computation and you could do this computation or you could have the network learn to do these types of computations early on in the first layer and so on. But if you don't need to—if you can just do this outside with a coding agent—that is so much easier. And so that's what the first thing we did is to write this TabPFN paper that did automated data science with LLMs. And so this is the type of relationship I absolutely see: for the data preparation, cleaning, etc., LLMs are fantastic and we're using them left and right. But in the end, they need to call a model and yeah, they can call XGBoost, but it's just not going to be as good as the tabular foundation models that are getting better by the day.

结合LLM与TabPFN Combining LLMs with TabPFN

Host

我想我们可以鱼与熊掌兼得,因为现在人们可以在家里告诉 Claude 或 Codex 去获取 TabPFN V3,因为你知道,对于非商业用途,它是免费的,他们实际上可以让智能体使用 TabPFN。这里就变得有点有趣了,对吧?因为我认为 TabPFN 的一大优点是我们的数据集和表格是可读的。但如果 Claude 做了奇怪的事情呢?因为你知道,有时为了获得更好的表示摩擦,你会做奇怪的高深莫测的特征工程之类的。也许 Claude 会那样做。也许 Claude 把我们的两个特征变成了 100 个特征,我们最终得到一团乱麻。那会是个问题吗?

I suppose we could have our cake and eat it because what folks can do now at home is they can tell Claude or Codex to go and grab TabPFN V3 because, you know, for non-commercial use, it's free and they could actually get the agent to use TabPFN. And here it gets a little bit interesting, right? Because I think one of the great things about TabPFN is that our data sets and tables, they're legible. But what if Claude did weird stuff? Because you know, sometimes to get better representational friction, you do weird inscrutable feature engineering and whatnot. Maybe Claude does that. Maybe Claude takes two of our features and it turns it into a 100 features and we end up with a spaghetti mess. Would that be a problem?

Frank

我的意思是,我不认为这会是个大问题,因为你会提示 Claude 实际上告诉你它到底在做什么,并给你代码和特征等等。这实际上和数据科学家会做的很相似,对吧?他们写某种特征工程管道,随着时间的推移可能会变得有点乱。但你知道,归根结底这只是简单的代码。它是一些简单的操作,你应用到不同的特征上。如果 Claude 做得好,它还会去访问各种在线特征库等等。例如,如果你有一个数据集,我不知道,其中一列告诉你数据点来自哪个国家,也许是一个购物数据集,然后你可以说,在这个国家,实际上这个特定日子是假日,所以需求会更大,因为人们有更多时间在网上购物等等,所以这将是一个非常重要的特征,你可以生成它,比如这个国家是否是假日,这需要外部信息。你不能只是从另一个上下文中看到这个,可能 Claude 并不真的想只是记住这个,而是实际上在一些知识库中查找,当然它可以做到这一点,现在你可以使用 Claude 和 TabPFN,因为我们有一个 MCP 服务器,与 Claude 集成,与 Gemini 集成,也与各种其他框架集成,很容易上手,然后基本上你可以说:“嘿,Claude,为我构建一个这个问题的数据集,然后使用 TabPFN 来预测。”实际上,是的,我的一个朋友用这个来预测世界杯。所以我们有点乐趣,为每场比赛做了预测,但那个朋友实际上使用 Claude 和这个来真正提示,嘿 Claude,请为我构建一个来自历史比赛的数据集,等等,任何你想要的方式,表格数据集,然后使用 TabPFN 来预测每场比赛,凭借这个,他们实际上在公司赢得了他们的竞猜。所以你现在可以拥有这个蛋糕,你绝对需要它,这对我来说真的是 AutoML 的融合点,因为我们可以将最先进的机器学习带给新手,他们甚至不需要知道如何编码,他们不需要真正知道什么是表格预测问题等等。而且摩擦非常低,但同时你还能获得最先进的性能,是的,我们实际上有各种不同的演示。我们有一个演示,它集成在 DataBricks 中,是的,在那里你实际上可以浏览你的仪表板等等,并向智能体提问,对于某些问题,你实际上想要进行表格预测,它就会直接做,它甚至不会告诉用户它正在使用 TabPFN,为什么要告诉呢,对吧?它只是在做出更好的预测,所以这绝对是几年后我们将身处其中的世界,你会看到表格基础模型被广泛使用,通常你并不知道。

I mean I don't think it would be too much of a problem because you would prompt Claude to actually tell you exactly what it's doing and to give you the code and give you the features etc. And that is actually quite similar to what a data scientist would do, right? They write some sort of feature engineering pipeline and that might get a little messy over time. But you know it's simple code in the end of the day. It's some simple operations that you apply to your different features. What Claude would also do if it's doing a good job is to go access all kinds of feature stores online etc. If you for example have a data set and I don't know one of the columns tells you about the country that the data point is from and maybe it's a shopping data set and then you can say well in this country actually this particular day is a holiday so there's going to be more demand because people have more time to actually go shopping online etc and so that is going to be a super important feature that you can then generate is like is this a holiday in this country or not and that you know requires external information. You can't just see this from another context and probably Claude doesn't really want to just remember this but actually look this up in some knowledge base and of course it can do that and you can now use Claude with TabPFN as we have an MCP server that's integrated with Claude, integrated with Gemini, also with all kinds of other frameworks and it is really easy to get going and then basically you can just say, "Hey, Claude, build me a data set for this problem and then use TabPFN in order to predict." And actually, yeah, a friend of mine used this for the World Cup. So we had a little bit of fun and made predictions for every game, but that friend actually used Claude with this in order to just really prompt, hey Claude, please build me a data set from historical games, etc. like any way you want, tabular data set and then use TabPFN in order to predict every game and with that they actually won their kicktip in their company and so you can now have this cake you need it absolutely and this is to me really sort of the convergence point of AutoML because we can give state-of-the-art machine learning to novices who really they don't even know need to know how to code they don't need to know really what a tabular prediction problem is etc. And the friction is just so low but still at the same time you get state-of-the-art performance and yeah we actually have a variety of different demos. We have one demo where it's integrated in DataBricks and yeah there you actually can look through your dashboards etc and ask the agent questions and for some of the questions you actually want to do tabular prediction and it will just do that it won't even tell the user that it's using TabPFN why should it right it's just making better predictions and so that is definitely the world that we will be in in a couple of years where you'll see tabular foundation models used left and right often without you knowing it.

超越预测:其他用途 Beyond prediction: other uses

Host

是的。不,我实际上真的很兴奋自己尝试一下,因为它是那种几乎在任何地方都有用的东西,特别是如果你能让编码智能体帮你做的话。所以那太棒了。还有一些我们没谈到的其他优势。例如,你可以用它们来生成数据,或者做密度估计,甚至嵌入。所以我想我们想摆脱它。它不仅仅是用于预测的东西。你实际上可以用它来做可解释性,或者作为构建其他预测架构的一部分。

Yeah. No, I'm actually genuinely excited to try it myself because it's one of those things that just has utility almost everywhere, especially if you can get a coding agent to help you do it. So that's amazing. There are some other advantages that we haven't spoken about. So for example, you can use them to generate data or do things like density estimation or even embeddings. So I guess we want to just get away from it. It's not just something which is only used for prediction. You can actually use it for explainability or as part of building some other predictive architecture.

Frank

是的。

Yeah.

超越分类与回归 Beyond Classification and Regression

Frank

我认为这非常重要,对吧?要超越单纯的分类和回归。它是一个基础模型,可以做各种事情。我们希望推动关系型数据。很多表格数据实际上来自关系数据库中两个不同表的合并。直接在关系型数据上工作会强大得多。还有时间序列,实际上非常相关。还有因果性。这是我们没怎么谈过的,但这非常核心,因为你确实想要区分因果和相关。能够以因果的方式理解世界非常有帮助。然后你会理解得更深入,这是一个巨大的挑战,你知道因果机器学习是一个大领域,我认为我们可以用表格基础模型在这方面取得坚实的进展。

And I think that's super important, right? To move beyond just classification and regression. It is a foundation model. It can do all kinds of things. We want to push towards relational data. A whole lot of data that is tabular actually comes from merging two different tables in a relational database. And working directly on the relational data would be so much more powerful. There's time series which is actually very related. There is causality. So that is something we haven't talked about much but this is really core because you do want to make distinctions between just causation and correlation there. It's so helpful to be able to understand the world in a causal manner. Then you understand it just much more deeply and this is a big challenge and you know causal ML is a big field and I think we can make a very solid dent in there with table foundation models.

Host

是的。我认为我们应该回到这一点,因为那非常非常重要。但所以是的,现在人们可以在 Python 中导入这个库,它非常类似于 scikit-learn 之类的。所以你可以说我有一些数据,我有我的信号和标签,拟合后开箱即用,它做回归和分类。现在如果我理解正确,你不能结合分类和回归,因为你在架构上决定专门化这两个模型。所以如果你想结合这些模态,我认为你需要创建两个模型。这是正确的吗?

Yes. And I think we should come back to that because that's very very important. But so yeah, right now folks can just in Python import this library and it's very similar to scikit-learn or something like that. So you know you can just say I've got some data here. I've got my signals and my labels fit and out of the box it does regression and classification. Now if I understand correctly you can't combine classification and regression because you've decided architecturally to specialize in those two models. So if you want to combine those modalities, I think you need to create two models. Is that correct?

Frank

有时为分类和回归分别建立模型是有帮助的,因为梯度传播方式略有不同,同时训练两者有点棘手。但如果你能同时训练两者,那当然好得多。是的,当然,我们也在研究这个。我们还想要一个因果头,一个时间序列头等等,并联合训练所有这些。那会好得多。

It helps to have separate models for classification and regression sometimes because the gradients propagate a bit differently and it's a bit tricky to train both at the same time. But if you can train both at the same time, that's of course so much nicer. And yeah, of course, we're also working on that. And we also want to have a causal head and we want to have a time series head, etc. and just train all of that jointly. That would be so much nicer.

模型架构与预测头 Model Architecture and Heads

Host

是的。你能解释一下吗?所以做回归时,你使用 MLP 头,而做分类时,你有一个非常有趣的,我猜它类似于注意力头。所以你有一个特殊类型的解码器头。告诉我背后的理由。

Yeah. Can you explain that? So when doing regression, you're using an MLP head and when you're doing classification, you've got this very interesting I guess it resembles an attention head. So you've got a special kind of decoder head for that. Tell me about the rationale there.

Frank

为不同任务使用不同类型的头不是问题,对吧?例如,对于因果性,我们需要与其他方法不同类型的头。我们过去做回归的方式是将回归作为分类,你有一个分箱分布,将空间分成比如 10,000 个类别,数据多的地方箱子很小,数据少的地方箱子很大。这样每个箱子中的样本数相同,然后你训练一个标准分类模型。这样做的理由是,实际上当时我的博士生 Sam Miller 做了这些实验,他也尝试了一个回归头,直接预测均值和方差,但用这种分类方式预测效果更好。其中一个原因可能是,Transformer 的所有超参数等一切都是为了分类做得很好而设计的,有 10,000 个类别根本不是问题,所以我们实际上就坚持了这种方式。一个非常好的副作用是我们可以做这些多模态预测。我们可以说,我不知道鸟是撞向杆子还是向左转或向右转。它们不会撞到杆子。所以你会有一个分布在这里,一个分布在那里,中间几乎为零概率。你可以在一次前向传播中再次进行这些预测,所以你可以得到非常校准的输出,具有这种类型的分布。

It's not a problem to have different types of heads for different tasks, right? We like for example for causality we would need a different type of head than that for the other methods and the way we used to do actually regression is a regression as classification where you just have this binning distribution where you bin the space into just like say 10,000 classes and where a lot of the data falls you have very small bins and where not so much of the data falls you have large bins. So that there's the same number of samples in each of the bins and then you just train a standard classification model. The rationale for that was well actually back then my PhD student Sam Miller did these experiments and he did try also yeah a regression head that actually just predicted the mean and the variance and it was just better to predict in this classification manner. One of the reasons for this might have been that actually well transformers were just like all the hyperparameters etc. Everything is just made in order to do classification really well and having 10,000 classes was not an issue at all and so we actually just stuck with that. One of the really nice side effects of that was that we could do these multimodal predictions. We could say well, I don't know if the bird flies against like a pole like they're going to turn left or they going to turn right. They're not going to hit against the pole. So, you're going to have a distribution here and a distribution here and pretty much zero probability mass in between. And you can do these predictions again in a forward pass and so you can have like really well calibrated outputs with this type of distribution.

表示与架构演进 Representation and Architecture Evolution

Host

你能告诉我你是如何做表示的吗?所以,我的意思是有一点传承,对吧?所以版本一更像一个普通的 Transformer。所以只是对事物进行标记化,如果我理解正确,现在在版本二和版本三中,它更加结构化,你有一种特定的方式在注意力中映射,因为 Transformer 是置换等变的。也许我们应该从那里开始。

And can you tell me about how you do the representation? So, I mean there's a bit of a lineage, right? So the version one kind of was it more resembled a normal transformer. So just kind of tokenizing things and if I understand correctly now in version two and three it's far more structured and you have a specific way of kind of mapping the in the attention because the transformer is permutation equivariant. Maybe we should start with that.

Frank

是的,所以 TabPFN v1 的架构非常像 Transformer,只是你去掉了位置嵌入,因为注意力已经对顺序不变,唯一让 Transformer 真正关注单词在序列中位置的是位置嵌入,我们不想要那个。我们想要不变性,所以我们直接去掉了位置嵌入。这是我们能做的最简单的事情。但这样要求实际上取一行并将其编码为一个嵌入,我们使用了一个超级简单的方法。在 TabPFN1 中,我们实际上只有连续特征,我们只使用了一个线性层,就这样。你知道这可能不是最好的,因为如果我们将其应用于分类值,我们只会将分类值编码为 1 2 3 4,然后放一个线性层,当然那不是最好的。所以在 tapfn2 中,我们实际上真正拥有一个知道行和列的架构,在那里我们对每个矩阵元素都有一个嵌入。所以对于每个值,你会有一个对行的注意力,一个对列的注意力,你会交替这些,然后你当然可以真正理解,啊,在这个列中数字是 1 2 3 4。可能这是分类的,你可以将其与数值 1 2 3 4 区别对待。而且,在 TabPFN 2 中,我们也会在先验中拥有更多的分类信息。所以对这些分类参数的处理越来越好。而且,在复杂性方面,我应该谈谈这个,因为复杂性方面 2 实际上比 TabPFN1 差得多,因为 FN1 只在行数上是二次的。但有点糟糕,因为将行编码为单个嵌入,这本质上是有局限性的,而在 topfn2 中,你实际上因为每个嵌入有这么多嵌入,你实际上会对所有行进行注意力,有时对所有列进行注意力,因为这一行有那么多元素。

Yeah, so the architecture for TabPFN v1 was very much like a transformer just that you drop the positional embedding because well attention is already invariant to the order and the only thing that makes a transformer actually pay attention to where the words are in the sequence is positional embedding and we didn't want that. We wanted to be invariant so we just dropped the position embedding. It was the simplest thing we could do. Then but what that required is to actually take a row and encode that into an embedding and we used a super simple method for that. In TabPFN1 we actually had only continuous features and we just used a linear layer and that was it. And you know that is maybe not the best because then we if we applied it for categorical values we would just encode the categorical values as one two three four and then put a linear layer on there and of course that is not the best. So in tapfn2 what we did instead is actually really have an architecture that knows about rows and columns and where there we did an attention over like so we had an embedding for each individual element of the matrix. So for each value then you would have an attention over the rows and an attention over the columns and you would alternate these and then you could of course really understand aha in this column the numbers are 1 2 3 4. Probably this is a categorical you can treat this differently than the numerical 1 2 3 4. And yeah, we would also in TabPFN 2 have much more categorical information in the prior already. So there's just better and better treatment of these categorical parameters. And yeah, also complexity wise I should talk about that because complexity wise 2 was actually much worse than TabPFN1 because FN1 was only quadratic in the number of rows. But it's sort of yeah bad because of the encoding of the row into a single embedding and then that is really inherently limiting whereas at topfn2 you actually since you had so many embeddings for each of the embedding you would actually have this attention over all the rows and you have at times all the columns because there's that many elements of this row.

TabPFN v2与v3的复杂度 Complexity of TabPFN v2 and v3

Frank

假设你有 n 行、m 列,那么复杂度就是 n²·m + n·m²——这就是 TabPFN v2 的复杂度。对于最多 1 万个数据点来说,这还可以,但如果我们想扩展到更大规模——10 万、100 万等等——就需要突破这个限制。其实对于 10 万,TabPFN 2.5,我们还是做到了,我们只是换用更强的 GPU 就撑过去了。但与此同时,Gwas 小组开发了 TabPFN,这真的很棒。有一种架构,最终基本上只用了 TabPFN v1 的架构——只对行数是二次的——但它有一个复杂得多的机制,通过一个不同的列变换器和行变换器来做嵌入。由于那个架构对于 TabPFN v3 想要处理的规模范围来说已经足够好了,而且在推理速度方面也够用,我们基本上就把它适配到了 TabPFN v3 上。当然,我们在很多方面做了创新,比如多分类、输出头等等。但基础架构我们其实不需要改动。我们有很多东西正在酝酿,但对于 TabPFN v3 的表格数据,我们其实并不需要它们。所以它们会出现在未来的版本里。

So you have n numbers of rows and m columns, and then you have n²·m + n·m² — that is the complexity of TabPFN v2. That was okay for up to 10,000 data points, but if we wanted to scale higher — for 100,000, a million, etc. — we needed to go beyond that. Well, actually for 100,000, TabPFN 2.5, we still did that; we got away with just going to beefier GPUs. But then at the same time, the Gwas group actually developed TabPFN, which was really nice. There was an architecture that basically in the end used the TabPFN v1 architecture only — only quadratic over the number of rows — but it had a much more complex mechanism to actually do this embedding through a different column and a row transformer. And since that architecture was actually sort of good enough for the size range we wanted to do for TabPFN v3, and also in terms of inference speed, we actually just basically adapted that for TabPFN v3. And well, innovated in many different ways — for example, for many classes and for the output head, etc. But the base architecture we didn't actually have to change from that. We have a lot of things cooking, but we didn't actually need them for our tabular data for TabPFN v3. So they're going to be in future versions.

Host

但说清楚一点,v3 的复杂度是多少?

But just to be clear, what is the complexity of v3?

Frank

v3 的复杂度和 TabPFN 其实一样。基本上对行数是 n²,但对预处理、对嵌入来说是 n·m²。所以它已经更快了,但仍然是二次的。有无数种方法可以把它做到次二次,我们当然也在研究这些。所以下一个目标是 1000 万,我们相当有信心。

Of v3, it's actually the same as TabPFN. So it's basically n² for the number of rows, but then it's n·m² for the pre-processing, for the embedding. But so it is already faster, but it's still quadratic. And there's a gazillion methods to make it sub-quadratic, and of course we're looking at those. So the next target is 10 million, and we're pretty confident.

归纳先验与局部性 Inductive priors and locality

Host

这太有意思了。我是说,就连行和列——这都是我们给这些模型注入的归纳先验的一个绝佳例子,对吧?因为我们对数据有某些假设,行与行、列与列之间存在某种关系,这看起来很合理。但快速注意力会是什么样子?比如,如果你要削减注意力的跨度,会不会是某种局部性先验?

That's super interesting. I mean, even the rows and the columns — that is a wonderful example of the kind of inductive priors that we put into these models, right? Because we have certain assumptions about the data, and it seems reasonable that there are kind of relationships row-wise and column-wise. But what would a fast attention look like? So for example, if you were going to cut down that span of attention, would it be some kind of locality prior?

Frank

是的。我是说,在语言里,你有这种局部性先验。在表格数据里,你没有,对吧?因为你想要对特征有这种不变性,但你可以在嵌入空间里有某种局部性。

Yeah. I mean, so in language, you have this locality prior. In tabular data, you don't, right? Because you want to have this invariance over features, but you could have some sort of locality in the embedding space.

测试时适应与直推学习 Test-time adaptation and transductive learning

Host

非常酷。现在,我对测试时自适应非常兴奋。你知道 o1 出来了,那很了不起——模型变得智能了,因为它们通过这种强化学习训练,学会了给自己提示。它们实际上在深思、在思考。而很多人在这些语言模型上做的就是测试时自适应,他们在测试时故意做一些结构化的推理。感觉这就是做这件事的完美机会。所以目前你刻意选择了一个归纳式模型,这有一些有趣的特性——当然是在缓存、校准、性能这些方面。我是说,如果是直推式的,那意味着测试样本实际上可以绑定到训练样本上,这可能不好,因为假设我在做医院预测,我可能有一个异常的病人,然后现在另一个病人影响了我对第一个病人的预测。所以这不一定好,但它可以是非常强大的东西。这对你来说有意思吗?

Very cool. Now, I am very excited about test-time adaptation. So you know o1 came out and that was remarkable — that the models became intelligent because they were kind of, through this RL training, learning to prompt themselves. They were actually deliberating, thinking. And what a lot of folks have been doing with these language models is test-time adaptation, where they actually deliberately do some kind of structured inference at test time. And it just feels like this is the perfect opportunity to do that. So at the moment you've made a deliberate decision to have an inductive model, and that has some interesting properties — certainly in terms of caching, calibration, performance, that kind of thing. I mean, if it were transductive, what that would mean is that the test samples could actually bind to the training samples, and that could be bad because, let's say I'm doing a hospital prediction thing and I might have some weird patient who's an outlier, and now this other patient is affecting my prediction of the first patient. So it's not necessarily a good thing, but it can be an incredibly powerful thing. Is that interesting to you?

Frank

我觉得我不会太往直推式的方向走。我可以往半监督的版本走,就是你有少量标注数据,然后有大量未标注数据,你知道这是你真正想要做好的一类数据。但最终,我想我会坚持一种接口,即我为某个数据点做的预测不影响其他数据点的预测,因为数据科学家根本不会接受那种情况。我觉得如果你在不同的批次里预测同一个数据点却得到不同的结果,那太奇怪了。会非常奇怪。

So I don't think I would go too much into the transductive direction. I can go into the semi-supervised version where you have some low numbers of labelled data and then you have a whole lot of unlabelled data, and you know this is a type of data you actually want to work well for. But in the end, I think I would stick with an interface where what I'm making predictions for doesn't affect the predictions for the other data points, because data scientists just wouldn't have that. I think that would be so weird if, if you predict the same data point in a different batch, you get different results. It would be super strange.

自适应思考与潜变量 Adaptive thinking and latent variables

Host

是的,这很有意思。我是说,比如在 ARC Prize 里,它之所以效果这么好,是因为你只有几个例子,而那另外两个例子给了你大量信息,我们拼命想让模型适应。但我想另一件事就是某种通用的自适应思考。所以我可以想象——我觉得你有类似的东西,但我不认为你们具体怎么做是公开的——只是在这里大声思考一下,我可能会这样处理:你可能有某种潜变量,你可以有一个潜变量,它能随着更多计算而演化。所以我做更多前向传播,实际上在演化某种表示。这是你会考虑的吗——就是现在我们有一个固定的计算量,但有很多任务我们不确定,我们可能需要思考,你知道,多搜索一下才能得到更好的预测?

Yeah, it's interesting. I mean, in the ARC Prize, for example, the reason it works so well is you only have a few examples of something, and those other two examples, they give you so much information, and we're desperately trying to adapt the model. But I suppose another thing is just some kind of adaptive thinking in general. So I could imagine — and I think you've got something like this, but I don't think it's publicly known what you do — but just to think out loud here, the way I might approach something like that is you might have some kind of a latent variable and you can actually have some latent which could evolve with more computation. So I do more forward passes and I'm actually evolving some kind of representation. Is that something you think about — that right now we have a fixed amount of computation, but there are many tasks that we're uncertain about, that we might need to think, you know, search around a little bit to get a better prediction?

Frank

是的,这里有很多类比。比如,我是说,你可以有——对某些情况你可能想非常快地做出预测,因为它是一个很简单的数据点,所以你可以让一个大网络提前退出。对更难的数据点,你想一路走到底。对更难的数据点,你可能想更用力地思考。而更用力思考有很多方式。比如你可以去做微调——你可以看看,嘿,这是我的数据集。让我们生成类似的其他数据集,在上面做微调。其实我的一个硕士生有一篇关于这个的论文,对于非常小的数据集,这确实能帮助提升。然后你可以做提示微调。所以如果你有一个非常大的数据集,你可以说,好吧,这个大数据集里哪些我应该真正喂进我的上下文?你可以选择——假设你有十亿个数据点,你可以选 10 万个数据点放进你的上下文。你可以用基于梯度的方式来做,但你也可以实际上“幻觉”出 10 万个数据点,更好地近似这十亿个数据点。我们在 NeurIPS 也有一篇关于这个的论文。一个有趣的事实是,你其实也可以把它用于可解释性。

Yeah, there's a whole lot of analogies here. So for example, I mean, you could have — for some cases you might want to make a prediction really quickly because it's a really simple data point, and so you could have early exit of a big network. For harder data points, you want to go all the way through. For even harder data points, you might want to think it a lot harder. And there's a whole lot of ways of thinking harder. Like you could go and do fine-tuning — like you could look at, hey, this is my data set. Let's generate other data sets like this and do fine-tuning on that. And actually, one of my master students had a paper about that, and that for very small data sets actually can help improve. Then what you can do is you can do prompt tuning. So if you have, for example, a very large data set, then you could say, well, what of this large data set should I actually feed into my context? You could select — say you have like a billion data points, you can select 100,000 data points to put into your context. And you could do that in a gradient-based manner, but you could also actually hallucinate 100,000 data points that approximate these billion data points better. We also had a paper at NeurIPS about that. One fun fact there is you can actually also use this for interpretability.

数据集蒸馏与模型选择 Dataset Distillation and Model Selection

Frank

我之前的一位博士生 Robin Shemis 拿了一个包含一千个数据点的医学数据集,把它压缩成两个数据点,作为提示词输入,结果预测性能与输入一千个数据点完全相同。这相当于给出了数据的原型,类似的事情还有很多可以做。你也可以用不同的先验训练不同的网络,然后对它们做标准的交叉验证。有研究表明,Transformer 实际上可以在一次前向传播中完成模型选择。所以你可以用大约 10 种不同的先验训练一个 Transformer,让它在测试时通过一次前向传播选出正确的那个。能做的事情太多了,我认为天空才是极限。

One of my previous PhD students, Robin Shemis, actually took a medical dataset of a thousand data points and condensed that into two data points that you would feed as a prompt, and that was actually the same predictive performance as feeding the thousand data points. So that sort of gave you the prototypes of the data, and there's a whole lot of things like that you can do. You could also have different networks that were trained with different priors and do standard cross-validation on them. There is work that shows that Transformers can actually do model selection in a forward pass. So you could basically train a Transformer with like 10 different priors and have it choose the right one at test time in a forward pass. There are so many different things you can do and I think the sky's the limit.

适应与智能体 Adaptation and Agents

Host

没错。我的意思是,还有采样和集成,你也可以有某种草稿本。不过总的来说我对这个方向非常兴奋,因为我在想这里是否存在一种张力,因为我们把它说成是一种基础模型,你知道,无需适应。但感觉适应总是严格更好。这就是为什么我对使用智能体非常兴奋,因为新信息不断进来,我们将其结晶、优化再优化,最终得到更好的东西。感觉这可以成为实现这一点的架构基础。数据集蒸馏,这真的很有趣。我记得我采访过 MIT 的一位叫 Andrew Ilyas 的人。他有一篇关于数据集蒸馏的出色论文,还有机器教学等等。基本思想是,如果你只是修剪和蒸馏数据集,就能获得显著更好的性能,而且这可以是一个在线自适应过程,因为当人们实施这项技术时,他们生活在他们所在的地方,他们有非常专业化的数据,情况在不断变化,这个适应过程可以成为预测架构的一部分。

Yeah, exactly. I mean, there's also like sampling and ensembles and you could have some kind of a scratch pad. I'm just really excited about this in general though because I wonder whether there's a tension there because we're kind of couching this as a foundation model, you know, without adaptation. But it feels that adaptation is always strictly better. I mean, this is why I'm quite excited about using agents because, you know, new information comes in and we crystallize and optimize and optimize and we end up getting something better. And it feels like this could become the base of an architecture to do that. The dataset distillation, that's really interesting. I think I interviewed a guy called Andrew Ilyas from MIT. He had a great paper on dataset distillation and there's machine teaching and so on. And you know the basic idea is that if you just kind of prune and distill the dataset you can get dramatically better performance, and this can be an online adaptive process because when folks are implementing this technology they live where they live, you know, they have very specialized data, things are changing, and this adaptation process could be part of the predictive architecture.

特征重要性与对抗检测 Feature Importance and Adversarial Detection

Frank

当然。正如我所说,有很多很酷的事情你可以做,而像 XGBoost 这样的方法做不到。例如,你可以在一次前向传播中做特征重要性分析,或者你可以做一次前向-反向传播,然后说,如果我改变这个特征,我的输出会变化多少?你可以得到梯度信号,所以前向-反向传播很简单。但你也可以对某个数据点做同样的事,比如这个数据点对我的模型训练有多重要?你实际上可以检测对抗攻击等等。你可以说,啊,这真的很奇怪。有一个数据点完全主导了我的模型。如果我去掉这个数据点重新训练,会不会好很多?如果是,那就迭代这个过程,等等。让我们找出哪些是重要特征,只用重要特征训练,丢掉其他特征。能做的事情太多了。

Yeah, absolutely. And there are so many cool things as I said that you can do that you couldn't do with methods like XGBoost, for example. You could do feature importance in a forward pass, or you can do a forward-backward pass and say, well, if I change this feature, how much would my output change? You get a gradient signal for that, and so forward-backward is trivial. But you could also do that for a data point, like how important is this data point here for training my model? You can actually detect adversarial attacks, etc. And you could say, ah, this is really strange. There's this one data point that completely dominates my model. How about I train again without that data point and then is that going to be much better? And if so, then let's iterate this process, etc. Let's figure out what are the important features, train with only the important features and drop the other ones. So many different things you can do there.

插值与外推 Interpolation vs Extrapolation

Host

我们来谈谈插值与外推。我的好朋友 Randall Balestriero 有一篇关于神经网络样条理论的论文,他当时经常谈到,决策树,你知道,MLP 就像决策树,但它们能外推。所以这确实是一个新现象,你研究过这个,因为我见过一些图表,但这些模型确实以普通模型做不到的方式进行外推。那里的行为是怎样的?

Let's talk about interpolation versus extrapolation. Now, my good friend Randall Balestriero had this paper on the spline theory of neural networks and he was actually talking a lot at the time about how, you know, decision trees, you know, MLPs are like decision trees but they extrapolate. And so this is genuinely a new phenomenon then, and you've studied this because I've seen some of the graphs, but these models actually extrapolate in ways that normal models don't. What's the behavior there?

Frank

外推非常重要,对我这样来自 AutoML 和贝叶斯优化领域的人来说也是如此。如果你只看到空间的一部分,而你想外推到空间的其他部分。例如,缩放定律。你只看到小型神经网络的点,但你实际上想预测你的大型网络会有多好。你需要能够泛化。是的,你可以直接把泛化放进你的先验里。就说,我想用这里的这些数据点训练,并且我希望能预测那边的。如果训练集和测试集不是独立同分布这一点在你的先验中,那么你也可以学习其他分布的预测。

Extrapolation is super important, also for me coming from AutoML and Bayesian optimization. If you only see parts of the space and you want to actually extrapolate to other parts of the space. For example, scaling laws. You only see points for like small neural nets, but you actually want to predict how good your large network is going to do. You need to actually be able to generalize. And yeah, I mean you can just put it in your prior that you want to generalize. Just say, well, I want to train on these data points from here and I want to be able to predict over here. If that's in your prior that the train and test is not i.i.d., then you can also learn other distribution predictions.

谷歌TabFM与扩展 Google's TabFM and Scaling

Host

我想大约三周前,是 Google 发布了一个叫 TabFM 的东西,我记了一些笔记。首先,模仿是最真诚的奉承。但你知道,他们基本上深受你所做工作的启发,将其规模扩大了约 30 倍,并在这个 Tabarina 基准测试上获得了显著更好的性能。那这到底是怎么回事?

So I think about three weeks ago, was it Google released something called TabFM and I took some notes about it. So the first thing is imitation is the sincerest form of flattery. But you know they basically were inspired very much by the work you've done and they scaled it up around 30 times and they got significantly better performance on this Tabarina benchmark. So what's the story with that?

Frank

是的,正如你所说,他们也加入进来,基本上采用相同的架构、相同类型的先验等等,并利用 Google 的算力来解决,只是为了得到更大更好的模型,这有点令人受宠若惊。这更好并不奇怪。我认为做到那么大可能有点为时过早,因为它也带来一些问题。模型非常大,因此也相当慢。它大约比一次前向传播慢 15 倍,这很酷。我的意思是,它仅通过一次前向传播就能做到这一点,获得这样的性能,这确实很棒,但这是一次非常慢的前向传播。它大约比我们的前向传播慢 15 倍。我们有这种测试时计算的 topfn 思考,实际上比我们的标准模型慢 10 倍。不过,我们的测试时计算实际上比 TabFM 的一次前向传播更快,并且在 ELO 分数等方面更强。但它是一次前向传播,所以这表明如果我们真正研究缩放定律——当然我们已经有缩放定律——并且投入那样的算力,我们会变得越来越好。他们确实在小数据上下了不少功夫。

Yeah, I mean it's as you said, it's sort of flattering that they're also jumping on this and basically taking the same architecture, the same types of prior etc., and you know solving it with Google compute to just get bigger and better models. It's not surprising that this is better. I think it's maybe a little premature to go that big because it also comes with some issues. The model is very large and sort of quite slow as a corollary. That's sort of like 15 times slower than just a forward pass, which is super cool. I mean, the fact that it can do this just in a forward pass, get this performance, that is really nice to see, but it's a very slow forward pass. It's like 15 times slower than our forward pass. So we have this test-time compute topfn thinking that's actually 10 times slower than our standard model. Still, our test-time compute is actually faster than a forward pass of TabFM and is stronger in terms of ELO scores etc. But it is a forward pass and so it shows that if we actually look at the scaling laws and of course have scaling laws already and if we invest that compute, we will get better and better. They did focus quite a bit on small data.

在TabArena上的表现 Performance on TabArena

Frank

所以特别是在 TabArena 中最小的数据集上,它比之前的方法好得多。对于极小和小的数据集,它仍然更好;对于中等规模,实际上差不多打平。中等规模最多到 10 万个数据点,而这正是我们之前模型关注的重点。在 TabPFN 3 中,我们其实已经转向了其他基准测试,我们去了 TabLib 基准,它最多到 100 万个数据点,还有 BeyondArena。我们其实没有太关注 TabArena,它最多到 10 万,这是一个很好的提醒,我们应该也多关注一下那个。当然,我们当前的基准测试也会确保正在酝酿的下一代模型在 TabArena 本身上也更强大。但我们不想过度关注它,因为我们非常关心扩展到 100 万到 1000 万个数据点、非独立同分布的分组数据等等。而 TabPFN 作为一个如此大的模型,我们实际上甚至无法在 BeyondArena 上运行它,对于超过大约 10 万个数据点的数据集,它就会崩溃,内存不足等等。但很酷的是他们现在参与其中,非常兴奋看到这个领域成长。

So in particular for the smallest data sets in TabArena it's much better than previous methods. For tiny and small it's still better, for medium it's actually sort of breaking even. Medium is up to 100,000 data points, and that was sort of our focus with the previous models. In TabPFN 3 we actually already went to other benchmarks, we went to the TabLib benchmark which is up to a million data points, and to BeyondArena. We haven't actually focused that much on TabArena, which is up to 100,000, and this is a good reminder that we should also focus a bit more on that one. And so of course our current benchmarks also will make sure that the next model generation that's cooking is also much stronger on TabArena itself. But we don't want to overindex on it because we do care a lot to scale to a million to 10 million data points, to non-IID group data, etc. And TabPFN being such a big model we actually just can't even run it on BeyondArena, like it just breaks, out of memory, etc., for data sets that go beyond roughly 100,000. But it's very cool that they're part of this now and super excited for the field to grow.

护城河:有原则的先验 The Moat: Principled Priors

Host

有一个关于护城河的问题。你们实际上做了研究,你们理解这个。所以想必你们可以把这个东西朝着有趣的方向演进,但有一件事我们还没谈到,而它是你们护城河的一部分,就是设计这些先验时融入了大量有原则的知识,对吧?我们还没有真正讨论过你们是如何做到的。你们是如何制作那些先验的?

There is the question of the moat thing. So you guys actually did the research, you understand this. So presumably you can evolve this thing in interesting directions, but one thing we haven't spoken about that is part of your moat is a lot of principled knowledge went into designing these priors, right? And we haven't really discussed how you did that yet. How did you make those priors?

Frank

先验的开发是一个相当迭代的过程。我们开始的方式基本上是,对于 TabPFN 2,我们有一组留出数据集,我们不会在上面训练,但我们会用它来说:如果我们对先验做一个改动,训练一个网络,看看它在这些数据集上表现如何,有哪些失败模式,然后实际看到,啊,这个数据集有一个失败模式,这个数据集中的哪些结构可能导致它失败,例如离群值、缺失值或无信息特征等等。然后我们也会把这些放入我们的先验中,然后随着时间的推移,我们基本上只是让它变得越来越复杂,以捕捉数据可能产生的越来越多可能的方式。因为如果你的先验是有限的,例如再次只是线性线,那么后验将基于线性回归而受到限制。这是一个完全没问题的模型,但如果你的数据实际上不是线性的,它就非常糟糕。所以我们会有一个非常复杂的先验,可以尽可能多地捕捉世界。

Prior development is a fairly iterative process. The way we started was basically, for TabPFN 2 we had this set of hold-out data sets that we wouldn't train on but that we would use in order to say if we make a change on the prior, train a network, see how well does it do on these data sets, where are some failure modes, and actually see, ah there's a failure mode in this data set, what are some of the structures in this data set that it might be failing because of, for example, outliers or missing values or uninformative features, etc. And then we would also put those into our prior, and then over time basically we just made it ever more complex to capture more and more possible ways that the data might have come about. Because if your prior is limited, for example again just linear lines, then the posterior is going to be limited based on linear regression. It's a perfectly fine model but it's just very bad if your data is actually not linear. And so we would actually have a very complex prior that can capture as much of the world as possible.

因果与相关 Causality and Correlation

Host

我们来谈谈因果关系。这是在这个背景下经常被谈论的事情之一。我的意思是,首先你能解释一下因果和相关的区别吗?我的意思是,我们在多大程度上可以期望这些模型具有因果性,那甚至意味着什么?

Let's talk about causality. So this is one of the things that's spoken about a lot in this context. I mean first of all just can you explain what is the difference between causation and correlation? I mean to what extent can we expect these models to be causal and what would that even mean?

Frank

好的,让我用一个例子来解释。假设有一个医学例子,你有一种疾病,一些患者患有这种疾病,这些患者会服用一种特定的药物,比如他们得到一种特定的药物来帮助他们对抗这种疾病,疾病越严重,他们服用的药物就越多。如果你有一个数据集,你只看到患者服用了多少或哪些药物,你想预测他们是否患有这种疾病,那么你基本上可以建立一个完美的模型,说如果他们服用了这种药物,那么他们患有这种疾病。你可能会忍不住说,哈哈,让我们停止给他们那种药物,他们就不会再有这种疾病了。但那是愚蠢的,对吧?因为那里的因果关系是反过来的。因为他们有病,所以他们服药,而不是反过来。所以这就是为什么当你开始进行这些干预时,区分因果和因果关系非常重要。所以我刚才说让我们停止给他们那种药物。这意味着你在因果机器学习中改变一个变量,叫做 do 算子。所以你 do 给患者这种药物,而不是观察患者正在服用这种药物,这是两件完全不同的事情。因为如果我观察到患者正在服用这种药物,那么就有某种关系,他们服用它的原因是他们患有这种疾病,但如果我只是决定他们将服用它,那么我对这种疾病一无所知。所以掌握这一点非常棘手,但这正是因果机器学习发挥作用的地方。

Yeah, let me explain that with an example. So if you think of a medical example where you have a disease that some patients have and these patients become a particular type of medicine, like they get a particular type of medicine that helps them against that disease, and the stronger the disease is the more of that medicine they get. If you have a data set where you just see how much or which medicines a patient gets and you want to predict whether they have that disease or not, then you can basically build a perfect model that says well if they get this medicine then they have this disease. And you might be tempted to say haha let's stop giving them that medicine and they won't have that disease anymore. But that would be foolish, right? Because there the causal relationship is the other way around. Because they have the disease, they get the medicine, not the other way around. And so that's where causality and causation is really important to distinguish when you start with these interventions. So I just said let's stop giving them that medicine. So that means you change a variable in causal ML that says do operator. So you do give the patient this medicine rather than observe the patient is taking that medicine, and those are two entirely different things. Because if I observe the patient is getting this medicine then there is some relationship, the reason that they're getting it is that they have this disease, but if I just decide that they're going to get this then I just know nothing about the disease. And so getting a hold of that is very tricky, but that's where causal ML comes in.

Host

这个 do 演算来自 Pearl,这个因果阶梯,请简要说明一下。

And this do calculus came from Pearl, this ladder of causation, just frame that up.

Frank

是的,do 演算确实来自 Pearl,在 Pearl 的因果学派中,你知道图,如果你知道图,那么你可以计算各种东西。但在实践中,你往往不知道图。比如,如果你是一家公司,你想对你的某些产品进行定价,你观察到客户的数百个不同特征等等,你想决定某些产品的价格,你并不确切知道这些特征中哪个导致哪个等等。但你仍然想在这个因果空间中做出决策,并意识到这里可能存在因果关系。所以你可以做两件事,或者解决两个不同的问题。一个是实际做出预测,知道这些因果关系,所以这基本上就是我们的 SEM 发挥作用的地方,我们实际上是在对所有可能的因果关系进行积分的情况下做出预测。我不知道正确的因果关系,但数据似乎更多地表明这一种,所以我会更多地加权这一种等等,然后我基本上对所有 SCM 的空间进行积分。还有这个问题,嘿,实际上告诉我因果关系,那要棘手得多。实际上还有第三个问题,就是告诉我该进行哪些干预。所以第一个只是在观察下的预测。第三个是干预的预测,或者干预的效果,那个你实际上也可以用这些基础模型来做,而好处是,我们只需要改变训练过程,元训练。

Yeah, so the do calculus does come from Pearl, and there is this in the Pearl school of causality, you know the graph, and if you know the graph then you can compute all kinds of things. But in practice, often you just don't know the graph. And we like, if you're a company and you want to actually do pricing on some of your products and you observe hundreds of different features of your customers and so on and you want to decide prices for some of your products, you don't know exactly which of these features cause which and so on. And you still nevertheless want to make decisions in this causal space and be aware that there might be causal relationships here. So there are two things you can do, or two different problems to solve. One is to actually just make predictions knowing about these causalities, and so that's basically where our SEMs come in, where we actually make predictions under like sort of integrating over all the possible causal relationships. I don't know the right causal relationship but the data is sort of suggesting this one more so I will weight this one more etc., and then I basically integrate over the space of all the SCMs. There's also this question of hey actually tell me about the causal relationships, that is much trickier. And actually there's also a third question which is well tell me about which interventions to do. And so that is the first one was just predictions under observations. This the third one is predictions of interventional, like or the effects of interventions, and that one you can actually also do with these foundation models, and the nice thing is that again all we need to change is sort of the training process, the meta training.

因果基础模型的元训练 Meta-training for causal foundation models

Frank

所以在元训练期间,我们可以控制——通常我们会采样一个因果图,观察其中的变量,然后我们会再次采样这个因果图并观察变量,但接着我们会取同一个样本,实际上在其中进行干预,并观察干预的效果。然后我们就有了配对数据:这是观测数据,这是干预预测。我们可以从先验中数亿个图里学习:如果我看到这类观测并做出这个干预,就会产生这个效果。因此,你实际上可以在一定限度内,从纯观测数据中元学习,对干预效果做出预测。存在一些可识别性的限制等等。但你会对所有可能的世界——所有与数据一致的可能结构因果模型——进行积分。

So during meta-training we can control—usually we sample a causal graph, observe the variables in there, and now we would again sample this causal graph and observe the variables, but then we would take the same sample and actually make an intervention in there and observe the effects of that intervention. And then we have pairs of this is observational data and this is an interventional prediction. And we can learn from hundreds of millions of graphs from our prior: if I see these types of observations and I make this intervention, this will be the effect. And so you can actually meta-learn to make predictions about interventional effects from purely observational data within some limits. There are some limits of identifiability, etc. But you would just integrate out over all the possible worlds—all the possible structural causal models that are in line with the data.

抽象与因果变量 Abstraction and causal variables

Host

这也是语言模型的一个问题,它与我非常感兴趣的“抽象”概念有关。语言学家提炼出了一种非常简约、可能不完整的生成语法或最简方案之类的东西,我们把它提炼出来,使其对我们来说清晰易懂。因为如果你仔细想想,因果性可能发生在多个分辨率层次上。它可能一直深入到光锥。而我们所做的就是提出这些因果变量。我们提炼这些因果变量的一部分原因在于我们是世界中的行动者。所以我们可以实际尝试事物。我们可以探索反事实,然后我想随着时间的推移,我们只是将其提炼出来,并可以将这些注入我们的 AI 模型。这有点像你说的——你可以拥有大量核心的结构化因果知识,然后如果你能将因果知识挂接到机器学习模型中,你就可以用因果技巧进行预测。目前,机器学习模型在没有我们提供的情况下无法从头获得因果知识,这是不言而喻的吗?

This is also a problem with language models, and it's related to this concept of abstraction which I'm very interested in. So linguists have distilled a very parsimonious, possibly broken kind of generative grammar or the minimalist program or something like that, and we have distilled that down so it's legible to us. Because if you think about it, causality could happen at multiple levels of resolution. It could be all the way down to the light cone. And what we do is we come up with these causal variables. And part of how we distill those causal variables is because we are actors in the world. So we can actually try things. We can explore counterfactuals and then I guess over time we just distill it down and we can imbue these into our AI models. This is kind of what you're saying—that you can have a load of core structured causal knowledge and then you can do this prediction with causal trick if you can kind of latch on the causal knowledge into the ML model. I think does it go without saying at the moment that the ML model cannot acquire de novo causal knowledge without us giving it to it?

从干预与观察中学习 Learning from interventions vs observations

Frank

我的意思是,为了了解干预的效果,模型需要以某种方式观察干预的效果。所以我们通常做的是随机对照试验、AB 测试等,我们实际进行干预,然后基于这些干预数据你可以拟合模型。但通常你只有大量观测数据,而没有干预数据。而因果基础模型这一新研究方向实际上在某种程度上允许你在测试时无需观察干预就能对干预进行推理。但你在元训练时观察它们,你说这是我的因果模型,我采样了它。如果在这个因果模型中我进行这个干预,那么就会发生这种情况,你只需从数亿个数据集中学习其效果,然后你就观察到了大量的干预效果,并且你可以对这些进行预测。有时你会遇到不可识别性。可能是 A 导致 B 或 B 导致 A,你不太确定是哪一个。但我们不需要说出是哪一个。我们只需要解决贝叶斯积分:这个模型会说什么,可能性有多大;那个模型会说什么,可能性有多大。当你实际上不确定是哪一个,并且它们做出完全不同的预测时,在极端情况下你想预测——我不知道——有 0.5 的概率这个是对的,这是我的结果;有 0.5 的概率那个是对的,这是我的预测。再次,你有这个多模型分布,实际上中间的概率为零。要么是这个,要么是那个。而我只是不知道正确的因果图。但在许多其他情况下,你实际上可以以更高的概率确定是这个因果模型,然后实际上,概率质量会转移到那个因果模型会做出的预测上。

I mean, in order to know something about the effect of interventions, the model needs to observe the effect of interventions somehow. So what we usually do is randomized control trials, AB tests, etc., where we actually do interventions and then based on this interventional data you can fit models. But often you just have a whole lot of observational data and you don't have the interventional data. And what this new line of work on causal foundation models actually allows you to some degree is to actually reason about interventions without observing them at test time. But you observe them during meta-train time where you say well this is my causal model that I sampled. If in this causal model I make this intervention then this is going to happen and you just learn the effects of that over hundreds of millions of data sets and then you have observed a whole lot of interventional effects and you can actually make predictions about these. Sometimes you have non-identifiability. There might be that A causes B or B causes A and you don't quite know which one it is. But we don't need to say which one it is. We just need to solve the basin integral over well what would this one say and how likely is it and what would this one say and how likely is it and when you're actually not sure which one it is and say they make completely different predictions then in the extreme you want to predict I don't know with 0.5 probability this is right and this is my result and 0.5 probability this one is right and this is my prediction and again you have this multi-model distribution where you actually have zero probability that it's in the middle. It's either that or that. And I just don't know the right causal graph. But in many other cases, you can actually figure out with a higher probability it's this causal model and then actually yeah, the probability mass shifts over to the predictions that that causal model would make.

高保真反事实模型与因果标准 High-fidelity counterfactual models and the bar for causality

Host

所以如果我理解正确的话,你是说我们可以有高保真度的模型,可以想象反事实。所以它们可以说我可以——我进行这个干预——我可以想象这种可能性和那种可能性,而且模型如此之好,以至于想象的可能性空间相当不错。因此,模型原则上可以从那里选择并创建一个因果模型。但我想我试图理解的是,因果模型的标准是什么?因为即使我们在现实世界中进行随机对照试验,它仍然有些统计上的任意性。所以你知道我们可能建立因果关系,但它是二元的东西吗?我们是否处于这样的领域:这些统计因果模型将几乎与具有因果地位一样好,但又不完全是这样?

So if I've understood you correctly, you're saying that we can have high-fidelity models that could imagine counterfactuals. So they can kind of say I can—I do this intervention—I could imagine this possibility and this possibility and the model is so good that the imagined space of possibilities is quite good. So the model could then in principle select and create a causal model from that. But I guess I'm trying to understand what is the bar for a causal model because even if we do RCTs in the real world it's still somewhat statistical arbitrary. So you know we might establish a causal relationship but is it a binary thing? Are we in the domain where these statistical causal models will be nearly as good as having causal status but not quite or something like that?

因果基础模型与RCT Causal foundation models and RCTs

Frank

是的。我的意思是,你可以进行随机对照试验,通常当你进行随机对照试验时,你会得到更好的预测,但实际上现在有了这些因果基础模型,在某些情况下你可以非常接近有随机对照试验时的性能。所以已经有论文可以减少你在随机对照试验中所需的数据点数量,同时仍然获得相同类型的性能。当然,你知道,我认为如果将来我们可以进行更短的研究,让我们的药物更早上市,并且仍然具有相同的置信度甚至更高的置信度,这将彻底改变医学和 AB 测试等等。所以我认为在那里做很多伟大的事情有很大的潜力,我对此非常兴奋。但当然,你知道,我们需要正确地做这件事,并且有理论基础等等,这是高风险的,但是的,潜在收益是巨大的,我对此超级兴奋。

Yeah. I mean so you can run RCTs and typically when you run RCTs then you get much better predictions but actually now with these causal foundation models you can in some cases get very close to the type of performance you get if you have RCTs. So there are already papers that then reduce the number of data points you need to have in your RCT in order and still get the same type of performance. And that of course you know I think will completely revolutionize medicine and AB tests and so on if in the future you know we can get away with much shorter studies and get our medicine to the market much earlier and still with the same confidence or even higher confidence and so I think there's a whole lot of potential to do a lot of great things there and I'm very excited about this but of course you know we need to do this properly and theoretically grounded and so on and this is high stakes but yeah the potential gains are huge and I'm super excited about that.

假设的下一个产品:因果模型 Hypothetical next product: causal model

Host

那会是什么样子?所以如果你假设创建了——你知道就像你的下一个产品是 Tab Causal 而不是 TabPFN,而是某种因果的东西——那会是什么样子?所以它会仅仅以纯观测数据的形式呈现,还是会有某种数据科学家以某种方式构建它,以便模型的目的是——就像 TabPFN 那样——基本上非常高效地抓住因果结构,然后这个因果模型的好处是它可解释、可靠、可解释,诸如此类。

And what would that look like? So if you hypothetically created the—you know like your next product was Tab Causal not Tab PFN but something causal—what would that look like? So would it be couched as just purely observational data or would there be some kind of data scientist who would be structuring it in a certain way so that the purpose of the model would be—you know just as TabPFN does—to essentially grab that causal structure very very efficiently and then the benefit of this causal model is it would be interpretable, it would be reliable, explainable, that kind of thing.

因果模型的现有工作 Existing work on causal models

Frank

我们已经有一些关于这个的论文。我们有这篇 DP PFN 论文,与本内特·奇尔科夫合作,他是我们的顾问,也是全球因果机器学习领域的领导者之一,世界上被引用最多的因果机器学习科学家。

We already have some papers on this. So we have this DP PFN paper together with Bennett Chilkov who is an adviser of ours and sort of yeah one of the leaders of causal ML worldwide, most cited scientists in causal ML in the world.

因果基础模型与开源研究 Causal Foundation Models and Open-Source Research

Frank

所以他绝对知道自己在做什么。我们还有其他一些团队也在做这个。Causal FM 来自 Layer 6 和多伦多大学。我们上周还迎来了那篇论文的第一作者加入。所以我们团队里有很多做因果性的优秀人才。基本上,你可以这样想象:你拿到一些观测数据,想预测如果我改变这个变量会发生什么。比如我给这个病人用这种药,另一个变量——比如疾病状态——会发生什么?会变好还是变坏?在某些情况下你确实能做出预测。在某些情况下,你总会做出预测,但预测可能完全是“不确定,我就是不知道”。但如果数据里确实有一些因果性的痕迹,指向某个方向,比如 A 更可能引起 B,那我们就可以隐式地给它更多权重。当然这是模型学到的,但我们可以给“A 确实引起 B”的模型更多权重,从而做出更好的预测。很酷的是,我们现在还有后续工作,发表在最近的……领域科学家如果知道 A 引起 B,他们可以明确指定这一点,然后我们构建结构因果模型,覆盖所有……抱歉,是在所有“A 确实引起 B 而非 B 引起 A、也非两者独立”的结构因果模型上构建后验。这样我们就能真正利用人们的领域专业知识。还有一些工作可以概率性地推断出因果结构,也就是说:给定这些数据,A 引起 B 的可能性有多大?这让你能做各种事情,比如根因分析等等。我觉得这非常令人兴奋。我们当前的模型还做不到这一点,但这绝对是我们未来要做的方向。我们刚刚启动了开源研究部门,与社区里的任何人合作攻克这类难题,因为我认为这些问题对人类来说太令人兴奋、太重要了,不能只靠我们自己来做。我们应该真正与世界上最优秀的研究者合作来解决这些问题。当然,我们会围绕它们构建产品,但核心问题必须解决。在这方面,我们不希望有任何红线阻碍,而是真正能够与任何人合作,开源。

So he definitely knows what he's doing. We also have some other teams working on this. Causal FM comes from Layer 6 and the University of Toronto. We actually have the lead author of that joining us last week as well. So we have a lot of great people on the team that work on causality. Basically, you can imagine this as: you're getting some observational data and you want to predict, if I change this variable here, what will happen. So if I give this patient this medicine, what will happen to another variable, for example the disease status? Is that going to get better or worse? In some cases you can actually make predictions about this. In some cases, well, you will always make a prediction, but the prediction might be just complete uncertainty: I just don't know. But if the data actually has some traces of causality that point in some direction, that hey, it's more likely that A causes B, then we can give more weight implicitly. Of course this is learned by the model, but then we can give more weight to the model where A actually causes B and then make better predictions. The cool thing is we now actually have a follow-up work of that as well, published at the last... where domain scientists, if they know that A causes B, they can actually specify that, and then we build the structural causal model over all... sorry, the posterior over all structural causal models where A actually causes B and not B causes A, and not where they're independent. So that allows us to really use this domain expertise of people. And then there's also work that can probabilistically get out some causal structure, so that can say: given this data, how likely is it that A causes B? And that allows you to do all kinds of things like root cause analysis, etc. And I think that's really exciting. Our current models do not do that yet, but that is definitely something that we are working on for the future. And we just actually started our open-source research arm, where we collaborate with anyone in the community on tough problems like this, because I think these problems are too exciting and too important for humanity that we just work on them by ourselves. We should really work together with the greatest researchers in the world in order to tackle these problems. And then of course we'll build products around them, but it's really important to solve these problems at the core. And there we don't want to have any red line in the way, but really be able to work with anyone, open source.

Host

是的,这真的很令人兴奋。关于这一点,最后一个问题:因果结构学习的潜力是什么?目前是我们决定变量是什么,给它们命名,它们是可读的、简约的等等。但我可以想象一个未来,AI 能够提出某种非常高分辨率但好得多的东西。所以我想,因果结构的保真度与它对我们的可理解性和可解释性之间,是否存在权衡?

Yeah, that's really exciting. And just a final question on that bit: what is the potential for causal structure learning? So at the moment we decide what the variables are and we give them names and they're legible, they're parsimonious and whatnot. But I can imagine a future where an AI would be able to come up with something which was very high resolution but much better. So I guess, is there a trade-off between the fidelity of the causal structure versus its intelligibility and explainability to us?

Frank

LLM 会提出的当然是世界知识和世界中的因果关系。我不认为它们能提出我们从数据中看到的特征之间的恰当因果关系。如果你有数万、数十万、数百万个数据点,这又是在数字里、在统计里,而这是 LLM 不擅长的。我看不出它们短期内会在这方面变强。

What LLMs will come up with is, of course, the world knowledge and causal relationships in the world. I do not think that they will come up with proper causal relationships between features that we see from the data. If you have tens of thousands, hundreds of thousands, millions of data points, again, it's in the numbers, in the statistics, and that's what the LLMs are not strong at. And I don't see them getting stronger at that anytime soon.

Host

就这一点,我有个题外问题:LLM 非常了不起。你知道,你读到那个 Erdős 单位距离猜想,但那是 125 页的思维链。它们似乎做不到的是——这也是我们参与这个过程的原因——我们可以看着那团乱麻说,啊,这基本上就是这个,我们可以把它抽象出来,我们是在用更少做更多。我们在压缩。为什么 LLM 不这么做?

Well, just on that, that's a tangential question, which is that LLMs are incredible. You know, you read that the Erdős unit distance conjecture, but it was 125 pages of chain of thought. What they don't seem to do, and the reason why we are part of the process, is we can kind of look at the spaghetti and we could say, ah, that's basically this, you know, and we can abstract it and we're doing more with less. We're compressing. Why don't the LLMs do that?

Frank

老实说,我认为这也是它们会学会做的事情。因为 LLM 确实会压缩知识。如果你读十亿个 token,你不可能全记在记忆里,你需要以某种方式压缩,而它们会学着做得越来越好。我不知道,它们会形成引理之类的东西然后记住。但这仍然是在语言的空间里,而不是在数据和数字的空间里。我认为挑战在于把两者结合起来,真正利用世界知识、利用惊人的思维链能力等等,同时也引入我们从数据中能看到的因果信息。

So honestly I think that is something that they will also learn to do. Because, you know, LLMs do compress their knowledge. If you read a billion tokens, you can't keep that in your memory. You need to compress it somehow, and they will learn to do this better and better. And I don't know, form lemmas and whatnot that they will then remember. But that is still sort of in the space of language and not in the space of the data and the numbers. And I think the challenge will be to bring the two together and really use the world knowledge, use the amazing chain of thought capabilities, etc., but also bring in the causal information that we can see from the data.

关系数据与TabPFN Relational Data and TabPFN

Host

我们之前承诺稍后会回到多模态。所以你现在关注的一些东西是文本数据、关系数据、图数据之类的。这在 TabPFN 里是怎么体现的?

So we promised we'd come back to multimodality later. So some of the things that you're looking at now is text data, relational data, graph data, that kind of thing. How does that play out in the TabPFN?

Frank

是的,关系数据目前是个有趣的方向。我们其实在这方面做得还很少。但我们从开源社区免费得到了这个,他们开发了一种关系数据库的嵌入方法。也就是说,在关系数据库查询中,你想预测某张表里的一些实体,并且想把所有关联的表作为上下文——那些能告诉你关于这个特定实体信息的其他表。于是你可以把那个关系数据库查询嵌入成一张单表。然后他们直接跑了 TabPFN——我记得是 TabPFN 2 或 2.5——我们拿过那个方法,在上面加了一点东西,跑了 TabPFN 3。我们还构建了更好的基准等等,确保我们优化的是正确的目标。TabPFN 3 开箱即用就是这些关系数据库上最好的基础模型。我们在 Kumo 原始的 RelBench 上做了评估,它比 Kumo RFM 更好,后者是他们自己的关系基础模型,最后也接了一个单独的表格式基础模型。然后我们还花了一些功夫做了一个更好版本的 RelBench,下周就会发布。我想等这个采访播出时,那大概是三周前的事了。

Yeah, so relational data right now, that's an interesting one. There we haven't actually worked on this really much at all. But we got this for free by the open source community, who just actually developed an embedding of a relational database. So queries in a relational database where you want to predict some entities in one table and you want to actually take as a context all the connected tables, where you have other tables that tell you something about this particular entity. And so you can actually have an embedding of taking that relational database query and putting that into a single table. And then they actually just ran TabPFN — in this case I think they ran TabPFN 2 or 2.5 — and we took that method and well, did put a little bit on top and ran TabPFN 3 with that. And what we did is actually built better benchmarks and so on to make sure that we're optimizing the right objective. And yeah, TabPFN 3 out of the box was actually just the best foundation model for these relational databases. We evaluated this on the original RelBench from Kumo. It was better than Kumo RFM, which is their own relational foundation model. That also has an individual tabular foundation model at the end. And yeah, then we also worked a bit more on making a better version of RelBench that we're actually releasing next week. I think when this interview airs it's going to be three weeks ago.

Host

太棒了。有意思的一点是,我们之前聊到模型里有一堆先验,对吧?比如,它可能有局部性先验。

Amazing. I mean, one interesting thing is, we were talking earlier about there's a bunch of priors in the model, right? So, for example, it might have a locality prior.

几何先验与关系数据 Geometric Priors and Relational Data

Host

它有列和行之类的。如果加入其他类型的几何先验,会意味着什么?比如,原则上你能不能直接放一个关系先验进去,它就能理解,这样你实际上可以放入多个数据表,它就能理解它们之间的关系,并且能以某种方式隐式地做连接之类的操作,或者处理图数据或树数据。这在原则上可能吗?

It's got columns and rows and whatnot. What would it mean to put in other types of geometrical prior? So for example, could you in principle just put a relational prior in there and it would just understand, so you could actually put in multiple data tables and it would just understand the relationships between them and it could somehow implicitly do joins and all of that kind of stuff, or maybe graph data or tree data. Is that possible in principle?

Frank

是的。

Yes.

Host

好的。

Okay.

Frank

所以基本上,这显然就是构建先验的关键所在。你只需看看你想擅长的数据类型,把它放入先验中,生成数据,然后模型就会擅长那个。当然,如果你有不同的关系表,你需要解决一些架构上的挑战。你究竟如何把它放入你的网络中?你如何确保它是高效的等等。但没错,那里天空才是极限。我认为这完全会发生。

So basically this is clearly the name of the game of building prior. You just look at the types of data you want to be good at, put that in your prior, generate your data, and then the model is going to be good at that. Of course, you need to solve some architectural challenges if you have different relational tables. How exactly do you put this into your network? How do you make sure that this is efficient and so on? But yeah, the sky's the limit there. I think this is totally going to happen.

企业采用与安全 Enterprise Adoption and Security

Host

那么,你认为这在企业中会如何发展?因为我认为企业中有一个小问题,就是存在安全边界,有时表格数据和运营数据通常是最难获取的。所以就个人而言,我迫不及待地想进入 Quickbooks,下载我所有的数据,然后把它塞进 TabPFN。但在企业中,你认为这会如何发展?

And how do you see this playing out in enterprise? Because I think there is a bit of an issue in enterprise that there are security boundaries, sometimes tabular data and operational data in general is the hardest to get. So speaking personally, I can't wait to go into Quickbooks and download all of my data and stick it into TabPFN. But in the enterprise, how do you see this playing out?

Frank

是的,表格数据是企业中最常见的模态,然而,历史上并没有这些表格基础模型,但人们使用传统模型,如 XG boost 等,现在又接触到编码智能体和 LLM 等。许多新手尝试在表格数据上使用 LLM,结果一败涂地,因为它们没有见过这种类型的数据,因为这些数据没有在互联网上,所以没有经过训练来很好地处理这种数据。但然后,是的,有一个巨大的机会,实际上,这就是需要表格基础模型的地方,而且表格基础模型越能了解组织中的数据,它们就会变得越好。所以你也可以采用一个表格基础模型,我们的模型是开放权重的,你实际上可以在组织的数据上对其进行微调,然后它会在这种特定数据上获得更好的性能,因为这样你就可以真正提取出那种数据中的复杂模式,并在这方面工作得更好。

Yeah, tabular data is the most common modality in the enterprise, and nevertheless, historically there aren't these tabular foundation models out there, but people use traditional models like XG boost etc., and now are being exposed to coding agents and LLMs etc. Many newbies sort of try the LLMs on their tabular data and see them fall flat on their nose because they haven't seen this type of data because this data hasn't been on the internet, so it hasn't been trained to work well on this data. But then yeah, there's this big opportunity to actually, yeah, that is where tabular foundation models are required, and the more the tabular foundation models will actually learn about the data in the org, the better they will get. And so you can also take a tabular foundation model and our models are open weight, and you can actually fine-tune it then on the data of the org, and it will get much better performance for this particular data because then you can actually just extract the type of intricate patterns in that type of data and work much better for that.

组织设计与数据治理 Organizational Design and Data Governance

Host

那么,你认为组织应该如何围绕这一点设计结构?例如,可能有某种数据平台。你认为组织中的个人应该构建这样的模型,还是应该有一个平台团队?你会如何进行数据治理等?因为如果你想一想,组织中有着丰富的表格数据。其中一些在 Office 365 中。一些只是来自运营数据。你知道,这里可能有一个零售系统,那里有一个金融系统。所以本质上,你想要做的是让这些数据可用,但也需要有控制。你会怎么做?

And how do you think an organization should design structures around this? So for example, there might be some kind of data platform. Do you think that individual folks in your organization should be building out models like this or should there be some kind of a platform team? How would you do the data governance etc? Because if you think about it, there's a wealth of tabular data in the organization. And some of it is in Office 365. Some of it is just from operational data. You know, might be a retail system over here, financial system over here. So essentially what you want to do is make this data available, but there also needs to be controls. How would you do that?

Frank

这看起来会和 LLM 非常相似,因为你不想让你的数据对每个人都可用。你也不希望个人训练自己的 LLM。那没有任何意义。相反,你希望组织中的某一部分真正理解这些模型,知道如何微调它们,知道它们在哪里仍然不足,哪里应该引入一些其他数据,哪里需要检查 XG boost 是否仍然更好,如果你有十亿个数据点之类的。所以我认为这是一个相当集中的服务。当然,作为 Prior Labs,我们想在这方面提供支持,但是的,需要与组织合作,如何实际实施这一点。

It's going to look pretty similar to LLMs because you don't want to make your data available to everyone. You also don't want individuals training your own LLMs. That doesn't make any sense. Rather, you want sort of one part of the org really understanding these models, knowing how to fine-tune them, knowing where they still fall short, where should I bring in some other data, where do I need to check whether maybe XG boost is still better if you have a billion data points or whatnot. And so I see this as a pretty centralized and services that are being offered. And of course, as Prior Labs we want to support in that, but yeah, the need to work with the orcs of how to actually play this out.

收入模式与商业产品 Revenue Model and Commercial Offerings

Host

那么,你们的收入模式是什么?因为据我了解,目前你可以直接使用你的 clawed 智能体,它可以使用 tapfn,你可以做很多事情,但你们正在构建什么样的商业产品?

And what is your kind of revenue model because as I understand at the moment you know you can just use your clawed agent it can use tapfn you can just do a bunch of stuff but what what are you building as as as a kind of commercial offering?

Frank

是的。所以,你可以以多种方式使用 TabPFN。你可以从开源版本使用,采用非商业许可证。你实际上可以尝试、测试,在非生产环境中完全免费测试。如果你真的想在生产中使用并赚钱,那么你需要支付许可费。所以,这是一种商业模式,即许可费。然后我们还有各种围绕它的东西。我们有这种思考模式、扩展模式,在大数据上表现更好。我们也有特别适合非常小数据的模型。我们当然也有一些服务,帮助人们起步,让他们在表格基础模型上获得更好的性能。我们提供微调即服务。所以基本上有各种层次,你可以将它们组合起来。但最终,你可以把它看作一个平台,只是让数据科学家极大地提高效率。

Yeah. So, you can use TabPFN in many different ways. You can use it from the open source with a non-commercial license. You can actually try it out, test it, it's entirely free to test it in a non-production setting. If you actually then want to use it in production and make money with it, then well, you need to pay license fee. So, that's one business model is this license fee. And then we have all kinds of things around it. We have this thinking mode, scaling mode working better for large data. We also have models that work particularly well for very small data. We of course also have some services to get people off the ground to get them better performance with tabular foundation models. We have fine-tuning as a service. So all kinds of layers on top basically that you can stitch together. But ultimately you can think of it as a platform that just makes a data scientist dramatically more efficient.

未来产品与服务栈 Future Offerings and Serving Stack

Host

好的。然后我只是在想象几年后的情况。所以你可能有一个前端工程团队。你最终可能会有一些模型太大,客户无法自己服务。所以想必你可能会提供某种服务栈。微调的事情很有趣。我们没有谈到那个。但我的理解是,你可以微调模型,但客户还不能微调他们自己的模型,或者

Okay. And then I'm just sort of like imagining a couple of years ahead. So you might have a forward engineering team. You might eventually have models that are too big for customers to serve themselves. So presumably you might offer some kind of a serving stack for that. The finetuning thing is interesting. We didn't talk about that. But do I understand that you can fine-tune the models but the customers can't yet fine-tune their own models or

Frank

模型还没有大到客户无法服务,但我们可能能够更快地服务它们。我们可能有不同的技巧。所以我们已经提供了一个 API,我们只是控制基础设施,我们可以确保我们可以批量处理不同的预测。我们可以确保 GPU 保持忙碌,如果客户要预留一个 GPU 来进行所有预测,那个 GPU 可能会一直闲置,实际上我们做服务可能比客户自己做服务更具成本效益。所以这个 API 也是商业化的一部分。然后我们有私有 VPC,对于不能离开组织的数据,但如果你已经有云合作伙伴,那么它可以在那里运行。

The models are not too large for the customers to serve yet but we might be able to serve them faster. We might have different tricks. So we do already offer an API where we just control the infrastructure and we can make sure that we can batch different predictions. We can make sure that the GPUs keep busy and if a client was to reserve a GPU in order to do all the predictions that GPU might idle all the time and it might actually be much more cost-efficient for us to do the serving than for the clients to do the serving. So this API is also something that is part of the commercialization. And then we have private VPCs where for data that cannot leave the org but if you already do have a cloud partner then it can run there.

微调商业化 Commercializing Fine-Tuning

Frank

比如 Azure 和 AWS,你已经可以在那里使用它。当然在 SAP 和 GenAI Hub 里,我们有 TabPFN。所以有多种不同的方式来实现微调商业化。我们特别擅长微调。也有一些微调封装工具。人们可以微调模型,但同样,如果他们针对自己组织内的数据微调模型,那么模型已有的许可证也适用于他们微调模型的情况。所以微调——我们可以做,或者客户可以做。两种方式都可行。

So for example, Azure and AWS, you can already use it there. And then of course in SAP and the GenAI Hub, we have TabPFN. So there's a variety of different ways to commercialize fine-tuning. We're particularly good at fine-tuning. There are also some fine-tuning wrappers. People can fine-tune the models, but again, if they fine-tune the models to the data in their org, then there's already the license for the model that also applies when they fine-tune the model. So the fine-tuning is—we can do it or the customer can do it. Both would work.

Host

如果 OpenAI 或 Anthropic 直接把它嵌入到他们的框架里,会发生什么?你会直接向他们收费,还是根据客户的使用情况来收费?

What would happen if OpenAI or Anthropic just embedded it in their harness? Would you just charge them straight away, or would it be based on how it was used by the customer?

Frank

是的。看看代币经济会如何发展会很有趣。我认为表格预测代币会比 LLM 自己尝试做要便宜得多。你可以有一个智能体去写一些代码,得到可能接近的结果,但那会花费大量时间和代币。相反,你可以直接调用一个表格基础模型,它应该做得更快、更好,而且用的代币更少,或者用更少成本的代币量。所以在这个未来的智能体经济中,我们正在为此开发方法。它们应该比其他任何方法都更好。你知道,LLM 会调用计算器,因为它们更擅长做数学。同样地,它们应该调用表格基础模型,因为它们在处理这类任务上比 LLM 自己做得更好、更便宜。

Yeah. That's going to be interesting to see how the token economy will work out. I think the tabular prediction tokens will be much cheaper than the LLM trying to do it themselves. You can have an agentic agent that goes and writes some code and gets something that's maybe close to that, but that takes a lot of time and tokens. Rather, you could just actually call a tabular foundation model and it should do it faster and better and for less tokens, or for maybe an amount of tokens that costs you less. So in this agentic economy of the future, that's what we're developing the methods for. They should be just better than any other method. And you know, LLMs call calculators because they're better at doing maths. And just like that, they should call tabular foundation models because they're better and cheaper at doing this than they can do that themselves.

Host

我的意思是,我只是在想,如果 Anthropic 的人听到这个,他们的第一个想法会是:我要更新系统提示,如果有人使用表格数据,就直接用 TabPFN。或者你知道,大多数普通人在家里听到这个,他们会想:我要更新我的 Claude 全局技能,任何时候使用表格数据都用 TabPFN。

I mean, I'm just imagining that if someone from Anthropic was listening to this, their first thought would be, I'm going to update the system prompt and just use TabPFN if anyone's using tabular data. Or you know, most normal people at home, if they listen to this, they're going to be thinking I'm going to update my Claude global skill to use TabPFN anytime I'm using tabular data.

Frank

是的,绝对可以这么做。我们有技能文件等等。

Yeah, absolutely do that. We have skills files and everything.

Host

是的,我的意思是,这似乎是不言而喻的。是的,另一件事是,你们正在启动一个研究部门。给我们讲讲吧。你们在做什么?

Yeah, I mean, it just seems like a no-brainer. Yeah, the other thing is, so you're launching a research arm. Tell us about that. What are you guys doing?

启动开放研究部门 Launching an Open Research Arm

Frank

是的,我们正在启动一个开放研究部门。我们已经在 Prior Labs 做了很多研究,我们还想做更多基础研究,也因为,嗯,我们刚被 SAP 收购,我们有很多资金可以做很多酷炫的事情。收购的主要动机一直是 DeepMind,所以我们想成为世界上最好的工作场所。我们想大量发表论文。我们想成为每个人为了做最伟大的工作而去的地方,好事自然会随之而来。在开放研究部门,将是完全开放的研究。我们可以与世界上任何人合作——与大学、与 Ellis 研究所、与其他合作伙伴,比如——我还在弗莱堡大学有一个研究小组。我们已经与多伦多的 Layer 6、多伦多大学、新加坡、图宾根的 Ellis 研究所等合作。所有这些都应该非常自由和容易地实现。所以基本上,我们将把在大学做的事情扩大规模,也在 Prior Labs 这样做,这对人们来说当然会超级协同,如果他们想产生影响,然后也看到他们的一些工作被转移到产品中。我们还想要攻克登月项目。所以真的去解决高风险、高回报的问题,如果我们解决了它们,就能真正推动人类进步。如果你在听,并且你有这方面的数据,如果你能更好地为这些数据做表格预测,也许是超级复杂的数据,你可以——我不知道——用这个帮助治愈癌症或帮助全球粮食生产或帮助任何对人类有益的问题,请联系我们。这绝对是我们想要真正关注并做一些酷炫事情的方向。

Yeah, so we're launching an open research arm. We're already doing a lot of research in Prior Labs, and we want to do a lot more fundamental research also because, well, we just did get acquired by SAP, and we have a lot of funds available to do a lot of cool stuff. The leading motive for the acquisition has been DeepMind, and so we want to be the best place in the world to work. We want to publish up a storm. We want to be the place where everybody goes in order to do the greatest work, and good things will happen from there. And in the open research arm, it's going to be fully open research. We can collaborate with anyone in the world—with universities, with the Ellis Institute, with other partners like—I also still have a university group at the University of Freiburg there. We already collaborate with Layer 6, for example, in Toronto, University of Toronto, with Singapore, with the Ellis Institute in Tübingen, etc., etc. And all of those should just be possible super freely and easily. So basically, we're going to take what we do at the university and scale that up, and also do that in Prior Labs, and it's going to be of course super synergistic for people if they want to have impact and then also see some of their work being moved into a product. And we also want to tackle moonshots. So like really take high-risk, high-gain problems that if we solve them could really move the needle for humanity. If you're listening and you have data for that, if you could do tabular predictions better for this data, maybe super complex data, and you could—I don't know—with this then help cure cancer or help global food production or help any type of problems that are good for humanity, please get in touch. This is definitely something we want to actually really focus on and do some cool stuff.

Host

太棒了。所以,明确一下,你们正在招聘研究员、实习生。我对此很兴奋,因为我们听到太多关于美国和中国主导 AI 领域的消息。我们在德国南部的弗莱堡,做着前沿 AI 研究。欧洲有大量有才华的人。所以给 Frank——他们给你一个邮箱吗,Frank?

Amazing. So yeah, to be clear, you are hiring so researchers, interns. I'm excited about this because we hear too much about America and China dominating the AI space. We're here in the south of Germany in Freiburg and doing frontier AI research. There are loads of talented folks in Europe. So give Frank—do they give you an email, Frank?

Frank

我只是——请去我们的网站申请。我们有超过 11,000 份申请。

I just—please go to our website and apply. We have over 11,000 applications.

Host

好的。好的。

Okay. Okay.

Frank

我已经雇用了其中 45 人。但如果你对在这种团队工作感到兴奋——我们很幸运已经有大量的人申请,我们能够挑选世界上最好的人。如果你对在这个领域工作感到兴奋——是的,无论如何,请申请:brain.ai/careers。我们几乎招聘你能想到的每一个职位。特别是现在还有研究科学家,做模型开发和前瞻性模型开发。你也可以考虑 6 到 9 个月后的情况,构建更好的架构等。但也有完全开源的部门,你可以与世界上任何人合作。

And I have hired 45 of them. But if you're excited to work in that type of team where we had the luxury of a whole lot of people already applying and us being able to pick the world's best folks in the space. If you're excited to work in this—yes, by all means, please do apply: brain.ai/careers. We hire for pretty much every role you can think of. In particular now also the research scientists that do model development and do sort of forward-looking model development. You can also think for like 6 to 9 months out and build better architectures, etc. But also the completely open source arm where you can collaborate with anyone in the world.

Host

太棒了。Frank,非常感谢你今天加入我们。非常精彩。非常感谢你的邀请。

Amazing. Well, Frank, thank you so much for joining us today. It's been great. Thank you so much for having me.

TabPFN 3.5更新 Update on TabPFN 3.5

Host

现在,正如你们许多人可能已经看到的,自从我们录制这次采访以来,Prior Labs 已经发布了 TabPFN 的新版本,即 3.5 版,它看起来非常非常好。所以,我联系了 Frank,他为我们录制了一个非常简短的更新,我现在将把它附加到视频中。

Now, as many of you may have seen since we recorded this interview, Prior Labs have now released a new version of TabPFN, which is version 3.5 and it's looking really, really good. So, I got in touch with Frank and he has recorded a very short update for us, which I'm going to tag on the video now.

Frank

我们刚刚发布了最新的模型 TabPFN 3.5,通过它我们在比 TabArena 更多的基准上进行了评估,我非常高兴地说,我们在七个不同的表格相关基准上获得了第一名,包括带有文本的表格、多模态数据、关系数据等,涵盖了许多数据科学家实际面临的问题。我们没有太关注 TabArena,但也在那里有所改进。所以 TabPFN 3.5 在帕累托意义上优于其他表格基础模型,例如在相同质量下前向传播比 TabPFN 快 20 倍,或者 ELO 分数高出 100 多分,同时仍然更快。但更重要的是,我们在提到的 beyond arena 基准上有了巨大改进,这更类似于数据科学家每天实际面临的数据复杂性,包括分组数据、时间数据、从微小到百万数据点的数据规模、表格中的文本、高基数特征等。因此,TabPFN 现在也与数据科学竞赛更加相关。

We just released our newest model, TabPFN 3.5, and with that we evaluate on much more than TabArena and I'm super happy to say that we took first place on seven different tabular related benchmarks including tables with text, multimodal data, relational data, etc., covering a lot of the problems that data scientists actually face. We didn't focus on TabArena that much but also improved there. So TabPFN 3.5 Pareto dominates other tabular foundation models, for example being 20 times faster in a forward pass than TabPFN for the same quality or also over 100 ELO points higher while still being faster. But more importantly, we improved dramatically on the beyond arena benchmark that I mentioned, which is much more similar to the types of data complexities that data scientists actually face every day with group data, temporal data, data scales from tiny to a million data points, text in the tables, high cardinality features, and so on. And as a result, TabPFN now is also much more relevant to data science competitions.

TabPFN 3.5在Kaggle Auto挑战赛 TabPFN 3.5 on the Kaggle Auto Challenge

Frank

举个例子,我们的研究科学家 Nick Ericson,他之前在 AWS 构建了 AutoGluon,他拿 TabPFN 3.5 去试了 Kaggle 2015 年的历史性 Auto Challenge。那场比赛有 3,500 名参赛者,奖金池 1 万美元,最终夺冠的是一个由 36 个不同模型组成的多层堆叠集成,还配了手工特征工程。Nick 这些年一直用这个比赛来改进 AutoGluon,而 AutoGluon 那篇论文其实已经在 96 个 CPU 上、24 小时内拿到了前 1% 的方案。现在他用 TabPFN 3.5,一行代码、在一块 RTX Pro 6000 GPU 上跑 1 分钟算力,就拿到了排名第一的方案。所以我现在对 TabPFN 3.5 作为数据科学家的实用工具感到非常兴奋。

For example, our research scientist Nick Ericson, who previously built AutoGluon at AWS, tried TabPFN 3.5 on the historical Auto Challenge from Kaggle from 2015, which had 3,500 competitors and a prize pool of $10,000, and was won by this multi-layer stack ensemble of 36 different models with handcrafted feature engineering. Nick had used this competition to improve AutoGluon over the years, and the AutoGluon paper had actually gotten a top 1% solution in 24 hours on 96 CPUs. Now he used TabPFN 3.5 and got the number one ranked solution in one line of code and 1 minute of compute on one RTX Pro 6000 GPU. So I'm super excited about TabPFN 3.5 as a practical tool for data scientists now.

Host

就是这样。以上就是 Frank 关于 TabPFN 3.5 的更新。我觉得能活在这个时代真的很令人兴奋,因为你可以直接拿这个基础模型去解决那么多现实世界的问题。我甚至已经把它用到了我们 Discord 服务器的垃圾信息过滤器上,这简直太棒了。当然,用起来需要一台相当强劲的机器,而且机器得有相当大的内存。但它真的、真的非常酷。所以希望大家也去玩玩看。我记得它对非商业用途是免费的。总之,希望你们喜欢这期节目。我们下期再见。

So there you go. That was the update from Frank on TabPFN 3.5. I think it's a genuinely exciting time to be alive because you can just take this foundation model and apply it to so many real-world problems. I've even applied it to the spam filter on our Discord server now, which is absolutely amazing. Admittedly, you need a fairly beefy machine to use it, and you need quite a lot of memory on the machine. But it really, really is cool. So I hope you folks have a play around with it. I think it's free for non-commercial use as well. Anyway, hope you enjoyed the show. See you on the next one.

互动版:逐字朗读 + 针对本期提问 →