HuggingFace 进军机器人领域:为物理 AI 构建开源社区

HuggingFace's Leap into Robotics: Building an Open Source Community for Physical AI

托马斯·沃尔夫 Thomas Wolf · Training Data · 2025-05-01 · 约 43 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

HuggingFace 联合创始人 Thomas Wolf 讨论公司新的机器人项目 LeRobot,以及他为何认为机器人技术正处于与几年前 Transformer 和语言模型相似的转折点。

Thomas Wolf, co-founder of HuggingFace, discusses the company's new robotics initiative, LeRobot, and why he believes robotics is at a similar inflection point to transformers and language models a few years ago.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 24)

全文 · Full transcript(中英对照)

Hugging Face 机器人入门 Introduction to Robotics at Hugging Face

Host

很多初创公司已经在机器人之上构建产品,他们想做一些东西,比如自动化手动测试,或者在物理世界中做点什么。他们用我们发布的机器人,一个非常简单的机械臂 S100,我们设计它就是为了成为最便宜的机械臂,只要 100 美元。他们已经在尝试围绕这个开展业务。这就是 Hugging Face 的理念:提供平台和基础模块,让人们在上面构建疯狂的东西。机器人领域也是同样的目标。本期节目,我们采访了 Hugging Face 的联合创始人兼首席科学官 Thomas Wolf。Hugging Face 是最大的 AI 开源社区。Thomas 一直能预测未来,正是因为他,Hugging Face 几年前大力投资 Transformer 和语言模型,推动了 LLM 浪潮。现在,他看到了机器人和物理 AI 的同样机遇,并正在引领 Hugging Face 的 LeRobot 项目,该项目整合了策略模型、数据集和物理实体,帮助开发者成为机器人专家,覆盖各种用例和形态。我们很高兴与 Thomas 聊聊机器人领域的爆发、物理 AI 的瓶颈、中美开源模型竞赛等话题。欢迎收听。Thomas,作为 Hugging Face 的首席科学官,你帮助公司提前一两年投资“登月计划”。上次我们聊时,你说你有种直觉,觉得机器人领域现在正处在几年前 Transformer 和语言模型的那个时刻。能跟我们说说你看到了什么吗?

Many many startups just already being built on top of the robot just you know they want to build something they have this idea of of a manual test they can automate or they have an idea of something they could do in the physical world and then they take the robot they take already like the basic building blocks we've shipped which is just a robotic a very simple robotic arm s 100 that we designed basically to be the cheapest robotic arm to be $100 and they're already like trying to start business around this at the bottom you know it's the uging face etos which is you bring all these platform, all these basic building blocks for people to build really crazy things on top. And robotics is the same goal for us. In this episode, we interview Thomas Wolf, co-founder and chief science officer of HuggingFace, the largest open source community for AI. Thomas has consistently predicted the future, and he's responsible for why HuggingFace invested heavily behind transformers and language models enabling the LLM wave a handful of years ago. Now he sees the same opportunity for robotics and physical AI and is helping to shepherd Hugging Face's low robot project which brings together policy models, data sets, and physical embodiment to help developers everywhere become roboticists across a diversity of use cases and form factors. We're excited to chat with Thomas about the robotics explosion, the bottlenecks ahead for physical AI, the US versus China in the open model race, and a lot more. Enjoy the show. Thomas, in your role at Hugging Face, you help Hugging Face invest behind Moonshots as the chief science officer, one or two years ahead of the field. Uh, you mentioned to me last time we chatted that, you know, you have a spidey sense that we're at the same moment for robotics today that we were for transformers and language models a handful of years ago. Tell us what you're seeing.

Thomas Wolf

嗯,说实话,我觉得这始于两年前。我们是在 18 个月前开始机器人相关活动的。当时,斯坦福等实验室取得了一些突破,他们展示的机器人能打结、叠衣服、做饭,比如把东西抛到空中再用锅接住。所有这些都用了很少的数据,但前景很好,可以借鉴世界模型等从互联网规模数据中受益的技术。这一切都指向一个不远的未来,机器人将真正工作。硬件已经存在很久了,但缺失的砖块是能适应、能动态变化的软件。我们开始看到这一点,所以大约 18 个月前启动了 LeRobot 项目。对我们来说,LeRobot 的成功,比如看到这个庞大的社区,我们的大赌注是:能否在机器人领域也建立一个大型社区?以前只有一些小众的爱好者社区,或者为工厂生产线认真造机器人的人,但在我看来那只是一个小垂直领域。问题是,能否把这个小垂直领域变成一个全面的横向领域?就像现在每个软件开发者几乎都成了 AI 研究者,他们都想了解 LLM 如何工作、如何训练。这是一个平滑的转变,1 到 2 亿软件开发者正在变得有 AI 意识。我认为未来可能有一个类似的转变,如果给这些人工具,他们也会成为机器人专家。这就是我们的目标。所以我们从软件库开始,它的成功促使我们也尝试硬件,这是一个大问题。今年我们收购了第一家硬件公司 Poland Robotics,并在 7 月中旬开放了第一批机器人的订单,那是一个月前的事,反响非常热烈。

Yeah, I think I think honestly it started two years ago. I mean, we started we started our activities in robotics 18 months ago. Uh and I think at this time there was a couple of breakthrough out of I mean these labs you know where right Stanford and basically these teams were starting to show robots that were able to tie notes, fold clothes uh you know cook like threw things in the air on a pan and grab them. And basically all of these thing uh with in in one way very little data but also with also good perspective to be able to leverage some of the world models and thing like that that we see you know really benefit from from internet size data. So all of this kind of pointed to uh a short future where robotics were going to work in a way like hardware was already there and in my opinion has been there for quite some time but the missing brick was really software that could adapt that could be dynamic all of that and we started to see that and that's why we started to work on the robot yeah a little bit more than than 18 months ago and I think for us the success of the robot like seeing like this huge community for us the big bet was can you build a big community in robotics as well. There were small community of kind of hobbies like or people very seriously building robots for like you know factory lines but it was kind of a like a tiny vertical in my opinion. The question was could you move this tiny vertical to like a full horizontal thing just like nowadays every software developer is kind of an AI researcher almost they all want to know how you know LLM works how you train them and there's this very smooth transition where the 100 200 million software developer are becoming kind of AI AI aware and I think there is a potential future transition where all of these people also become roboticist in a way if you give them the tools so that was kind of our goal so we started with with the software library and the success of that kind of brought us to also try to go in hardware which is which is a big question as well and that's why we acquired first hardware company this year Poland Robotics and and we shipped our at least we opened the orders for our first robots that was mid July so that's one month ago really and this went crazy

Host

能跟我们说说 LeRobot 是什么吗?

Can you tell us what LeRobot is?

Thomas Wolf

当然。LeRobot 是我们试图在机器人领域复制 Transformers 库成功的尝试。这个想法是拥有一个核心库,每个人都能用,以非常简单易用的方式提供所有最新技术、高效训练机器人的算法、最新数据集,并连接到执行器,即机器人的硬件部分。这是三个方面的交叉:策略模型、数据集和硬件,我们试图在 LeRobot 中把它们结合起来。

Of course. So LeRobot is our attempt to reproduce the success of the Transformers library in the robotics field. The idea is to have a central library that everyone would use, bringing in a very simple and accessible way all the latest technology, algorithms to train robots efficiently, the latest datasets, and also connecting this to actuators, the hardware part of robotics. It's the intersection of three aspects: the policy models, the datasets, and the hardware, which we try to combine in LeRobot.

Host

Hugging Face 在机器人领域的角色有何变化?对于在物理世界构建的人来说,Hugging Face 扮演的角色与在数字世界用 LLM 构建时相同还是不同?

How does the role of Hugging Face change in robotics? For people building in the physical world, does Hugging Face play the same role or a different role than it has played for people building in the digital world with LLMs?

Thomas Wolf

我们的目标是扮演同样的角色,即建立社区,让人们接受 AI 可以开源,可以调整、训练、控制,并托管在任何地方。在机器人领域,托管在何处甚至更重要,因为未来机器人无处不在,你希望很多模型在本地运行。如果机器人失去 Wi-Fi 连接,撞到墙或你的孩子,那比 LLM 产生幻觉要严重得多。所以机器人领域的安全问题是一个很好的理由,让你不依赖远程 API,而是让模型尽可能靠近硬件。我认为我们的角色在安全性和机器人未来方面可能比在 LLM 领域更重要。

Our goal is to play the same role, which at a very high level is building communities and bringing people into the idea that AI can be open source, something you can tweak, train, control, and host where you want. Hosting where you want is even more important in robotics because in a future with robots everywhere, you want many models to run locally. If your robot loses Wi-Fi connection and runs into a wall or your kid, it's much more dramatic than an LLM hallucinating. So safety in robotics is a good reason to not depend on a distant API but have models as close to the hardware as possible. I would say our role is maybe even more important for safety and the future of robotics than it is in LLMs.

Host

能说说 LeRobot 社区的规模吗?有多少人在构建?有多少人在贡献数据集等?

Could you say a word on the size of the community you have at LeRobot? How many people are building? How many people are contributing data sets and things like that?

Thomas Wolf

当然。我应该查一下最新数字,因为它在指数级增长。有几千人,大概六千到一万人。几个月前我们举办了一个全球黑客马拉松,在六大洲有 100 个地点。所以虽然还没有达到百万级别,但已经远超几千人了。

For sure. I should have checked the latest number because it's exponentially growing. It's several thousand people. I would say six to ten thousand. One event we did a couple of months ago, a worldwide hackathon, had 100 locations across six continents. So I would say it's still not in the millions, but it's way above several thousand.

社区成长与硬件演进 Community growth and hardware evolution

Thomas Wolf

对我们来说,主要的指标是我们在 Hub 上测量的数据集数量,我们看到社区成员或数据集的数量都在指数级增长,我认为这很好地表明我们走在正确的轨道上。而且你要知道,目前可用的硬件仍然非常像是爱好者的硬件,比如 3D 打印的机械臂,到处都还连着线。所以从今年夏天开始,我们想推出更面向大众市场的硬件,不仅能吸引黑客和习惯到处插线的人,也能吸引家庭中的每个人——外观更精致。

And the main indicator for us is we can measure the number of datasets on the hub, and we see in all these things—number of community members or datasets—exponential growth, which I think is a very good indication that we're on the right track. And you have to keep in mind that the hardware available right now is still very much hobbyist hardware, like 3D-printed arms, still wired everywhere. That's why starting this summer, we wanted to bring much more mass-market hardware, something that would interest not only the hacker and people used to plugging cables everywhere, but also everyone in families—something that looks much more polished.

机器人社区三类开发者 Three types of developers in the robot community

Host

你的机器人社区里的开发者是什么样的人?我很好奇他们与传统构建经典控制系统的开发者有哪些相同或不同之处。

What's the persona of developers in your robot community? And I'm curious how it's the same or different from the people that have traditionally been building classical control-based systems.

Thomas Wolf

我认为有三种类型的人。第一种是传统的机器人专家——他们肯定想用 AI。对他们中的很多人来说,他们知道如何构建硬件,知道能用什么,但一直被软件栈的限制所困扰,所有最优控制模型都非常限制你能做的事情。所以这些人都很高兴地加入了这股潮流,我们看到了与 Transformer 中相同的效应:许多学术实验室开始使用机器人,因为这对学生来说是一个很好的切入点。所以这部分增长非常强劲。第二种社区在我看来更有趣。他们是那些原本不真正涉足机器人领域的人,但因为对 AI 感兴趣,而机器人看起来是 AI 的物理体现,所以他们想进入机器人领域。这些人包括软件开发者,甚至只是对机器人感兴趣的人。一个很好的例子——这很有趣——很多投资者实际上买了 100s 机械臂,只是为了从物理上理解这个机器人是什么以及它能做什么。因为它看起来非常容易上手,你拿到机械臂,软件只是 Python 代码,现在通过一点 vibe coding,你实际上可以很容易地调整或控制它。我们看到有些人可能不是纯粹的技术人员,但想了解机器人领域正在发生什么,他们通过低成本机器人这个切入点进入。所以你可以对机器人进行 vibe coding。

I would say there are three types of personnel. One is the traditional roboticists—they definitely want to use AI. For a lot of them, they know how to build hardware, they know what they can use, but they've been frustrated by the limitations of the software stack, all the optimal control models, and these are very limiting in what you can do. So all these people have really happily joined the bandwagon, and we see the same effect we saw in transformers: many academic labs starting to use robots because it's a very nice entry point for all the students. So this has been growing very strongly. The second community is much more interesting in my opinion. It's people who were not really into robotics, but because they're into AI and robotics looks like a physical manifestation of AI, they kind of want to go into robotics. These people cover software developers, but even people who are just interested in robotics. A good example—it's interesting—a lot of investors have actually bought a 100s arm just to try to understand physically what this robotic thing is and what it can do. Because it seems so accessible, you get the arm and the software is just Python code, and now with a bit of vibe coding, you can actually tweak it or control it quite easily. We see people who maybe are not purely technical but want to understand what's happening in robotics, and they use this entry point which is low-cost robots. So you can vibe code the robot.

氛围编码与可及性 Vibe coding and accessibility

Host

是的,这确实是我们的目标。所以你已经可以做到一点了,但对于新机器人 Rich Mini,我绝对希望它成为最容易使用的方式之一。我希望我的孩子们能够对机器人进行 vibe coding 来设定行为。

Yeah, that's really my goal. Yeah. So you can already do that a little bit, but for the new robots, Rich Mini, I definitely want this to be one of the easiest ways to use. I would love my kids to be able to vibe code behavior on the robots.

机器人市场成熟与 ChatGPT 时刻 Maturity phase of robotics market and ChatGPT moment

Host

你认为整个机器人市场处于什么样的成熟阶段?什么时候我们会在机器人领域迎来 ChatGPT 时刻?

What sort of phase of maturity do you think we're in for the robotics market at large? When will we have a ChatGPT moment in the world of robotics?

Thomas Wolf

是的,这就是我在寻找的东西。有人称之为 iPhone 时刻。也许第一个用例,第一个让所有人或很大一部分人觉得‘我想要一个机器人’的时刻,是在消费领域。我认为企业市场相当复杂。在某些地方,某些行业已经有很多机器人——汽车制造是最好的例子。然后是第二部分,机器人开始进入,这里有很多关于可靠性的挑战:这些机器人能否可靠地部署在零售店并真正有用?但我更感兴趣的是第三部分:娱乐、趣味、演示、教育。在这些领域,我认为‘我需要一个 3000 美元的机器人因为需要可靠性’这个问题不那么紧迫。所以你可以选择一个非常容易获得的机器人——比如 Rich Mini,定价 300 美元,绝对可以冲动购买。你把它作为礼物买下来,不确定它能不能用。但在这个价位,我们想探索的是,在娱乐、趣味、教育、通过物理交互学习 AI 方面是否有很大潜力,而不仅仅是在聊天机器人或屏幕上编码。我认为这完全没有被探索过。过去有过一些尝试——比如 MIT 媒体实验室的 Cynthia 做了 Jibo——但过去它们通常定价偏高,我认为超过一千美元,更重要的是,软件非常有限。你买一个有趣的机器人,但可能只有五到十个行为,一旦你试过所有,就结束了。而这里,Rich Mini 的目标是让它几乎像智能手机一样:它自带一些行为,但正因为你可以调整它,人们可以构建新行为并分享,并接入所有新的 VLN、语音模型、聊天模型,可能性几乎是无限的。这基本上是为重建 iPhone 的 App Store 打开了一扇门。这就是我非常兴奋的地方。老实说,最后这部分仍然是一个很大的赌注,因为那里什么都没有——没有真正的证据。主要的迹象是社区的指数级增长,这使它看起来相当可行。

Yeah, that's the thing I'm looking for. Some call it the iPhone moment. Maybe what will be the first use case, the first moment where everyone, or a large fraction of the population, will think 'I want a robot' in terms of consumer. I think the enterprise market is quite complex. In some places, there are already a lot of robots in certain industries—car manufacturing is the best example. Then there is a second part where there is an entry of robots, and here there are a lot of challenges around reliability: will these robots be reliable enough to be deployed in retail stores and be really useful? But the third part I'm much more interested in is entertainment, fun, demo, education, where I think the question of 'I need a $3,000 robot because I need reliability' is less pressing. So you can take a robot that's really accessible—Rich Mini, for instance, is priced at $300, something that can definitely be an impulsive buy. You buy it as a gift, and you're not sure if it's going to work or not. But for this price, what we want to find is whether there is a lot of potential in entertainment, fun, education, learning AI through physical interaction instead of just coding on a chatbot or on a screen. I think that's something that has not been explored at all. There were a couple of tries—I mean the MIT Media Lab, Cynthia did Jibo, for instance—but in the past, they were usually priced a bit high, I think above a thousand, and more importantly, the software was very limited. You would buy a robot that would be fun, but you had maybe five or ten behaviors, and once you've tried them all, that's it. Here, the goal for Rich Mini is really to make it almost like a smartphone: it comes with a couple of behaviors, but just because you can tweak it and people can build new behaviors and share them, and plug in all the new VLN, speech models, chat models, the possibilities are kind of endless. It's an open door to rebuilding the app store of iPhone, basically. So that's what I'm very excited about. To be honest, this last part is still very much a bet because nothing exists there—there is no real proof. Major signs are all this exponential growth of community, which makes it quite plausible.

Rich Mini:机器狗重生与创业平台 Rich Mini as a reincarnation of robot dogs and a platform for startups

Host

所以你把 Rich Mini 看作是 90 年代机器狗的转世,人们可以真正地玩耍、实验,并在家庭中拥有机器人伴侣。

So you see Rich Mini as the reincarnation of the robot dogs of the 90s, where people can actually play and experiment and have robot companions in households.

Thomas Wolf

公平地说,这是一个很大的赌注,但昨天我在 Tech Barbecue 讨论机器人技术时,有人告诉我,作为投资者,他们看到很多初创公司已经在机器人基础上建立起来。那些想构建东西的人,他们有一个想法,比如自动化一个手动测试,或者在物理世界中做某事。然后他们拿起机器人,拿起我们已经发货的基本构建模块——一个非常简单的机械臂 S100,我们设计它是为了成为最便宜的机械臂,售价 100 美元——他们已经试图围绕它开展业务。Rich Mini 在某种程度上也是为此设计的。

I mean, to be fair, this one is a big bet, but something I was discussing just yesterday on robotics at Tech Barbecue. Someone was telling me, as an investor, how they see so many startups already being built on top of robots. People who want to build something, they have an idea of a manual test they can automate, or an idea of something they could do in the physical world. Then they take the robot, they take the basic building blocks we've shipped—just a very simple robotic arm S100 that we designed to be the cheapest robotic arm at $100—and they're already trying to start businesses around this. Rich Mini is also in a way designed for that.

Hugging Face 在机器人数据中的角色 Hugging Face's role in robotics data

Host

我想谈谈数据作为瓶颈的问题。我认为语言和机器人技术之间的一大区别是,公共互联网上有数万亿的 token 可以用来训练大语言模型。这种动态在机器人领域并不存在,我认为这正是 Hugging Face 在生态系统中可以发挥更有趣作用的地方,尤其是在去中心化的数据集策划和创建方面。谈谈机器人数据集方面正在发生的事情。

I want to talk about data as a bottleneck. I think one of the big differences between language and robotics is you have trillions of tokens out in the public internet to train LLMs. That dynamic doesn't exist in robotics, and I think that's where Hugging Face's role in the ecosystem could be much more interesting in terms of decentralized dataset curation and creation. Talk about what's happening on the dataset side of robotics.

Thomas Wolf

是的,这非常有趣。机器人技术面临几个挑战,主要挑战是数据。就是没有足够的数据。有一些方法可以利用互联网上的视频作为训练数据,但非常有限。在某些方面,我们可以使用模型,但在其他方面,如果你想自动化一个任务,除了记录某人或机器人执行任务之外别无他法。我认为这里有一个可能性和一个限制。主要限制是你可以自己记录很多任务,但通常你非常缺乏多样性。所以你基本上可以训练一个机器人在你的房间里做得很好,当一切看起来都一样时,但一旦你把它放到隔壁房间,墙壁可能是绿色而不是红色,机器人就很难泛化。所以这是主要限制。我们 Hub 的想法是每个人都可以记录数据集,如果我们设法激励他们分享数据,那么我们就可以建立一个非常多颜色、多位置的数据集,这将非常多样化,希望也非常大。所以我会说这是一个长期目标,我们希望这能有所帮助。我们尝试做的另一件更直接的事情是直接与社区的参与者合作。所以我们发布了一些数据集,试图帮助他们发布一些数据集。我认为在机器人技术中,一个很好的方面是很多人最终想卖硬件,所以他们可以负担得起将一些软件作为开源分享,如果它能推动整个领域发展,因为最终这不是他们直接销售的东西。所以我正在试图说服很多机器人公司这样做,令人惊讶的是,很多公司似乎对此感兴趣。

Yeah, it's super interesting. There are a couple of challenges in robotics, and the main challenge is data. There's just not enough data. There are some ways to use video on the internet as training data, but it's very limited. In some ways, we may be able to use models, but in other ways, if you want to automate a task, there's no way around just recording someone or the robot possibly doing the task. I think here there is one possibility and one limitation. The main limitation is you can record a lot of tasks yourself, but usually what you will lack a lot is diversity. So you will basically be able to train a robot to do something very well in your room when everything looks the same, but once you put it in the next room where maybe the walls are green instead of red, the robot has a lot of trouble generalizing. So this is the main limitation. Our idea with the Hub was that everyone could record datasets, and if we managed to incentivize them to share the data, then we could maybe build a very multi-colored, multi-location dataset that would be extremely diverse and hopefully also very big. So I would say that's a long-term goal, and we hope this can help. Another more direct thing we try to do is work directly with the actors of the community. So we release a couple of datasets to try to help them release some datasets. I think in robotics, one nice aspect is that a lot of people in the end want to sell the hardware, so they can afford to share a bit of the software as open source if it brings the whole field forward, because in the end that's not directly what they sell. So that's what I'm trying to convince a lot of robotics companies to do, and surprisingly, a lot of them seem interested.

世界模型及其对机器人的影响 World models and their impact on robotics

Host

你前几天发推文提到了世界模型。我想你和我见过同一位世界模型创始人。世界模型开源领域发生了什么?这对机器人技术的发展有帮助还是没有帮助?我也可以问一下吗?为什么现在会出现世界模型?因为感觉它们最近开始涌现。

You tweeted about world models the other day. I think you and I met the same world model founder. What's happening in the world model open source space, and how does that help or not help what's going to happen in robotics? And can I ask on that too? Is there a 'why now' for world models at the moment? Because it feels like they've started to pop up recently.

Thomas Wolf

有趣的是,感觉有几个团队独立工作了几个月,然后恰好现在发布了。当你和他们交谈时,他们并不是在互相模仿。我想一件事是出现了非常酷、非常好的图像生成,最终理解了如何修复六指问题,基本上为图像获得了更可靠、更连贯的世界模型,这自然被转置到了视频。所以我们也看到了一些非常酷的视频模型,这只是下一步。我与之交谈过的这个领域的许多创始人还说,他们得益于开源视频模型生成和开源图像生成的进步。他们采用这个视频生成模型,进行微调,然后训练它能够对某些输入做出反应,这也是我们在机器人技术中做的事情。这两者之间有很多共同点,而且似乎效果很好。所以你开始拥有这种全新的体验,你实际上拥有一个可控的、照片级真实的电影,并且以连贯的方式对你输入的动作做出反应,无论是移动还是要求它添加一些东西,比如一个骑士城堡汽车驾驶。你会看到这个东西反应得很好,我认为有很多潜在的应用,既在娱乐领域——某种可能全新的娱乐形式,我们从未见过的东西,也许是我们第一次创造一种新的虚拟娱乐形式——也在商业领域,以及如何拥有交互式事物。这些下游应用之一是为机器人生成更多数据。生成数据只有两种方式:一种是在现实世界中记录,我认为这仍然非常有趣;另一种是模拟。令人惊讶的是,在模拟方面,我们最近没有看到太多发展,但也许这是我在相当长一段时间内看到的第一个模拟生成数据的突破。

What is interesting is it feels like a couple of teams have been working on that independently for a few months and they just happen to release this right now. When you talk to all of them, they are not really copying each other. I guess one thing was the advent of really cool, really good image generation and finally understanding how to fix the six-finger issue and basically get a more reliable and coherent world model for images, which naturally was transposed to video. So we see some really cool video models as well, and this is just one next step. A lot of the founders I've talked to in this field also say they are helped by the advance of open source video model generation and open source image generation. They take this video generation model, fine-tune it, and then train it to be able to react to some inputs, which is also what we do in robotics. There are a lot of common points between these two things, and it seems to work quite well. So you start to have this kind of totally new experience where you actually have a film that's controllable, photorealistic, and reacting in a coherent way to the action you input, either just moving around or asking it to add something like a rider castle car driving. You see this thing react very well, and I think there are a lot of potential applications, both in entertainment—some form of entertainment that might be totally new, something we've never seen, maybe the first time we create a new form of virtual entertainment—and also a lot of applications in business and how you can have interactive things. One of these downstream applications is generating more data for robots. There are just two ways to generate data: one is to record it in the real world, which I think is still very interesting, and the other is to simulate it. Surprisingly, on simulation, we have not seen a lot of development recently, but maybe this is the first breakthrough I've seen on simulated generated data in quite some time.

人形机器人及替代形态 Humanoid robots and alternative form factors

Host

我很兴奋地看到 DeepMind 用 Genie 训练他们的具身机器人。非常令人兴奋。人形机器人:你相信人形机器人是终极形态吗?

I was very excited to see some of the even what DeepMind's doing with Genie to train their embodied robots. Super exciting. Humanoids: do you believe in humanoids as the kind of ultimate form factor?

Thomas Wolf

是的,大辩论。可以肯定的是,我现在对尝试其他形态更感兴趣。人形机器人的主要问题,我认为有两个主要问题。第一个是它总是相当昂贵,因为你需要很多电机,而机器人中的所有成本都在执行器上。那总是大约 70% 的价格标签。所以当你有 60 个执行器时,那就是你的账单。所以很难将人形机器人的价格降到汽车以下。我认为汽车的价格已经是一个相当高的要求了。如果你以汽车的价格买东西,你确实期望从中获得很多价值。所以这就是为什么我们正在探索更小的机器人,比如只有一只手臂或一个移动的头,以及这类东西。

Yeah, big debates. What is sure is I'm quite more excited about trying other form factors right now. The main problem with humanoids, I think there are two main problems. The first one is it's always quite expensive just because you need a lot of motors, and all the price in the robots is the actuator. That's always like 70% of the price tag. So when you have 60 actuators, that's just your bill. So it's really hard to drive humanoids below the price of a car. I think the price of a car is still already quite a high requirement. If you buy something at the price of a car, you do expect to have a lot of value out of it. So that's why we're exploring smaller robots, like just one arm or a moving head, and this type of thing.

人形机器人成本与采用 Humanoid robot cost and adoption

Thomas Wolf

有可能在某个时候我们能得到更便宜的人形机器人。KC Scale 试图做单元树,他们一直在努力降价,很多公司也在朝这个目标努力,但要降到 1 万美元以下真的很难。人形机器人的好处当然是,一旦你解决了人形,你就同时解决了很多任务。所以如果你解决了人形,你可以做人类能做的一切,这非常令人兴奋。主要问题是:你需要解决人形吗?就我而言,我更希望看到一系列不同的形态。我也认为其中一些比人形可爱得多。对于社会接受度,我认为人形也对人们要求很高。它直接处于恐怖谷中,看起来很像你,动作也很像你。我曾认为这会是社会接受度的重大限制。但老实说,我见过很多单一功能的机器人,你在某个时候会忽略它们。所以我也更有信心,人们会说,'好吧,也许我们对机器人恐怖谷过于担心了。'也许在某个时候,一旦我们开始看到几个机器人,人们就会非常容易地接受它们。

There is some possibility that we can get cheaper humanoids at some point. KC Scale was trying to do unit trees, they've been trying to cut the price, and there are a lot of companies trying to aim for that, but it's going to be really hard to get it under $10,000. The nice thing about the humanoid, of course, is that once you've solved the humanoid, you solve a lot of tasks at the same time. So if you solve the humanoid, you can do everything a human does, which is very exciting. The main question is: do you need to solve the humanoid? On my side, I would like to see a galaxy of different form factors. I also think some of them are much more cute than the humanoid. For social adoption, I think the humanoid is also asking a lot from people. It's directly in this uncanny valley where it looks a lot like you and moves a lot like you. I thought this would be a big limit for social adoption. But to be honest, I've seen a lot of unitary robots, and you kind of ignore them at some point. So I'm also much more confident that people will just say, 'Yeah, let's just... maybe we're too worried about the uncanny valley in robotics.' Maybe at some point, once we start to have seen a couple of robots, people will just accept them very easily.

Host

好的。那么,我们很快就会看到人形机器人了。

Okay. So, we're going to see the robot humanoid soon.

Thomas Wolf

我的意思是,如果我们的 Chimney Mini 和小型机器人运行得很好,那么在某个时候我们会重新爬升到人形形态。我会说逐步地,就像我们一直做的那样,带着社区一起前进。

I mean, the goal would be if our Chimney Mini and our small robots work really well, that at some point we'll climb back to make the humanoid form factor. I would say progressively, as we've done, bringing the community along with us.

Host

当你想象 10 年后的世界,你认为我们中间有多少机器人?你认为其中 80% 是人形,20% 是硬件和用例的漫长尾部,还是你认为世界会如何发展?

As you imagine the world in 10 years, how many robots do you think there are among us? Do you think 80% of them are humanoids and 20% are this long tail of diversity of hardware and use cases, or how do you think the world plays out?

Thomas Wolf

是的。我很乐意看到第二种选择,因为我认为那是一个我们生活中有更多机器人的选项。我真正不太兴奋的未来是机器人成为一种精英事物,因为它们要 10 万美元。所以基本上,如果你有钱,你家里有三个机器人;如果你没有,就没有。Hugging Face 一直关注大社区。所以我们关心这一点。出于这个原因,我更兴奋看到许多基本上对很多人可及的形态。其中一些更便宜,一些比这个昂贵的人形更贵。如果你能买,那很好;如果你不能,那太糟了。所以在 Hugging Face,这是我们试图推动的未来。我认为这也更有趣,因为在某种程度上,你也在限制自己。就像大语言模型,如果你只是试图让它们模仿人类,那是一回事。但如果你试着思考也许它们能做人类不能做的事情,那也有趣得多。

Yeah. I would love to see the second option because I think that's an option where we have much more robots in our life. What I would really not be super excited about is a future where robots are kind of an elite thing because they cost $100,000. So basically, if you're rich, you have three robots at home, and if you're not, you don't. Hugging Face has always been about the big community. So we care about that. For this reason, I'm much more excited to see a lot of form factors that are basically accessible to a lot of people. Some of them are cheaper, some are more expensive than this single humanoid that costs a lot. And if you can buy it, that's nice, and if you cannot, too bad for you. So at Hugging Face, that's the future we try to nudge and push toward. I think it's also much more fun because in a way, you're also restricting yourself. Just like LLMs, if you just try to make them copy humans, it's one thing. But if you try to think maybe they can do something that humans cannot do, it's also much more interesting.

基础模型 vs 专用模型 Foundation models vs specialized models

Host

你认为我们正在走向一个由大型基础模型组成的世界,这些模型几乎可以做任何事情,然后只需几个提示就能快速适应任何新领域?还是你认为你社区中的开发者会从一个小型基础模型开始,然后进行大量自己的数据收集和定制来适应他们的领域?

Do you think we're heading towards a world of big foundation models that can kind of do everything and then be adapted quickly to any new domain with just a few prompts? Or do you think that developers in your community are going to start from a small base model and then do a lot of their own data collection and customization to adapt to their domains?

Thomas Wolf

我认为我们会越来越多地看到两者。随着领域的发展,我们开始看到一个很长的尾部。例如,如果我们看 Hugging Face 上的下载量,我们看到非常大的最先进模型被下载,这些模型通常太大,无法在本地笔记本电脑上运行。但我们也看到一些下载量最大的模型实际上大小正好适合在笔记本电脑上快速运行。所以我们看到这两种模式。随着领域的成熟,我们会越来越多地看到这一点。这不是你选择其中一个或另一个的问题;而是取决于你需要什么,你可能会在本地使用一个,也可能不会。我认为带有路由器的 GPT-5 是一个很好的例子。也许最大的模型或最长的推理链并不是万能的。你实际上需要智能地选择你想要的模型。所以它可以放在路由器后面,但也可以只是在本地运行一些可能非常有用的模型。我们越来越知道如何训练实际上非常有用的模型。但是当你需要更复杂的东西,当你需要长时间反思时,那么你会转向更大的模型。

I think we'll see more and more of both. As the field evolves, we start to see a really long tail. For instance, if we take the downloads on Hugging Face, we see both very large state-of-the-art models being downloaded, which are usually too large to run on a local laptop. But we also see some of the most downloaded models are actually just the right size to run quickly on a laptop. So we see these two modalities. And as the field matures, we'll start to see this more and more. It's not like you choose one or the other; it's just depending on what you need, you might use one locally or not. And I think GPT-5 with the router is a good example of this. Maybe the largest model or the most reasoning, the longest reasoning chain, is not the answer to everything. You actually need to smartly select the one you want. So it can be behind a router, but it can also be just locally you run some models that might be extremely useful. And we know better and better how to train models that are actually extremely useful. But when you need something much more complex, when you need reflection for a very long time, then you will turn to much larger models.

开放与封闭模型及 OpenAI 在 HF Open vs closed models and OpenAI on Hugging Face

Host

过去几年非常流行的一个叙事是开放与封闭之间的战斗。封闭模型与开放模型,谁会赢?就在最近几周,OpenAI 现在出现在 Hugging Face 上。所以我很好奇这怎么理解,以及这可能意味着开放与封闭的未来,或者它们如何合作。

One of the narratives that's been really popular over the last few years is this narrative of the battle between open and closed. Closed models versus open models, who's going to win? And just in the last few weeks, OpenAI is now present on Hugging Face. So I'm curious what to make of that and what it might imply about the future of open versus closed, or maybe how they work together.

Thomas Wolf

我们非常高兴欢迎他们回来。他们曾经在那里。我参与的第一个模型,也是我们从游戏公司转型为开源平台的原因,是 GPT-1。没有多少人记得,但很有趣,因为它主要是在小说和言情小说上训练的。所以当你把两个角色放入续写中,他们总是会相爱。在某种程度上,我还有点怀念这个。然后谷歌采纳了这个想法,也在维基百科上训练,维基百科有很多世界知识,然后扩展到 GPT-2、GPT-3 等等。但在那个时候,他们非常支持开源。我认为开源,就像在软件中一样,两种解决方案共存。有些公司同时开源两者,或者两者都做,我的意思是谷歌已经是一个例子,有 Gemma 系列和 Gemini 系列。有一些有趣的时刻,我听说一个 Gemma 模型实际上非常好,以至于比闭源模型还好,所以他们不得不不开源它。所以前沿在某种程度上目前相当封闭。而挑战者主要是中国的,但我认为我们也会在美国看到一些新的基础模型团队。所以我认为我们可能会在美国看到一些挑战。前沿会保持相当封闭,我认为,两者在性能上会有微小差异。主要原因,老实说,在这个时间点,我认为我们并不完全处于 AI 的成本节约时期。

We're super happy to welcome them back. They were there. The first model I worked on, and the reason we switched from being a game company to an open source platform, was GPT-1. Not a lot of people remember, but it was very funny because it was trained mostly on novels and romance novels. So when you would put two characters in the continuation, they would always fall in love. In some way, I still miss this one a little bit. Then Google took this idea and trained it also on Wikipedia, which had a lot of world knowledge, and then it expanded to GPT-2, GPT-3, all of that. But at that time, they were very pro open source. I think open source, just like in software, both solutions just coexist. Having companies that open source both, or that do both, I mean Google has been an example for quite some time with the Gemma line and Gemini line. There were some interesting moments where I heard that one Gemma model was actually so good that it was better than closed source models, so they had to not open source it. So the frontier is quite closed at the moment in a way. And challenging new players, mostly in China, but I think we'll start to see some new foundation model teams in the US as well. So I think we might see some challenge in the US. The frontier will stay quite closed, I think, and both will see a tiny difference in performance. The main reason right now, to be honest, at this exact point in time, I think we're not exactly in a kind of cost-saving time of AI.

开源采纳的动机 Motivations for Open Source Adoption

Thomas Wolf

所以这意味着,对很多参与者来说,我认为转向开源以节省成本并不是最重要的。他们现在转向开源通常是因为想要数据隐私,想要能够适配模型,或者他们可能有一个新想法,比如一个新的行动模型,他们想做些不存在的东西。这就是我们目前看到的。所以我们看到很多新的探索以开源方式涌现。我预期的是,随着市场走向成熟,成本、在更快或不同硬件上运行的能力、拥有模型以及拥有模型运行的整个技术栈,会变得越来越重要。所以我认为,就像软件领域一样,长期来看开源是许多应用和用例的制胜方案,但我们仍处于动荡期。

So which means that for a lot of actors, I think moving to open source because it saves cost is not the most important thing for them. So usually they move to open source right now because they want data privacy. They want to be able to adapt the model. They have maybe a new idea or a new action model, for instance. They have a new idea of something that does not exist and they want to do that. So that's usually what we see right now. So we see a lot of emergence of new exploration in an open source way. What I do expect is as we go to a more mature market, then the cost and being able to run it on faster hardware or other hardware, and being able to own the model and also own the full stack of where the models run, becomes actually more and more important. So I think just like in software, I think in the long term, open source is kind of a winning solution for many applications, for many use cases, but we're still in the turbulence phase.

Host

是的。

Yeah.

Hugging Face 角色演变 Hugging Face's Evolving Role

Host

你认为随着这些模型推动前沿并出现闭源模型,Hugging Face 在 LLM 生态系统中的角色是如何演变的?我记得以前你可以在 Hugging Face 上下载小的 BERT 模型并在本地运行,对吧?那是很多的使用场景。现在模型变得太大,无法在消费级硬件上运行,你的业务是如何演变的?你如何看待 Hugging Face 角色的演变?

How do you think Hugging Face's role in the LLM ecosystem has evolved as these models have pushed the frontier and there are closed models? I remember back when you could download the small BERT model on Hugging Face and run it locally, right? That was a lot of the usage. How has your business evolved now that we're going towards models too large to run on consumer hardware, and how do you see Hugging Face's role evolving?

Thomas Wolf

令人惊讶的是,我去年年底做了统计,BERT 模型仍然被大量使用。所以开源一个令人惊讶的有趣方面是韧性:一旦你有了在生产中运行良好的东西,你可能不想被迫迁移到新的 GPT,对吧?我的意思是,这是 GPT-5 周围的一些反弹:人们实际上想继续在许多事情上使用 GPT-4。也许他们爱上了它,或者它是他们的主要朋友,就像一些 Reddit 帖子所说的那样,但也可能只是他们的应用程序运行得很好,他们不想重新设计。我认为开源,对我们的长期利益来说,也是提供这个非常稳定的基础:你构建了一些东西,你知道它会存在,你可以把它作为一个非常稳定的基础。总的来说,我认为在社区中,我们的角色已经逐渐从自己推动很多东西、推动我们的库、推动我们的早期产品,转变为更多地赋能整个社区。所以我们现在与社区中的许多参与者合作:我们与 llama.cpp 合作很多,与 vLLM 合作很多,与所有大玩家合作,试图看看整个生态系统如何能高效运作。比如,一个模型发布后,你希望它能直接在 vLLM 中使用,直接在 llama.cpp 中使用。所以我们越来越多地扮演这种元社区建设者的角色,试图协调所有参与者,让他们步调一致,以相同的方式前进。所以从某种意义上说,我们比几年前更关注社区和 Hub。

I mean, surprisingly, I was doing these stats at the end of last year, and the BERT model is still really used a lot. So a surprisingly interesting aspect of open source is also resiliency: once you have something that works really well in production, you may not want to be forced to move to the new GPT, right? I mean, that was a little bit of the backlash around GPT-5, you know, people actually wanted to keep using GPT-4 for many things. Maybe they fell in love with it, or it was their main friend, as some Reddit posts were around this, but also maybe they just had their application which worked really well and they don't want to redesign it. I think open source, the long-term interest for us is also to provide this very stable base: you build something, you know it will exist, and you can keep it as a very stable base. In general, I think in the community, our role has switched progressively from maybe pushing a lot of things ourselves, pushing our library, pushing our early product, to more enabling the community in general. So we now work a lot with many actors in the community: we work a lot with llama.cpp, we work a lot with vLLM, we work a lot with all the big players to try to see how this whole ecosystem can be very efficient and work really well. So like, one model is released, you want to be able to use it directly in vLLM, you want to be able to use it directly in llama.cpp. So we try to have more and more this kind of role of meta-community builder, where we try to align and bring all the players at the same pace and help them move in the same way. So in a way, we are much more focused on the community and the hub than we were maybe a couple of years ago.

中国开源领导力 China's Open Source Leadership

Host

太酷了。你怎么看正在发生的事情?你提到中国最近有很多开源模型。你认为为什么会这样?西方的开源模型发展状况如何?

Really cool. What do you think about what's happening? You mentioned China has had a lot of open models recently. Why do you think that is happening? And what is the state of open model development in the West?

Thomas Wolf

是的,这是过去两年最令人惊讶的事情,对吧?中国成为开源的冠军。谁会在 2020 年代预测到呢?我两周前刚去拜访了他们,试图更好地了解当地的情况。事实是,这是一个内部竞争非常激烈的市场。那里有很多非常优秀的团队,在某种程度上让我想起了硅谷。人们工作非常努力,这些模型提供商之间相互竞争。而他们竞争的一个方面,令人惊讶的是,就是最开放:开源方面。所以他们为非常开放而自豪。有些公司,当他们不再开放时——比如一家叫智谱的公司——他们决定不开源,然后立即看到了反弹,我认为主要是招聘方面:人们不想再去那里工作了。所以他们又回到了开源。所以现在这很强大,我认为是智力上的强大。所以我预计这将继续。我也预计会有更多团队加入,因为我看到很多——我的意思是,我们也看到了,对吧?当 GPT-5 发布时,很多人实际上做了研究,其中一些在清华大学,对吧?我们知道团队也在那里,部分成员是中国人。所以他们有非常强大的人,他们都想训练最好的模型。我认为有趣的是,西方最近也开始回归开源,老实说,就在今年夏天,对吧?但这次开源的呼声:OpenAI 决定回归。现在,我们只是在等待 Anthropic 可能开源他们的第一个模型。所以我认为是时候尝试让他们参与了。是的,我会说现在开源的情况相当不错,但从来不是——就像《星球大战》中的绝地武士。从未胜利。我们必须继续推动这一点。我们必须继续高举开放的旗帜。

Yeah, this is the most surprising thing that happened, I think, in the last two years, right? The fact that China would become a champion of open source. Who would have predicted that in the 2020s, right? I've been visiting them two weeks ago to try to understand a bit better on the ground how it's happening. And the thing is, it's a very competitive market internally. There are a lot of teams there that are extremely good, and it reminded me in some way of Silicon Valley. People are working extremely hard and they compete with each other, all these model providers. And one part on which they compete, which is surprising, is being the most open: the open source aspect. So they're extremely proud of being very open. And some of these companies, when they stopped being open—like one called Zhipu—they decided not to open source it, and they saw an immediate backlash, I think mostly on hiring: people didn't want to come work there anymore. So they went back to open sourcing. So it's quite strong now, I would say strong in brain. So I would expect this to continue. I would expect also quite more teams to come, because I see a lot of—I mean, we see that as well, right? When you have the presentation of GPT-5, a lot of people actually did the study, some of them at Tsinghua University, right? We know the teams are there also, partly with Chinese members. So they have extremely strong people, and they all really want to train the best model. What I think is interesting is to see the West kind of coming back to open source very recently, to be honest, just over the summer, right? But this call for open sourcing: OpenAI decided to come back. Now, we're just waiting for Anthropic to maybe open source their first model. So I think it's time to try to ask them to participate. Yeah, I would say right now the situation for open source is pretty good, but it's never like—it's like the Jedi and Star Wars. It's never won. We have to keep pushing this. We have to keep pushing our flag of openness.

西方开源复兴 Resurgence of Open Source in the West

Host

是什么推动了西方开源的复兴?

What's driving the resurgence of open source in the West?

Thomas Wolf

是的,我认为一件事是,当你没有什么可失去的时候,开源总是你团队的一个好解决方案。例如,你创建了一家新公司,想迅速崛起:你开源你的模型,对吧?这就是 Mistral 的秘诀:如何迅速成为大玩家。但对于中国公司来说,在西方几乎没有人会使用中国的 API,所以他们反正也不在西方销售 API。所以从某种意义上说,他们通过开源模型在西方市场没有什么可失去的。所以我认为存在这样一种情况:作为参与者,其后果是,当没有人开源时,就像一个市场。有人有兴趣占据这个空间,对吧?说我们要成为开源玩家。所以当其他人都停止开源时,Meta 就成了这个开源玩家。

Yeah, I think one thing is, when you have nothing to lose, open source is always a good solution when you're on your team. So it can be, for instance, you create a new company and you want to quickly rise to the top: you open source your model, right? That's the Mistral recipe: how you can very quickly become a great player. But for the Chinese, for instance, it's also—in the West, almost nobody will use a Chinese API, so they don't sell APIs in the West anyway. So in a way, they have nothing to lose from the Western market by open sourcing their model. So I think there is this thing: as players, the consequence of that is also that when nobody is open source, it's like a market. There is an interest for someone to take the room, right? To say we're going to be the open source player. So Meta was this open source player when everyone kind of stopped open sourcing.

开放模型与商业信任 Open models and business trust

Host

Thomas,你提到西方公司不会用中国模型而非中国 API。那开源模型呢?当权重托管在美国服务器上时,西方公司愿意使用中国的开源模型吗?是否仍有犹豫,这种犹豫有道理吗?

Thomas, you mentioned that Western companies won't use a Chinese model over a Chinese API. What about open models? Are Western companies willing to use Chinese open models when the weights are hosted on US servers? Is there still hesitation, and is it well-founded?

Thomas Wolf

说实话,我没看到很多这种情况。这是个好问题。我经常做调查,问人们的看法,因为这确实可能是个担忧。比如 DeepSeek 出来时,Perplexity 就有一个不错的无审查模型。但在很多商业场景中,我觉得人们其实没太注意到区别。所以我认为大家普遍希望有更好的方式来理解模型的安全性。人们有点担心模型在某些情况下行为异常。很多公司问:你能保证这个模型始终表现良好吗?我们知道这很难——即使是 GPT,有时你问它草莓里有多少个 R,它也会答错。所以这是一个普遍需求,有几个团队正在研究。

I don't see that a lot, to be honest. It's a good question. I try to do regular polls and ask people what they think, because it can be a concern. When DeepSeek came out, there was a nice uncensored model from Perplexity, for instance. But in many business cases, I don't think people really notice anything. So I think there's a general appetite for a better way to understand the safety of a model. People are a bit worried about a model that might behave strangely in some cases. Many companies ask: can you guarantee this model will always behave well? We know that's really hard—even with GPT, sometimes you ask the number of R in strawberry and it behaves badly. So this is a general need, and several teams are working on it.

AI for Science 与超人类能力 AI for science and superhuman capabilities

Host

我们能谈谈开放科学吗?

Can we talk about open science?

Thomas Wolf

我们像人类一样构建大语言模型。但如果 AI 模型能看到红外线或辐射,这些人类做不到的事呢?那已经是超人类了。对科学来说,这非常有趣。许多科学 AI 模型在某种程度上已经是超人类的,因为它们能感知人类无法触及的模态或预测事物。这是一个跳出人类局限思考的好基础。

We build LLMs like humans. But what if an AI model could see infrared or radiation, things humans cannot? That's already superhuman. For science, it's super interesting. Many AI models for science are already superhuman in a way because they can see modalities or predict things inaccessible to humans. It's a good ground to think outside human limitations.

开放科学热情 Passion for open science

Host

你一直对开放科学充满热情。能谈谈什么是开放科学,Hugging Face 扮演什么角色,以及你的热情从何而来吗?

You've been passionate about open science for a while. Can you say a word about what open science is, what role Hugging Face plays, and where your passion comes from?

Thomas Wolf

这要从很久以前说起。在我当律师之前,我是一名物理学研究员,研究超导材料。令人惊讶的是,超导领域的许多伟大研究是由苏联人完成的。他们发明理论的方式与西方截然不同,有很多绝妙的想法,但我得在苏联的《实验与理论物理期刊》中追踪它们,有些还是俄文的。从那时起,我意识到获取知识很难,如果我能让它变得更容易,就能解锁很多酷东西。当我进入计算机科学领域时,我发现了 arXiv 和开源——一切都免费、共享、用英文写,每个人都能读。我很兴奋,直到我试图复现一篇 DeepMind 论文,发现了局限:人们只发表他们想发表的,但不给所有诀窍。所以对我来说,开放科学是一种延伸:提供开放模型很好,但更好的是解释如何训练模型。授人以鱼不如授人以渔。长期来看,AI 是一项基础技术,应该像物理学一样——每个人都能通过书本学习。训练智能体的配方应该人人皆知。短期来看,如果我们教人们如何训练好模型,他们就会把好模型带到 Hub 上,我们就有更多内容。例如,我们写很长的博客文章,有些成了书。今年夏天我们出版了一本关于在千块 GPU 上训练、负载均衡和并行化的书。另一篇长文是关于制作高质量数据集——我们制作了 FineWeb,被 Qwen 等模型使用。我们还写了如何构建它、如何过滤、什么重要。所有这些都为 Hugging Face 带来了更好的开源 AI 模型。

It started a long time ago. Before I was a lawyer, I was a researcher in physics, working on superconductive materials. Surprisingly, a lot of great research on superconductivity was done by the Soviets. They had very different ways of inventing theory, with great ideas, but I had to track them down in Soviet JETP letters, some still in Russian. From that time, I realized accessing knowledge is hard, and if I could make it easier, it would unlock a lot of cool stuff. When I joined computer science, I discovered arXiv and open source—everything free, shared, in English, everyone can read it. I was excited until I tried to reproduce a DeepMind paper and found limits: people publish what they want but don't give all the tricks. So open science for me is an extension: it's nice to give open models, but even better to explain how to train a model. Give a fish vs. teach to fish. In the long term, AI is a fundamental technology that should be like physics—everyone can learn from a book. The recipe to train an intelligent artifact should be known by everyone. In the short term, if we teach people how to train great models, they bring great models to the hub, and we have more content. For example, we write long blog posts, some become books. This summer we published a book on training on a thousand GPUs, load balancing, parallelism. Another long post was on making high-quality datasets—we made FineWeb, used by Qwen models. We also wrote how we built it, how to filter, what's important. All this brings better open-source AI models to Hugging Face.

AI 对科学发现的影响与开源角色 AI's impact on scientific discovery and open source's role

Host

我想回到你关于物理和超导的评论。很多 AGI 实验室认为 AI 颠覆科学并不遥远。在数学、物理、材料科学方面已有令人兴奋的发现。你认为我们会看到这些模型带来科学发现的转折点吗?开源将扮演什么角色?

I want to go back to your physics and superconductivity comments. Many AGI labs believe AI disrupting science is not far off. There have been exciting discoveries in math, physics, material science. Do you think we'll see an inflection point in scientific discovery from these models, and what role will open source play?

Thomas Wolf

一如既往,有些炒作会推动人们,但有时我们高估了正在发生的事情。数学是一个很好的例子:有人认为 AI 正在为定理做新证明,发明新科学。但我们需要谨慎。开源将至关重要,因为它使这些工具和方法的获取民主化,让更多研究人员在此基础上构建并加速发现。

As always, there's some hype that drives people, but sometimes we overestimate what's happening. Math is a good example: there's this idea that AI is doing new proofs for theorems, inventing new science. But we need to be careful. Open source will be crucial because it democratizes access to these tools and methods, allowing more researchers to build on them and accelerate discovery.

科学中提出正确问题 Asking the right question in science

Thomas Wolf

我本人就是科学家,所以我认为这种看法是错误的。原因是我曾是个糟糕的科学家,我可以讲讲这个。我过去是个非常好的学生。你给我一个问题,我总能找到证明。我知道这个问题有解,所以我只需要填补空白,把我已知的一些东西组合起来。但当我成为研究者后,我发现我是个很糟糕的研究者,因为我不会提出正确的问题。如果有人让我证明一个定理,我能做到;但如果有人问‘数学里现在有什么值得探索的?’,我就完全没主意了。在科学中,要取得重大突破,关键是要提出正确的问题,找到一个没人问过的问题,一个能开启全新研究领域的问题。诺贝尔奖得主通常就是开创了一个新领域的人,因为他们问对了问题。比如,也许光速应该是常数,我们来探索这意味着什么,结果就引出了广义相对论和黑洞。我认为现在的 LLM 在这方面仍然非常糟糕,不擅长这种有品味地提出正确问题的方式。这并不意味着我们不能用它们做很酷的事情,但我现在更多把它们看作非常有用的助手。一旦人类研究者说‘这个值得研究’,你就可以用它们把预测能力放大 10 倍、100 倍甚至 1000 倍,快速完成对这个分子或蛋白质已有研究的全面综述,或者找出检验某个假设最合乎逻辑的方法。但我仍然认为这是科学研究的加速器和助手。我真正想看到的是一个 AI 说:‘嘿,我有个关于如何超光速的想法。’但要做到这一点,你不能直接写出答案,你必须提出正确的问题:我们今天理论中什么需要被挖掘,或者我们应该重新考虑什么,才能发明出像那样具有突破性的东西?

I think as a scientist myself, that's really the wrong way to view science. The reason is I was a bad scientist, so I can tell about that. I was a very good student. When you give me a problem, I'm always pretty sure I can find the proof. I know this thing has a solution, so I just have to fill the gap, grab a couple of things I know, and combine them together. When I became a researcher, I discovered I was a pretty bad researcher because I was not able to ask the right question. If somebody asked me to demonstrate a theorem, I could do it, but if someone said, 'What is interesting to explore now in math?' I had no idea. In science, the main thing you need to do for a big breakthrough is to ask the right question, find a way to ask a question nobody has asked before, a question that will open a whole new field of research. That's typically a Nobel Prize—someone who just opened a new field because they asked the right question. Maybe the speed of light should be constant, and let's explore what that means, and it leads to general relativity and black holes. I think LLMs right now are still extremely bad at this kind of tasteful way to ask the right question. That doesn't mean we cannot do really cool stuff with them, but the way I see them nowadays is more as very useful helpers. Once a human researcher says, 'This is interesting to study,' you can use them to multiply by 10, 100, or a thousand the predictions you can do, quickly do a full survey of what has been done on this molecule or protein, or say what would be the most logical way to test this hypothesis. But I still see this as an accelerator and assistant of scientific research. What I would love to see is an AI that says, 'Hey, I have an idea on how to go faster than light.' But for that, you cannot just write the answer; you have to ask the right question: what should we trench in today's theory, or what should we reconsider to invent something as groundbreaking as that?

AI 中有趣问题与分歧 Interesting questions in AI and disagreement

Host

关于你提到的提出正确问题,你认为现在 AI 领域有哪些有趣的问题,或者有哪些人们没问但应该问的问题?

To your point on asking the right questions, what do you think are the interesting questions in the world of AI right now, or maybe the questions that people are not asking that they should be asking?

Thomas Wolf

我认为这是一个问题,而且它和我们经常讨论的一个现象有关:谄媚,即 AI 模型总是同意你的倾向。我认为一个好的研究者实际上是一个经常不同意别人的人。我以前的教授是诺贝尔奖得主,他在讨论时非常不友好,但我认为这是其中的一部分。你必须非常有主见。所以找到一种方法让模型有更强的观点,或者有品味的观点,我认为对科学来说将是关键。当然,这将基于深度学习和 LLM,但可能涉及其他训练方式,其他思考方式。我认为这是个大问题。有少数人在探索,但不多。

I think this is one question, and it's related to something we talk a lot about: this sycophancy, the tendency of AI models to always agree with you. I think a good researcher is actually a good example of someone who disagrees with a lot of people. My former professor, who was a Nobel Prize winner, was very unfriendly in how he would discuss, but I think that's part of it. You have to be extremely opinionated. So finding a way to push this model to have stronger opinions, or a taste in their opinions, I think for science will be key. And of course, this will be based on deep learning and LLMs, but it may involve other ways to train them, other ways to think about it. I think that's one of the big questions. There are a couple of people exploring that, but not a lot.

Hugging Face 十年角色 Hugging Face's role in 10 years

Host

好的。当你展望 10 年后的世界,Hugging Face 在其中扮演什么角色?你认为你的社区中有多少人在用 LLM 和机器人技术构建东西?我知道 10 年时间跨度很难想象,但你觉得 10 年后的世界会是什么样子?

Okay. When you see the world in 10 years, like what is Hugging Face's role in it? How much of your community do you think is building with LLMs with robotics? I know it's hard to think in 10-year time spans, but what do you think the world looks like in 10 years?

Thomas Wolf

是的,10 年后会非常不同。我希望能看到一个世界,基本上每个人都觉得他们可以用 AI 来构建东西,而不仅仅是消费 AI,他们觉得自己可以成为这件事的参与者。有点像过去我们有很多媒体是为我们生成和创造的,然后我们进入了每个人都能创造媒体的时代。这催生了新一代的 YouTuber 和网红,他们制作了极其有趣的内容。我希望 AI 也能如此:一个像软件开发者社区一样庞大的社区,每个人都能用 AI 创造东西,他们觉得这只是工具箱里的另一个工具。他们可以写代码,也可以训练模型,甚至调整模型。好的一面是,我坚信社区的创造力和自然发明。这是很美的景象。所以 10 年后,我希望人们不只是消费 AI 内容什么都不做,而是真正发挥创造力,用身边的 AI 工具构建很棒的东西。老实说,我认为我们现在就在建设这样的东西。所以我相当乐观。这件事将给整个社会带来很多改变,因为很多工作会变得不同。

Yeah, 10 years is very different. What I would love to see is a world where basically everyone feels like they can build with AI, not just consume AI, but they feel like they can be actors in this thing. A little bit like the difference between when we used to have media generated and created for us, and then we moved to the current era where everyone is able to create media. That created a whole new generation of YouTubers and influencers making extremely interesting content. I would love AI to be the same: a very big community like the software developer community where everyone can create things with AI, and they feel like it's just another tool in their box. They can code stuff, but they can also train a model and adapt it. The nice thing about that is I'm a big believer in the creativity and natural invention of the community. It's something beautiful to witness. So in 10 years, I hope people are not just consuming AI content and doing nothing, but they're actually exerting their creativity to build really nice things with a lot of AI tools around them. To be honest, that's something I think we're building right now. So I'm quite optimistic. This thing is going to change a lot for society in general because a lot of jobs will just be different.

结束语 Closing remarks

Host

这是一个美好的愿景。Thomas,非常感谢你今天参加我们的节目。我们真的很享受这次对话。

It's a beautiful vision. Thomas, thank you so much for joining us today. We really enjoyed this chat.

Thomas Wolf

谢谢。这是我的荣幸。

Thanks. It was a pleasure.

互动版:逐字朗读 + 针对本期提问 →