构建物理智能:从失望到突破

Building Physical Intelligence: From Disappointment to Breakthrough

卡罗尔·豪斯曼 Karol Hausman · The Generalist · 2026-03-17 · 约 74 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Karol Hausman 分享了他从童年对机器人的迷恋到创立 Physical Intelligence 的历程,这家公司正在为现实世界的机器人构建 AI 大脑。

Karol Hausman shares his journey from childhood fascination with robots to founding Physical Intelligence, a company building an AI brain for real-world robots.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 30)

全文 · Full transcript(中英对照)

童年对机器人的痴迷 Childhood Fascination with Robots

Host

今天能进行这场对话,我真的很兴奋。在研究你们用 Physical Intelligence 构建的东西时,自然大部分讨论都围绕技术展开,但我越深入了解你的故事和这家公司,就越觉得你正在构建的东西非常个人化,并且与你哲学等方面的背景有关。所以,我想先从这些开始。你最早对机器人感到兴奋的记忆是什么?

I'm so excited to have this conversation today. In studying what you're building with Physical Intelligence, naturally much of the conversation is around the technology, but the more I looked into your story and the company, it feels like what you're building is really personal and connects to some of your background in philosophy and all these other things. So, I'd love to start, you know, maybe with some of that first. One of your earliest memories of being excited by robots.

Karol Hausman

可能是小时候看电影吧。我是《星球大战》的忠实粉丝。我读了所有的书,看了很多遍电影。我真的很喜欢那个世界里机器人的描绘方式,它们的多样性、不同的功能、与人类互动的方式,还有好机器人和坏机器人。但我不确定能否指出某个具体时刻。我只记得在成长过程中,我对能够建造出如此像人类的东西这一概念着迷。我一直认为,如果你能做到这一点,它就能让我们更好地理解自己。我对那些我认为人们在成长过程中都会着迷的事情感到好奇。你知道,就是那些关于现实本质、大脑如何运作、思考意味着什么的深刻问题。我认为机器人可能会为这些问题提供一些答案。因为阅读其他类型的书籍或学习哲学并没有给出好的答案。

Probably with watching movies as a kid. I think I was a big Star Wars fan. I was reading all the books, watching all the movies multiple times. And I really loved the way that robots were portrayed in that world, the diversity of them, different functionalities, the way they interacted with people, the way that there are good robots and bad robots. But, I don't know if I can pinpoint a specific point. I remember just growing up being fascinated by that concept that you can build something that is so human-like. And I always thought that if you were able to do that, it would allow us to understand ourselves a little bit better. And I was fascinated by the things that I think people are just fascinated as they're growing up. You know, asking deep questions about the nature of reality, how the brain works, what does it mean to think? And I thought that robots probably are going to have some answers to these questions. Because reading other kinds of books or studying philosophy didn't have good answers to these questions.

Karol Hausman

然后,更重要的是,当我开始研究机器人学并学习它时,我的第一印象是非常失望,因为每当你看到一个机器人,基本上感觉就像没有智能机器人存在,也没有人在研究它们。你看到的每一个机器人都是这种预先编程的机器,只是从 A 点走到 B 点,重复执行,没有任何智能。所以,我认为那可能是我最具决定性的时刻,因为那就像一个警钟,你知道,有这么多人对我们从电影和书中看到的未来感到兴奋,但似乎没有人真正在研究它。

Then, more importantly, as I started looking into robotics and started studying it, the first impression I had was that it was very disappointing in that every robot you would see would basically seem like there were no intelligent robots out there and no one was working with them. Every robot you would see would be this pre-programmed machine that just goes from A to B and does it repeatedly and has no intelligence whatsoever. So, I think that was maybe the most defining moment for me when I saw that because that was like a wake-up call of you know, there's so many people excited about this future that we saw in the movies that we read about in the books and it seems like no one was working on it.

Host

你说过,在某种程度上,你问的是每个孩子都会问的关于宇宙本质和人类思维的经典问题。我不知道这是否像你想象的那么普遍。我想可能有些孩子确实对这个神话着迷,但也许不是全部。

You said that, you know, in some ways you were asking the classic questions of every child about the nature of the universe and the human mind. I don't know if that's as universal as you might imagine it to be. I think probably there are some kids who really become fascinated by that myth, but maybe not all.

可乐罐实验:突破 The Coke Can Experiment: A Breakthrough

Karol Hausman

我们做了一个特别的实验:一个可乐罐放在机器人面前,前面还有三张不同名人的照片。我们给构建的模型下达的指令是“把可乐罐放在泰勒·斯威夫特的照片上”。机器人捡起可乐罐,然后慢慢移向泰勒·斯威夫特。这一切都来自互联网数据。那一刻我们恍然大悟,你可以从 LLM 和互联网中引入大量先验知识,并将其与机器人动作连接起来。感觉就像打开了一扇新的大门。也许,只是也许,如果我们做对一切,将其与互联网知识结合,进行规模扩张,完成所有需要完成的步骤,它就可能成功。那时我们清楚地认识到,实现这一目标的方法是创建一个以解决物理智能为唯一目标的组织。

We had this one particular experiment with a Coke can in front of a robot and three pictures of different celebrities in front of it. And the prompt we gave to the model we built was put a Coke can on picture of Taylor Swift. The robot picked it up and then slowly moved it towards Taylor Swift. All from internet data. That was the moment where it clicked for us that you can bring in a lot of prior knowledge from LLMs, from the internet and connect it to robot motions. And it felt like it opened another door. Maybe, just maybe, if we do everything right, if we combine it with internet knowledge, if we scale it up, if we do all of the pieces that need to be done, it might work. And at that point it became clear that the way to accomplish this is to create an organization whose sole purpose is to solve physical intelligence.

机器人悖论与物理智能使命 The Paradox of Robotics and Physical Intelligence's Mission

Host

早在教会机器人可靠地叠毛巾或端咖啡杯之前,我们就已经教会了机器击败世界上最优秀的棋手。这个悖论是机器人行业的核心,也是多年来许多虚假曙光的原因。Karol Hausman 相信他找到了前进的道路。Physical Intelligence 的 CEO 已经筹集了超过十亿美元的风险投资,旨在为现实世界构建一个 AI 大脑。这个大脑能帮助机器人跨不同任务和环境准确泛化。他采取的方法与许多资金最充足的竞争对手截然不同,专注于从实际交互中收集数据,而不是通过模拟。在今天的节目中,Karol 和我讨论了从《网球内心游戏》中学到的经验教训、为什么强化学习正在回归,以及创造真正物理智能的挑战。我是 Mario,这里是 The Generalist。

We taught machines to beat the world's best chess players long before we taught them to reliably fold a towel or carry a coffee cup. That paradox is at the heart of the robotics industry, responsible for many of its false dawns over the years. Karol Hausman believes that he's found the way forward. The CEO of Physical Intelligence has raised more than a billion dollars in venture funding to build an AI brain for the real world. One that helps robots generalize accurately across different tasks and environments. He's doing so by taking a very different approach from many of his best funded competitors, focusing on gathering data from actual interactions rather than through simulations. In today's episode, Karol and I discuss lessons learned from The Inner Game of Tennis, why reinforcement learning is making a comeback, and the challenges of creating true physical intelligence. I'm Mario and this is The Generalist.

机器人越多,模型越好 The More Robots Deployed, the Better Models Get

Karol Hausman

你部署的机器人越多,模型就应该变得越好。因为我认为模型会不断改进,它们将能够吸收越来越多的数据。

The more robots you deploy, the better the models should get. Because I think the models will keep on getting better. They will be able to absorb more and more data.

早期影响与成长 Early influences and upbringing

Host

你的父母是科学家吗?他们有没有做过什么事情,鼓励了你那种天生的好奇心?

Were your parents scientists? Were there things that they did that sort of encouraged that natural curiosity that you must have had?

Karol Hausman

其实不是。我爸爸是机械师,妈妈是企业家。所以家里并没有太多科学氛围。我觉得主要是看电影让我着迷。我来自波兰的一个小镇,上的高中在当地人才密度很高。我记得那是一个鼓励提出各种问题的环境。我找到了一群同样对这些事情着迷的朋友,这大大激发了我的兴趣。在高中,这就是我们会聊的话题。

Not really, no. My dad was a mechanic and my mom was an entrepreneur. So, not really. There wasn't a lot of science at home. I think it was mostly coming from watching movies and being fascinated by it. I also come from a small town in Poland. I went to high school that had very high talent density for that area. I remember that being an environment where a lot of these questions were being asked and it was very encouraged. So, I found a group of friends that were really fascinated by similar things and I think that's what really increased my interest in that area. It was just a thing that you would talk about in high school.

Host

你发过一条推文,谈到李飞飞的传记,说移民故事和一路上有相信你的老师,这些与你的生活有相似之处。高中是你遇到第一批看到你早期才华的老师的时候吗?

You had a tweet where you were talking about Fei-Fei Li's biography and how it spoke to you in some sense, the immigration story and having teachers that believed in you along the way were parallels to your own life. In high school, was that when you had that first batch of teachers who saw some of that remarkable early talent?

Karol Hausman

我觉得贯穿始终,不只是高中。有很多人,现在回想起来我非常感激,他们要么看到了什么,要么鼓励我深入探索,要么激励我,要么给了我一些本不会有的机会。每个阶段我都能指出某个人。读李飞飞的书让我想起了这一点。我记得读完后我感谢了其中一些人,感谢他们为我所做的一切。所以,是的,从波兰到德国再到美国,在美国多个地方,这是一段漫长的旅程,感觉每个阶段都有人在为我加油、帮助我。

I think it was throughout. It wasn't just high school. There were multiple people that now looking back I'm really grateful to, that either saw something or just encouraged me to dig deeper or motivated me or gave me some opportunities that I wouldn't have otherwise. At every single stage there's someone I could point to. Reading Fei-Fei's book just reminded me of that. I remember after I read it I thanked some of them for everything they did for me. So, yeah, it was a long journey from Poland through Germany to the US, going through multiple places in the US, and it felt like at every single stage of the journey there was someone rooting for me and helping me.

早期兴趣与机器人之路 Early interests and path to robotics

Host

在你意识到自己想从事机器人领域之前,你本科学习的是计算机科学和哲学,但我好奇在此之前你是否有其他兴趣。比如小时候,你有没有想过'我要成为天体物理学家'或者神经科学家?

Before you realized that you wanted to work in robotics in particular, as an undergrad you studied computer science and philosophy, but I'm curious if there were interests before that. Like as a kid were you thinking, 'I'm going to be an astrophysicist' or a neuroscientist?

Karol Hausman

是的,我其实非常喜欢物理。高中毕业后我申请了很多物理和数学项目。我以为自己会成为天体物理学家或理论物理学家。我对此非常着迷,读了很多书,高中时非常喜欢物理。但后来很多朋友决定攻读工程学位,我们都要去华沙。我想也许我可以两者都学,或者先学工程,如果觉得无聊再转回物理。当时这并不是一个刻意的选择,更像是:这些人非常聪明,我想和朋友们在一起。而且机器人真的很酷。所以让我试试,看看情况如何。于是我去了华沙,学习当时所谓的机电一体化,基本上是电子、机械工程和计算机科学的混合。后来我也学了哲学。但我觉得搬到华沙后,我才开始意识到机器人是一个领域,但非常令人失望。我记得在华沙学习机器人学,你能得到的梦想工作是在华沙附近的雅芳工厂做机器人工程师。那是每个人都竞争的工作。如果你能在雅芳的生产线上工作,那就算在机器人行业成功了。我记得我去过一次,看到了那条生产线,那是我见过的最无聊的东西。机器人只是从 A 点移动到 B 点,毫无智能可言。我想那时我开始意识到,我真的想做超越这个的事情。似乎没有人在这方面努力,这反而激发了我更大的兴趣。于是我开始思考:谁在做智能机器人?这真的存在吗?这定义了之后的旅程:我试图找到那些人,从一个地方搬到另一个地方,最终自己开始做这件事。

Yeah, I was really into physics actually. Right after high school I applied to a lot of physics and math programs. I thought I was going to be an astrophysicist or a theoretical physicist. I was really fascinated by it. I was reading a lot of books. I was really into physics in high school. But then a lot of my friends decided to go for engineering degrees and we were all going to go to Warsaw. I thought that I could maybe try to study both or I could start with engineering and then go back to physics if it turns out to be boring. It wasn't like a very intentional choice at that time. It was more like these are really smart people, I really want to stay with my friends. And then I'm, robots are really really cool. So let me try that and then see how it goes. So I went to Warsaw to study what was called mechatronics at the time, which was basically a mixture of electronics and mechanical engineering, computer science. Then I also studied philosophy. But I think once I moved to Warsaw, that's when I started realizing that robotics is a thing, but it's a thing that is very disappointing. I remember studying robotics in Warsaw. The dream job you could get was being a roboticist at an Avon factory right next to Warsaw. That was like the job that everybody was competing for. If you could work for Avon on the production line, that's how you made it in the robotics industry. I remember I went there once and saw that production line and it was the most uninspiring thing I've ever seen. The robots, as I said before, just moving from A to B, not being intelligent whatsoever. I think that's when it started clicking for me that I would really like to do something beyond that. It seemed like no one was working on it, which sparked my interest even more. So that's where I started thinking who is doing intelligent robots and is it even a thing? That then defined the rest of the journey of trying to find those people and move from place to place as I try to find them. And then eventually start working on it myself.

哲学研究与反思 Philosophy studies and reflections

Host

回顾那段学习哲学的经历,也许现在它只是简历上的一行,但我好奇,当你开始从事关于意识本质或那些让我们更人性化或更少人性化的感觉的工作时,是否有你越来越频繁回顾的作家?

This sort of study of philosophy, as you've looked back on that time, maybe it's a line on the CV at this point, but I'm curious if there are writers that you've ended up coming back to more and more as you start to do this sort of work around the nature of consciousness or the certain senses that make us more or less human.

Karol Hausman

那是一个非常有趣的时期。我觉得在学习过程中有几个领悟。一是很多哲学其实是哲学史,而不是哲学本身。它是在谈论多年前别人怎么想。从外部看,总觉得柏拉图之类的人是天オ,在时代之前就搞清了一切。但当你真正阅读他们的作品时——这是学习哲学的一部分——你会发现只有很小一部分听起来是对的,很多他们说的绝对是错的。这很令人大开眼界。哲学史表明,很多人一路上犯了很多错误,认为世界以某种方式运作,而现在很清楚并非如此。他们有一些非常有趣的观察,我认为我们常常过度强调这些。但我们错了很多。我仍然记得一些我认为非常相关且我非常喜欢的哲学家。其中之一是斯宾诺莎。回顾我所学的内容,他是我最喜欢的哲学家。但也有一些科目,非常少,更关注当下的哲学思考而不是历史。那些对我来说最有趣。有一门课叫本体论,研究事物本身,研究它们是什么。我觉得甚至很难理解这意味着什么,但那是那种你坐下来思考现实的科目,非常有趣。尽管它非常无结构,不清楚什么是对什么是错,但它似乎触及了某种真实的东西。

It was a really interesting period. I think there were a few realizations I had while studying it. One was that a lot of philosophy is the history of philosophy. It's not philosophy itself. It's talking about what others thought about years ago. From the outside looking in, it always sounds like Plato or people like that were these geniuses that figured everything out way before their time. But when you actually read their works, which was part of the deal of studying philosophy, you realize that there's a very small percentage of things that sound right, and there are a lot of things that they said that are absolutely wrong. That was quite eye-opening. The history of philosophy shows that there are people who made many mistakes along the way and thought that the world worked a certain way when it's now very clear that it doesn't. They had some really interesting observations, and I think we often over index on those. But we were wrong a lot. There are still some philosophers that I remember studying that I think were very relevant and I was a big fan of. One of them is Spinoza. That was my favorite philosopher when I look back at what I studied. But then there were some subjects, very few, that were more about philosophizing today rather than looking at the history. Those were the most fun for me. There was this one subject called ontology, which was a study of things, of how they are. I think it's even difficult to comprehend what this means, but it was one of those subjects where you just sit and contemplate reality and it was very fun. Even though it was very unstructured and it wasn't clear what is right and what isn't, it seemed to touch something real.

斯宾诺莎与实在结构 Spinoza and the structure of reality

Host

而且可能是我印象最深的科目,也是我最有乐趣的地方。比哲学史有趣多了。所以很有意思,我喜欢探索这些,因为至少对我来说,这与你试图实现的一些事情有着深刻的相似之处。为什么是斯宾诺莎?他身上的什么东西让你觉得深刻或真实?

And probably does the subject I remember the most. That's where I had the most fun. Much more so than the history of philosophy. So interesting and these things I love to explore because there is to me at least such deep parallels between some of the things that you're trying to build to fruition. What why Spinoza? What was it about him that seemed deep to you or true?

Karol Hausman

真正让我产生共鸣的一点是,当时很多哲学家都在思考上帝的本质和现实的本质。通常有一种观点认为,上帝是一个独立的实体,要么是万物的创造者,要么以某种方式控制着宇宙中发生的一切。而斯宾诺莎的有趣之处在于,他说我们周围的一切,整个现实,就是上帝。就是这样。它有一种底层结构,而这种结构本身就是美和智慧所在。我觉得这非常发人深省。当你研究物理学以及机器人或人工智能之类的东西时,你会越来越意识到事物存在某种底层结构。在某种程度上,这种结构相当反直觉。比如它为什么存在。为什么你可以用一个简单的方程来描述世界的所有复杂性,而不是用一百万个不同的方程。而且这个方程越优美、越简短,它似乎就越真实,或者说越准确。我认为这最接近当今科学状态下的感受。

The one thing that really resonated with me was this idea that, at that time a lot of philosophers were contemplating the nature of God and the nature of reality. And there was often this idea that God is this separate entity that either was the creator of everything or had some kind of way of controlling everything that is happening in the universe. And I think with him what was so interesting was that he said that everything around us, like all of reality, that is God. That is what it is. And there is some underlying structure to it and that structure in itself is where the beauty lies and where the intelligence lies. And I thought that was very thought-provoking. As you study physics and things like robotics or AI, I think more and more you realize that there's some underlying structure to things. And in some ways that structure is quite counterintuitive. Like why it is there. Why is it so that you can describe all of complexity of the world in a simple equation rather than in a million different equations. And that the more beautiful and short that equation is, the more real it seems, or the more accurate it is. And I thought this was kind of the closest to what it feels like today given the current state of science.

网球内心游戏与无意识学习 The Inner Game of Tennis and unconscious learning

Host

你发过一条推文,谈到一本书,我为了准备这次访谈尽量多读了一些,叫《身心合一的奇迹力量》。你提到,虽然表面上是在讲网球,但似乎与机器人技术有深刻的相似之处。你指出,书中谈到,如果你有意识地专注于改进击球动作,你可能会达到一个局部最优,也许正手击球会好一点,但要达到最优击球,实际上需要在某种程度上让无意识接管。你如何看待在机器人技术中模拟这种近乎无意识的方面,以及它在流畅运动、本体感觉等中的作用?你读那本书时是在想这个吗?

You had a tweet where you were talking about a book that I tried to read as much of it as I could in preparation for this called The Inner Game of Tennis. And you sort of mentioned how you were really talking about tennis, but it seemed like some deep parallels to robotics. You sort of make the point that they talk about the fact in this book that if you really focus with your conscious mind on improving your stroke, you could sort of get to this local maximum where maybe you are doing a little bit better at hitting a forehand, but really what you need to do to get to an optimal stroke is sort of allow the unconscious mind to take over in some respect. How do you think about emulating this aspect of almost the unconscious and the role it plays in fluid motion and proprioception and all these sort of things when it comes to robotics? Was that what you were thinking about when you were reading that book?

Karol Hausman

很多。是的,绝对有。我认为这不仅仅适用于运动。在我们如何学习、我们以为如何学习、以及我们如何教别人之间有很大区别。比如,就拿学习语言来说,我们以为应该教的方式,或者我们在学校里想教的方式,就是定义所有规则。你先学语法、各种概念、时态如何运作、什么词跟在什么词后面。你学所有这些结构,以为如果你知道了所有结构,就能遵循这些规则,然后像母语者一样说话。但这不是我们学习语言的方式,对吧?你只是沉浸其中,听到每个人都在使用他们直觉上知道的结构,然后突然之间你就完全理解了语言如何运作,尽管你无法明确指出任何一条规则。有很多母语者说得完美,完全遵循所有规则,但他们根本不知道这些规则是什么,底层结构是什么,什么是变格,什么是时态。他们对此一无所知,但仍然完美地遵循它们。那本书展示了这一点,但在运动方面,你可以读所有你想读的关于如何打网球的资料,或者你可以尝试描述正手击球时肘部应该精确在什么位置,或者如何调整挥拍,但真正学会它的唯一方法就是去做,去沉浸其中。我认为这与机器人世界和人工智能世界有很多相似之处,我们曾以为机器应该如何思考或学习,与它们实际如何学习之间存在差异。

A lot. Yeah, definitely I was. I don't think it just applies to motion. I think there's a big difference between how we learn things and how we think we learn things or how we teach things to others. For example, even if you look at something like learning a language, the way we thought you're supposed to teach it or the way we would want to teach it to others in school is by defining all the rules. You first learn about grammar and all the different concepts and how tenses work and what word follows what other words. So you learn all of that structure so that the thinking is that if you know all of the structure, you can follow those rules and then speak like a native speaker would. But it's not how we learn language. Right? You are just immersed in it, you just hear everybody using the structure that they kind of intuitively know, and all of a sudden you emerge with a full understanding of how language works, even though you can't really pinpoint any of these rules. There are many native speakers that speak perfectly, follow all the rules exactly, but have no idea what these rules are, what the underlying structure is, what declension is, what tenses are. They have no idea of any of that, but they still follow them perfectly. With that book, I think that was shown, but in the motion aspect of things, where you could read all you want about how to play tennis, or you can try to describe where your elbow should be exactly when you hit a forehand, or how to adjust your swing, but the only way to really learn it is by just doing it and kind of immersing yourself in it. And I think there are a lot of parallels to the robotics world and to the world of AI, how we thought that the machines should think, or how they should learn things versus how they actually learn them.

机器人视觉的转折点 A turning point in robotics vision

Host

你,我想是在 2023 年,说你有一个时刻,第一次真正觉得能看到机器人行业未来的一点轮廓,事情开始变得明朗。但我觉得你在这个领域已经待了相当长一段时间了。所以我想知道那个时刻之前的时期是怎样的。在一个你不知道结果会如何的领域工作,是不是相当孤独?

You, I think in 2023, say that you had a moment where for the first time you really felt like you could see a bit of the future of the robotics industry, and things were really clicking into place. But you had really been in that field for quite a while already, it strikes me. So, I wondered what the period before that moment was like. Was that a quite a lonely time to be working in a space where you don't know how it's going to play out?

Karol Hausman

所以,也许我可以回顾一下,多讲讲这段旅程。那种感觉很有趣。搬到华沙后,如我所说,看到机器人技术的现状非常令人失望。很长一段时间,感觉没有人正在研究我真正想做的智能机器人。我记得当时我在寻找任何东西、任何人,只要在做一些更接近我想象中机器人样子的事情。很长一段时间,感觉没有人做。然后我记得我终于找到了一个人。我找到了一位德国教授,他写的论文似乎更接近我想象中的机器人。慕尼黑有一个硕士项目叫机器人与机器智能。我想,如果我申请这个项目并被录取,而他们不做智能机器人,那我基本上就完了。如果一个名字那样的项目不做我想象中的任何事情,那肯定大错特错。所以我申请了,搬到了慕尼黑。我记得第一天上课,我去上最早的一节课,在地铁里迷路了,到得非常晚,所以错过了整节课。我到了那里,一个人也没有。课已经结束了。但我看到一个清洁工,我在找洗手间,就问他洗手间在哪里。他指了楼的三楼。我记得非常清楚。当我上楼时,我看到一扇很大的门,后面传来一些机械的声音。我不知道为什么,但我就是决定打开门看看里面发生了什么。那个声音非常吸引人。

So, maybe I can go back and tell you a little bit more about the journey. It's very interesting that kind of the feeling. After I moved to Warsaw, as I mentioned, it was very disappointing to see the state of robotics. And for a long time it felt like no one is working on the intelligent robots that I really wanted to work on. And I remember at the time I was searching for really anything, anybody who is doing something that is a little bit closer to what I imagined robots to be. And for a long time it felt like no one was. Then I remember I finally found someone. I found a professor in Germany that was writing papers that seemed a little bit closer to what I would imagine robots to be. And there was this master's program in Munich called Robotics and Machine Intelligence. And I thought if I apply for this program and get in and they don't do intelligent robots, then I'm basically lost. Like if a program with that title doesn't do anything that I imagine, then something is very very wrong. So I applied and moved to Munich. And I remember the first day of classes, I went to one of the very first classes and I got stuck in a metro somewhere and then was super late, so missed the entire class. And I showed up. No one was there. The class was over at that point. But I saw a janitor and I was looking for a bathroom, so I asked him where the bathroom is. He pointed me to the third floor of the building. I remember it very vividly. And as I went upstairs, I remember seeing these big big doors and there was some mechanical sound coming from behind them. I don't know why, but I just decided to open them and see what's going on there. The sound was very intriguing.

寻找智能机器人 Finding intelligent robots

Karol Hausman

当我推开那扇门,看到两个类人机器人,一个在做爆米花,另一个在给吐司抹黄油。那一刻,我知道这就是我多年来一直在寻找的东西。那是我那段旅程中最震撼的时刻之一——我终于找到了。有看起来智能、能做类人动作的机器人,还有人在研究这个。我立刻离开那个房间,因为里面基本没人。我敲了隔壁的门,问能不能在那里工作。我完全不够格,不知道怎么编程这些机器人,也不熟悉任何工具。但那里的人又给了我机会,让我周一过来试试看。我加入了那个实验室,得到了那份工作,就这样我终于开始研究智能机器人了。所以,我真的很感激那个人。之后我搬到了美国,攻读智能机器人方向的博士。当时我兴奋极了,终于能做这件事了,终于找到了一小群志同道合的人。即使机器人还不太灵光,我也不在乎。我只是很激动能遇到这些人,向他们学习,推动这个领域前进,做我一直想做的事。整个博士期间我都是这么做的。

And as I opened these doors, I see these two basically humanoid robots doing one of them was making popcorn and the other one I believe was spreading butter on a toast. And that was the moment where, you know, this is something I've been searching for at this point for years. So that was the moment that was probably one of the most powerful moments I had during that journey where I finally found it. There's robots that look intelligent, that do something that is very human-like. There are people who work on this. So, I immediately left the room not because there was basically no one there. And I knocked at the room next door and asked if I could work there. And basically I was completely unqualified. I had no idea how to program these robots. I was not familiar with any of the tools. But the person who was there it was another one of those people that gave me a chance and asked me to just come in on Monday and see what I could do. Well, I got involved, got the job at that lab. That's how I got into finally working on intelligent robots. So, yeah, I think this is like yet another person that I'm really grateful to. And then afterwards I moved to the US. Got into a PhD program studying intelligent robots. And at that point I was so stoked that like I finally get to do this and I finally found a small group of people that work on those things. It didn't matter that, you know, it doesn't work that well. I was just so excited that I finally get to, you know, meet the people that are working on this. I can learn from them. I can push it forward and I can finally work on the thing that I always wanted to work on. So, I did this throughout my PhD.

转向深度学习 Switching to deep learning

Karol Hausman

博士期间还有另一个决定性时刻。我当时在研究一个叫主动感知的问题。如果你感兴趣,我可以深入讲讲。但那时感觉虽然有办法写论文、推进博士进度,但没有什么能真正解决这个问题。世界太复杂了,你可以把技术推到一定程度,但永远到不了终点。这些方法不奏效,看不到出路。后来有个博士后候选人到我们实验室参观,作为博士生,我需要向他展示我的工作,作为面试的一部分。我给他看了我的研究,他听完后说,我应该放弃这一切,转向深度学习,那才是解决我问题的办法。我当时想,他根本不知道自己在说什么,可能没认真听,因为我已经在这个课题上花了两年时间。但那天晚些时候,我去听了他的讲座,他展示了他的工作。那又是一个大开眼界的时刻,因为我第一次看到了真正可能奏效的东西。不再是这里一篇小论文、那里一篇小论文,但拼不起来。而是把所有碎片整合在一起的东西。那和我在慕尼黑的经历类似,我决定放弃博士课题,完全改变方向,开始和那个人合作,尽一切努力推动机器人领域的深度学习。那个人就是 Sergey Levine,他现在是我在 Physical Intelligence 的联合创始人,也是机器人深度学习的先驱之一。从那以后,我再也没有回头。这回答有点长,但我想说,我并不觉得孤独,因为能研究这个问题我真的很兴奋。

And I had another moment like this during my PhD, this kind of defining moment where I was working on a certain sort of problems where it was referred to us as active perception. And I can go a little bit deeper into this if you're interested. But at that point it kind of felt like there were ways to write papers to progress in your PhD, but there isn't anything that felt like it could actually solve this problem. It all felt like the world is too complex. You can push it to some extent this technology, but you won't really get to the finish line. They're not kind of adding up. You don't really see a path out of it. And there was another moment where a post-doc candidate stopped by our lab and at that time, when a post-doc candidate stops by, you are supposed as a PhD student to show them your work and kind of with them as part of the interview. So, I showed to that post-doc candidate stuff that I was working on and he listened to all of it and then at the end of it he said that I should drop all of this and switch to deep learning. That's the way to solve the problem that I was actually working on. At that point, I just thought to myself that he had no idea what he's talking about and he probably didn't listen to me because at that point I think I spent already 2 years or something like this working on the topic that I was working on. But then later that day I went to his lecture where he showed what he was working on. And that was another one of those moments where it was extremely eye-opening because for the first time I saw something that like could actually work. Where it was less of like a you know, here's this a little paper over there and a little paper over there, but they don't add up. But it's something that kind of brought all the pieces together and yeah, that was another one of those moments like the similar one to that to the one I've experienced in Munich where at that point I decided to drop my PhD topic, change it completely, start collaborating with that person and do everything I can to push deep learning in robotics. And that person was Sergey Levine who is now my co-founder at Physical Intelligence and one of the pioneers of deep learning in robotics and yeah, I never looked back after that. But to kind of this is a long way of answering your question, but it didn't feel lonely in that you know, I was just really excited to be working on that problem.

信念确立的时刻 The moment of conviction

Host

这些故事太棒了。第一个就像巫师被霍格沃茨录取一样,你终于找到了你的魔法师群体和他们施展的魔法。Sergey 的故事也很有趣,你似乎从怀疑到深信不疑转变非常快。这个周期有多长?就在同一天、同一周吗?

Such good stories there. The first one is almost like you know, when a wizard is accepted to Hogwarts or something. It's like you finally found your collection of magical people and the magic that they were doing. The Sergey story is also so interesting to me because you seem to have gone from skeptical to convinced very very fast. Like what was the cycle on that? Was that like literally the same day, the same week?

Karol Hausman

就在那场讲座中。讲到一半的时候,我就想,对,这说得通。就是它了。我之前做的都是错的,这才是正确的方向。我想尽一切努力研究这个课题。我经常有这样的感觉。我觉得每个研究人员时不时都会有这种感觉,看到某个东西,突然就通了。那是一种非常强烈的感觉,你会想放下手头的一切,因为你找到了比你之前认为的更接近真理的东西。后来我们在将大语言模型与机器人学习结合时,也有过类似的时刻,那又是一个一切开始变得清晰的强大瞬间。

I think it was within that lecture. I think like in the middle of that talk, I was like, yep, this makes sense. This is 100% it. Whatever I was doing is wrong. This is the right way. I want to do everything I can to work on this topic. And I had feelings like this. I think every researcher has a feeling like this every now and then where they just see something and it clicks. And at that point it's a very powerful feeling. You kind of want to drop everything you've been doing and you just feel like you found something that is much closer to truth than what you thought before. I think we all had a similar moment afterwards when we're working on combining large language models with robot learning. And that was I think one more of these another of these powerful moments where it's like things start to click.

泰勒·斯威夫特演示 The Taylor Swift demo

Host

就是那个泰勒·斯威夫特的演示?没错。也许你可以讲讲那个故事,非常有意思。

That was the Taylor Swift demo? That's right. Yeah, maybe you could tell that story because that is a fascinating one.

Karol Hausman

之后我开始和 Sergey 合作。我们几个人真正进入了那个领域。感觉我们就像机器人社区里的一小群叛逆者,因为当时深度学习有很多问题,有些问题至今仍然存在。比如它不可解释,无法真正模块化,似乎没人完全理解它的工作原理,样本效率也不高。问题一大堆。所以当时写任何关于机器人深度学习的论文都非常不受欢迎。我们夹在两个世界之间:机器学习界觉得任何机器人论文都不适合那个场合,而机器人界当时也从未完全接受深度学习。但和一小群人一起推动这些方法还是很酷的。后来唯一真正接受这个观点的地方是 Google Brain。所以我博士一毕业就决定加入 Google Brain,继续和 Sergey 等人合作,基本前提是:要让这些方法真正奏效,就得扩大规模。

Yeah, so what happened afterwards is I started working with Sergey. There was a few of us who were really getting into that field. It kind of felt like we were a small group of renegades within the robotics community because there were a lot of problems with deep learning at the time. It was and some of these problems remain. Like it was not interpretable. There is no way of really making it modular. It seemed like nobody fully understood how it worked. It wasn't very sample efficient. And there were all of these problems. So like at that time it was extremely unpopular to write any deep learning in robotics paper. And it was kind of like between these two worlds of machine learning world where any robotics paper seemed kind of like the wrong paper for that venue. And the robotics world that never fully embraced or didn't embrace deep learning at that time, but it was still really cool to work with a small group of people and push these methods forward. So then afterwards, the only place that was really embracing that view was Google Brain. So I decided to join Google Brain right out of my PhD, continue working with Sergey and others on those set of methods with the basic premise being that the way to really get it to work is to scale it up.

机器人结合大语言模型 Combining robotics with LLMs

Karol Hausman

所以我们正在扩大规模。我们试图弄清楚如何真正让它大规模运作。但同样,我开始觉得世界上有太多复杂性,如果机器人必须亲身经历一切才能学习,那将很难扩展。如果它们必须完全靠自己来学习逻辑和如何分解任务,那将非常困难。

So we're scaling it up. We're trying to figure out how to really get it to work at scale. But again, it started feeling like there was so much complexity in the world that if robots need to experience all of it first hand to learn, it'll be very difficult to scale. It will be very difficult to have them learn about logic and how to break down a task if they had to do it all by themselves.

Host

是的。

Yes.

Karol Hausman

所以我想大概在——我需要回想一下具体是哪一年——我们开始与我的另一位联合创始人 Brian Ichter 合作。我们开始将这些机器人方法与大型语言模型结合起来。这发生在 ChatGPT 时刻之前。那时人们刚开始尝试使用大型语言模型,并越来越深入地理解它们。有一个特别的演示。我们对此非常兴奋,因为我们认为这将是一条将大量机器人未曾亲身经历的、从互联网学到的先验知识引入机器人世界的途径。这将在某种程度上解决必须收集大量数据才能理解世界如何运作的问题,因为这种理解已经嵌入在大型语言模型中。所以如果你能找到一种方法将这两者结合起来,那将非常有帮助。

So then I think around—I would need to look back to see what year it was exactly—we started combining it together with Brian Ichter, my other co-founder. We started combining these robotic methods with large language models. This was before the ChatGPT moment. This is where people just started experimenting with large language models and understanding them better and better. There was this one particular demo. We were really excited about this because we thought that this would be a path to bringing a lot of prior knowledge that robots didn't experience first hand—that we learned from the internet—into the robotics world. And that would kind of solve this problem of having to collect so much data to understand how the world works, because that understanding is already embedded in large language models. So if you figure out a way to combine these two, that should really, really help.

Karol Hausman

我们做了一个特别的实验,测试从互联网规模的知识到机器人行为有多少迁移。通常,当你研究机器人时,你会专注于一个非常具体的任务,然后测试那个具体任务。所以你很少感到惊讶。大多数时候是令人失望的,因为你非常希望这个任务能成功。你已经在这个任务上工作了很长时间,但通常它不成功,不过如果你足够努力,最终你会让它成功。但你很少会遇到这样的情况:你原本没指望这个任务能成功,或者它不是你正在研究的任务,但它却成功了。所以你很少会感到惊喜。

And we had this one particular experiment where we were testing how much transfer you get from this internet-scale knowledge to robotic behavior. And usually when you work on robots, you work on a very specific task and then test that specific task. So you're very rarely surprised. It's mostly disappointing because you really want this task to work. You've been working on this task for a very long time, and then it usually doesn't work, but if you work hard enough, eventually you get it to work. But very rarely are you in a situation where you didn't expect the task to work, or it's not the task that you're working on and it works. So very rarely are you positively surprised.

Karol Hausman

所以我们设置了一些实验来测试有多少知识可以迁移。其中一个实验是:机器人面前放着一个可乐罐,前面有三张不同名人的照片。我们给结合了 LLM 和机器人模型的模型输入的提示是“把可乐罐放在泰勒·斯威夫特的照片上”。其中一张照片就是泰勒·斯威夫特。你可以在网上看到这个视频。那是一个相当不起眼的机器人能力演示。但机器人捡起了可乐罐,然后慢慢移向泰勒·斯威夫特。那是另一个让我们极度兴奋的时刻,尽管如果你看视频,它完全不起眼,因为机器人模型的数据中从未有过任何泰勒·斯威夫特的信息。它必须理解泰勒·斯威夫特的概念,将其与泰勒·斯威夫特的图像联系起来,然后将其与正确的动作联系起来,从而将可乐罐移到泰勒·斯威夫特的照片上,这一切都来自互联网数据。所以那一刻我们恍然大悟,这真的可行。你可以从 LLM、从互联网引入大量先验知识,并将其与机器人动作联系起来。尽管演示本身非常不起眼,但仅仅是将这两种知识来源结合起来这一事实,对我们来说就非常非常令人印象深刻,感觉就像打开了一扇新的大门。

So we set up a few of these experiments trying to test how much knowledge transfers. And one of them was this experiment with a Coke can in front of a robot and three pictures of different celebrities in front of it. And the prompt we gave to the model we built that combined LLM and robot models was 'put the Coke can on picture of Taylor Swift.' And one of the pictures was Taylor Swift. You can see this video on the internet. It's a pretty pathetic demonstration of what robots could do at the time. But the robot picked it up and then slowly moved it towards Taylor Swift. And that was another one of these moments of huge excitement, even though if you watch the video it's totally unimpressive because the robot models had never had any of Taylor Swift in their data. It had to understand the concept of Taylor Swift, connect it to the image of Taylor Swift, and then connect it to the right motion that would move the Coke onto the picture of Taylor Swift, all from internet data. So that was the moment where it clicked for us that that actually worked. You can bring in a lot of prior knowledge from LLMs, from the internet, and connect it to robot motions. And even though the demonstration itself was very unimpressive, just the fact that you can combine these two knowledge sources was really, really impressive for us and it felt like it opened another door.

Host

是的,感觉就像——我不知道你是否熟悉,我肯定你熟悉。我不敢相信我竟然在问你是否熟悉。你比我懂得多得多。但我想在 70 年代有一个实验,叫 Shrdlu 之类的,整个方法就是试图通过编程教给 AI 世界上所有的规则,比如,这是鸟的类别。它有这些特征,然后这是蝙蝠的类别。它有一些重叠的特征,诸如此类。但本质上,人们现在因为使用 LLM 的方式而熟悉这个概念,但就好像机器人突然继承了所有这些规则和知识,通过与 LLM 连接,使得泰勒·斯威夫特的演示成为可能。

Yeah, it feels like—I don't know if you're familiar, I'm sure you are. I can't believe I'm asking if you're familiar. You know much more about this than I do. But there was some experiment in the 70s, I think, called Shrdlu or something like that, where the entire approach was to try and teach AI all the rules of the world programmatically and to say, you know, here is the category of what a bird is. It has these characteristics, and then here's the category of a bat. It has some overlapping characteristics, all of these sorts of things. But essentially, people are familiar with this concept now because of the way we use LLMs, but it was almost as if the robot suddenly inherited all of these rules and pieces of knowledge from tying them up with the LLMs that allows the Taylor Swift demo to happen.

Karol Hausman

是的,完全正确。我认为这里可能有两点。一是我们在机器人领域犯了同样的错误。我们想写下所有这些规则。我们以为只要有足够多的规则,机器人就能遵循它们并做正确的事。但就像我们说的网球的内在游戏,你不能仅仅写下所有规则。你必须实际去做。有一些底层结构,但你无法完全指出来。你需要从数据中学习。我认为人们在语言方面也有类似的想法。他们以为只要我们能写下所有规则,那就够了。但事实证明,有数万亿或数十亿条这样的规则,有时我们甚至无法完全用语言表达它们。你只需要学习它们,如果你学了,那么你就能遵循它们,即使你仍然不完全理解它们是什么。所以我认为我们正在反复学习这一课。

Yeah, that's exactly right. And I think there's maybe two points there. One is that we made the same mistake in robotics. We wanted to write all of these rules. We thought that if we only had enough of these rules, the robots would be able to follow them and do the right thing. But kind of like we said about the inner game of tennis, you know, you can't just write all of the rules. You kind of have to do it. And there is some underlying structure, but you can't fully put your finger on it. You need to learn it from data. And I think people thought about it similarly in language as well. They thought that if only we could write all the rules, that would be enough. But it turns out that there are trillions or billions of these rules and sometimes we can't fully even express them in language. You just need to learn them, and if you do, then you would be able to follow them even though you still don't fully understand what they are. So I think we're learning this lesson over and over again.

Host

是的。

Yes.

Karol Hausman

所以,你知道,你几乎有了第三个顿悟时刻,技术又向前迈进了一步。到那时,听起来你身边已经有了 Physical Intelligence 的大部分联合创始人,但你如何把最后几位成员拉进来,并决定迈出那一步呢?

And so, you know, you have this almost third eureka moment for yourself where the technology is taking yet another jump. By that point, it sounds like you have most of the Physical Intelligence co-founders around the table, but how do you sort of pull the last few members aboard and decide to make that leap?

Karol Hausman

是的。我认为在那一刻,事情开始变得清晰,这是可能的。你知道,如果你在某个事情上工作了这么久——那基本上是我整个成年生活,大约 15 年——很长一段时间里,你认为这个问题没有解决方案。有时感觉更具体一些,但从未感觉它真的能被解决。然后你有了这个时刻,第一次看到了隧道尽头的曙光。也许,只是也许,如果我们做对了一切,如果我们把它与互联网知识结合起来,如果我们扩大规模,如果我们完成所有需要完成的环节,它可能会成功。如果你看到了隧道尽头的曙光,你不能——你想为此做点什么。然后在那一刻,变得很清楚,实现这一目标的方法是创建一个组织,其唯一目的是解决物理智能,解决这个问题。它不能作为另一个组织中的第 20 优先级来解决。这个组织存在的唯一理由就是解决物理智能。

Yeah. I think at that point it started becoming clear that it could be possible. And you know, if you worked on something for so long—that was basically my entire adult life, 15 years or so—and for a long time you thought that there was no solution to this problem. There were sometimes where it felt a little bit more tangible, but it never felt like it could actually be solved. And then you have this moment where for the first time you see the light at the end of the tunnel. Like maybe, just maybe, if we do everything right, if we combine it with internet knowledge, if we scale it up, if we do all of the pieces that need to be done, it might work. If you see that light at the end of the tunnel, you can't—you want to do something about it. And then at that point it became clear that the way to accomplish this is to create an organization whose sole purpose is to solve physical intelligence, to solve this problem. It can't be solved as like priority number 20 in another organization. The only reason for this organization to exist would be to solve physical intelligence.

创业条件 Conditions for Starting the Company

Karol Hausman

那么下一个问题就是,我们到底要怎么做,需要什么条件才能实现?很快我们就清楚了一点:你必须拥有真正最优秀的人才,真正最顶尖的研究人员。因为在这些领域,第二好的团队对这样的公司来说根本行不通。你需要的是领域内最顶尖的人。这一点其实并不难,因为我们彼此已经合作了很长时间。我碰巧已经和这个领域最优秀的人成了朋友。所以只需要确保我们都想做这件事,并且准备好了。另一个条件是,你需要获得大量资金。这需要非常长期的押注,以及完全认同这一点的投资者——这需要时间,而且公司起步时是一家研究公司,不以收入或短期营收为导向。这是第二个要求。第三个要求就是打造一家非常出色的公司,有正确的焦点、正确的人,以及我们在其他领域(比如硬件或运营等)所不具备的专业知识。我大部分时间都在琢磨这两点:我们如何找到合适的投资者,他们能真正帮助我们这次冒险,并且完全认同我们的做法;然后我们如何找到合适的人来填补所有缺失的部分。

So then the next question is how do we actually do it and what are the conditions to make it happen? And what became clear immediately is you need to have truly the best people, truly the best researchers to make it happen. Because the second best team in those areas doesn't really work with a company like that. You need to have the top top people in the field. So that one actually wasn't that hard because we were already working with each other for a very long time. So I happen to be friends with the best people in the field already. So it was just a matter of making sure that we all wanted to do it and we were ready for it. The other condition was that you need to get access to a lot of funding. It would require a very long-term bet and investors that are fully aligned with this taking some time and this starting as a research company and not being oriented around revenue or around the short-term revenue. So it was the second requirement. And the third requirement was just to build a very incredible company with the right focus, with the right people, with the expertise in all the other areas that we didn't have expertise in like hardware or operations and things like this. I spent most of my time figuring out these two points. How can we find the right investors that can really help us in this adventure and they're fully aligned with how we want to do this. And then how can we find the right people to fill all the missing pieces.

Host

Lachy Groom 是怎么加入团队的?因为他显然作为运营者和投资者有着非常令人印象深刻的职业生涯,但并不是人们会自然而然想到的那种人——比如,这个人会把人生的下一篇章奉献给机器人技术。

How did Lachy Groom join the crew? Because he obviously has had a very impressive career as an operator, as an investor, but isn't someone who one would automatically think, you know, this is someone who's going to devote the next chapter of their life to robotics.

Karol Hausman

在我们开始考虑创办公司之前,我并不认识 Lachy。但后来——他应该自己来讲这个故事,而不是由我来说——但我相信当时的情况是,他那时已经投资了几年。在他投资的这些年里,他一直觉得自己不想永远做投资。他想找点别的事情,可以全身心投入,他想创造东西。而他一直着迷的一个领域就是机器人技术。过去几年,他看到这个领域有很多创新。有那么一个时刻,机器人技术迎来了它的高光时刻。他尤其印象深刻的是,他一次又一次看到来自 Chelsea Finn 和 Sergey Levine 实验室的论文,以及我们来自 Google 的论文。所以,他跟他很多朋友说,如果 Chelsea、Sergey、Karol 或者那些团队里的任何人考虑创办公司,请一定介绍给我。我至少想投资,或者至少跟他们聊聊。就这样,我们通过一位朋友介绍认识,我们向 Lachy 做了推介,希望他能投资。在那次推介结束时,很明显他想做的远不止投资。这就是他多年来一直在寻找的机会。他真的很想投身其中。所以,下一步就是我们尽快互相了解,尽可能多地待在一起。在这个过程中,我也越来越清楚,他就是我一直在找的那个人,能在商业、融资、运营方面帮助我们。他填补了我们真正需要的那些缺失环节。

So I didn't know Lachy before we started thinking about starting a company. But then, and he should tell the story rather than me, but I believe what was happening is that he had been investing at that point for a few years. And as long as he's been investing, he always thought that he doesn't want to be investing forever. He wants to find something else that he can fully devote himself to and he wants to build things. And the one area he was always fascinated by was robotics. And he was seeing over the past few years that there is a lot of innovations happening there. There was some kind of moment. Robotics was having its moment. And particularly he was impressed. He was seeing over and over papers coming from Chelsea Finn's and Sergey Levine's lab, as well as our papers from Google. So, he talked to a lot of his friends that if ever Chelsea, Sergey, or Karol, or anybody from those teams are thinking of starting a company, please connect me to them. I would at least want to invest or at least talk to them. So, that's when we got introduced by a friend, and we pitched to Lachy as a way to have him invest. And at the end of that pitch, it became clear that he would like to do much more than just invest. This is the opportunity he's been looking for for all the years he's been investing. And he really wanted to dive in. So, then the next step was us trying to get to know each other as quickly as possible, spend as much time together as we could. And then as we were doing that, it also became clear that this is the person that I've been looking for to help us on the commercial side, on fundraising side, on operation side. They kind of filled the missing links that we really needed.

Host

在你之前没有共事过的人身上,你需要找到某些哲学上的一致性吗?你在他身上看到了哪些特质,让你觉得“是的,这是一个我可以带入这个已经建立了深厚历史的信任圈的人”?

Were there certain philosophical alignments that you needed to find in someone who you hadn't worked with before? What were the traits that you saw in him that made you feel, 'Yes, this is someone I can bring into this very trusted group where we've already built all this history together.'

Karol Hausman

有几个方面。我想最大的一个就是他立刻就理解了。那时,我和很多投资者、其他人、其他商业人士聊过。要让他们理解那个想法非常困难——你必须先做好研究,先构建技术,不能被短期收入分心。如果我们做对了,这将彻底改变世界,成为有史以来最有价值的企业,但你需要有耐心让我们以正确的方式去做,而不是走捷径、限制天花板。而他,我想,大概在头一分钟就理解了。这非常令人安心。除此之外,在我们的对话中,我特别喜欢的一点是,在之前和很多其他人的对话中,我总觉得自己是那个在推动项目野心的人,推动它如果设置得当能走多远。而我认为这是第一次有人为我们设定了更高的目标。这真的很酷,因为感觉我们都在朝同一个方向努力。他会让我们变得更好、更有野心。所以,我很高兴找到了一个能确保我们不走捷径的人。他完全认同要以最大的方式来做这件事。他让我们更有野心。而且他正是我们在那些真正需要更多专业知识的领域所需要的人。他会确保那些领域与我们发展公司的愿景完全一致。所以,基本上那已经是不用多想的事了。

There were a few. I think maybe the biggest one was that he immediately got it. At that point, there were a lot of investors or other people that I talked to, other business people that I talked to. And it was very hard to get that idea across, that idea that you have to do research right. You need to build a technology first, and you cannot be distracted by short-term revenue. And if we do this right, this is going to completely change the world and it's going to be the most valuable business of all time, but you need to have the patience to let us do it the right way, rather than short-circuit that path and cap the ceiling. And he got it, I think, within like the first minute. So, that was very reassuring. And then, on top of that, what I really liked about our conversations was that in many of those previous conversations with other people, I always felt that I'm the one pushing the ambition of this project, of how far it could go if we set it up right. And I think this was the first time where I had somebody else set up even higher ambitions for us. That was really cool to see because it felt like we're all pushing in the same direction. He will make us even better and more ambitious. So, I was really glad to find someone who will make sure that we are not going to short-circuit this. He's fully aligned in doing this in the biggest way possible. He makes us more ambitious. And he's the person in the areas that we really need more expertise. And he'll make sure that those areas are fully aligned with how we want to develop this company. So, it was basically a no-brainer at that point.

Host

你洞察到这是解决这个问题的正确时机,但至少从外部来看,人们可以想象许多不同的形态。比如,这个公司的另一个版本可能会说:“嘿,我们要自己尝试制造最好的人形机器人。”或者,我们会采取某种方法来获取数据,以便非常有效地运行这些模型,而不是另一种方法。你们是如何最终确定今天这种物理智能的版本的?

You had this insight that this was the right moment to go and solve this problem, but at least from the outside, one could imagine many different form factors that might take. Like another version of this company might be saying, 'Hey, we're going to try and build the best humanoid robot ourselves.' Or, you know, we're going to take this approach to getting the data we need to run these models very effectively versus another approach. How did you sort of land on the version of physical intelligence as it is today?

Karol Hausman

我认为从一开始,我们就有一个公司理念,这个理念基于我们到目前为止所做的所有研究——就像语言领域发生的那样,真正解决这个问题的不会是专门化模型。

I think from the get-go, we had a thesis for the company and that thesis was around all the research that we've done up until this point that similarly to what happened with language, it's not going to be specialist models that really solve this problem.

机器人通用模型 Generalist Models for Robotics

Karol Hausman

未来的方向是通用模型,它们能处理各种不同的任务、环境和机器人。这类似于我们在语言领域看到的:你可能认为最好的翻译器应该专门做翻译,最好的程序员应该只训练代码,但结果证明,要成为所有这些领域的最佳,就是训练一个通用模型。这个通用模型同时处理诗歌数据、编程数据和翻译数据,最终在那些专业任务上比所有专家模型都强得多。我们在机器人学习中也开始看到类似的现象。只要我们能收集足够多的数据,只要数据非常多样化,并且用正确的方式处理,我们就能构建出最好的通用模型。一旦解决了这个问题,我们就能拥有一个充满各种形态和高度智能机器人的世界。所以,我们从一开始就有两个想法:第一,智能一直是机器人技术的瓶颈,但我们不想创办一家专注于特定机器人的公司,而是想直接解决智能问题;第二,解决智能的方法是从视觉、语言等领域汲取经验,采用基础模型的方法,包括跨本体学习、大规模多样化数据、大量真实世界数据,并进行必要的研究来构建这些模型。简单来说,就是为不同类型的机器人打造一个 AI 大脑。

It's going to be generalist models that work across all kinds of different tasks, all kinds of different environments and all kinds of different robots. It was similar to what we've seen in language where you would think that the best translator would be specialized for translation or the best coder would be specialized or just trained on code, but it turned out that the way to be best in all of those fields is to train one generalist. The generalist that takes in poetry data and coding data and translation data turns out to be much better than all of those specialists at those specialist tasks. And we started seeing something similar in robot learning. If only we could collect enough data, if that data was very diverse and if we do it the right way, we should be able to build the best generalist. If we solve that, it's really going to allow us to have this world of many diverse form factors and robots that can be very intelligent. So I think the two thoughts we had from the beginning are: one, intelligence has always been the bottleneck for robotics, but rather than trying to start a robotics company that focuses on a specific robot, how can we tackle this problem head-on and just focus on the intelligence? And the second thought being that the way to solve intelligence is to take the lessons from vision and language and other fields and really take the foundational model approach, which includes things like cross-embodiment learning, large diversity of data, a lot of real-world data, and do the research necessary to figure out how to build these models. So you sort of land on this idea of the AI brain for these different robot types, to put it simply.

Host

你是如何开始思考收集数据的正确方式的?这显然是很大的一部分,而且显然有不同的参与者采取不同的方法。我想你一定经过了无数种方式的推理才找到了你现在的方法。

How did you start to think about the right way to gather the data? Because that's clearly such a big piece of it, and obviously there are different players that take different approaches there. I imagine you must have had to reason through that in a million different ways to land on the one you have.

Karol Hausman

我认为我们仍在推理中。我们还没有所有答案。关于如何思考物理智能,重要的一点是我们并不教条。我们不是坐下来苦思冥想然后得出解决方案并把它当作赌注。我认为更好的方式是,我们真正追求真理,进行实验。我们不知道解决方案,而且我们知道我们不知道。所以我们想遵循科学方法,真正找到真相,找到什么有效、什么无效。因为我认为存在一个真正的答案。我们只需要在寻找过程中保持谦逊。这就是我们得出当前答案的方式。我们做了很多实验,尝试了很多不同的想法,看看哪些有效,然后加倍投入。现在,根据我们看到的证据,我们认为它会如何运作:这些模型需要能够吸收非常多样化的数据集。这与其说是选择正确的数据或正确的收集方式,不如说是构建一个能够吸收各种数据的引擎——无论是人的视频、机器人的遥操作数据、手持设备的数据、视频数据还是模拟数据,任何数据都可以。它们能吸收的数据越多,模型就越好。我认为我们正处于机器人学习的这个阶段:我们试图把任何能用的数据都扔给这些模型,以它们能尽可能多地吸收的方式构建它们,让它们达到可部署的门槛——你可以真正将它们部署到世界中,让机器人实际收集数据,执行有经济价值的任务。一旦你达到那个阶段,一旦它们能真正工作并执行有价值的任务,你就进入了下一个阶段:大规模部署机器人,在多个垂直领域和多样化环境中实际交付价值。我认为第二个阶段实际上是我们获得最多数据的阶段。你部署的机器人越多,模型就应该变得越好,你就能部署更多——这是一个自然的飞轮。今天令人兴奋的是,我相信我们非常接近这个门槛。我们已经超过了那个门槛。这非常令人兴奋,因为模型会不断变得更好,它们能吸收更多数据,而且它们还将拥有可持续的数据来源,这些数据非常有价值,因为它们是最真实的,最接近你实际部署机器人的方式。

I think we're still reasoning through it. I don't think we have all the answers yet. The important piece about how to think about physical intelligence is we're not very dogmatic. It's not that we sit down and think very hard and then come up with the solution and this is our bet. I think the better way to think about it is that we're really truth-seeking and we run experiments. We don't know the solution, and we know that we don't know the solution. So we want to follow the scientific method and really try to find the truth and what works and what doesn't. Because I think there is a true answer out there. We just need to be very humble in finding it. So that's how we arrive at the current set of answers. We run a lot of experiments, try many different ideas, see which ones stick, and then double down on them. Now, based on the evidence we've seen, how we think it's going to work: these models will need to be able to absorb very diverse data sets. It's less about picking the right data or the right way of collecting data, more about building the engine that allows you to absorb all kinds of data—whether it's videos of people, teleoperation data from robots, data from handheld devices, video data, simulation data, really anything. And the more data they can absorb, the better the model will be. I think we're at this stage of robot learning where we try to throw anything we can at these models, build them in a way that they can absorb as much as possible, and get them to the threshold of being deployable—where you can actually deploy them in the world and have robots out there collecting data for real, doing economically valuable tasks. Once you're at that stage, once you're at that threshold where they can actually work and do valuable tasks, that's when you enter the next stage: deploying robots at scale in multiple verticals and diverse environments, actually delivering value. I think that second stage is actually going to be where we get most data from. The more robots you deploy, the better the model should get, the more you can deploy them—there's a natural flywheel. And what's exciting about the moment today is that I believe we're very close to this threshold. We're already above that threshold. That's really exciting because the models will keep getting better, they'll be able to absorb more data, but they'll also have this sustainable source of data that is very valuable because it's the most real, closest to how you actually want to deploy the robots.

Host

对于那些可能没有花太多时间深入研究或者刚接触这个领域的人来说,我想说 Physical Intelligence 所做的实验之一,可能比其他公司做得更多,就是这种真实世界数据的方法。为什么这种方法如此重要和宝贵,以至于进入你所说的第二阶段并启动飞轮如此有价值?

For folks that maybe haven't spent as much time digging into this or are coming to it fresh, I would say that one of the experiments that Physical Intelligence has done maybe more than others is this real-world data approach. Why is that so important and valuable to get, and so valuable to enter that phase two that you talked about that can start the flywheel?

Karol Hausman

我们有几种理论解释为什么这很重要,但我想元观点是,如果存在另一种效果更好的路径,比如模拟或视频或其他方法,我们会很乐意选择那条路。再次强调,我们并不是坐下来认为这是最好的路径,并且只做这一件事。我们做了很多实验,尝试了很多想法,而这是似乎效果非常好的一个。为什么我们认为这很重要?这很难绝对地描述。我认为与其他替代方案比较会更容易一些。一个流行的替代方案是模拟。这个领域取得了许多成功,尤其是在运动控制用例中,比如机器人行走或做后空翻等特技。大多数这些方法首先在模拟中训练。对于这类任务,主要复杂性在于如何移动自己的身体。它更多是关于如何正确移动腿以便行走或跑步,而不是与世界互动。从这个意义上说,只要你准确建模了自己的身体,就应该没问题。

We have a few theories why this is really important, but I think the meta point here is that if there was a different path that worked much better, like simulation or from videos or something else, we'd happily pick that path. Again, it's not that we sat down and thought this is the best path and the only thing we're going to do. We ran a lot of experiments, tried out a lot of ideas, and this is the one that seems to be working very well. Why do we think this is important? It's kind of difficult to describe in the absolute. I think it's a little easier to compare it to other alternatives. One popular alternative is simulation. This is an area that had a lot of successes, especially in locomotion use cases, where you have robots walking around or doing stunts like backflips. Most of these methods train in simulation first. For those kinds of tasks, the main complexity is about how you move your own body. It's less about interacting with the world and more about how do I move my legs correctly so that I can walk or run. In that sense, as long as you model your own body accurately, you should be good.

运动与操作的仿真 Simulation for Locomotion vs Manipulation

Karol Hausman

如果你把你的机器人本体建模得非常好,并且能从仿真迁移到现实世界,那你就可以在仿真中学习行为,然后直接应用到现实世界。但到了操作问题,也就是操控周围世界的时候,难点就不再是建模自己的身体,比如手臂从 A 点移到 B 点,而在于建模世界会如何响应。所以操作问题更多是关于你与世界的交互。我认为在这种情况下,模拟周围世界比模拟自身身体要困难得多。这不再只是模拟单个机器人,你需要模拟一切。而我们目前还不知道如何大规模地模拟一切,如何做到足够精确和可扩展,使其适用于每一个任务或物体。要把所有摩擦参数、仿真行为都调对,需要很长时间,而且根本不可扩展。这是我们目前的发现。所以,仿真可以用来做一些更粗粒度、更独立的大动作,但一旦你开始尝试拿起咖啡杯或毛巾之类的东西,你就需要依赖一个必须非常精确的仿真,实际上你得创造一个与现实完全相同的仿真,这时真实世界数据就变得至关重要了。你必须模拟整个外部世界。任何一个这样的任务都太昂贵、太困难了。但只要你能只模拟自己的身体,那就完全没问题,我们在运动、后空翻、舞蹈等方面已经看到了这一点。

If you model that very, very well, your own particular robot, and that transfers from simulation to the real world, you should be able to learn that behavior in simulation, and then that's good enough to work in the real world. Now, when it comes to the problem of manipulation, where you're manipulating the world around you, the difficulty is less about modeling your own body, like how you move your arm from A to B, but it's more about modeling how the world will react to it. So the problem of manipulation is more about this interaction with the world. And I think in this case, it's just much harder to simulate the world around you than it is to simulate your own body. It's no longer about just a single robot that you need to simulate; you need to simulate everything. And we don't really know how to simulate everything at scale, how to do this accurately enough and scalably enough so that it would work for every single task or object. It takes really long to get it exactly right, to get all the friction parameters right, to get the simulation behaviors right, and it's just not as scalable. That's been our finding so far. So it's sort of the case that with simulation you can do some of these coarser, larger actions that are more self-contained, but once you start trying to implement picking up the coffee cup or the towel or whatever, you're starting to rely on a simulation that would have to be so good you'd effectively have to create a simulation that is perfectly similar to reality, and that's where real-world data starts to become so important. You're having to simulate all of the external world. Any one of those tasks is just too costly, too difficult to do. But as long as you can get away with just simulating your own body, that works perfectly fine, and that's what we've seen in locomotion, backflips, dances, or things like that.

强化学习回归 Reinforcement Learning Comeback

Host

大约一年前,你提到强化学习正在卷土重来,这后来成为物理智能与 recap 协作中一个非常有趣的部分。当时你看到了什么让你觉得这项技术可能有第二春?为什么它对你如此有用?

There was a moment about a year and a little bit ago where you talked about reinforcement learning making a comeback, and that has since become a really interesting part of the way that physical intelligence seems to work with recap. What were you seeing at the time that made you think this technique might have a second wind, and why has that been so useful for you?

Karol Hausman

将强化学习应用于机器人已有很长的历史。强化学习的问题在于,你需要从自己的经验中学习,而在收集经验的过程中,你需要遇到一些成功。然后你希望增加导致成功的好的行动的概率,减少导致失败的行动的概率。这意味着,当你探索世界并收集自己的经验时,你需要有一些成功,因为如果你看不到任何成功,你基本上就不知道往哪里走,你会迷失。如果你从零开始,只给机器人随机指令,你很难遇到成功。例如,如果你试图教机器人抓取一个物体,给手臂的所有电机随机指令,这些指令中很难有哪一个能成功抓取。这就是我们常说的探索问题:我们没有办法引导机器人如何探索才能遇到成功。一旦你遇到一次成功,你就进入了飞轮效应;你可以增加那些行动的概率。但获得第一次成功非常困难。这长期以来一直是强化学习的问题。但现在,随着我们构建的模型,探索问题变得容易多了,因为机器人不再从零开始。它们从像 Pi Zero、Pi O Five 或 Pi O Six 这样的基础模型开始,这些模型已经对运动如何工作以及如何探索周围世界有了一些直观的理解。所以即使你在机器人面前放一个新物体,如果你让它去抓取,它要么成功抓取,要么以一种有趣的方式失败。因此,你实际上做一些有用的事情的概率变得非常高,至少比以前高得多。现在你有了探索世界的方法,这使得你可以更容易地应用强化学习方法,因为它们可以更快地遇到成功。这就是强化学习如何随着时间的推移而演变,也是我们能够重新开始应用它并看到其成功的原因。从另一个角度来看,我们认为强化学习重要的原因是,目前大多数这些模型主要是基于模仿训练的。你收集大量数据,无论是来自仿真、人类视频还是真实协作数据,然后这些模型的目标是复制你在数据集中看到的动作。一个重要的注意事项是,这实际上并不是我们关心的目标。我们并不关心是否精确地执行每一个动作,就像它被呈现的那样。我们真正关心的是任务的成功,即你是否完成了任务。我们没有一种方法可以用这些模型来编码这个目标。它们都是为模仿而优化的,因此很难提高成功率,因为这不是它们关心的目标。它们只关心尽可能接近地模仿动作。强化学习提供了一种方法来编码这个目标,让它们优化以实际成功完成任务。这两点结合起来让我们相信,我们应该开始应用强化学习,以便提高可靠性,让模型不是 70%的时间有效,而是 99.99%的时间有效,因为它们优化了正确的目标。同时,这变得可能,因为现在这些模型可以以比以前更智能的方式探索。

There's been a long history of applying reinforcement learning to robots. The problem with reinforcement learning is that you need to learn from your own experience, and while gathering that experience, you need to encounter some successes. Then you want to increase the probability of good actions that lead to those successes and decrease the probability of actions that lead to failures. This means that as you explore the world and collect your own experiences, you need to have some successes, because if you don't see any, you basically don't know where to go. You're lost. If you start from scratch and just command random commands to a robot, it's very unlikely that you'll encounter a success. For example, if you're trying to teach a robot to grasp an object and you command random commands to all the motors of the arm, it's very unlikely that some of these commands will lead to a successful grasp. That's what we often refer to as the exploration problem: we don't have a way of guiding the robot on how to explore so that you can encounter a success. As soon as you encounter one, you're on the flywheel; you can increase the probabilities of those actions. But getting to that very first success is very hard. That's been the problem of reinforcement learning for a very long time. But now, with the models we've been putting together, the exploration problem became much easier because robots now don't start from scratch. They start from a foundational model like Pi Zero, Pi O Five, or Pi O Six, where they already have some intuitive understanding of how motions work and how you can explore the world around you. So even if you put a new object in front of the robot, the chances are that if you ask it to grasp it, it will either grasp it or fail in an interesting way. So the probability of you actually doing something useful becomes very high, or at least much higher than before. Now you have a way to explore the world, and that allows you to apply reinforcement learning methods much more easily because they can encounter successes much quicker. This is how reinforcement learning has evolved over time, and that's what allowed us to start applying it again and seeing its successes. From the other perspective, the reason we thought reinforcement learning would be important is that most of these models today are trained mostly based on imitation. You collect a lot of data, whether from simulation, human videos, or real collaboration data, and then the objective for these models is to try to replicate the actions you've seen in the dataset. The important caveat is that it's not actually the objective we care about. We don't care about executing every single action exactly the same as it was presented. What we really care about is the success of the task, whether you accomplished the task or not. We don't have a way of codifying this objective with these models. They're all optimized for imitation, and therefore it's very difficult to drive the success rate because it's not the objective they care about. They only care about imitating the actions as closely as possible. Reinforcement learning provides a way to codify this objective, to have them optimize for actually accomplishing the task successfully. Both of these things combined led us to believe we should start applying reinforcement learning so that we can improve reliability and have models that work not 70% of the time but 99.99% of the time because they optimize for the right objective. And at the same time, it became possible because now these models can explore in ways that are more intelligent than before.

耐心与商业化压力 Patience vs. Commercialization Pressure

Host

回到你之前提到的一个话题,你说当你刚开始和 Lucky 交谈时,他理解商业化需要耐心,你提到如果公司太急躁,可能会破坏长期愿景。这种情况通常如何发生?你如何保持警惕?

To circle back to a topic you alluded to earlier, you mentioned that when you first started chatting with Lucky, he understood the need for patience on commercialization and you sort of referenced the fact that there's a version of this company that if you're too impatient, maybe you can short-circuit the long-term vision. What are the ways that that happens and how do you have to be alert to them?

Karol Hausman

我认为机器人公司有很长的历史会这样做,或者只是发生在它们身上。我们不是第一家以宏大愿景和广泛机器人愿景起步的公司。但分析过去,通常的情况是:你从那个愿景开始,开发技术,然后因为外部压力,你试图在技术尚未完全成熟时将其商业化,并深入某个特定应用,可能是 TAM 最大的应用。一旦这样做,你开始在技术上偷工减料,破坏了它成为真正通用技术的愿景,而通用技术本应带来最大价值。但现在你只是试图为那个客户或特定垂直领域提供最大价值。这完全合理,对吧?你所有的激励都与之挂钩:收入、客户满意度、估值、员工想法。所以你有所有理由开始偷工减料,试图做出更专用的解决方案。通常就是这样。你从解决通用机器人的宏大愿景开始,但很快变成了一家应用公司,比如仓库拣选公司。这对我来说会是一个令人心碎的结果。我认为我们真的有机会解决大问题——通用物理智能的问题。如果我们做到了,你不仅会实现最初可能专注的那个应用,我相信你能解决所有问题。你能让任何机器人做任何任务。反直觉的是,这也会是更好的商业结果。关键在于你心中要有正确的时间跨度。

I think there's a long history of robotics companies doing this or just happening to them. We're not the first company that starts with a big vision and broad vision for robotics. But if I analyze what happened in the past, what often happens is you start with that vision, you start developing the technology, and then because there's usually some external pressure, you try to commercialize this technology at that point. When it's not fully ready yet and you dive into a particular application, maybe the application that has the biggest TAM or something like that. As soon as you do that, you start cutting corners on the technology itself. And you short-circuit that vision of it being a very generalized general-purpose technology that would actually deliver the most value, but now you're just trying to deliver the most value for that customer or for that particular vertical that you chose. And that's perfectly reasonable, right? Like all of your incentives are tied to that. Your revenue is tied to that, the customer satisfaction is tied to that, your valuation, how employees think about this. So you have all incentives in the world to start cutting corners and trying to make a more special-purpose solution. That's usually what happens. So you start with this broad vision of how general-purpose robots are going to be solved, but you very quickly end up becoming an application company, like a warehouse pick and place company. And this would be a heartbreaking outcome for me. I think we really have a chance to solve the big problem, the problem of general physical intelligence. And if we do that, you won't just enable that one application that you could have focused on initially, but I believe you'd be able to solve all of it. You'll be able to solve it for any robot to do any task. And counterintuitively, that would be the much better commercial outcome as well. It's all just about trying to have the right time span in your head.

优化学习速率 Optimizing for Rate of Learning

Host

这如何影响你对业务合作伙伴的选择?你是优化一系列用例以获取数据多样性吗?还是关于规模,尽可能从实际部署中获取数据?

How does that influence how you think about the right partners for the business? Like are you optimizing for a range of use cases in that case so that you're getting the diversity of data? Is it about scale and getting as much of that data from real-world deployments as you can?

Karol Hausman

目前我们优化的是学习速度。所以现在还为时过早。不像你在推特上可能看到的,机器人并不会马上敲你的门,出现在你家做所有事情。我不认为一切都已经解决,只是扩大现有配方的问题。我认为我们可以把它们推得比现在更远,但还有很多研究要做,很多问题没有答案。所以现在我们主要优化速度,以及如何尽可能多地了解问题。这很大程度上归结为收集什么样的数据,如何将数据整合到模型中,以及如何让这个循环尽可能紧密。这说起来很复杂,但我们优化的是尽可能多地学习,以便找出可扩展的配方,然后尽可能扩大规模。

Right now we're optimizing for the rate of learning. So it's still quite early. Unlike maybe what you can see on Twitter today, it's not that robots are about to knock at your door and show up at your home and do everything. I don't think everything is solved yet and it's just a matter of scaling up the existing recipes. I think we can push them much further than where they are today, but there's still a lot of research to be done and a lot of questions that are unanswered. So right now we're mostly optimizing for speed and how we can learn as much about the problem as possible. A lot of it does come down to what kind of data to collect, how to integrate the data into the model, and how to have that loop be as tightly closed as possible. It's a long way of saying that it's a fairly nuanced question. But what we're optimizing for is to learn as much as possible so that we can figure out the scalable recipe that then we can just scale as much as possible.

最陡学习率论点 Theses on Steepest Learning Rate

Host

你现在有没有关于什么能带来最陡峭学习速度的假设?是非常精细的任务,还是完全不同的东西?

Do you have theses at this point on what gives you the steepest rate of learning, whether that's very fine-tuned tasks or something totally different?

Karol Hausman

我认为我们学到了一些东西。我们知道数据的多样性非常重要。我们知道数据的质量非常重要。我们知道与模型闭环非常重要。我觉得这些都是相当宽泛的说法。如果你一年前问我,我可能会说类似的话。但我们确实学到了这些术语实际意味着什么。我认为人们经常谈论数据多样性或数据质量,但我们并没有真正很好的定义:数据到底是什么?它意味着什么?数据多样性到底意味着什么?你如何衡量它?所以我认为这些都是相当深刻的问题,我们越来越了解数据到底是什么,如何衡量它,如何优化它。但这是一个持续的拉锯:你在多大程度上扩展现有事物?你能在多大程度上更好地理解它?你能在多大程度上提高改进的斜率?所以我认为我们会随着规模扩大继续沿着这条轨迹前进。

I think there are a few of those things that we learn. We know that the diversity of data is really important. We know that the quality of data is really important. We know that closing the loop with the models is very important. I just feel like these are fairly broad statements. If you had asked me this a year ago, I would probably say something similar. But we did learn a lot about what these terms actually mean. I think very often people talk about the diversity of data or quality of data or something like that. But we don't really have very good definitions of what data actually is. What does it mean? Or what diversity of data actually means? Or how do you measure it? And so I think these are fairly deep questions and we are getting more and more of a feel of what data actually is and how we can measure it and how we can optimize for it. But it's a constant push and pull on how much do you scale the existing thing? How much can you better understand it? How much can you increase the slope of improvement there? So I think we'll just continue on the trajectory as we scale it.

无需新研究的扩展 Scaling Without New Research

Host

你提到这个领域仍然需要真正的研究才能实现最终目标,但也有很多可以优化的地方。在没有任何新研究的思维实验中,你认为这能带我们走多远?你对能产生什么样的机器人有粗略的启发式判断吗?

You mentioned that there's still a need for real research in this space to achieve the ultimate end goal, but that there's also plenty to optimize. In the thought experiment where we get no new research, how far do you think that takes us? Do you have some rough heuristics of what kind of a robot that is able to produce?

Karol Hausman

这是我们经常思考的问题。因为我认为语言模型工作中的一个有力洞见是:很长一段时间人们认为我们仍然缺少一些想法,或者需要不同的架构或其他东西。但结果证明它已经足够好了。只是规模扩张的问题。我们只是没有足够的远见来预测规模扩张将如何解决那些在小规模下看似重大的问题。所以我认为配方已经存在的可能性不小。我们今天拥有的配方会起作用,会解决一切。我们只需要以正确的方式扩大规模并很好地执行。当然,我对此并不确定,但这是一个重要的问题,我们思考了很多。所以我们试图验证它。开始扩展现有配方,看看它的表现和扩展情况,同时继续研究,至少能改善规模扩张的斜率。

It's a question that we think about a lot. Because I think one powerful insight from the language model work is that for a long time people thought we still have a few ideas that are missing or we need a different architecture or different this or that. But it turned out that it was good enough. It was just a matter of scaling. And we just didn't have enough foresight to predict how scaling is going to resolve some of the issues that seem like big issues of today at small scale. So I think there is a non-trivial chance that the recipe is already there. And the recipe that we have today would work and would just solve everything. And we just need to scale it in the right way and execute on it really well. Well, I'm not certain about this at all, but I think that's one question that is an important question that we're thinking about quite a bit. So what we're trying to do is to verify it. So start scaling the existing recipe and see how it performs, how it scales, and at the same time continue the research that at the very least could improve the slope of that scale.

评估用仿真 Simulation for evaluation

Host

英伟达可能是最典型的例子,投入大量资金提升仿真引擎的保真度。你对仿真技术快速进步、缩小差距或提供大规模数据以显著改善效果有多乐观?

Nvidia is probably the best example of a company putting a ton of money into improving simulation engines to higher fidelity. How optimistic are you that simulation will improve fast enough to close the gap or provide data at scale to meaningfully improve things?

Karol Hausman

正如我之前所说,我对此持开放态度,并不固执于某一条路径。仿真在过去十多年里进步神速,现在更加逼真且可扩展。我相信它首先会在评估环节产生影响,而不是数据收集。在 Physical Intelligence,随着模型越来越强大,评估它们需要更多时间。你需要更多样本来区分 99% 和 95% 的成功率,而且任务范围也在扩大。因此需要在更多机器人、更多环境中评估更多任务。如果仿真能如我们所愿,我们就可以在多样化场景中评估模型,只要足够逼真,就应该能与真实世界评估相关联。

As I mentioned before, I'm quite open-minded and not dogmatic about the path. Simulation has been improving like crazy over the past decade plus. It's much more realistic and scalable now. I believe the first impact will be on evaluation, not data collection. At Physical Intelligence, as models become more powerful, evaluating them takes more time. You need more samples to distinguish between 99% and 95% success rates, and the task repertoire broadens. So you need to evaluate on more tasks across more robots and environments. If simulation works as hoped, we can evaluate models across diverse scenarios, and if realistic enough, it should correlate with real-world evaluations.

进展速度惊人 Surprising speed of progress

Host

我们聊过你职业生涯中的转折点。Physical Intelligence 最大的转折点可能还在后面,但过去几年里,最让你惊讶或印象深刻的时刻是什么?

We talked about inflection points in your career. Physical Intelligence's biggest ones are likely ahead, but what have been the biggest surprises or most impressive moments over the past couple of years?

Karol Hausman

最主要的是这一切进展之快,而且还在加速。公司成立前,我们以为部署机器人需要大约 5 年,但实际上 18 个月就做到了。要么是我预测能力太差,要么就是进展比预想快得多。这样的时刻有好几个。第一个版本 Pi Zero 就是其中之一:我本以为自主折叠各种衣物需要很多年,但我们在公司成立头 7 个月内就实现了。下一个模型 Pi 05 也是:我原以为让机器人在从未见过的环境中执行任务需要很长时间,但我们让机器人进入了从未去过的陌生家庭。家庭环境是最多样化的,也是难度最高的版本。到了 Pi 06,我们看到机器人以极高的成功率执行任务。现在我们有视频显示,机器人能连续 13 小时不间断地制作浓缩咖啡,反复操作并自行清理。每一次发布都让我惊讶。我原本以为问题要难得多。做 LLM 的人可能也有类似感受。

The main thing is how fast everything is moving and picking up speed. Before starting the company, we thought it would take about 5 years to deploy robots, but we deployed our first robots 18 months in. So either I'm bad at predictions or it's moving much faster. There have been multiple moments. Our first release, Pi Zero, was one: I thought folding diverse laundry autonomously would take many years, but we achieved it within the first 7 months. The next model, Pi 05, was another: I thought it would take very long to get robots to perform in unseen environments, but we got robots to new homes they'd never seen. Home is the most diverse environment, the hardest version. Then with Pi 06, we saw robots performing tasks at very high success rates. Now we have videos of robots doing extremely dexterous tasks like making coffee on a proper espresso machine for 13 hours straight without cuts, over and over, cleaning up. Every release has been surprising. I thought problems were much harder. People working on LLMs probably have similar answers.

机器人的可靠性与自我意识 Reliability and self-awareness in robotics

Host

当机器人连续 13 小时制作浓缩咖啡时,什么样的失败率算好?你对此感到兴奋的是什么?另外,有位朋友提到,失败往往很微妙——机器人意识不到自己卡住了,找不到出路。你常见到这种情况吗?

When running a robot for 13 hours straight making espresso, what does a good failure rate look like? What excites you about that? Also, a friend mentioned that failures are often subtle—the robot doesn't realize it's stuck and can't find a way out. Is that something you see?

Karol Hausman

人们对可靠性关注不够。这是语言模型或聊天机器人与机器人之间的重大区别。聊天机器人犯小错没关系,因为你可以纠正或理解。但物理世界毫不宽容:任何小错误,比如拿手柄时稍有偏差或没插好,都会导致灾难性失败。性能门槛要高得多,因为没有人帮忙。关于自我意识,强化学习通常需要一个价值函数来预测离成功有多近。我们已经训练出相当好的价值函数——有时即使一切看起来正常,它们也能预知机器人即将失败,因为成功期望开始下降。

People don't pay enough attention to reliability. That's a big difference between a language model or chatbot and a robot. With a chatbot, small mistakes are fine because you can correct or interpret it. The physical world is unforgiving: any small mistake, like grabbing the portafilter slightly off or not inserting it correctly, leads to catastrophic failure. The performance bar is much higher because there's no one to help. Regarding self-awareness, reinforcement learning often needs a value function that predicts how close you are to success. We've trained these value functions to be quite good—sometimes they know the robot is about to fail even when everything looks okay, as the expectation of success starts dropping.

AI 自我评估优于人类专家 AI's Superior Self-Assessment vs Human Experts

Host

这其实也是我们在 AlphaGo 等案例中看到的现象——一个计算机程序学会了如何下围棋。如果你观看这个 AI 系统与围棋世界冠军李世石的一些对局,非常有趣的是,在许多对局中,围棋界的顶尖专家们都在解说和评论。他们都认为比赛基本上势均力敌,很难说清是 AI 还是李世石占优。但如果你看价值函数的预测,比赛其实已经结束了。AI 系统 100%知道比赛已经结束,它已经赢了。至少李世石已经没有机会了。但全世界所有的专家,可能包括李世石本人,都认为这是一场非常接近的比赛。所以我认为,我们很可能会在机器人领域看到类似的模型,它们能够比我们更好地预测自己的表现或成功的可能性。

This is actually something that we've also seen with words like AlphaGo where we had a computer program learn how to play Go. If you watch some of the games that this AI system had against Lee Sedol, the world champion at Go, what was quite intriguing is that in many of these games, you have world experts at Go talking about and commentating on the game. They all think the game is basically head-to-head. It's very unclear who is winning, whether it's the AI or Lee Sedol. But then you look at the value function predictions and it's basically over. The AI system knows 100% that the game is over. It's already won by AI. At least Lee Sedol has no chance. But the entire world, all the experts, including probably Lee Sedol himself, think that this is a very close game. So, I think we will probably get to similar models in robots where they will be able to predict how well they're doing or how likely the success is much better than we will.

物理智能的本体感觉与补偿 Proprioception and Compensation in Physical Intelligence

Host

我和儿子最喜欢做的一件事是,他喜欢看我的书架,然后抽出任何他觉得封面有趣的书。他 15 个月大,所以只是根据封面上有趣的脸来选。他对奥利弗·萨克斯的自传《在运动中》非常感兴趣。上周我和他一起翻阅时,翻到了一页,奥利弗·萨克斯谈到他对本体感觉的痴迷,以及这如何是第六感,实际上是最重要的感觉,因为你可以剥夺人类的其他五种感觉,他们或多或少还能应付。但如果你失去了这种对自己身体位置的感觉,生活就会变得极其困难。他谈到与一些患者合作,这些患者可能因为病毒而失去了对自己身体的感觉。有一个案例让我对你的工作以及最终可能的样子思考了很多。这个人感染了病毒,失去了本体感觉,但他通过视觉来补偿。如果房间里的灯灭了,他就无法移动。但只要他能看到自己在哪里,他就能勉强走路,但不能边走边说话等等。所以这基本上是大脑补偿这种感觉缺失的方式。这是否就是我们目前在物理智能方面所处的阶段——我们没有真正的本体感觉,所以我们在用其他感觉来补偿?你认为未来是否有一个点,计算机能够模拟出某种真正等同于本体感觉的东西?

One of my favorite activities to do with my son is he loves to look at my bookcase and just pull out whatever covers he finds interesting. He's 15 months old, so he's just doing it based on interesting faces on the front. He's gotten really interested in Oliver Sacks's biography called On the Move. I was paging through with him last week and it lands on this page where Oliver Sacks talks about being obsessed with proprioception and how this is the sixth sense and it's actually the most important sense because you can strip humans of all the other five and they can more or less get along. But if you get rid of this sense of where your body is, life becomes extremely hard. He talks about working with these patients who maybe have a virus and lose this sense of their own body. He has this one case which made me think a lot about what you're doing and maybe what the final state of it might look like. This guy gets a virus and loses the sense of proprioception, but he compensates with visuals. If the lights go off in a room, he can no longer move. But as long as he can see where he is, he can sort of walk, but not walk and talk, etc. So it's basically the way the brain compensates for the loss of this sense. Is that sort of the stage we're at the moment with physical intelligence where we don't have true proprioception, and so we're sort of compensating with these other senses? And is there a point in the future, do you think, where however a computer is able to emulate that, we have something that really feels equivalent to it?

Karol Hausman

是的,有很多神经科学实验展示了类似的现象,你可以禁用一种感知模态,然后用另一种来补偿。我认为总的来说,正确的感知模态应该是什么并不清楚。我们的直觉常常是错误的,因为这些机器学习算法能够从对我们来说似乎不足的信号中提取出正确水平的信息。我经常被问到的一个问题是关于机器人的触觉传感。我们反而装了腕部摄像头。结果发现它们可以很好地补偿触觉的缺失。它们可能从数据中找到了某种规律,比如你看到夹爪的变形或其他东西,这指示了你施加了多少压力。这足以完全补偿触觉。也许我们会发现有些路径并非如此,你确实需要其他感知模态来提供信号。但到目前为止,我认为整个领域都惊讶于仅用简单的摄像头就能走多远。不过总的来说,智能很大程度上是关于补偿的。我的意思是,有点像预测将要发生的事情,然后试图弄清楚实际发生的事情与你的预测有多大差异。这不仅适用于感知模态,也适用于精度。例如,作为人类,你在运动方面非常不精确。如果我让你连续 100 次把手指放在一个特定的点上,并精确测量,会有相当大的差异,尤其是与现代工业机器人相比,它们可以达到亚毫米精度和非常好的重复性。但如果你将我们的灵巧性与那些机器人的灵巧性相比,那是天壤之别。我们灵巧得多,尽管我们比机器不精确得多。我认为原因是我们有物理智能来补偿。我们有足够的反馈,如果我们只是看着手指的位置,或者只是感觉它,我们就可以立即补偿。我认为我们也会看到这一点被应用到机器人设计中,也许未来的机器人不需要像我们想象的那样精确。它们可以适应很多事情,比如电机中的一点间隙或噪声,因为有了正确的智能,你可以补偿它。这就是我们在 Physical Intelligence 所看到的。我们展示了机器人上有史以来最令人印象深刻的演示之一。但这些机器人如果与任何最先进的工业机器相比,都是非常糟糕的机器人。它们非常非常糟糕。它们非常不精确,不太可靠,有很多间隙和其他问题,但这并不重要,因为正确的智能可以补偿它。

Yeah, there are many neuroscientific experiments showing something like this, how you can disable one sensing modality and then compensate for it with another one. I think in general it's quite unclear what the right sensing modality should be. I think our intuitions are often wrong about this because these machine learning algorithms can extract the right level of information from signals that to us seem insufficient. A common question I get is about touch sensing on the robots. Instead we put wrist cameras. And it turns out that they can compensate for the lack of touch just fine. They probably find some kind of regularities in the data where maybe you see the deformation on the gripper or something that indicates how much pressure you're applying. And that's enough for it to fully compensate for the sense of touch. It might be that we'll find some paths where that's not the case and you really need some other sensing modality to give you that signal. But so far I think it's been quite surprising to the entire field how far you can push it with just simple cameras. I think in general though, a lot about intelligence is about compensating. What I mean by this is kind of like predicting what's going to happen and then trying to figure out if what actually happened is how different it is to what you predicted. It doesn't just apply to sensing modalities. I think it also applies to precision. For instance, as a human, you're very imprecise when it comes to movements. If I ask you to put your finger on a specific point 100 times in a row, and I measure this exactly, there'll be quite a bit of variance, especially compared to modern industrial robots that can go to sub-millimeter accuracy and very good repeatability. But if you compare our dexterity to the dexterity of those robots, it's night and day. We're way more dexterous, even though we're way less precise than a machine. I think the reason for this is that we have physical intelligence that can compensate for it. We have enough feedback that if we just look where our fingers are, or if we just feel it, we can compensate for it immediately. I think we'll see that being translated to robot design as well, where maybe the robots of the future don't need to be as precise as we had thought. They can accommodate for a lot of things, like a little bit of backlash in the motors or just noise, because with the right intelligence, you can compensate for it. That's what we've seen here at Physical Intelligence. We show some of the most impressive demos ever on robots. But these robots are really bad robots if you compare them to any state-of-the-art industrial machine. They're really, really bad. They're very imprecise, not very reliable, a lot of backlash and other problems, but that doesn't really matter because the right intelligence can compensate for it.

书籍推荐:伟大不可计划 Book Recommendation: Why Greatness Cannot Be Planned

Host

嗯,你用一种非常有趣且优雅的方式回答了我这个非常宽泛的问题。所以,我很高兴我问了这个问题,因为这太迷人了。作为最后一个总结性问题,我总是喜欢问大家,如果你有机会给地球上的每个人指定一本书去读,并且知道他们都能理解,你会给大家推荐哪本书?

Well, that was a very interesting and gracious way of answering a very baggy question from me. So, I'm glad I asked it because that was fascinating. As a final wrap-up question, I always like to ask folks that if they had the chance to assign a book to everyone on Earth to read, and know that they would understand it, what is a book you would love to give everyone?

Karol Hausman

我最近重读了一本我非常喜欢的书。我不确定这是否是我会推荐给所有人的书,但我认为这本书真的非常非常好。它是肯·斯坦利的《为什么伟大不能被计划》。是的。我认为这是一本非常美妙的、反直觉的书,它展示了多个例子,说明我们如何在没有计划的情况下达到真正非凡的成就。

I reread recently a book that I'm a big fan of. I'm not sure if this is the book I would recommend to everyone, but this is the book I think is really, really good. It's Why Greatness Cannot Be Planned by Ken Stanley. Yes. I think it's just a wonderful counterintuitive book that shows multiple examples of how we arrive at something really spectacular without planning for it.

结束语 Closing remarks

Host

我觉得这相当鼓舞人心,也很有激励性,只是方式可能比往常稍微不那么直接。所以,我认为这是一本非常有见地的书,能帮助很多人重新思考创新和成就非凡事业的方式。

I find it quite inspiring and quite motivational in a way that is maybe a little less straightforward than usual. So, I think it's just a very insightful book that would help a lot of people to shed a new light on how to think about innovation and achieving something spectacular.

Host

好了,这正好是一个完美的结尾。非常感谢你,卡尔。谢谢。

Well, that's a perfect place to end. Thank you so much, Carl. Thank you.

Host

本期节目到此结束。感谢收听《通才》播客。请在 Apple Podcasts、Spotify 或你喜欢的播客应用上订阅。评分和评论能帮助更多人发现这些讨论。如果你喜欢本期对话,我将不胜感激,如果你能花点时间留下评论。查看往期节目及更多内容,请访问 thegeneralist.substack.com。下次再见,我们将继续探索未来。

That's it. Thank you for listening to this episode of The Generalist podcast. Please subscribe on Apple Podcasts, Spotify, or your preferred podcast app. Ratings and reviews help others discover these discussions. So, if you enjoyed the conversation, I'd be grateful if you could take a moment to leave one. For all past episodes and more, visit us at the generalist.substack.com. See you next time as we continue to explore the future.

互动版:逐字朗读 + 针对本期提问 →