伯克利教授 Jitendra Malik 谈 AI 与人类智能的差距

Berkeley Professor Jitendra Malik on AI's Limits vs Human Intelligence

吉滕德拉·马利克 Jitendra Malik · The Robot Brains · 2023-08-16 · 约 92 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Jitendra Malik 探讨莫拉维克悖论、智能进化,以及为何当今 AI(如 GPT-4)尽管成就斐然,仍远不及人类智能。

Jitendra Malik discusses the Moravec paradox, the evolution of intelligence, and why today's AI, despite impressive feats like GPT-4, still falls short of human-like intelligence.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 26)

全文 · Full transcript(中英对照)

Jitendra Malik 介绍 Introduction of Jitendra Malik

Host

我们今天的嘉宾是伯克利教授、Meta 兼职研究科学家总监 Jitendra Malik。Jitendra 在计算机视觉研究前沿已有数十年,指导过 70 多名博士生和博士后,其中许多人进入了顶尖工业研究实验室,或成为 MIT、伯克利、卡内基梅隆、加州理工、康奈尔等大学的教授。他的论文被引用超过 20 万次,是所有工程学科中引用最高的研究者之一。值得注意的是,他于 2007 年和 2008 年获得 Longuet-Higgins 奖,2015 年两次获得 Helmholtz 奖(表彰经过十年考验的贡献),并当选为美国国家工程院、美国国家科学院和美国艺术与科学院院士。在我个人对计算机视觉的观察中,我一直将 Jitendra 视为计算机视觉社区的旗手,他不仅贡献了广泛的算法,还推动该领域成为更以基准驱动的学科,极大地加速了整个领域。最近,Jitendra 将注意力转向了实现更接近人类智能的问题,其中机器人研究成为核心——这当然是我特别兴奋的。非常高兴您能来。欢迎来到节目。

Our guest today is Berkeley professor and part-time research scientist director at Meta, Jitendra Malik. Jitendra has been at the forefront of computer vision research for multiple decades. He has mentored over 70 PhD students and postdocs; many have gone on to top industry research labs as well as become professors at MIT, Berkeley, Carnegie Mellon, Caltech, Cornell, and so forth. Jitendra's works have been cited over 200,000 times, making him one of the most highly cited researchers across all engineering disciplines. Notably, he was awarded the Longuet-Higgins Prize in 2007 and 2008, and the Helmholtz Prize twice in 2015 for contributions that have stood the test of time, awarded to papers after 10 years of publication. He has been elected into the National Academy of Engineering, the National Academy of Sciences, and the American Academy of Arts and Sciences. In my personal outsider interview on computer vision, I've always seen Jitendra as the flag bearer of the computer vision community, both through a wide range of algorithmic contributions as well as through making computer vision into a more benchmark-driven discipline, which dramatically accelerated the entire field. Lately, Jitendra has turned his attention to the problem of achieving more human-like intelligence, a quest in which robotics research has become central—something I'm of course particularly excited about. So great to have you with us. Welcome to the show.

Jitendra Malik

谢谢你,Peter。很高兴来到这里,期待我们的对话。

Thank you, Peter. It's a pleasure to be here, and I'm looking forward to our conversation.

Host

我非常期待我们的聊天,Jitendra。但在深入今天的对话之前,我想感谢我们的播客赞助商:Index Ventures 和 Weights & Biases。Index Ventures 是一家风险投资公司,投资于从种子轮到 IPO 各个阶段的杰出企业家,在旧金山、纽约和伦敦设有办事处。该公司支持包括 AI、SaaS、金融科技、安全、游戏和消费在内的多个垂直领域的创始人。就个人而言,Index 是 Covariant 的投资者,我极力推荐他们。Weights & Biases 是一个 ML Ops 平台,通过实验跟踪、模型和数据集版本管理以及模型管理,帮助您更快地训练更好的模型。OpenAI、NVIDIA 以及几乎所有发布大型模型的实验室都在使用它。事实上,我在伯克利的学生和 Covariant 的同事中,许多人(如果不是全部)都是 Weights & Biases 的重度用户。Jitendra,让我们开始对话吧。

I'm really looking forward to our chat, Jitendra. But before diving into today's conversation, I'd like to thank our podcast sponsors: Index Ventures and Weights & Biases. Index Ventures is a venture capital firm that invests in exceptional entrepreneurs across all stages, from seed to IPO, with offices in San Francisco, New York, and London. The firm backs founders across a variety of verticals including AI, SaaS, fintech, security, gaming, and consumer. On a personal note, Index is an investor in Covariant, and I couldn't recommend them any higher. Weights & Biases is an ML Ops platform that helps you train better models faster with experiment tracking, model and dataset versioning, and model management. They are used by OpenAI, NVIDIA, and almost every lab releasing a large model. In fact, many if not all of my students at Berkeley and colleagues at Covariant are big users of Weights & Biases. Jitendra, let's dive into our conversation.

当前 AI 与人类智能的局限 Limitations of Current AI Compared to Human Intelligence

Host

ChatGPT 和其他大型语言模型现在风靡一时,当然,伯克利毕业生 John Schulman 在 OpenAI 的这些努力中处于核心地位,我们对此感到非常自豪。但尽管有这么多兴奋点,在您最近的演讲中,您非常明确地指出,今天的 AI 与人类智能相比仍然非常非常有限。您能详细谈谈您的想法吗?

ChatGPT and other large language models are all the rage right now, and of course we have a lot to be proud of with Berkeley graduate John Schulman being at the center of these efforts at OpenAI. But despite all the excitement, in your recent talks you've made it very clear that today's AI is still very, very limited compared to human intelligence. Can you say a bit more about your thinking here?

Jitendra Malik

当然,Peter。对我来说,这其实是 AI 社区中至少一部分人已经知道一段时间的事情,即莫拉维克悖论。我们可以将其与整个智能的历史联系起来。如果我们把智能看作是通过进化产生的生物构造,追溯到 5.5 亿年前,当时出现了第一批能够移动的动物,然后它们能够看见——视觉对移动很重要,因为这样你才能知道去哪里找食物等等——再到,比如说,过去七百万年,人类从其他灵长类动物进化而来。化石中有非常好的进化数据表明,大脑的进化跟随了手的进化,即具有对生拇指的手。这发生在我们成为两足动物之后,我们开始用双脚行走,这样手就可以自由地制造工具和操纵物体,然后对生拇指帮助了我们。当然,工具制造使人类相对于其他物种获得了优势,尽管与其他许多动物相比,人类本身相当弱小。然后,当然,语言出现了。粗略计算一下:如果你把整个进化史看作 24 小时,那么语言就是最后的两三分钟,语言、符号思维以及许多我们与复杂推理相关的东西。所以从某种意义上说,GPT-4 能在律师资格考试或各种测试中达到 90% 的水平,这非常了不起。我的意思是,这令人难以置信,从公众的角度来看,这是高度智能的标志。我想把这与 20 年前 AI 的成就联系起来,当时我们有了第一个能在国际象棋中击败人类的程序——我想那是在 90 年代末——然后,当然,还有 DeepMind 的工作,他们展示了他们的程序能在围棋中击败人类。这些成就非常了不起,对公众来说,在这些游戏中的表现是智能的终极成就。这确实了不起;我并不是要贬低这些成就。但对我来说,所有那些先于这种成就水平的东西同样重要,这包括基本的感知运动能力:看的能力、移动的能力、操纵物体的能力、规划的能力等等。这就是我的立场,我觉得我们只有征服了智能的这些方面,才算真正征服了智能。

Certainly, Peter. To me, this is something which has been known at least in a segment of the AI community for a while, as Moravec's paradox. And we can tie it to the history of intelligence as a whole. So if we think of intelligence as the biological construct which emerged through evolution, going back like 550 million years ago when we have the first animals that can move, and then they can see—seeing is important to moving because that's how you get to know where to go to find food and so on—and then on to, let's say, the last seven million years when you have the evolution of hominins from other primates. And there is this very nice evolutionary data from fossils which is that the evolution of the brain followed the evolution of the hand as an opposable hand with a thumb which is opposable. This comes after we became bipedal, we started to walk on two feet, so the hand was free to build tools and manipulate objects, and then the opposable thumb helped us. And then, of course, tool making led to humans getting an advantage over other species, even though they are as such quite small and frail compared to many other animals. So then, of course, we have the emergence of language. So by a rough calculation, something like this: if you think of all of evolutionary history as like 24 hours, then language is the last two or three minutes, and language, symbolic thinking, and many of the things that we associate with sophisticated reasoning. So in a sense, it's remarkable that GPT-4 can do like the 90th percentile in the bar exam or in various tests of various sorts. I mean, this is incredible, and from the perspective of the general public, this is a sign of great intelligence. I would like to connect this back to achievements of AI even 20 years ago, when we had the first programs that could beat humans at chess—I think that's in the late 90s—and then, of course, we had the DeepMind work where they showed that their programs could beat humans at Go. These are remarkable, and to the general public, performance at these games is the ultimate achievement in intelligence. And it is remarkable; I'm not trying to downgrade that accomplishment. But to me, all the stuff which precedes that kind of level of accomplishment is equally important, and this includes basic sensory-motor competence: the ability to see, the ability to move, the ability to manipulate objects, the ability to plan, and so on. So that's where I am, and I feel that we will not have conquered intelligence until we have conquered all those aspects of intelligence as well.

大脑发展紧随手部发展 Brain Development Followed Hand Development

Host

我有很多后续问题,Jitendra。首先,您说了一句让我印象深刻的话:大脑的发育跟随了手的发育。当您谈到大脑发育时,这意味着什么?我们的头骨作为物种变大了吗?出现了具有不同能力的新区域吗?到底发生了什么?

I have a lot of follow-up questions here, Jitendra. So first, you said something that really stood out to me: the development of the brain followed the development of the hand. What does that mean when you talk about development of the brain? Did our skulls get bigger as a species? Did new regions emerge that have different capabilities? What exactly happened?

Jitendra Malik

这里,我的意思是化石数据是关于头骨大小的,还有关于手的数据。所以我们有人们发现的与手部结构相对应的化石。在化石记录中,我们可以了解这一点,并且可以了解容纳大脑的头颅腔的体积。我们可以用这些来联系,但大脑内部的软组织在这些化石中从未保存下来,所以我们无法回答一些非常具体的问题。但有一个争论,实际上可以追溯到希腊哲学,那就是:是因为我们聪明所以能很好地操纵物体,还是因为我们能很好地操纵物体所以才变得聪明?我喜欢希腊哲学家阿那克萨戈拉的一句话:正是因为有了手,人才是其他动物中最聪明的。所以这可能是一个争论,对吧?但从化石证据来看,我们实际上知道……

So here, I mean the fossil data is about the size of the skull, and there's data on the hand. So what we have are fossils which people have found which are corresponding to the structure of the hand. So in the fossil record, we can get an idea of that, and we can get a sense of the volume of the cavity of the head in which the brain is. And so we can use those to connect up what the soft tissue inside the brain is never preserved in these fossils, so we can't answer some very specific questions. But there was this debate which actually goes back to Greek philosophy, which is: is it because we are intelligent that we can manipulate objects well, or is it because we manipulate objects well that we became intelligent? And there's a quote that I like from a Greek philosopher called Anaxagoras, which is that it is because of his hand that man is the most intelligent of other animals. And so it could be like a debate, right? But from the fossil evidence, we actually know the...

具身智能与语言基础 Embodied intelligence and language grounding

Host

关于先后顺序,似乎手的发育在某种程度上促进了大脑的发育,因为更大的大脑可以利用这一点。回想早期,人类生活远没有今天复杂,需要应对的事情少得多。我理解你的意思是,你希望先看看我们能否开发出一种智能,能够完成那些基本能力——比如追逐或躲避其他动物、在野外寻找东西、帮助其他人做事——而这些能力如今明显缺失。我好奇的是:你认为这样做也会让语言模型变得更好吗?如果有了物理交互能力的基础,语言模型会变得更鲁棒、更不脆弱吗?

Sequencing which came first, and it appears to follow the order where the development of the hand led to the development of the brain in some ways, because a bigger brain could exploit that. Now I'm thinking back to the early days. Human life wasn't nearly as complex in many ways as today's life is; there was much less to navigate. And as I understand what you're saying, you'd like to first see if we can develop an intelligence capable of doing those things—basic abilities like chasing or evading other animals, finding things in the wild, helping other humans do things—and that is clearly missing today. What I'm curious about is: do you think that doing that will also make language models better? Will language models become more robust, less brittle, if there is a foundation of physical interaction abilities?

Jitendra Malik

这是个好问题,在某种程度上是个实验性问题,我们的领域会随时间给出答案。但让我说说我的倾向。以儿童的发展为例,早期阶段是感觉运动阶段。孩子玩物体,学习爬行、走路,当然也从父母或照顾者那里接收语言输入。但那种语言输入是非常具身的。母亲可能说“把球给我”,孩子面前有一个明确的物体,他知道那是球。“给”这个词在动作上有明确含义。所以孩子早期习得的词汇深深植根于感觉运动经验。“球”这个词唤起视觉印象,也唤起运动印象——我看到一个球,就想扔它。这就是吉布森所说的可供性。椅子有视觉形象,并告诉我如何坐下,所以有运动动作。孩子词汇中的早期词汇就是这种性质。然后孩子习得动词-主语-宾语三元组等,这是语言的开始。后来孩子上学开始读书,遇到“正义”、“和平”、“公平”等词。许多这些词在语境中获得意义。我不一定对公平有直接的视觉画面,但我有与其他概念相关的概念。那个阶段的学习——我们不会去查字典里每个词的意思;我们通过阅读和语境获取意义。这接近 GPT-3 和 GPT-4 内部发生的事情。一个词被删除,你试图从语境中预测它。我认为这个过程相当接近人类通过读书获取知识的方式。但还有更早的阶段,那是非常直接具身的。对于成年人,也许 80%的词汇是通过语境获取意义的,但我认为那 20%是先来的,而且非常重要。为了智能的全面发展,我们需要捕捉这一点。不具身的语言模型会错过这个方面。我不认为这是根本性的失败,因为显然现在有来自谷歌和 Meta 等地的视觉-语言多模态模型,这方向是对的。但在我看来,所有这些都需要具备。

That's a great question, and to some extent it's an experimental question that our field will answer over time. But let me tell you my bias on this. If we take developmental life—the life of a child—the early stage is very much a sensory-motor stage. The child plays with objects, learns to crawl, then to walk, and of course receives linguistic input from parents or caregivers. But that linguistic input is very grounded. The mother may say 'give me the ball,' and there is a clear object in front that the child knows is the ball. The word 'give' has a clear meaning in terms of an action. So the early words a child acquires are very grounded in sensory-motor experience. The word 'ball' evokes a visual impression and also a motor impression—I see a ball that I feel like throwing. This is what Gibson called affordances. A chair has a visual image and tells me how to sit down, so there's a motor action. The early words in a child's vocabulary are of this nature. Then the child acquires verb-subject-object triplets and so forth, and this is the beginning of language. Later, the child goes to school and starts reading books, encountering words like 'justice,' 'peace,' 'fairness.' Many of those words acquire meaning in context. I don't necessarily have an immediate visual picture of what fairness is, but I have the concept related to other concepts. The kind of learning we do at that stage—we don't look up individual words in dictionaries; we read and acquire meaning by context. This is close to what's happening inside language models like GPT-3 and GPT-4. A word is deleted, and you try to predict it from context. I think that process is fairly close to how we acquire knowledge when we read books as humans. But there was that earlier stage, which was very directly grounded. For a grown adult, maybe 80% of vocabulary was acquired through words whose meaning was captured in context, but I argue that the 20% came first and is very important. For the full development of intelligence, we need to capture that. Language models that are not grounded will miss out on this aspect. I don't view this as a fundamental failing, because clearly there are now vision-and-language multimodal models coming from places like Google and Meta, and that's in the right direction. But in my mind, all of these need to be there.

Host

只是研究头脑风暴——它们可能以任意顺序出现。也许在人类中,具身先出现,基本的具身词汇先出现。我们看到的一些研究中,人们只是拿大型语言模型,希望之后附加具身。对人类来说,那可能不现实。但有可能——我好奇你的看法:AI 能否以那种方式构建,先有语言模型,然后具身?还是你坚信应该先具身,先有基本词汇,然后从那里扩展?

Just research brainstorming—they could come in either order. Maybe in humans, grounding comes first, a basic grounded vocabulary comes first. What we see in some research is people just take large language models and hope to later attach grounding. For humans, that would probably not be practical. It is possible maybe—I'm curious about your take: is it possible that AI could be built that way, language model first and then grounding next? Or do you have a strong belief it should be grounding first, basic vocabulary, and then expand from there?

Jitendra Malik

我没有确凿的科学证据,所以这关乎个人的信念或判断。我属于具身优先派,但我愿意被证明是错的。让我给出一些推测性的证据支持我的观点。语言中的许多词是作为隐喻演化的。我们谈论“攀登成功的阶梯”。我们使用时空世界、物体运动和智能体导致物体发生变化的世界。隐喻对语言至关重要。伯克利语言学家乔治·莱考夫写过关于这个主题的书,强调隐喻对语言的重要性,本质上是为了反驳乔姆斯基更形式化的语法方法。语言的使用是通过这种隐喻的阶梯构建的。这让我觉得那样构建语言可能更容易。此外,我们知道没有语言也能构建技能——乌鸦有智能行为,大猩猩和黑猩猩也有。构建运动技能——打开门把手、将钉子插入孔、穿针——这些不需要语言。所以作为纯粹的科学问题,其发展应该不需要语言的叠加。但时间会证明。我把科学生涯押在物理智能优先、语言随后上。但我们都应该尝试自己喜欢的方法,然后就会知道什么有效。

I don't have hard scientific evidence on this, so it's a question of one's belief or judgment call. I'm of the grounding-first school, but I'm willing to be proved wrong. Let me give some speculative evidence in favor of my beliefs. Many words in language evolved as metaphors. We talk about 'climbing the ladder of success.' We use the spatial-temporal world, the world of movement of objects and agents causing things to happen to an object. That is very much something from which metaphors are central to language. George Lakoff, a linguist at Berkeley, has written books on this topic emphasizing the importance of metaphor for language, essentially to counter Chomsky's more formalistic grammar approach. Language in usage is built up with this laddering of metaphor. That would argue to me that it might be easier to build language that way. Also, we know we can build skills without language—crows have intelligent behavior, gorillas and chimpanzees have intelligent behavior. Building motor skills—the ability to open a door handle, insert a peg into a hole, thread a needle—these don't require language. So the development of that, as a purely scientific matter, should proceed without needing the layering of language. But time will tell. I'm betting my scientific career on physical intelligence first, followed by language. But we should all try our favorite approaches, and then we'll know what works.

Host

这确实让我共鸣,但我也愿意像你一样对冲赌注。另一种可能性目前当然不能排除。

That definitely resonates with me, but I'm willing to hedge my bets just like you. The other possibility is certainly not to be ruled out at this point.

五岁前的挑战 Challenges before age five

Host

你在最近一次与 Tendra 的谈话中提到,我们面前的大挑战,或者说至少是人工智能领域的许多大挑战,本质上是一个孩子在五岁之前学到的东西。你能详细说说吗?

You said in a recent talk to Tendra that the big challenges ahead of us, or at least many of the big challenges ahead of us in artificial intelligence, are essentially what a child learns acquires before age five. Can you say a bit more about that?

Jitendra Malik

当然可以。这些本质上都是感觉运动协调的挑战。也就是说,孩子操控物体、捡起物体、识别物体、扔球、拼乐高积木的能力。所有这些都不容易。我们认为它们简单,是因为作为成年人我们能轻松做到,但如果你观察两三岁的孩子,你会发现他们需要大量的练习。事实上,有数据支持这一点,因为我们的心理学同事正在研究儿童如何获得这些能力。例如,纽约大学的 Karen Adolph 研究了儿童如何学习走路,她尽可能以生态效度高的方式,在更自然的环境中观察儿童。结果发现,他们经常摔倒。这是学习走路过程的一部分。他们摔很多次。但摔得不太疼,因为他们的身体柔软,脂肪多,个子矮,重心低,摔倒的冲击不大,而且他们总能爬起来再试。她分享了摔倒次数的数据,数量相当大。同样,尝试操控物体:孩子们似乎只是尝试各种他们感兴趣的活动,从而获得这些技能。我几乎认为你提到的五岁这个年龄实际上是由此决定的,因为如果孩子不能握住铅笔并操控它,你就不能送他们去学校学习写字。所以精细运动控制必须达到一定的能力阈值,才能让孩子接受学习写字的训练。因此,你选择的五岁这个门槛实际上与某些专业技能的获得有关。当然,不仅仅是感觉运动,还有社交技能。孩子们获得了建立其他智能体模型的能力,比如他们的照顾者和其他孩子。他们学会认识到别人有目标。他们学会如何共同注意。所以五岁之前已经获得了许多复杂的能力。

Yes, certainly. These are essentially challenges of sensory-motor coordination, if you will. So these are the ability of a child to manipulate objects, to pick up objects, to recognize objects, to throw balls, to assemble Lego pieces together. All of these are not easy. We think of them as easy because as adults we can do them easily, but if you observe your child at the age of two and three, you will find that they have to do an enormous amount of practice for this. It turns out that there is data on this because our colleagues in psychology are studying how children acquire these abilities. For example, Karen Adolph at NYU has studied how children learn to walk, and she did it in a way where she was trying to be as ecologically valid as possible, observe children in a more natural setting. It turns out that they just fall a lot. That's part of the process of learning to walk. They fall a lot. It doesn't hurt them so much because their bodies are soft, there's a lot of fat, they're relatively short so the center of gravity makes the fall not too big, and they somehow just get up and start trying again. She shares these numbers on how many falls it is, and it's quite a lot. Similarly, the attempt to manipulate objects: children seem to be just trying to do various activities which are of interest to them and acquire these skills. I almost think of the age you mentioned, age five. I think that age five is actually set by this because you can't send a child to school to start to learn to write if they cannot hold a pencil and manipulate it. So fine motor control has to come to a certain threshold level of capability before you subject a child to the discipline of learning to write. So in fact, that five threshold that you pick is actually related to the acquisition of certain expertise. And of course, it's not just sensory-motor; there is also social expertise. Children acquire the ability to build models of other agents, their caregivers, other children. They learn to figure out that they have goals. They learn how to work with mutual attention. So there is a lot of sophisticated capabilities that have been acquired by age five.

Host

这一点特别有共鸣,Jitendra,正如你所说,孩子们经常摔倒。他们学习东西确实需要很长时间,对吧?你知道,我不久前刚当了爸爸,人们会问我,有没有什么启发我的研究。我的主要回答是:我认为我们没有给强化学习智能体足够的 rollout。因为那些看似简单的任务,正如你所说,实际上需要很长时间才能掌握,我们可能太没有耐心了,没有给智能体足够的 rollout 来获得这些基础技能,然后其他东西才能在此基础上建立。

One thing here really resonates, especially Jitendra, as you said the kids fall very often. It actually takes a long time to learn things for them, right? And as you know, I became a dad not too long ago, and people will ask me if anything is inspiring my research. My main response has been: I think we're not giving our reinforcement learning agent enough rollouts. Because this task that seems simple to us, as you said, actually takes a long time to acquire, and we might just be too impatient in terms of how many rollouts we give to our agents to acquire these foundational skills that then other things can be built on top of.

Jitendra Malik

现在你的研究有了很大变化。作为一个常量,我稍后想谈谈计算机视觉,因为直到最近那都是你主要的研究领域,但现在你花了很多时间在机器人学和机器人学习上。你现在在做什么?最让你兴奋的是什么?

Now your research has changed quite a bit. As a constant, I want to get to computer vision later because that's where you spent most of your time until very recently, but now you're spending a lot of time on robotics, robot learning. What are you working on right now? What are you most excited about?

Jitendra Malik

我最感兴趣的高层问题是我所谓的技能获取。我用技能这个词来指某种感觉运动行为。让我举例说明。走路是一种技能,四肢爬行是一种技能,在手中旋转物体是一种技能,扔球是一种技能,接球是一种技能,切苹果是一种技能。那么技能包括什么?技能包括感觉方面,即使用视觉,还有触觉,还有本体感觉。所有这些对于驱动动作都很重要。然后会发生一些物理变化:物体状态改变,或者物体被切开,等等。所以对我来说,这就是我思考的核心问题。我认为技能:我们有一些原子技能,然后我们将它们组合起来产生更复杂的技能,然后越来越复杂的技能,以此类推。作为成年人,我们基本上已经拥有了这个技能包,因此我们现在可以学习相当复杂的任务。假设你被一家公司聘为洗衣机修理工。你需要学习一些具体的东西,但你是从一套所有人类都拥有的基本技能开始的。另一个例子可能是开车:作为一个 16 岁的青少年,你学习开车。你大概用 10、15、20 个小时学会开车,但这实际上掩盖了很多东西,因为它建立在你在 0 到 60 岁之间获得的技能之上。

The high-level problems that I'm most interested in is what I would call skill acquisition. I'm using the term skill as some kind of sensory-motor behavior. Let me explain with examples. The ability to walk is a skill, the ability to crawl on four legs is a skill, the ability to twirl an object in your hand is a skill, the ability to throw a ball is a skill, the ability to catch a ball is a skill, the ability to slice an apple is a skill. So what does a skill include? A skill includes a sensory aspect, so the use of vision but also touch, also proprioception. All of these are important for driving the action. And then there is some physical transformation that happens: an object state is changed, or an object is sliced, or what have you. So to me, that's how I think of a central problem. I think that skills: we have some atomic skills, and then we combine them to produce more complex skills, and then ever more complex skills, and so forth. As grown-ups, we sort of have this wrapper with us, and therefore now we can learn a fairly complex task. Suppose you are hired by a company to become a washing machine repair person. You'll have to learn some specifics, but you start with a set of basic skills which all humans have. Another example might be driving: as a teenager at 16, you learn to drive. You learn to drive in like 10, 15, 20 hours of driving, but that's really concealing a lot of stuff because it builds on skills that you acquired between 0 and 60.

研究议程与技能习得 Research Agenda and Skill Acquisition

Jitendra Malik

好的,这就是高层次的东西。我的研究议程就是关于如何获取技能,然后显然我们可以把它看作一个元问题:我想为机器人获取特定技能,然后我想有一种方法论,使得我获取下一个技能时能更快、更高效。在我们组里,我们首先选择的是行走,这是有原因的。首先,行走是每个人都能做的事情,两条腿行走、四条腿行走,问题的两个版本都很有趣。它在生物学中很常见,因为它与运动相关,而运动很重要。这也是机器人学中研究了数十年的问题。波士顿动力公司的视频有数百万人观看,但他们用更经典的控制设计方法处理这个问题。所以这里的挑战是,我们如何用学习框架来处理它,并非常关注感官输入。我觉得当我们进入这个问题时,我们带来的这两个角度并不是标准做法。我想,人类婴儿有一个发展故事,所以也应该有一个适用于机器人学的学习故事。作为一个职业生涯大部分时间花在视觉上的人,我觉得感知必须从一开始就存在。不要认为感知是后来添加的东西,或者你做了之后得到一些干净的状态,然后把它扔过墙,希望它能工作。它必须从一开始就是系统设计的一部分。这就是我们开始研究行走的原因。我也很幸运有非常好的学生和合作者。Ashish Kumar 是我的学生,Deepak Pathak 当时是 Merita 的博士后,和我一起工作,我们玩得很开心,也取得了很大进展。是的,巨大的进展。

So okay, so that's the high-level thing. I mean, that's what my research agenda is. I feel it's about how to acquire skills, and then obviously we can think of it as a meta problem: I want to acquire specific skills for robots, and then I want to have a methodology such that the next skill I acquire, I can do it more quickly and more efficiently. In our group, the one we picked on first was walking, and there's a reason for that. First of all, walking is something everyone can do, I mean, walking on two legs, walking on four legs, both versions of the problem are interesting. It's obviously common in biology because it's connected to movement, and movement is important. And it's a problem which has been studied in robotics for decades. I mean, Boston Dynamics' videos are seen by millions of people, but they approached the problem in a more classical control design sort of way. So the challenge here was how do we approach it in a learning framework, paying a lot of attention to sensory inputs. I feel that when we got into this problem, those were the two angles that we were bringing in, which are not, shall we say, standard issue. I sort of thought, okay, there has to be a developmental story for human babies, so therefore there has to be a learning story which should work for robotics. And as a person who has spent most of his career in vision, I feel that sensing has to be there from the get-go. Don't think of sensing as something which you just add later, or you do it and then you get some clean state and then you throw it over the wall and hope that will work. It has to be part of the system design from the beginning. So that's how we got into walking. And I was also lucky to have very good students and collaborators. Ashish Kumar was my student, Deepak Pathak, who was at that time a postdoc working with me at Merita, and we had a real fun time and we were able to make a lot of progress. Yes, I mean, a ton of progress.

快速运动适应概念 Rapid Motor Adaptation Concept

Host

是的,巨大的进展。也许我们可以稍后播放一些结果视频,因为它们不言自明,但当然也要解释一下,让只听音频的人也能理解发生了什么。你和你的合作者一起开发了快速运动适应的概念。什么是快速运动适应?它让这些机器人能够做什么?

Yes, I mean a ton of progress. And maybe at some point we can queue up some of the videos of the results because I think they kind of speak for themselves, but of course talk to them such that people who are just listening can also understand what's happening. You, together with your collaborators, developed this concept of Rapid Motor Adaptation. What is Rapid Motor Adaptation and what is it enabling for these robots to do?

Jitendra Malik

让我先解释一下适应这个词。在机器学习中,一个核心概念是泛化。我们在计算机视觉中取得进展,是因为我们决定要识别所有椅子时,无法写出椅子的数学定义,而是提供例子,然后系统学习椅子的概念。这就是泛化。椅子有很多种,但它们有共同点。对我来说,泛化的对应术语是适应,针对任何运动活动。当你走路时,你需要在平地上走,在楼梯上走,在室内走,在雨后泥泞的徒步小径上走等等。我们能做到这一点。如果我们有一个固定的程序,它就不会工作得那么好。你需要适应地形。因此有了适应这个术语。这个想法实际上与经典控制领域从 50 年代和 60 年代就知道的东西有关。他们有一个领域叫自适应控制,但框架略有不同,模型大多是线性的且假设已知。我们在学习范式中所做的类似于那个目标,但使用了深度学习提供的更现代、更灵活的工具。所以我已经讲到了适应部分,我们需要实现运动适应,因为行走是一种运动动作。为什么需要快速?我们需要快速,因为脚下的地形变化很快。如果我们走路,前面有滑的东西,你会开始打滑,但我们会恢复。时间尺度大约是一秒。在行走中,步态周期大约是一秒,如果你踩错一步,你会踉跄并恢复,这应该也在那个时间尺度内。如果花 10 秒,那就没用了;你会摔倒并可能严重受伤。不幸的是,这种情况确实会发生。老年人摔倒是一个重大的健康危机。所以我们的目标是快速运动适应:在一秒或更短的时间内适应,这能让你避免严重摔倒或损伤。那么如何实现呢?这实际上就是改变你走路的方式。走路在这里指的是你向不同腿的不同关节发出的指令,以适应地形:沙子、硬地等。我们开发了一种技术,可以非常快速地推断地形的某些方面,并相应地改变行为。这就是我们所说的快速运动适应,简称 RMA。它基于一个非常简单的想法:如果我向身体发出相同的指令,实际反应会取决于地形。如果我在硬地上走,假设我是盲人,不看不想,像个心不在焉的教授,只是在硬地上走,然后我走到沙滩上,现在在沙子里。会发生什么?我会用几秒钟前同样的力放下脚并抬起,但在沙子里我的脚会下沉。所以当我抬起时,用同样的力它不会抬到同样的程度。这被我的本体感觉系统感知到,它告诉我这里有些不同。所以我们有一个期望的效果,我们发出指令,但实际发生的情况取决于地形。这种差异就是我们可以用来适应的信号。你只需在上面加一点数学,就能让它工作。我认为直觉就是我刚才解释的,然后还有一些技术细节,但我觉得我已经给了你 80%的想法。

So let me maybe first explain the term adaptation. In machine learning, a central concept is generalization. We made progress in computer vision when we decided that to recognize all the chairs, we were not going to be able to write down a mathematical definition of a chair, but we would have examples, and then the system would learn what the concept of a chair is. That's generalization. There are many different kinds of chairs, but there's something common about all of them. The counterpart of that term generalization is adaptation for me, for any motor activity. When you are walking, what you need to do is you need to walk on flat ground, you need to walk on stairs, you need to walk inside, you need to walk in very slushy mud just after the rains on a hiking trail, etc. We are able to do that. If we had a fixed, exactly a fixed program, it would not work so well. You want to adapt to the terrain. Hence the term adaptation. This idea actually connects to something that people in classical control have known about from the 50s and 60s. They have a field called adaptive control, but it was in a slightly different framework where by and large the models were linear and assumed to be known. What we are doing in the learning paradigm is something akin to that goal, but with much more modern and flexible tools which deep learning provides. So I've got to the adaptation part, which we need to achieve motor adaptation because the motor action walking is a motor action. Now why rapid? We need to be rapid because the terrain can change very quickly under our feet. If we are walking and there is something slippery in front of you, then you will start to slip, but we recover. The time scales are on the order of a second. In walking, a gait cycle is on the order of a second, and if you take a misstep, you stumble and you recover, that should take on that order as well. If it takes 10 seconds, it's no good; you will have fallen down and could hurt yourself badly. Sadly, this happens. Old people falling is a major health crisis. So that's why our goal was rapid motor adaptation: adaptation on the scale of a second or less, which enables you to recover without severe falls or damage. Now how do we achieve it? It's really about changing how you are walking. Walking here means the commands you issue to the motors of the different joints in the different legs, in a way that adapts to the terrain: sand, hard ground, etc. What we developed was a technique for very rapidly inferring some aspects of this terrain and changing our behavior accordingly. This is what we call Rapid Motor Adaptation, or RMA for short. It's based on a very simple idea: if I issue the same commands to my body, how it actually reacts will depend on the terrain. If I'm walking on hard ground, let's say I'm blind, not looking, not thinking, being the absent-minded professor, just walking on hard ground, and then I start and walk onto the beach, now I'm in sand. What will happen? I will put my foot down and lift it up with the same force that I was doing a few seconds ago, but in the sand my foot sinks. So when I lift it up, with the same force it won't come up to the same extent. This is being sensed by my proprioceptive system, and it tells me that something is different here. So we have a certain desired effect, we issue a command, but what actually happens is different depending on the terrain. This discrepancy is the signal that we can use to adapt. You just throw a little bit of math on top of it, and we can make it work. I think the intuition is basically what I explained, and then there are a few technical details, but I think I've given you 80% of the idea right here.

机器人适应性与学习 vs 工程 robot adaptability and learning vs engineering

Host

是的,应对并适应如此多样地形的能力——就像这里实验一下,那里实验一下,地形每次都在变化,而机器人却因为不同的摩擦力、与之互动的物体的不同惯性,产生了完全不同的交互,但它总能自我稳定并继续前进。我觉得它的能力真的令人惊讶。在我看来,这很像早期波士顿动力的一些视频,那些机器人的能力同样令人惊讶,但区别在于,这是通过学习实现的,而不是花一整年工程时间去攻克下一个地形或发布下一个视频。所以这里有种魔力——学习居然能实现这一切。我想知道这让你作何感想,既然学习能有效达到甚至可能超越波士顿动力工程实现的水平,你会据此推断接下来可能实现什么吗?

Yeah, the ability to deal with and adapt to such a wide variety of terrains—it's like experiment here, experiment there, terrain changes every time, and somehow the robot has this completely different interaction because of different friction, different inertia of objects it's interacting with, and somehow it just stabilizes itself and keeps going. I think it's really surprising how capable it is. In my mind, it's very reminiscent of some of the old Boston Dynamics videos in the sense that they were also very surprising how capable those robots were, but the difference is that this is just done by learning, not a whole year of engineering to get to the next terrain or the next video you can release. So there's something magical here—the fact that learning can achieve all this. I wonder what it makes you think, given learning can effectively achieve as much as was done with engineering at Boston Dynamics and possibly more. Would you extrapolate that to what is going to become possible next?

Jitendra Malik

我想也许现在对于能观看的读者来说,我想展示一些视频,然后再回答你的问题。

I think that maybe at this point for the readers who can watch the thing, I would like to show some videos and then I'll answer your question.

Host

好的。

Sounds good.

Jitendra Malik

我来边放边解说。这里是我们机器狗,它在伯克利码头附近河床边的岩石区,在岩石间攀爬。它卡住了,但随后移动了脚,成功脱困。这是一个例子,它在一条徒步小径上走下台阶,台阶被落叶覆盖,但它成功恢复并继续前进。这里机器人试图在一个建筑工地上操作新的模块化试点。所有这些例子中,落脚点都不太稳,但它都成功通过了。这是一个室内例子,在几块木板上。我的学生阿什什在做这个,当时是疫情期间,他把机器人带回家,基本上就睡在机器人旁边。这些是在他家里做的实验,机器人在一堆木板上攀爬。同样,在这种场景下,人类落脚也会很不稳。这些都是它能应对的各种变化的例子。我先停在这里,回到你的问题。

So I'll just talk over what's happening. Here we have our robot dog, and it's in a rocky area next to a riverbed on the Berkeley Marina, and it's scrambling among these rocks. It gets stuck, but then it moves the foot over and manages to do that. Here's an example where it's going down some stairs on a hiking path, and the stairs are sort of concealed by a bunch of leaves, but yet it manages to recover and not pause. Here the robot is trying to work on a new modular pilot at a construction site. In all of these, the foothold is a bit unsteady, but it manages to make it through. Here's an example indoors on some planks. My student Ashish, who was working on this, this was in the era of the pandemic, and he had the robot at home and basically went to sleep with the robot next to him. These are experiments in his house where the robot is scrambling on a bunch of planks. Again, this is a setting where humans would be quite unsteady in their foothold. So these are all examples of the kind of variability it can deal with. Let me stop here and return to your question.

Host

等等,等等。在你回到问题之前,我想强调一下,你展示的所有不同案例,都是同一个神经网络在控制机器人,对吧?同一个神经网络知道如何适应所有这些场景。

Hold on, hold on, hold on. Before you return to the question, I just want to emphasize that all the different cases you showed, it's the same neural network controlling the robot, right? The same neural network knows how to adapt to all of these scenarios.

Jitendra Malik

完全正确。谢谢你澄清这一点。是的,所有案例中都是完全相同的策略。不像老式计算机编程那样有 if-then-else 语句:如果这样,就做这个;如果那样,就做别的。不,这里都是完全相同的程序。现在,在幕后发生的是,在这些不同的地形中,对外部条件的一些估计正在被估算。我们使用术语“潜在外在因素”,然后这本质上引导机器人的行走行为做出适当改变。所以,彼得,在我开始展示视频之前,你问过这种风格相对于更经典的风格(比如波士顿动力所代表的,你坐下来分析并写出控制律)意味着什么或有什么影响。让我在这里做一点历史评论。我认为如果你回顾物理学史,看看 19 世纪的物理学,那时没有计算机。所以聪明的物理学家要理解一个现象,本质上就是写下微分方程。牛顿、麦克斯韦——物理学的本质被这些描述现象的微分方程所捕捉。这种思维方式进入了控制理论:你有一个物理系统,所以你用一些微分方程来描述动力学。现在,实际问题是这些微分方程可能需要知道质量、摩擦力等信息。所以物理学是正确的,但物理学有所有这些未知参数,而这些参数实际上会不断变化,当我在楼梯、树叶、泥地、沙地上行走时——这些参数一直在变化。所以尽管理论上物理学家有几百年前的理解方式,但在实际实现中,我被困在不知道所有这些参数的问题上。然后是接触的复杂性:当你有腿时,有些腿在地上,有些腿在空中,每一种情况都需要不同的模型。所以所有这一切,当你试图用分析风格——老式风格——来做时,你写下方程,设计控制器,用纸笔完成,然后在计算机上实现。这种风格适用于已知参数、已识别的简单系统。这是经典控制成功的原因。我的意思是,人类登上了月球——我想到了阿波罗任务——你需要确保绕月轨道准确,这些能力是由经典控制理论的技术提供的。现在在我们的当前设置中,我们拥有这个非常复杂的系统,其模型可能无法提前写出,而学习方法的好处是它们可以在学习过程中处理这些未知数。所以我们同时识别系统和学习控制律,这非常强大。回到我的物理学类比:19 世纪的经典物理学家以某种方式运作,但 20 世纪的物理学家做了什么?1960 年或 1952 年之后,我们可以使用计算机,所以你不必依赖微分方程的解析解;你可以模拟系统。你可以非常精确地模拟它,而不必做简化近似。在统计学中也是如此:早期一代统计学家必须做简化假设——n 趋于无穷,存在某种高斯分布——因为这使得数学可行。但如果你用计算机采样,我们就不必做那些近似。所以在我看来……

Exactly. Thank you for making that clear. Yes, it's exactly the same policy in all the cases. It's not like in the old style of computer programming where you have these if-then-else statements: if something, do this; if something else, do something else. No, here it's all exactly the same program. Now what's happening underneath the hood is that in these different terrains, some estimate of the external conditions is being estimated. We use the term 'latent extrinsics,' and then essentially that guides the robot's walking behavior to change appropriately. So, Peter, before I started to show the videos, you were asking about what this means or what is the implication of this style relative to the more classical style, as represented by Boston Dynamics, where you sit down and figure it out and write a control law. So let me make a little historical remark here. I think that if you look at the history of physics over time, if you look at physics in the 19th century, there were no computers then. So what a clever physicist had to do to understand a phenomenon was essentially to write down differential equations. Newton, Maxwell—the nature of physics was captured by these differential equations which described the phenomena. That's the mindset which came into control theory: you have a physical system, so you describe dynamics by some differential equations. Now, the practical issues here are that these differential equations might require knowing something about the mass, something about the friction, and so forth. So the physics is correct, but the physics has all these unknown parameters, and these unknown parameters will actually keep changing as I'm walking on stairs, leaves, mud, sand—I mean, these parameters keep changing. So even though in theory the physicists have a way of understanding this going back a couple of hundred years, in practical implementation I am stuck with the issue that I don't know all these parameters. Then there is the complexity of contact: when you have legs, then some legs are on the ground, some legs are above the ground, and every one of these regimes requires a different model. So all of this, when you try to do it in an analytical style—the old style—you write down the equations, you design a controller, you do it with pen and paper, and then you just implement it on the computer. This style works well for simple systems which are identified, where we know these parameters. This is responsible for the successes of classical control. I mean, man went to the Moon—I think of the Apollo Mission—and you need to make sure that the orbits around the Moon were accurate, and these kinds of capabilities were provided by techniques derived from classical control theory. Now in our current setup, what we have is this very complex system for which the model may not be written out in advance, and learning approaches have the benefit that they can deal with these unknowns as part of the learning process. So we are dealing with identifying the system at the same time as learning a control law, and this is very powerful. To return to my analogy of physics: classical physicists of the 19th century operated in a certain way, but what did physicists of the 20th century do? Post-1960 or 1952, we have access to a computer, so you don't have to rely on an analytic solution of a differential equation; you can simulate the system. You can simulate it very accurately without having to make simplifying approximations. In statistics, the same thing: an earlier generation of statisticians had to make simplifying assumptions—that n goes to infinity, there's some distribution which is Gaussian—because that enabled you to make the math go through. But if you take samples with a computer, we don't have to make those approximations. So in my view...

机器人学中的仿真 Simulation in Robotics

Jitendra Malik

计算机模拟的力量在于,我们不再需要像前几代统计学家和控制工程师那样做出简化的假设。从这个意义上说,模拟就是答案。模拟依赖于底层的物理模型——这并非黑魔法——但它受益于摩尔定律:计算机每年都在变快,模拟技术也越来越精确。如果你依赖模拟,你就在乘着这条曲线上升,而依赖人类智慧则变化缓慢。这就是为什么我对机器人领域的模拟非常乐观。并非所有同事都同意;有个老笑话是模拟注定会成功。但我认为你必须巧妙地使用模拟,了解其局限性,并掌握从模拟迁移到现实的技术。我相信模拟在推动机器人技术前沿方面将发挥重要作用。

The power of computer simulation is that we no longer have to make the simplifying assumptions that were needed for previous generations of statisticians and control engineers. Simulation is the answer in that sense. Simulation relies on having an underlying physical model—there's no black magic—but it benefits from Moore's Law: every year computers get faster and simulation technology gets more accurate. If you rely on simulation, you're riding that curve, whereas relying on human ingenuity doesn't change that rapidly. That's some intuition for why I'm very bullish on simulation in robotics. Not all colleagues agree; there's an old joke that simulations are doomed to succeed. But I think you have to use simulations artfully, know their limits, and have the right technique for transferring from simulation to reality. I believe simulation has a major role to play in advancing the state of the art in robotics.

Host

我个人同意模拟将发挥非常大的作用,至少用于原型设计,甚至有望学习到可以零样本或少样本迁移到现实世界的东西。我在自己的模拟工作中经常遇到的挑战是,模拟器往往比现实世界的多样性有限。很难将现实世界的那种通用经验带入模拟器。

I personally agree that simulation will play a very big role, at least for prototyping everything, and hopefully even for learning things that can then be transferred zero-shot or few-shot to the real world. The challenge I tend to run into in my own work with simulation is that simulators tend to have limited diversity compared to the real world. It's hard to have the same kind of general experience that the real world brings into a simulator.

Jitendra Malik

这是个合理的观点,但我认为这只是时间问题。对我来说,理想的模拟器不是一次性过程,而是一个反馈循环。当我们构建模拟时,我们是在构建外部世界的模型。然后我们在其中训练系统,将其放入现实世界,如果出了问题,那就告诉我们模拟中需要改变什么。这是一个循环,我认为这就像科学过程:模拟器是理论,现实世界是实验。科学是一个循环,不是单向的。所以你所寻求的多样性会随着我们投入更多努力而出现。在某些情况下,我们可以通过扫描真实公寓并将其放入计算机来捕捉现实。例如,在伯克利和 Meta 的同事一起,我们创建了像 Gibson 和 Habitat 这样的模拟环境,扫描真实公寓并用它们来研究移动机器人的导航策略。这样,你就能获得现实世界的统计数据。模拟依赖于多种技术,所有这些技术都在进步。想想好莱坞电影——它们有非常好的图形和物理效果,足以在视觉上欺骗我们。这表明前景是好的。但这并非在所有情况下都有效;有些同事相信在现实世界中训练,我也相信一点。我是在对冲风险。我团队中的人既做现实世界训练也做模拟训练。我认为未来属于两者。

That's a valid point, but I view this as a matter of time. To me, the ideal simulator is not a one-shot process; it's a feedback loop. When we build a simulation, we are building a model of the external world. Then we train a system in it, put it into the real world, and if something doesn't go right, that tells us something to change in our simulation. This is a cycle, which I think of as the scientific process: a simulator is a theory, and the real world is an experiment. Science is a loop, not one-way. So the diversity you seek will emerge as we put more effort into this. In certain settings, we can capture reality by scanning real-world apartments and putting them inside the computer. For example, at Berkeley and with colleagues at Meta, we created simulation environments like Gibson and Habitat, where we scanned real apartments and used them to study navigation strategies for mobile robots. That way, you get the statistics of the real world. Simulation rests on a variety of technologies, all of which are getting better. Think of Hollywood movies—they have very good graphics and physical effects, good enough to fool us visually. That tells us the prospects are good. It doesn't work in all settings; some colleagues believe in training in the real world, and I believe a bit in that too. I'm hedging my bets. People in my group do both real-world and simulation training. I think the future belongs to both.

盲视 vs 视觉行走 Blind vs. Vision-Based Locomotion

Jitendra Malik

我展示的那些通过快速运动适应训练的机器人可以导航各种地形,但仍有一些地形它们无法处理。到目前为止,我展示的例子都是针对盲机器人,这个机器人无法处理爬楼梯。它可以下楼梯,因为基本上就是摔下去并避免跌倒。但上楼梯它做不到。为此,我们发现有必要引入视觉。如果你不介意,我给你看一些视频。这是一个带有摄像头的机器人。右边是摄像头看到的——它是一个 RGB-D 摄像头。它没有地形的先验知识,正在尝试在一堆工具上行走。如果犯错,它就会摔倒。这是一个它在公园里尝试爬楼梯的例子。它很挣扎,但相当有效地完成了。令人印象深刻的是,这是一个相对较小的机器狗,所以它的腿高很小,而楼梯相对于机器人来说相当大。家里有小型犬的人知道,如果狗很小,它在家里爬楼梯会有困难。这里机器人挣扎着,快要摔倒,但设法保持了平衡。这些例子说明了视觉的用武之地:当地形非常困难时。如果我必须过河,河中间有石头,我可以把脚放在一块石头上,再放另一块,然后过河。我可不想盲目地做这件事。我可以在海滩上盲目行走,但那里不行。我们的机器狗大多数时候可以盲目行走,就像人类一样。盲人可以相当有效地行走,但他们在楼梯上有困难。他们会用棍子戳来感知高度。我们的机器狗借助视觉系统可以处理这些例子。

The robots I showed that were trained with rapid motor adaptation can navigate a wide range of terrains, but there are still terrains they cannot handle. The examples I showed so far were for a blind robot, and this robot could not handle climbing stairs. It could handle climbing downstairs because you sort of just fall down and avoid falling. Climbing upstairs it couldn't do. For that, we found it necessary to introduce vision. If you don't mind, I'll show you some video. This is a robot with an onboard camera. On the right, you see what the camera sees—it's an RGB-D camera. It has no advanced knowledge of the terrain and is trying to walk on a bunch of tools. If it makes a mistake, it will fall. Here is an example where it's trying to climb a bunch of stairs in the park. It's struggling but manages quite effectively. What's impressive is that this is a relatively small robot dog, so the height of its leg is very small, and the stairs are fairly big compared to the robot. People who have very small dogs at home know that if a dog is very small, it has trouble with stairs. Here the robot is struggling, about to fall, but manages to retain its balance. These examples illustrate where vision comes in: when the terrain is really difficult. If I have to cross a river with stones in the middle, I can put a foot on one stone, then another, and cross. I would not want to do that blind. I can walk at the beach blind, but not there. Our robot dog has the behavior that most of the time it can walk blind, like humans. Blind humans can walk quite effectively, but they have trouble with stairs. They poke with a stick to get a sense of the height. Our robot dog with a vision system can manage these examples.

四足 vs 双足行走 Quadrupedal vs. Bipedal Locomotion

Jitendra Malik

你问我我们不能做什么。我认为四足运动已经接近解决。双足运动更难。两条腿对四条腿,四条腿更容易,因为平衡更容易。我们已经为双足机器人做了我们模型的版本,它们能工作,但不知何故这个问题似乎还没有解决。有趣的是,最近似乎对人形机器人的兴趣又复兴了。

You asked me what we cannot do. I think quadrupedal locomotion is pretty close to solved. Bipedal locomotion is harder. With two feet versus four feet, four feet are easier because balance is easier. We have done versions of our model for bipedal robots with two legs, and they work, but somehow it doesn't seem like that problem is solved. Interestingly, there seems to be a revival of interest in humanoid robots these days.

机器人学与运动 Robotics and Locomotion

Host

机器人需要行走能力,当然,为了让机器人有用,手臂和手必须能够举重和做其他事情。这个领域有很多公司,我认为开发双足运动仍然是一个未解决的问题。

Robots will need the ability to walk, and of course for the robot to be useful, the arms and hands have to be able to lift weights and do stuff. There are a number of companies in that space, and I regard developing locomotion, bipedal locomotion, as still not a solved problem.

计算机视觉现状 State of Computer Vision

Host

我想换个话题。你职业生涯的大部分时间,我想今天仍然有一部分时间,都花在计算机视觉上。我很好奇:你怎么看计算机视觉的现状?我们处于什么阶段?只是再扩大规模一点的问题吗?我们甚至应该扩大什么?做无监督学习然后事情就会自然出现,这够吗?你对这个领域的现状有什么个人想法和直觉?

I want to switch gears for a moment. You spend most of your career, and still today I imagine some of your time, in computer vision. I'm curious: what do you think of the current state of affairs in computer vision? Where are we at? Is it just a matter of scaling up a bit more? And what should we even scale up? Is it enough to do unsupervised learning and then things will just emerge? What is your personal thinking and intuition on the state of the field?

Jitendra Malik

有那些我称之为核心视觉的经典问题。我有一个口号:我称之为视觉的三个 R。第一个 R 是识别:你有一张图像,你必须说出它是否有椅子、狗或猫。第二个 R 是重建:进入 3D,恢复物体的 3D 形状或场景的空间布局。第三个 R,为了凑成三个 R,我称之为重组,但有些人会称之为分割或分组:将像素集合分解成单个物体,例如,这些像素属于椅子,这些像素属于人类,等等。在这些问题上,我们看到了显著的进展。识别:我认为有很多来自 Meta 的工作,我在 Meta 的同事们,以及过去几十年的前期工作,表明构建用于识别猫、狗和椅子的计算机程序——我们确实可以做得非常好。我们可以用数千个类别做到这一点。这个领域的工作实际上就是关于扩大规模:我们可以做得很好,如果你给我更多数据和更多例子、更多标签或例子,我就能让它工作。当然,实际应用场景仍然存在,例如,想想在医学图像分析领域工作的人,他们想诊断某人 X 光片中的问题;那将需要使用这种方法,只需再训练一些,提供一些例子,训练一个系统。在将图像分解成物体方面,这也是我研究多年的问题,但最近,例如,Meta 出了一个系统,Segment Anything,它在数百万甚至数十亿个例子上训练,有了这个,那个问题就可以解决。3D 重建的问题,我认为部分解决但尚未完全解决。有很好解决方案的版本是如果你有一个物体的多个视图;这曾经被称为运动恢复结构或 SLAM,现在首选的技术是 NeRF,神经辐射场,它能产生惊人的图像。但我认为这个问题还没有达到人类能做的水平,即从单张图像就能解决。所以我认为从单张图像恢复 3D 的问题仍然——我的意思是,我们取得了进展,但大部分尚未解决。但这只是指核心视觉问题。核心视觉本身并不是目的;视觉是为了别的东西。我认为视觉连接到两个相邻领域。一个是机器人学,即用于指导行动的视觉,这也是我过去五年选择追求的方向。视觉连接的另一个领域是视觉到认知,对于认知,我会把语言作为代理。所以同时处理视觉和语言的模型——这个领域非常开放。我认为已经取得了实质性进展,但还有很多工作要做。我绝不是——我的第一篇视觉论文是四十年前,1983 年,我写了第一篇计算机视觉论文。视觉取得了巨大进步,但我不认为它已经解决了。还有这些领域,比如视觉和语言的交互,长视频理解:我看一部电影,想理解并回答关于电影的问题,这将是语言和视觉推理的结合。显然,机器人学的应用是真实的。如何用更少的数据管理:人类不知何故能够用比我们在 AI 和机器学习中少得多的例子来学习概念。所以那里一定有一些我们遗漏的东西。在我看来,计算机视觉还有很多工作要做。

There are what I would call the classic problems of core vision. I had a slogan for this: I used to call them the three Rs of vision. One R is recognition: you have an image and you have to say, does this have a chair, or a dog, or a cat? The second R is reconstruction: going to 3D, recovering the 3D shape of an object or the spatial layout of the scene. The third R, to make it three Rs, I called it reorganization, but some people would call it segmentation or grouping: taking the collection of pixels and breaking them up into individual objects, for example, these are the pixels that belong to a chair, these are the pixels that belong to a human, and so on. On these problems, we have seen remarkable progress. Recognition: I think there's a lot of work that came out of Meta, my colleagues at Meta, and preceding work over the decades, which shows that building computer programs for recognizing cats and dogs and chairs—we can really do remarkably well. We can do this with thousands of categories. The work in this area is really about scaling up: we can do quite well, and if you give me more data and more examples, more labels or examples, I can make it work. Of course, the practical settings remain, for example, think of someone who works in medical image analysis and wants to diagnose some problem in someone's X-ray; that will require using this methodology, just training some more, giving some examples, and training a system. In terms of breaking up the image into objects, that's again a problem I've worked on for many years, but recently, for example, you have a system that came out of Meta, Segment Anything, which was trained on millions or even a billion examples, and with that, that problem can be solved. The problem of 3D reconstruction is somewhat solved but not fully solved, I would say. The version for which there are very good solutions is if you have multiple views of one object; this is what used to be called structure from motion or SLAM, and now the preferred technology is something called NeRF, neural radiance fields, and that gives remarkable pictures. But I don't think the problem is solved at the level of what humans can do, which is even from a single image we can solve it. So I regard the problem of recovering 3D from a single image as still—I mean, we have made progress, but large parts of it are not solved. But this is referring just to the core problems of vision. Core vision doesn't exist for its own sake; vision exists for something. I think vision connects to two neighboring fields. One is robotics, which is vision for guiding action, and this is what I have chosen to pursue in the last five years. The other field to which vision connects is vision to cognition, and for cognition I would include language as a proxy. So models that deal with vision and language together—this area is very open. I think there has been substantial progress, but there's a lot more to do. I'm by no means—my first vision paper was four decades ago, 1983, when I wrote my first computer vision paper. Vision has seen tremendous strides, but I don't think of it as solved. There are these areas like the interaction between vision and language, long-form video understanding: I watch a movie and I want to understand and answer questions about the movie, which will be a combination of language and visual reasoning. Obviously, the applications to robotics are real. How to manage with much less data: humans somehow manage to learn concepts with far fewer examples than we can in AI and machine learning. So there must be something there we are missing. In my view, there is still plenty to do in computer vision.

对抗样本 Adversarial Examples

Host

我想到的另一件事是计算机视觉中对抗样本的概念。你能解释一下它们是什么以及你目前对它们的看法吗?

One thing that also comes to my mind is this concept of adversarial examples in computer vision. Can you maybe expand on what they are and what are your current thoughts on them?

Jitendra Malik

这些例子引起了大量公众关注。比如说,你有一张停车标志的图像,然后你稍微修改它,添加一些难以察觉的噪声,但当你把它输入你的计算机程序、你的分类器时,它说这是一条蛇或类似的东西。这相当引人注目。我们可以很容易地解释这一点,因为在这些神经网络中,它们是分类器,所以它们有一个决策边界。那些难以察觉的噪声所做的是在某个方向上修改图像,从而欺骗分类器;它只是越过了决策边界的另一边。但人类不会被它欺骗。看那张图像的人类不会被欺骗。这告诉我,人类解读图像的过程比那个简单的分类器更复杂。所以这是一个值得我们研究的话题。我自己没有研究过这个问题,但我会在这里随便说几句。我觉得理解任何图像的正确方式既有自下而上的成分,也有自上而下的成分。自上而下的成分是我们在做梦时运用的。当我们睡觉做梦时,发生的是大脑某处的概念被触发,神经元被激活,它们实际上激活了视觉皮层,几乎类似于实际看到图像时被激活的方式。所以有自上而下和自下而上的成分。自上而下的成分有时与生成模型相关,这里有很多技术定义。我不确定哪个版本会成功,但我觉得那可能是方向。

These are examples that have attracted a lot of popular attention. You have, let's say, an image of a stop sign, and then you modify it a little bit, adding some imperceptible noise, but then you feed it to your computer program, your classifier, and it says this is a snake or something like that. It's kind of remarkable. We can explain these fairly easily because what's happening is that in these neural networks, they are classifiers, so they have a decision boundary. What that imperceptible noise is doing is modifying the image in a certain direction such that it fools the classifier; it just goes over the other side of the decision boundary. But humans are not fooled by that. A human looking at that image is not fooled. That tells me that the human process of interpreting an image is more sophisticated than just that simple classifier. So that's a worthwhile topic for us to study. I have not worked on this problem myself, but I'll give a few throwaway remarks here. I feel that the right way to understand any image has both a bottom-up and a top-down component. The top-down component is what we exercise when we dream. When we are sleeping and dreaming, what's happening is concepts somewhere in our brain go and there is some triggering of neurons, and they in fact activate the visual cortex almost similar to what it would be activated by actually seeing an image. So there is a top-down and a bottom-up component. The top-down component is sometimes connected to generative models, and there are many technical definitions here. I'm not sure what version will work out, but I feel that that's probably the direction.

计算机视觉基准测试动机 Motivation for Benchmarking in Computer Vision

Host

你在计算机视觉领域除了许多算法洞见和贡献之外,还早在二十多年前就开始强调基准测试的重要性。你是如何决定开始做这件事的?为什么这么做?显然它产生了很大影响。你如何看待这种影响随着你开始做这件事而演变?

One of the things you've done in computer vision, in addition to many algorithmic insights and contributions, is you started to put emphasis on the importance of benchmarking now over two decades ago. How did you decide to start doing that? Why did you do that? And obviously it's had a lot of impact. How do you see that impact evolve as you started doing this?

Jitendra Malik

这要追溯到大约 2000 年甚至更早。背景是在一次会议的讨论中,我与一位同事 Olivier Faugeras 进行了一场小组辩论。他是研究计算机视觉几何的人。他认为运动恢复结构和 SLAM 是值得尊敬的问题,因为它们有清晰的数学公式。但他认为分割和识别这类问题质量可疑。他特别说:'分割,好吧,每个人都写论文,提交,然后他们会有一些图像,他们说这些是我的系统找到的组,然后我们必须接受它,我们很高兴。这不是做科学的方式。'他深受法国传统影响,认为需要数学形式化等等。所以那次旅行回来后,我决心证明他是错的。思考方式是:其他没有精确数学模型的科学是如何做的?比如心理学等,它们用人类做实验。所以自然的事情是考虑用人类做实验。我们可以向人们展示图像,让人们标记物体的边界。如果人类在这方面是一致的,那就意味着我们有一个科学有效的问题可以研究。这是一个主要动机。另一个有趣的动机是,我当时有一个学生 David Martin,他之前从事计算机架构工作。他是 Dave Patterson 的学生,Dave Patterson 是伯克利的著名计算机架构师和获奖者。在计算机架构领域,他们非常相信基准测试,因为当任何新硬件出现时,你运行相同的代码片段,尝试分类,得到性能数字。所以这些趋势以某种方式结合在一起。我们做的是:我们说,让我们收集一组图像。这是 2000 年的一堆 CD,这是 Corel 集合。那时,你无法从网上获取图像;网上没有那么多图像,互联网仍处于初期。所以我们买了这套 CD。每张 CD 包含一定数量的图像,你可以把它们加载到计算机上。然后我们可以让本科生来标记物体的边界,并观察一致性。我们用它来实际研究自然图像的统计,并为寻找物体边界这个任务定义人类真实标准。即使不是 100%一致,也有 90%一致。这成为训练机器学习算法的数据。所以我们做了几件事:它促进了数据的使用,促进了基准测试,并促进了对图像作为图像集合、图像流形等的理解。当然,我不想把这种思维模式的功劳全归于自己。逐渐地,这种思维模式在计算机视觉中变得越来越普遍。我也非常感谢 Pietro Perona,他在加州理工学院,以前是我在伯克利的学生。Pietro 领导了从网络收集的第一批图像集合,称为 Caltech 101,用于物体识别。然后英国的 Andrew Zisserman 做了这个,还有 PASCAL 项目,MIT 的 Antonio Torralba,Alyosha Efros,他们也推动了很多数据,强调了数据的重要性。李飞飞带来了 ImageNet。所以在 2000 年到 2010 年之间,范式改变了。过去发表一篇带有数据集的论文非常困难,而到 2010 年,这成为我们做科学的方式。

This takes me back to about 2000 or something like that, or even earlier. The context was a discussion at a conference. I had a little debate at a panel with one of my colleagues called Olivier Faugeras, who was one of the people who studied the geometry of computer vision. He regarded structure from motion and SLAM as worthwhile, respectable problems because they have a clean mathematical formulation. But he regarded problems like segmentation and recognition as kind of of dubious quality, shall we say. In particular, he said: 'Segmentation, okay, everybody writes up, submits a paper, and then they'll have some images and they'll say these are the groups that my system found, and then we have to accept it and we are happy. This is not a way of doing science.' He was very much in the French tradition where you need to have mathematical formalisms and so on. So I came back from that trip and I was determined to prove him wrong. The way to think about it was: how do other sciences which have no precise mathematical models do this? Like psychology and so on, they do experiments with humans. So the natural thing here was to think about experiments with humans. We can show images to people and we can have people mark the boundaries of objects. If humans are consistent in this, that means we have a scientifically valid problem to work on. So this was one big motivation. Another motivation interestingly was that I had a student at that time, David Martin, who had previously worked in computer architecture. He had been a student of Dave Patterson, a famous computer architect and award winner from Berkeley. In the field of computer architecture, they believed a lot in benchmarking because when any new hardware comes, you run the same pieces of code and you try to classify, get the number of drops. So these trends sort of came together. What we did was: we said, let's get a set of images together. Here is a stack of CDs from 2000, this is a Corel collection. At that time, you couldn't get images on the web; there were not that many images on the web, the internet was still in its infancy. So we bought this collection of CDs. Each of these contains a certain number of images, but you can load them on the computer. Then we could have undergrads come and mark the boundaries of objects and see consistency. We used that as a way of actually studying the statistics of natural images and to define for this task of finding boundaries of objects what was human truth. Even if it was not 100% consistent, it was 90% consistent. This became data for training machine learning algorithms. So we did a number of things: it promoted the use of data, it promoted benchmarking, and it promoted understanding the image as a collection of images, the manifold of images, and so forth. Of course, I don't want to take the only credit for this mode of thinking. Gradually, this mode of thinking became more and more common in computer vision. I give also a lot of credit to Pietro Perona, who was at Caltech and formerly my student at Berkeley. Pietro led the collection of the first collection of images from the web, which was called Caltech 101, used for object recognition. Then there was Andrew Zisserman in Britain who did this, and there was the PASCAL program, Antonio Torralba at MIT, Alyosha Efros, who also pushed a lot of data, the importance of data. Fei-Fei Li came in with ImageNet. So sometime between 2000 and 2010, the paradigm changed. It used to be that publishing a paper with a dataset would be really difficult, and by 2010 it became that this is the way we make science now.

ImageNet 上的深度学习突破 Deep Learning Breakthrough on ImageNet

Host

如果你想想现代 AI 时代,一切都由深度学习驱动,而深度学习的重大突破是在 ImageNet 挑战赛上的图像识别,这个基准测试从早期视觉基准测试中发展而来。我听说你是那个故事的一部分。为什么杰夫·辛顿和他的学生在 ImageNet 上运行?你能多说一点吗?

If you think about modern era AI, it's all driven by deep learning, and the big breakthrough in deep learning was in image recognition on the ImageNet challenge, a benchmark that kind of found its way from these early days of starting benchmarking in vision. I hear you've been part of the story there. Why Jeff Hinton and his students ran things on ImageNet? Can you say a bit more about that?

Jitendra Malik

这是个有趣的故事。大约在 2011 或 2012 年,我坐在伯克利校园的办公室里,电话响了。我接了电话;现在通常不接电话,因为通常是随机骚扰电话。我接了电话,是杰夫·辛顿。我从 80 年代就认识他;我们都是该领域的老前辈。杰夫没有以谈论天气开始对话。他的第一个问题是:'你为什么不喜欢深度学习?' 他很直接。所以我想,好吧,我需要直接回应。我说:'哦,因为你没有证明你的观点,你没有拿出证据。' 然后他说:'哦,但我们在 MNIST 和 CIFAR 等上有这些结果。' 我说:'那些是小数据集,我不会被说服。我们需要具有自然图像复杂性的东西。' 所以他说:'哦,好吧。' 我说:'例如,我们有 PASCAL 物体检测挑战。如果你在那上面取得好结果,那会令人印象深刻。' 他说:'好吧,让我想想。' 然后他挂了电话。第二天他回电话说:'PASCAL,我知道你们喜欢 PASCAL,但我认为它没有足够的图像来支持我们的技术。另一方面,ImageNet 出现了,它有更多的图像,我觉得那更有希望。你觉得呢?' 我说:'是的,我会对 ImageNet 的结果印象深刻。' 所以他说:'好的,我回去做。'

This is a fun story. This is somewhere around 2011 or 2012 in this era. I was sitting in my office in Berkeley campus and the phone rang. I picked up the phone; normally you don't pick up phones these days because these are usually some random spam caller. I picked up the phone and it was Geoff Hinton. I know him from ages from the 80s; we are both old-timers in the field. Jeff didn't start the conversation with asking about the weather. Jeff's first question is: 'Why don't you like deep learning?' Right, he was very direct. So I thought, okay, I need to be direct back. So I said: 'Oh, because you haven't made the point, you haven't made the case.' So then he said: 'Oh, but we have these results on MNIST and CIFAR and this and that.' I said: 'Those are little baby datasets, and I'm not going to be convinced by that. We need something with the complexity of natural images.' So he said: 'Oh, okay.' I said: 'For example, we have this PASCAL challenge for object detection. If you get good results on that, then that would be impressive.' So he said: 'Okay, let me think.' Then he hung up. He called me back the next day and said: 'PASCAL, I know you guys like PASCAL, but I don't think it has enough images for the techniques that we have. On the other hand, ImageNet, which has come in, has a lot more images, and I feel that that's more promising. What do you think of that?' I said: 'Yeah, I'd be impressed by ImageNet results.' So he said: 'Okay, I'm going to go back.'

AlexNet 与 GPU 因素 AlexNet and the GPU factor

Jitendra Malik

我会让几个学生研究这个问题,如果有成果就联系你。嗯,俗话说,剩下的就是历史了。所以那是亚历克斯·克里热夫斯基,当然这里有一个隐藏变量,那就是 GPU 已经登场,并且发挥了作用。所以这个成就——我的意思是,有 GPU,有数据,有深度学习将它们整合在一起。这种方法的结果远优于传统的计算机视觉技术。这些结果在佛罗伦萨的 ECCV 研讨会上展示,我记得我们都在那个房间里,那真是见证历史的时刻。亚历克斯展示了他的结果,然后我们进行了一番争论。扬·勒昆在房间里,我也在,阿廖沙·埃夫罗斯也在,我们争论着‘这是真的吗?这意味着什么?’当然,历史已经清楚地说明了后来发生的事。

And I'll make a couple of students work on this problem and if we have something to show I'll call you. Well, as they say, the rest is history. So that was Alex Krizhevsky, and of course there's a hidden variable here, which is that GPUs had come on the scene and they played their role in this. So this accomplishment — I mean, there were GPUs, there was data, there was deep learning putting it all together. And the results of this methodology were much better than the classical computer vision techniques. These results were presented at ECCV in a workshop in Florence, and I remember we were all there in the room, and the times when you say history being made right there. Alex presented his results, and then we had a little argument. Yann LeCun was in the room, I was in the room, Alyosha Efros was in the room, and we were arguing about 'Is this real? What does it mean?' And then, of course, history is very clear about what happened.

未来 AI 应用:医疗与养老 Future AI applications: medicine and elder care

Host

谈谈历史和未来走向。您认为人工智能在不久的将来会走向何方?您认为可能会出现哪些令人兴奋的应用,真正对人类产生巨大影响?

Talk about history and where things are going. Where do you see AI going in the near future? Where do you see maybe exciting applications that could emerge that could really impact humanity in a great way?

Jitendra Malik

我认为显然遍布各个领域,但让我说说我最喜欢的一些。我认为医学、医疗保健——这些非常重要。有很多,这些不会完全交给 AI,而是人与机器协同使用,例如在诊断 X 光图像方面。我选一个最接近我自己能力的例子;我以前做过类似的工作。这些可以非常显著且具有民主化意义。想想血液样本,你可以拍照——现在很容易,相机无处不在。每个人,即使在欠发达国家、贫穷国家,人们也能用上手机和相机。你能以某种方式利用它们吗?我知道印度有一个项目,在资源不足的医院里,给刚出生的婴儿拍照,然后从中可以估计孩子是否发育迟缓等等。同样在美国,18%的 GDP 用于医疗,所以这是一个大领域。我想到了老年人护理。我自己有年迈的父母。我们知道现在退休后人们寿命很长,如果在家中得到帮助,他们的生活质量会大大提高。不是每个人都能负担得起全天候的护理人员;经济上不可能支持。但机器人辅助可以发挥重要作用。所以我选的是贴近我兴趣的视觉和机器人学例子,但当然还有很多其他领域。

I think all across the place, obviously, but let me say some of my favorite ones. So I think medicine, healthcare — I mean these are very important. There are many, and these will not be completely transferred to AI, but sort of the use of a human and a machine together, for example in diagnosing X-ray images. I'm picking an example which is closest to my own competence; I've worked on things like that before. And these can be very remarkable and democratizing. So think of a blood sample which you can take a photo of — that's easy now, cameras are ubiquitous. Everybody, even in a less developed country, in a poor country, people have access to phones and cameras. And can you use them in some ways? There was a project I knew about in India where in underserved hospitals, you take a photo of a kid who's just being born, and from that you can estimate aspects of whether the child is stunted in growth, and so on. And then equally in America, where 18% of our GDP goes to health, so that's a big area. I think of elder care. I have very old parents myself. We know that now post-retirement people live long lives, and the quality of their lives would be improved considerably if they had assistance at home. Not everybody will be able to afford to have attendants all the time; it's just not possible for the economy to support that. But robot assistance could play a major role. So I'm picking examples which are close to my interests of vision and robotics, but of course there are ones all over the place.

AI 与就业替代:历史视角 AI and job displacement: a historical perspective

Jitendra Malik

我想说一件事,因为每当讨论 AI 时,总会讨论 AI 取代工作等等。我个人并不担心。原因是,我认为我们看到了计算机化随时间推移的影响,但系统适应需要很长时间。当过程是渐进的时候,它会经历一个过程:最初是人与机器一起工作,然后机器只是辅助人类,然后人与机器协同工作,然后部分工作完全外包给机器,人类开始做更有创意和有趣的事情,等等。想想编程:曾几何时,人们必须用汇编语言和机器码编程。我在 1970 年代通过按某些按钮来引导程序。然后我们有了更高级的语言,甚至更进一步。现在我们用 GPT 来帮助我们准备大量标准代码。那么什么时候出现过大规模失业?我认为我们从未见过,也不会见到。生产力所做的是,我们现在可以将这些技术应用于更多场景。在 1950 年代,考虑到编程的工作量,这样做是不值得的;有很多应用你根本不会尝试计算机化,但现在我们愿意去做。所以我不是专业经济学家,但作为 AI 研究者,我必须思考我工作的后果,我倾向于认为这一切都将与人类协同,不必以那种方式担心。我看到的好处远多于坏处。

I want to say one thing because whenever there's a discussion of AI, there's a discussion of AI taking jobs and so on. I'm not worried about that personally. The reason is that I think we have seen the effect of computerization over time, but it takes a lot of time for the systems to adapt. When the process is gradual, it goes through a process where initially you have humans and machines working together, then the machine is just assisting the human, then the human and machine work together, then there is some work part which is completely outsourced to the machine, the human then starts to do something more creative and fun, and so on. If you think about programming: once upon a time, people had to program in assembly language and machine code. I have programmed a machine in the 1970s by pushing certain buttons to bootstrap the program. Then we got to higher-level languages, and even more so. Now we use GPT to help us with getting a lot of the standard verbiage of the code ready. So when was there ever this mass unemployment? I don't think we have seen that, and we won't. What the productivity has done is that we could now apply these techniques across many more settings. It would not have been worthwhile to do it in the 1950s given the effort of programming; there were many applications you would simply not have tried to computerize, but we are willing to now. So I'm not a professional economist, but as an AI researcher I have to think about the consequences of my work, and I tend to think that it will all be in conjunction with humans, and it's not something to worry about in that same way. I see a lot more upside than downside.

给 AI 博士生的建议 Advice for PhD students in AI

Host

当一位新的博士生来找您,或者一位潜在的博士生在会议上与您交谈时,如果他们想开启人工智能研究的职业生涯,您会给他们什么建议?

When a new PhD student comes to you or a prospective PhD student chats with you at a conference, what kind of advice do you give them if they want to embark on their AI research career trajectory?

Jitendra Malik

我一直认为人们需要拥有一个技能组合。他们应该在某些方面广泛,在某些方面深入,就像 T 形。我意思是,有时人们称之为 T 型模型。我认为你需要掌握一些技术技能。如果你编写 PyTorch 模型或 TensorFlow 模型,你需要理解机器学习的基础数学、基本编码能力等等。所以有一个基本的技术技能核心是必不可少的。但还有一定的知识广度,我也很欣赏,我总是告诉我的学生要尝试跳出框框思考。如果你在做计算机视觉,也许也读读感知方面的文献,读读神经科学方面的文献,也许和计算机图形学的人聊聊,甚至艺术。所以我认为,因为我们的领域与许多其他学科相连——人工智能广泛地连接——我认为思想更开阔的人能够在更长的时间尺度上做出更多贡献。而且我认为适应性非常重要。如果你坚持‘这是我曾经学过的一项技术,我很擅长,可以用它写 15 篇论文’,如果你一直固守于此,你注定会过时。所以适应性不仅对行走机器人有用,对研究人员也有用。在我的职业生涯中,我不得不多次适应,放弃曾经认为很棒、我真的很擅长并花了多年学习的思维模式,比如微分几何,然后它变得无关紧要。没关系,你知道,你继续做别的事。最重要的是,在任何研究事业中,我认为激情很重要。在我看来,最优秀的研究者,他们热爱自己所做的事情,你必须感受到那种热情。

I always believe that people need to have a portfolio of skills. There is something where they should be broad and some ways in which they should be narrow, like a T. I mean sometimes people call it the T-model. I think there are some technical skills you need to master. If you code a PyTorch model or TensorFlow model, you need to understand the basic math of machine learning, basic coding abilities, and so on. So there is a basic core of technical skills that are essential. But then there is a certain breadth of knowledge which I also appreciate, and I always tell my students to try to think outside the box. If you are doing computer vision, maybe also read the literature on perception, maybe read the literature on neuroscience, maybe talk to people in computer graphics, even art. So I think because our field connects to so many other disciplines — AI broadly does — I think people who are more broad-minded can contribute more over a long scale of time. And I think it's very important to be very adaptable. If you insist that 'here is a technique which I once learned and I was good at it and I could write 15 papers with it,' and if you just stay with that, you're doomed to extinction. So being adaptable is useful not just for walking robots but also for researchers. I have in my career had to adapt many times, given up modes of thought which I regarded as great at once, where I was really good at and I spent years studying differential geometry, and then it became irrelevant. Fine, you know, you do one thing. And above all, in any research enterprise, I think passion is important. I think the best researchers, in my view, they sort of love what they're doing, and you sort of have to feel that.

印度早期生活与教育 Early Life and Education in India

Host

我们来聊聊你的经历。据我所知,你在印度长大。你在印度哪里长大的?

Let's talk about your trajectory. As I understand it, you grew up in India. Where in India did you grow up?

Jitendra Malik

我在印度中部一个叫贾巴尔普尔的城市长大。我在那里读完高中,然后去了坎普尔的印度理工学院,位于德里附近。我本科专业是电气工程。那是 1975 年到 1980 年。当时计算机还不常见。我因为读到关于计算机的文章而兴奋,还上了几门课。我开始思考人工智能,因为我在《科学美国人》上读到文章,心想,哇,这是一个宏伟的目标。然后我申请了研究生院,意外地被斯坦福录取了。我很惊讶他们录取了我。

I grew up in a city called Jabalpur, which is in central India. I finished high school there and then went to IIT Kanpur, which is in the north near Delhi. I was an electrical engineering major as an undergrad. This was from 1975 to 1980. At that time, computers were still not so common. I got excited by reading about computers and took a few classes. I got into thinking about AI because I read articles in Scientific American and thought, wow, this is a grand goal. Then I applied to graduate school and got into Stanford by some fluke. I was very surprised they admitted me.

斯坦福博士:从逻辑到视觉 PhD at Stanford and Switch from Logic to Vision

Jitendra Malik

我刚到斯坦福时学的是编程语言和体系结构,但一个学期后我转到了人工智能。我跟约翰·麦卡锡一起工作,他创造了“人工智能”这个词。他是我的导师。我阅读了经典人工智能的文献,斯坦福当时非常注重基于逻辑的人工智能。麦卡锡认为,我们将通过使用一阶逻辑表示世界事实,然后进行逻辑推理来解决智能问题。有一个问题叫做“资格问题”,麦卡锡认识到了这一点:在常识推理中,新事实会改变你对旧事实的信念。例如,你可能会说鸟会飞,但如果你知道它是企鹅,你就会改变信念。这对概率推理来说不是问题,但在经典逻辑中却是问题。麦卡锡让我研究这个问题。这非常困难,几年后我感到沮丧。然后我遇到了罗德·布鲁克斯,他在喝酒时告诉我所有这些逻辑的东西都是胡说八道,我应该研究视觉或机器人学。他告诉了我莫拉韦克悖论。这让我很震惊,因为我一直相信逻辑和推理很重要,但他说重要的是猫能做什么。我被说服了,几个月内我放弃了麦卡锡的论文,转向了汤姆·宾福德的计算机视觉。我在两年半后放弃了博士学位,但这是我做出的最好的决定之一。

When I started at Stanford, I was in programming languages and architecture, but after a quarter I switched to AI. I worked with John McCarthy, who coined the term AI. He was my advisor. I read the literature of classical AI, which at Stanford was very much logic-based. McCarthy believed we would solve intelligence by representing facts about the world using first-order logic and then doing logical inference. There was a problem called the qualification problem, which McCarthy recognized: in common sense reasoning, new facts change your belief in old facts. For example, you might say a bird can fly, but if you learn it's a penguin, you change your belief. This is not a problem for probabilistic reasoning, but in classical logic it is. McCarthy assigned me to work on this problem. It was very hard, and after a few years I got frustrated. Then I met Rod Brooks, who told me over a beer that all this logic stuff is nonsense, and I should work on vision or robotics. He told me about Moravec's paradox. It shook me because I believed logic and reasoning were important, but he said what matters is what a cat can do. I got convinced, and within a few months I abandoned my thesis with McCarthy and switched to computer vision with Tom Binford. I abandoned my PhD two and a half years in, but it was one of the best decisions I made.

转至伯克利与建设 AI 项目 Move to Berkeley and Building AI Program

Host

据我所知,你从斯坦福来到了伯克利。你是伯克利最早的人工智能教授之一,至少在当时是现代人工智能教授。

From Stanford you came to Berkeley, as I understand. You were one of the first AI professors at Berkeley, or at least modern AI professors at the time.

Jitendra Malik

是的,没错。当时有尚卡尔·萨斯特里,一位转向机器人学的控制理论家,还有罗伯特·维伦斯基和洛特菲·扎德等人代表经典人工智能。我是被聘用的年轻人。第二年斯图尔特·罗素被聘用,再下一年约翰·坎尼被聘用。那是伯克利人工智能和机器人学的开端。这很有趣,因为我们都是年轻的助理教授,试图在一个之前在这个领域没有太多活动的系里创建一个领域。

Yes, that's right. At that time, there was Shankar Sastry, a control theorist who switched to robotics, and others like Robert Wilensky and Lotfi Zadeh representing classical AI. I was the young kid hired. The next year Stuart Russell was hired, and the following year John Canny was hired. That was the beginning of AI and robotics at Berkeley. It was fun because we were all young assistant professors trying to create a field in a department that didn't have much activity in this area before.

工业界与学术界协同 Industry vs Academia Synergy

Host

你培养了超过 70 名博士生和博士后,他们去了顶尖大学和行业实验室。你也在工业界担任过职务。你怎么看待工业界和学术界在人工智能研究上的协同作用?

You have graduated over 70 PhD students and postdocs to top universities and industry labs. You've also taken positions in industry. How do you see the synergy between industry and academic research for AI?

Jitendra Malik

这是一个严肃的问题。在最近的 CVPR 上,我们有一个小组讨论,一些学术同事担心学术研究正在失去相关性,因为像 GPT-3 或 GPT-4 这样的模型无法在学术实验室中训练。我认为学术界和工业界并不冲突,而是互补的。最明显的方式是,学术界吸纳原始人才——一个刚开始的博士生——在五年内把他们培养成知识渊博、能够出去做大事的人。工业界离不开学术界培养的人才。工业界的研究有多种模式。我在谷歌和 Meta(当时是 Facebook AI)待过。

This is a serious question. At the last CVPR, we had a panel discussion where some academic colleagues worried that academic research is losing relevance because models like GPT-3 or GPT-4 cannot be trained in an academic lab. I believe academia and industry are not in conflict but complementary. The most obvious way is that academia takes raw talent—a beginning PhD student—and over five years makes them into someone who knows a lot and can go out and do great things. Industry cannot survive without the people academia trains. The kind of research that can happen in industry has various modes. I spent time at Google and at Meta (then Facebook AI).

工业研究模式 Modes of Industrial Research

Jitendra Malik

工业研究有不同的模式。一种模式像学术界:独立研究者,可能两三个人一起合作,他们有个想法,实现它,做实验,写论文。但还有另一种可能性,就是更大的团队。这是 DeepMind 和 OpenAI 采用过的模式,一个大型团队共同攻克一个项目,集中大量资源,形成临界质量,从而构建出更宏大的成果。在我看来,这两种模式都很重要。我喜欢把它们称为小科学和大科学。这跟物理学有类比之处:有些人可能看过电影《奥本海默》。其中一个重要角色是欧内斯特·劳伦斯,他是伯克利的物理学教授,建造了回旋加速器,开创了大科学的传统。在 1930 年之前,物理学论文从来没有多个作者——可能一个或两个。而现在,高能物理的论文有上千个作者。这始于劳伦斯在伯克利的工作,因为要建造能足够加速粒子的设备,人们必须合作;他们无法再独自工作。我认为我们的领域也是如此。我们需要承认,AI 既有小科学的成分,也有大科学的成分,两者都需要推进。这不是非此即彼。小科学更容易进行探索,大科学更容易进行利用。根据目前的架构,工业界会有更多小科学,公司里会有更多大科学,但这并非必然。在物理学中,大型强子对撞机的工作完全是学术研究,但它是大科学,由大学联盟完成。所以我认为它们是互补的,我两者都重视。只关注其中一个会是错误的。

There are different modes of industrial research. One mode is like academia: individual investigators, maybe two or three people working together, they have an idea, implement something, do experiments, write a paper. But you also have the possibility of bigger teams. This is a mode that DeepMind and OpenAI have had, where a large team works on one project, putting a lot of wood behind that arrow, achieving critical mass to build something much bigger. In my view, both modes are important. I like to think of it as small science and big science. There's an analogy to physics: some of you may have seen the movie Oppenheimer. A prominent character is Ernest Lawrence, a physics professor at Berkeley who built a cyclotron and started the tradition of big science. Before 1930, physics papers never had multiple authors—maybe one or two. Now, papers in high energy physics have a thousand authors. It started with Lawrence at Berkeley because to build equipment to accelerate particles fast enough, people had to get together; they could no longer work alone. I think the same is true for our field. We need to acknowledge that AI has both a small science component and a big science component, and both need to be pursued. It's not either/or. Small science can do exploration more readily; big science can do exploitation more readily. Given the current structure, there will be more small science in industry and big science in companies, but it's not necessary. In physics, work on the Large Hadron Collider is all academic research but it is big science, with a consortium of universities. So I believe they are complementary, and I value both. Focusing on just one or the other would be a mistake.

研究热情与放松 Passion for Research and Relaxation

Host

你谈到对研究的热情。我觉得我从未见过你不充满激情、不忙于研究或教学。但你会休息吗?你做什么来放松?

You talk about passions for research. I feel like I've never encountered you not being passionate and pretty busy getting research done or teaching. But do you ever take time off? What do you do to relax?

Jitendra Malik

我喜欢散步,所以有时就出去走走。我喜欢去博物馆。我喜欢旅行,虽然不常有时间。我喜欢去罗马、伦敦、巴黎等地的博物馆。我喜欢阅读各种东西。这些可能是我最主要的户外活动:散步、徒步、旅行、阅读。可惜我不擅长任何运动。我可以听音乐,但不会创作。

I like walking, so sometimes just going on a walk. I like going to museums. I like traveling, though I don't often have time for that. I like going to museums in Rome, London, Paris, and things like that. I like reading about lots of things. These are probably my biggest outside activities: walking, hiking, traveling, reading. I'm not good at any sports, unfortunately. I can listen to music but cannot create it.

Host

如果你擅长运动,也许你就不会在这里做 AI 研究了,所以这可能是件好事。你阅读时,是读小说,还是其他学科的科学书籍?你读什么?

If you had been good at sports, maybe you wouldn't have been here doing AI research, so it might have been a blessing. When you're reading, are you reading fiction, scientific books in other disciplines? What are you reading?

Jitendra Malik

我相当杂食。我什么都读。我读过很多小说,比如推理小说。福尔摩斯——我能随口引用福尔摩斯里的句子。我非常喜欢读历史,觉得很有趣。我也喜欢科学史。让我为此做个论证,以及为什么我鼓励学生读历史。历史给了我们最近邻。在生活的许多方面,你永远不会遇到完全相同的情况重复,但你可以找到最近邻。所以如果你身处某个情境,你可以想:‘在某个时间点有过类似的情况。’这在人类历史中成立——王国、帝国、战争、革命——在科学史中也成立。我们看着物理学这样的领域,觉得它如此优美和成熟,但 300 年前的物理学并非如此。生物学在 50 或 100 年前也不是这样。那时事情非常不确定,有争论,有完全错误的想法。看到这些,给了我关于 AI 的灵感和信心。今天 AI 人人都在谈论;我们在报纸上看到它。但我在 80 年代开始进入这个领域时,它只是一个小学科,在计算机科学系里被容忍,因为大多数人认为那只是空谈,永远不会成真。你如何有勇气在一个尚未能交付成果的领域坚持下去?通过历史类比:说,‘哦,这些其他领域也曾是萌芽状态,也不成功。’当我试图预测 AI 的未来时,我也这样做。有时我们担心我们过于快速外推,认为突然就会发生。或者有这样一种说法:短期内我们总是高估事情完成的速度,长期内我们低估,因为发现来自我们从未梦想过的地方。这种历史视角支撑了我的研究;我从中获得灵感。此外,历史很有趣——很多酷故事,而且它们真的发生过。

I'm pretty omnivorous. I read everything. I've read a lot of fiction, like mystery novels. Sherlock Holmes—I can quote lines from Sherlock Holmes readily. I like reading history a lot; I find it a lot of fun. I also like scientific history. Let me give an argument for that and why I encourage students to read history. History gives us nearest neighbors. In many aspects of life, you never have exactly the same situation repeat, but you can find nearest neighbors. So if you are in a situation, you can think, 'Here was a similar situation faced at this point in time.' This is true in human history—kingdoms, empires, wars, revolutions—and also in the history of science. We look at a field like physics and think it's so beautiful and mature, but physics wasn't like that 300 years ago. Biology wasn't like that 50 or 100 years ago. Things were very uncertain, there were debates, absolutely wrong ideas. Looking at that gives me inspiration and confidence about AI. AI today is on everybody's tongue; we see it in newspapers. But when I started in the field in the 80s, it was a little discipline, kind of tolerated in computer science departments because most thought it was just talk and would never be real. How do you have the confidence to persist in a field when it's not yet able to deliver? By historical analogy: saying, 'Oh, these other fields were also embryonic and didn't work.' When I try to project the future of AI, I do the same. Sometimes we worry we extrapolate too quickly, thinking suddenly this will happen. Or there's this line that in the short term we always overestimate how quickly things will get done, and in the long term we underestimate because discoveries come from somewhere we never even dreamed of. That historical perspective grounds my research; I get inspiration from it. Besides, history is fun—lots of cool stories that actually happened.

Host

我真正喜欢读的是传记。那是历史,但我把它归为同一类。好了,Jitendra,非常感谢你抽出时间。我真的很享受这次对话。

The thing I really enjoy reading is biographies. That is history, but I put that in the same category. Well, Jitendra, thanks so much for making the time. I really enjoyed this conversation.

Jitendra Malik

彼此彼此。不,这太棒了。Peter,谢谢你抽出时间。我非常开心。

Likewise. No, it was wonderful. Peter, thanks for taking the time. I enjoyed myself immensely.

互动版:逐字朗读 + 针对本期提问 →