Yann LeCun:LLM 的智能不如家猫

Yann LeCun: LLMs Less Intelligent Than a House Cat

杨立昆 Yann LeCun · AI Inside 播客 · 2025-04-09 · 约 46 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Meta 首席 AI 科学家 Yann LeCun 探讨为何当前 LLM 的智能不如家猫,发展理解物理现实的世界模型仍是 AI 最大挑战,以及 Meta 开源 Llama 如何赋能数千家公司而仅颠覆少数几家。

Meta's chief AI scientist Yann LeCun discusses why current LLMs are less intelligent than a house cat, the need for world models to understand physical reality, and how Meta's open-source Llama approach enables thousands of companies while disrupting only a few.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 13)

全文 · Full transcript(中英对照)

引言与 LLM 局限 Introduction and LLM limitations

Host

Meta 首席 AI 科学家、图灵奖得主 Yann LeCun 加入我们,讨论为什么当前的 LLM 还不如一只家猫聪明。如何开发理解物理现实的世界模型仍然是 AI 最大的未解挑战。以及为什么 Meta 对 Llama 的开源方法正在赋能数千家公司,同时只颠覆了三家。这些内容即将呈现。大家好,欢迎收听 AI Inside,一档关注科技世界中无处不在的 AI 的节目。我是主持人之一 Jason Howell。本周的节目确实有点不同,我想直接进入正题,但在此之前,衷心感谢那些在 Patreon 上直接支持我们的人。那是 patreon.com/inside。Corky Garco,你知道你是谁。你是我们的重要支持者,我们非常感谢你。好的,我和联合主持人 Jeff Jarvis 有机会与 AI 界非常著名的人物 Yann LeCun 进行了交谈。这发生在上周五,也就是这次采访录制的时间,非常精彩。所以我不再浪费时间,直接开始。非常激动地欢迎 AI Inside 的嘉宾 Yann LeCun,Meta 首席 AI 科学家、图灵奖得主,被许多人称为 AI 之父。欢迎来到节目,Yann。很高兴见到你。

Meta's chief AI scientist and touring award winner Yann LeCun joins us to talk about why current LLMs are less intelligent than a house cat. How developing world models that understand physical reality remains AI's biggest unsolved challenge. And why Meta's open-source approach to Llama is enabling thousands of companies while disrupting just three. That's coming up right after this. Hello and welcome to AI Inside, the show where we take a look at the AI that's layered throughout so much of the world of technology. I'm one of your hosts, Jason Howell. And this week's episode is definitely a little different, and I want to get right to it, but just real quick before we do, huge thank you to those of you who support us directly on Patreon. That's patreon.com/inside. Corky Garco, you know who you are. You're a huge supporter and we appreciate you. So, thank you so much for that. All right, Jeff Jarvis, my co-host, and I had the chance to chat with Yann LeCun, very notable figure in the world of AI. This happened last Friday. That's when this interview was recorded, and it was pretty incredible. So, I'm not going to waste any time. Let's just jump right into it right now. Thrilled to welcome to AI Inside Yann LeCun, chief AI scientist at Meta, touring award winner, known by many as the godfather of AI. Welcome to the show, Yann. It's really nice to meet you.

Yann LeCun

谢谢邀请。

Thanks for having me on.

Host

听到别人介绍你是 AI 之父,会不会觉得腻了?你可能会想,好吧,又来了。

Does it ever get old hearing someone introduce you as the godfather of AI? You're kind of like, yeah, here we go again.

Yann LeCun

我捂住耳朵,这样就不会脸红了。但到了这个地步,你可以接受它,因为这是事实。

I shut my ears so I don't turn red. But you can accept it at this point because it's the truth.

Host

我要开始的问题是这样的:我们深深地嵌入了当前的人工智能领域,这似乎就是 LLM 时代,而且可能还有新事物即将出现,但我们仍然深陷其中。你对 LLM 的局限性发表了相当多的看法,而与此同时,我们看到像 OpenAI 这样的公司获得了创纪录的融资,这很大程度上建立在其 LLM 技术的成功之上。所以我看到一边是收益递减,另一边是公司把一切都押注在生成式 AI 和 LLM 上。我很好奇你的想法,为什么他们可能没有看到你所看到的这项技术的问题?或者也许他们看到了,只是处理方式不同。你怎么看?

The kind of question that I have to kick things off is that we are so firmly implanted into the current realm of artificial intelligence, which really seems to be the LLM generation, and there's probably something on the horizon around that, but we're still firmly implanted in there. And you've been pretty opinionated on the limits of LLMs at a time when we're also seeing things like OpenAI securing a record-breaking round of funding largely built on its success in LLM technology. And so I see diminishing returns on one side, on the other companies betting everything on generative AI and LLMs. I'm curious to know what you think as far as why they might not be seeing what you're seeing about this technology. Or maybe they are, they're just approaching it differently. What are your thoughts there?

Yann LeCun

哦,也许他们看到了。毫无疑问,LLM 是有用的,特别是对于编程助手之类的东西。未来可能用于更通用的 AI 助手工作,人们正在谈论智能体系统,但它仍然不完全可靠。对于这类应用,主要问题——也是 AI 和计算机技术更普遍的一个反复出现的问题——是你可以看到令人印象深刻的演示,但当真正部署一个足够可靠的系统,交到人们手中并让他们日常使用时,还有很大的距离。让这些系统足够可靠要困难得多。我的意思是,10 年前我们看到汽车在乡村道路上自动驾驶的演示,大约 10 分钟后你就得干预,我们取得了很大进展,但仍然没有达到像人类一样可靠的自动驾驶汽车,除非我们作弊,这正是我们和其他人正在做的。所以在过去 70 年里,AI 领域反复出现这样的历史:人们提出一个新范式,然后声称‘好了,就是它了,这将在 10 年内带我们达到人类水平的 AI,地球上最智能的实体将是一台机器。’而每一次都被证明是错误的,因为新范式要么遇到了人们没有看到的局限性,要么被证明非常擅长解决某一子类问题,而不是通用智能问题。所以一代又一代的 AI 研究人员、实业家和创始人都在做出这些声称,而他们每次都错了。所以我不想贬低 LLM。它们非常有用。应该对它们进行大量投资,也应该对运行它们的基础设施进行大量投资,实际上大部分资金都流向了那里——不是用于训练,而是用于运行,可能服务数十亿用户。但就像其他计算机技术一样,即使不是人类水平的智能,它也可以是有用的。现在,如果我们想要追求人类水平的智能——我认为我们应该这样做——我们需要发明新技术。我们离达到那个水平还差得远。

Oh, maybe they are. There's no question that LLMs are useful, particularly for coding assistants and stuff like that. And in the future probably for more general AI assistant jobs, people are talking about agentic systems, but it's still not totally reliable yet. It's a bit like, for this kind of applications, the main issue, and it's been a recurring problem with AI and computer technology more generally, is the fact that you can see impressive demos, but when it comes time to actually deploy a system that's reliable enough that you put it in the hands of people and they use it on a daily basis, there's a big distance. It's much harder to make those systems reliable enough. I mean, 10 years ago we were seeing demos of cars driving themselves in countryside streets for about 10 minutes before you had to intervene, and we made a lot of progress, but we're still not to the point of having cars that can drive themselves as reliably as humans, except if we cheat, which is what we and others are doing. So there's been a repeated history over the last 70 years in AI of people coming up with a new paradigm and then claiming, 'Okay, that's it, this is going to take us to human-level AI within 10 years, the most intelligent entity on the planet would be a machine.' And every time it's turned out to be false because the new paradigm either hits a limitation that people didn't see or turned out to be really good at solving a subcategory of problem that didn't turn out to be the general intelligence problem. So there's been generation after generation of AI researchers and industrialists and founders making those claims and they're being wrong every time. So I don't want to poo-poo LLMs. They're very useful. There should be a lot of investment in them, and there should be a lot of investment in infrastructure to run them, which is where most of the money is going actually—not to train them, but to run them, serving billions of users potentially. But like every other computer technology, it can be useful even if it's not human-level intelligence. Now if we want to shoot for human-level intelligence, and I think we should, we need to invent new techniques. We're just nowhere near matching that.

Host

非常感激你能来,Yann,因为我在这个节目和其他地方经常引用你的话,因为你是 AI 领域现实主义的声音。我没有听到你像其他地方那样大肆宣传。你非常清楚地说明了我们现在所处的位置。我想你曾把我们比作可能达到了一只聪明的猫或一个三岁小孩的水平。甚至还不完全是。所以你说过我们已经触及了 LLM 所能做的极限。所以下一个范式、下一个飞跃,我想你谈到过更好地理解现实。但你能谈谈你认为研究下一步应该走向哪里,我们应该把资源投向哪里以从 AI 中获得更多吗?

I'm really grateful you're here, Yann, because I quote you constantly on this show and elsewhere because you are the voice of realism in AI. I don't hear you spouting the hype that I hear elsewhere. And you've been very clear about where we are now. I think you've equated us to maybe getting to the point of a smart cat or a three-year-old. Not even exactly. And so you've talked about we've hit kind of the limits of what LLMs can do. So there is a next paradigm, a next leap, and I think you've talked about understanding reality better. But can you talk about where you think research should go next, where we should be putting resources next to get more out of AI?

Yann LeCun

三年前我写了一篇长论文,解释了我认为 AI 研究在未来 10 年应该走向何方。那是在世界了解 LLM 之前,当然我知道 LLM,因为我们之前就在研究它,但这个愿景没有改变。它没有受到 LLM 成功的影响。事情是这样的。我们需要能够理解物理世界的机器。我们需要能够推理和规划的机器。我们需要具有持久记忆的机器。我们需要这些机器是可控且安全的。这意味着它们需要由我们给它们的目标驱动。我们给它们一个任务,它们完成它,或者给我们问题的答案,就这样。它们不能逃脱我们要求它们做的任何事情。

So I wrote a long paper three years ago where I explained where I think AI research should go over the next 10 years. This was before the world learned about LLMs, and of course I knew about it because we were working on it before, but this vision hasn't changed. It's not been affected by the success of LLMs if you want. So here's the thing. We need machines that understand the physical world. We need machines that are capable of reasoning and planning. We need machines that have persistent memory. And we need those machines to be controllable and safe. Which means that they need to be driven by objectives we give them. We give them a task, they accomplish it or they give us the answer to the question we ask, and that's it. And they can't escape whatever it is that we're asking them to do.

世界模型概念 World Model Concept

Yann LeCun

所以我在那份文件中解释的是,我们如何可能达到那个点,核心概念是世界模型。我们头脑中都有世界模型,动物也有。它是让我们能够预测世界将要发生什么的心理模型,无论是世界自身的变化还是我们可能采取的行动。如果我们能预测行动的后果,那么我们可以设定一个目标、一个任务,利用我们的世界模型想象特定行动序列是否真的能实现那个目标。这让我们能够规划。规划和推理实际上就是操纵我们的心理模型,来判断特定行动序列是否能完成我们设定的任务。这就是心理学家所说的系统二,一种关于如何完成任务的深思熟虑的过程。我们并不真正知道如何做到这一点;我们在研究层面取得了一些进展。该领域许多最有趣的研究是在机器人学的背景下进行的,因为当你需要控制机器人时,你需要提前知道向手臂发送扭矩的效果。在控制理论和机器人学中,这种想象行动序列后果然后通过优化搜索满足任务的行动序列的过程有一个名称:模型预测控制(MPC)。这是最优控制中一个非常经典的方法,可以追溯到几十年前。主要问题是,在机器人学和控制理论中,世界模型是由工程师编写的一组方程。你想控制机器人手臂或火箭,你可以写下动力学方程。但对于 AI 系统,我们需要这个世界模型从经验或观察中学习。这似乎是动物和人类(婴儿通过观察学习世界如何运作)头脑中发生的过程。这部分似乎非常难以复现。

So what I explained in that document is how we might potentially get to that point, centered on a central concept called a world model. We all have world models in our head, and animals do too. It's the mental model that allows us to predict what's going to happen in the world, either because the world is being the world or because of an action we might take. If we can predict the consequences of our actions, then we can set ourselves an objective, a goal, a task to accomplish, and using our world model, imagine whether a particular sequence of actions will actually fulfill that goal. That allows us to plan. Planning and reasoning really is manipulating our mental model to figure out if a particular sequence of actions is going to accomplish a task we set for ourselves. That is what psychologists call system two, a deliberate process of thinking about how to accomplish a task. We don't know how to do that really well; we are making some progress at the research level. A lot of the most interesting research in that domain is done in the context of robotics, because when you need to control a robot, you need to know in advance what the effect of sending a torque on an arm is going to be. This process in control theory and robotics of imagining the consequences of a sequence of actions and then by optimization searching for a sequence of actions that satisfies a task has a name: model predictive control, MPC. It's a very classical method in optimal control going back decades. The main issue is that in robotics and control theory, the world model is a bunch of equations written by an engineer. You want to control a robot arm or a rocket, you can write down the dynamical equations. But for AI systems, we need this world model to be learned from experience or observation. This is the kind of process that seems to take place in the minds of animals and maybe humans, infants learning how the world works by observation. That part seems really complicated to reproduce.

自监督学习与 LLM Self-Supervised Learning and LLMs

Yann LeCun

这可以基于一个非常简单的原则,人们已经尝试了很久但没有太大成功,称为自监督学习。自监督学习在自然语言理解和 LLM 方面取得了巨大成功。事实上,它是 LLM 的基础。你拿一段文本,训练一个大型神经网络来预测文本中的下一个词。基本上就是这样。有一些技巧可以让它更高效,但这就是 LLM 的基础。你训练它预测下一个词,然后在使用时,让它预测下一个词,将预测的词移入它的视野窗口,然后预测第二个词,再移入,预测第三个。这就是自回归预测。这就是 LLM 的基础。所有的技巧都关于你能花多少钱雇人来微调它,以便它能正确回答问题。这就是目前大量资金投入的地方。

This can be based on a very simple principle which people have been playing with for a long time without much success, called self-supervised learning. Self-supervised learning has been incredibly successful in the context of natural language understanding and LLMs. In fact, it's the basis of LLMs. You take a piece of text and train a big neural net to predict the next word in the text. That's basically what it comes down to. There are tricks to make this efficient, but that's the basis of an LLM. You train it to predict the next word, then when you use it, you have it predict the next word, shift the predicted word into its viewing window, and then predict the second word, then shift that in, predict the third. That's autoregressive prediction. That's what LLMs are based on. All the tricks are about how much money you can afford to hire people to fine-tune it so it can answer questions correctly. That's what a lot of money is going into right now.

视频自监督学习 Applying Self-Supervised Learning to Video

Yann LeCun

你可以想象利用自监督学习的原则来学习图像的表示,学习预测视频中将要发生的事情。如果你给计算机播放一个视频,训练一个大型神经网络来预测视频中接下来会发生什么,如果系统能够学会并很好地完成预测,它可能已经理解了很多关于物理世界的基本性质。比如物体按照特定规律运动,有生命的物体可以以更不可预测的方式运动,但仍然满足一些约束。你不会看到没有支撑的物体因为重力而掉落。人类婴儿需要九个月来学习重力。这是一个漫长的过程。年幼的动物学得更快,但它们最终对重力的掌握程度不同。猫和狗在这方面非常擅长。那么我们如何复现这种训练呢?如果我们做天真的做法,也就是我研究了 20 年的,拿一个视频训练系统预测接下来会发生什么,这实际上行不通。如果你训练它预测下一帧,它学不到任何有用的东西,因为太容易了。如果你训练它预测更长时间,它真的无法预测会发生什么,因为有很多可能发生的事情。在文本的情况下,这是一个简单的问题,因为字典中的单词数量有限,所以你永远无法精确预测下一个单词,但你可以预测所有单词的概率分布。这已经足够了。你可以在预测中表示不确定性。但你不能用视频做到这一点。我们不知道如何表示所有图像或视频帧或视频片段的适当概率分布。这实际上是一个数学上难以处理的问题。这不仅仅是计算机不够大的问题;它本质上是难以处理的。

You could imagine using this principle of self-supervised learning for learning representations of images, learning to predict what's going to happen in a video. If you show a video to a computer and train some big neural net to predict what's going to happen next in the video, if the system is capable of learning this and doing a good job at that prediction, it will probably have understood a lot about the underlying nature of the physical world. Things like objects move according to particular laws, animate objects can move in more unpredictable ways but still satisfy some constraints. You're not going to have objects that are not supported fall because of gravity. Human babies take nine months to learn about gravity. It's a long process. Young animals learn this much quicker, but they don't have the same kind of grasp of gravity in the end. Cats and dogs are really good at this. So how do we reproduce this kind of training? If we do the naive thing, which I've been working on for 20 years, of taking a video and training a system to predict what happens next in a video, it doesn't really work. If you train it to predict the next frame, it doesn't learn anything useful because it's too easy. If you train it to predict longer term, it really cannot predict what's going to happen because there are many plausible things that might happen. In the case of text, it's a simple problem because you have a finite number of words in the dictionary, so you can never predict exactly what word follows, but you can predict a probability distribution over all words. That's good enough. You can represent uncertainty in the prediction. You can't do this with video. We do not know how to represent an appropriate probability distribution over the set of all images or video frames or video segments. It's actually a mathematically intractable problem. It's not just a question of not having big enough computers; it's intrinsically intractable.

联合嵌入预测架构 Joint Embedding Predictive Architecture

Yann LeCun

直到五六年前,我对此没有任何解决方案。我认为没有人有任何解决方案。我们想出的一个解决方案是一种改变我们做法的架构。我们不预测视频中发生的一切,而是训练一个系统学习视频的表示,并在那个表示空间中进行预测。这种表示消除了视频中许多不可预测或无法确定的细节。这种架构称为联合嵌入预测架构。可能令人惊讶的是,它不是生成式的。每个人都在谈论生成式 AI。我的直觉是,下一代 AI 系统将基于非生成式模型。

Until maybe five or six years ago, I didn't have any solution to this. I don't think anybody had any solution. One solution we came up with is a kind of architecture that changes the way we would do this. Instead of predicting everything that happens in the video, we basically train a system to learn a representation of the video and we make the prediction in that representation space. That representation eliminates a lot of details in the video that are just not predictable or impossible to figure out. That kind of architecture is called a joint embedding predictive architecture. What may be surprising about this is that it's not generative. Everybody is talking about generative AI. My hunch is that the next generation AI system will be based on non-generative models.

AGI 时间线与定义 AGI timeline and definition

Host

本质上,听你谈到我们当前真正的局限性时,我想到的是:我们正处在 AGI(通用人工智能)的边缘。原因如下——这取决于你问谁,对吧?有些人说‘它就在眼前’,另一些人说‘哦,它已经在这里了,看看这个,多神奇?’但当你谈论时,其他人又说它永远不会到来。我觉得在这个节目上我们经常带着怀疑讨论这个话题,而你刚才说的让我更加确信 AGI 并非近在咫尺,它其实是一个遥远的理论。

Essentially, what occurs to me in hearing you talk about the real limitations of where we're at. We're right on the precipice of AGI, artificial general intelligence. And here's the reason why. It depends on who you ask, right? Some people are like, 'It's right around the corner.' Other people are like, 'Oh, it's already here. Take a look at this. Isn't that amazing?' But when you're talking and others say it'll never be here. I think often on this show we talk about this topic a little bit in disbelief and I think what you just said kind of punctuates that for me. When you put it that way, it really makes me more confident that AGI is not right around the corner. That AGI is really this distant theory.

Yann LeCun

首先,我毫不怀疑在未来某个时刻,我们会拥有至少在人类擅长的所有领域都像人类一样聪明的机器。这不是问题。很多人对此有巨大的哲学疑问。很多人仍然相信人性是不可捉摸的,我们永远无法将其简化为计算。我在这方面不是怀疑论者。我毫不怀疑在某个时刻我们会拥有比我们更聪明的机器。它们在狭窄领域已经做到了。那么问题就是:AGI 到底意味着什么?你是指像人类智能一样通用的智能吗?如果是这样,你可以用这个词,但它非常误导人,因为人类智能根本不通用。它极其专门化。我们被进化塑造,只做那些值得为生存完成的任务。我们认为自己拥有通用智能,但我们根本不通用。只是所有我们无法理解的问题,我们都能想到它们,这让我们相信自己拥有通用智能,但我们绝对没有。所以我认为这个词是胡说八道,非常误导人。我更喜欢我们在 Meta 用来指代人类水平智能概念的词:AMI,高级机器智能。这是一个更开放的概念。我们实际上发音为‘ami’,在法语中是朋友的意思。但如果你愿意,我们可以称之为人类水平智能。毫无疑问它会实现。它不会在明年发生,也不会在两年后发生。它可能在未来十年内达到某种程度。所以并不遥远。如果我们目前正在做的所有事情都成功,那么也许十年内我们就能很好地判断是否能达到那个目标。但这几乎肯定比我们想象的要难,而且可能难得多,因为在 AI 历史上它总是比我们想象的要难。所以我乐观。我不是那些说我们永远达不到的悲观主义者。我不是那些说我们现在做的一切都无用的悲观主义者。那不是真的。它非常有用。我不是那些说我们需要量子计算或全新原理的人。不,我认为它基本上会基于深度学习,而且这个底层原理会伴随我们很长时间。但在这个领域内,我们需要发现和实现的东西,我们还没有达到。我们缺少一些基本概念。说服自己这一点最好的方法是:我们有系统可以回答互联网上任何有答案的问题。我们有系统可以通过律师资格考试,这很大程度上是信息检索。我们有系统可以缩短文本并帮助我们理解,可以批评我们写的文章,可以生成代码——但生成代码相对简单,因为语法很强且很多是正确的。我们有系统可以解方程,可以解决它们被训练过的问题。但如果它们遇到一个全新的问题,当前系统就是找不到解决方案。最近有一篇论文显示,如果你用所有最好的大语言模型测试最新的数学奥林匹克竞赛题,它们基本上得零分,因为那是它们没被训练过的新问题。所以我们有能操纵语言的系统,这让我们误以为它们很聪明,因为我们习惯于聪明人能巧妙地操纵语言。但我的家用机器人在哪里?我的五级自动驾驶汽车在哪里?能像猫一样做事的机器人在哪里?甚至模拟的能像猫一样做事的机器人?问题不在于我们造不出机器人。我们能造出具有物理能力的机器人。只是我们不知道如何让它们足够聪明。处理现实世界和产生行动比理解语言难得多。这又回到了我之前提到的问题:语言是离散的,有很强的结构。现实世界是一团乱麻,不可预测,非确定性,高维,连续。所以让我们尝试建造一个能像猫一样快速学习的东西吧。

First of all, there is absolutely no question in my mind that at some point in the future we'll have machines that are at least as smart as humans in all the domains where humans are smart. That's not a question. People have big philosophical questions about this. A lot of people still believe that human nature is kind of impalpable and we're never going to be able to reduce this to computation. I'm not a skeptic on that dimension. There's no question in my mind at some point we'll have machines that are more intelligent than us. They already are in narrow domains. So then there is the question of what does AGI really mean exactly? Do you mean intelligence that is as general as human intelligence? If that's the case, you can use that phrase but it's very misleading because human intelligence is not general at all. It's extremely specialized. We are shaped by evolution to only do the tasks that are worth accomplishing for survival. We think of ourselves as having general intelligence but we just are not at all general. It's just that all the problems that we're not able to apprehend, we can think of them, and that makes us believe that we have general intelligence but we absolutely do not. So I think this phrase is nonsense. It's very misleading. I prefer the phrase we use at Meta to designate the concept of human-level intelligence: AMI, advanced machine intelligence. This is a much more open concept. We actually pronounce it 'ami', which means friend in French. But let's call it human-level intelligence if you want. No question it will happen. It's not going to happen next year, it's not going to happen two years from now. It may happen to some degree within the next 10 years. So it's not that far away. If all the things we are working on at the moment turn out to be successful, then maybe within 10 years we'll have a good handle on whether we can reach that goal. But it's almost certainly harder than we think, and probably much harder than we think because it's always been harder than we think over the history of AI. So I'm optimistic. I'm not one of those pessimists who say we'll never get there. I'm not one of those pessimists who say all the stuff we're doing right now is useless. It's not true. It's very useful. I'm not one of those who say we're going to need quantum computing or some completely new principle. No, I think it's going to be based on deep learning basically, and that underlying principle is going to stay with us for a long time. But within this domain, the type of things we need to discover and implement, we're not there yet. We're missing some basic concepts. The best way to convince yourself of this is to say: we have systems that can answer any question that has a response somewhere on the internet. We have systems that can pass the bar exam, which is basically information retrieval to a large extent. We have systems that can shorten text and help us understand it, that can criticize a piece of writing, that can generate code—but generating code is relatively simple because the syntax is strong and a lot of it is right. We have systems that can solve equations, that can solve problems as long as they've been trained to solve those problems. If they see a new problem from scratch, current systems just cannot find a solution. There was a paper recently that showed that if you test all the best LLMs on the latest math olympiad, they basically get zero performance because they are new problems they haven't been trained to solve. So we have systems that can manipulate language, and that fools us into thinking they're smart because we're used to smart people being able to manipulate language in smart ways. But where is my domestic robot? Where is my level-five self-driving car? Where is a robot that can do what a cat can do? Even a simulated robot that can do what cats can do. The issue is not that we can't build a robot. We can build robots that have the physical abilities. It's just that we don't know how to make them smart enough. It's much harder to deal with the real world and to produce actions than to understand language. Again, it's related to the problem I mentioned before: language is discrete, it has strong structure. The real world is a huge mess and unpredictable, not deterministic, high-dimensional, continuous. So let's try to build something that can learn as fast as a cat.

Host

我有很多问题要问你,但我会再停留在这个话题上一分钟。

I've got so many questions for you but I'm going to stay on this for another minute.

人类与机器智能 Human vs Machine Intelligence

Host

人类级别的活动或思维是否应该成为模型?这是否有局限性?几年前亚历克斯·罗森伯格有一本很棒的书叫《历史如何出错》,他认为他解构了心智理论,我们并没有那种推理过程,实际上我们有点像大语言模型那样运作——我们脑子里有一堆录像带,遇到情况时就找最近的录像带播放,然后以此决定是或否。这听起来确实有点像人类思维,但我们通常认为的人类思维模型是推理和权衡等等。而且正如你所说,我们并非通用智能,但机器可能做到我们——它现在已经能做到我们做不到的事。它可以做得更多。那么当你思考成功和那个目标时,那个模型是什么?达到猫的水平会是一个巨大的胜利,但你的更大目标是什么?是人类智能还是别的什么?

Should human-level activity or thought even be the model? Is that limiting? There's a wonderful book from some years ago by Alex Rosenberg called How History Gets Things Wrong, arguing that he debunks the theory of mind, that we don't have this reasoning we go through, that in fact we're kind of doing what an LLM does in the sense that we have a bunch of videotapes in our head and when we hit a circumstance we find the nearest videotape and play that and decide yes or no in that way. And so that does sound like the human mind a bit, but the model we tend to have for the human mind is one of reasoning and weighing things and so on. And also as you say, we're not generally intelligent, but the machine conceivably could do things that we—it right now does things we cannot do. It could do more. So when you think about success and that goal, what is that model? Is it a cat would be a big victory to get to the point of being a cat, but what's your larger goal? Is it human intelligence or is it something else?

Yann LeCun

这是一种与人类和动物智能相似类型的智能。当前的 AI 系统很难解决从未遇到过的新问题,对吧?所以它们没有这种心智模型,我之前提到的世界模型,让它们能够想象自己行为的后果等等。它们不以那种方式推理。我的意思是,大语言模型肯定不,因为它唯一能做的就是产生词语、产生 token。所以,你让大语言模型在复杂问题上花更多时间思考的一个技巧是让它逐步推理,结果它产生更多 token,然后花更多算力来回答那个问题,但这是一个糟糕的技巧。这是一个 hack。这不是人类推理的方式。另一个例子是大语言模型写代码或回答问题。你让大语言模型生成大量 token 序列,这些序列都有一定的概率。然后你有第二个神经网络试图评估每个序列,然后选出最好的那个。这就像产生很多答案,然后让一个评论者告诉你哪个答案最好。有很多 AI 系统是这样工作的,在某些情况下有效,比如你想让计算机系统下棋。这正是它的工作方式。它生成一个树,包含你和对手所有可能的走法,然后你和对手再走,如此反复。这个树呈指数增长,所以你无法生成整个树。你必须有一些聪明的方法只生成树的一部分,然后你有一个所谓的评估函数或价值函数,选出最有可能导致获胜位置的最佳分支。现在所有这些都经过训练。它们基本上是神经网络,生成好的分支并选择它。这是一种有限的推理形式。为什么有限?顺便说一句,这是一种人类非常不擅长的推理。一个你在玩具店花 30 美元买的小玩意能在象棋上打败你,这表明人类完全不懂这种推理。我们真的很不擅长。我们没有记忆容量、计算速度等等。所以我们在这方面很糟糕。但我们真正擅长的是猫、狗和老鼠非常擅长的推理:在现实世界中规划行动,并以分层的方式进行规划。例如,在人类领域,但动物任务中也有类似的。你看到猫学会打开罐子、跳上门开门、打开门锁等等。它们学会如何做到这一点,并规划一系列行动以达到目标,比如到另一边去获取食物。你看到松鼠这样做。它们实际上很聪明,学会了如何做这类事情。这是一种我们不知道如何用机器复现的规划。很多是完全内在的。与语言无关。我们人类认为思考与语言有关,但并非如此。动物能思考。不会说话的人也能思考。大多数类型的推理与语言无关。如果我让你想象一个立方体漂浮在我们面前的空中,然后绕垂直轴旋转 90 度。可能你假设立方体是水平的,底部是水平的。你没有想象一个侧着的立方体,然后你旋转 90 度,你知道它看起来和开始时一样,因为它是立方体。它有 90 度对称性。这个推理中没有语言参与。只是图像和对情况的抽象表征。我们如何做到这一点?我们有这些抽象的思想表征,然后我们可以通过想象进行的虚拟动作来操纵这些表征,比如旋转那个立方体,然后想象结果。这就是让我们能够在抽象层面完成现实世界任务的原因。立方体是什么做的、有多重、是否漂浮在我们面前都不重要。所有这些细节都不重要,表征足够抽象,不在乎这些细节。如果我计划明天去巴黎,我可以尝试用我能采取的基本动作来规划我的行程,这基本上是毫秒级的肌肉控制,但我无法做到,因为那是几个小时的肌肉控制,而且依赖于我不知道的信息。比如,我可以上街拦一辆出租车。

Well, it's a type of intelligence that is similar to human and animal intelligence in the following way. Current AI systems have a very hard time solving new problems that they've never faced before, right? So they don't have this mental model, this world model I was telling you about earlier that allows them to kind of imagine the consequence of their actions or whatever. They don't reason in that way. I mean, an LLM certainly doesn't because the only way it can do anything is just produce words, produce tokens. So one way you trick an LLM into spending more time thinking about a complex question versus a simple question is you ask it to go through the steps of reasoning, and as a consequence it produces more tokens and then spends more computation answering that question, but it's a horrible trick. It's a hack. It's not the way humans reason. Another example that LLMs do is for writing code or answering questions. You get an LLM to generate lots and lots of sequences of tokens, all that have some decent level of probability or something like that. And then you have a second neural net that tries to evaluate each of those and then picks the one that is best. It's like producing lots of answers to a question and then having a critic tell you which answer is best. There are many AI systems that work this way, and it works in certain situations, like if you want a computer system to play chess. This is exactly how it works. It produces a tree of all possible moves from you and then from your opponent and then from you and then from your opponent. That tree grows exponentially, so you can't generate the entire tree. You have to have some smart way of only generating a piece of the tree, and then you have what's called an evaluation function or value function that picks out the best branch that results in a position most likely to win. All of those things are trained nowadays. They're neural nets basically that generate the good branch and select it. That's a limited form of reasoning. Why is it limited? And by the way, it's a type of reasoning that humans are terrible at. The fact that a $30 gadget you buy at the toy store can beat you at chess demonstrates that humans totally suck at this kind of reasoning. We're just really bad at it. We don't have the memory capacity, the computing speed, and everything. So we're terrible at this. What we are really good at, though, is the kind of reasoning that cats, dogs, and rats are really good at: planning actions in the real world and planning them in a hierarchical manner. For example, in the human domain, but there are similar ones in animal tasks. You see cats learning to open jars, jump on doors to open them, open the lock of a door, and things like that. They learn how to do this and plan that sequence of actions to arrive at a goal, like getting to the other side perhaps to get food. You see squirrels doing this. They're pretty smart actually in the way they learn how to do this kind of stuff. This is a type of planning that we don't know how to reproduce with machines. A lot of it is completely internal. Has nothing to do with language. We think as humans that thinking is related to language, but it's not. Animals can think. People who don't talk can think. Most types of reasoning have nothing to do with language. If I tell you to imagine a cube floating in the air in front of us, and then rotate that cube 90° along a vertical axis. Probably you made the assumption that the cube was horizontal, that the bottom was horizontal. You didn't imagine a cube that was kind of sideways, and then you rotate it 90° and you know that it looks just like the cube you started with because it's a cube. It has 90° symmetry. There's no language involved in this reasoning. It's just images and abstract representations of the situation. How do we do this? We have those abstract representations of thought, and then we can manipulate those representations through virtual actions that we imagine taking, like rotating that cube, and then imagine the result. That is what allows us to accomplish tasks in the real world at an abstract level. It doesn't matter what the cube is made of, how heavy it is, whether it floats in front of us or not. All those details don't matter, and the representation is abstract enough to not care about those details. If I plan to be in Paris tomorrow, I could try to plan my trip in terms of elementary actions I can take, which basically are millisecond-by-millisecond controls of my muscles, but I can't possibly do this because it's several hours of muscle control and it will depend on information that I don't have. Like, I can go out on the street and hail a taxi.

分层规划与世界模型 Hierarchical Planning and World Models

Yann LeCun

我不知道出租车多久会来。我不知道灯会是红的还是绿的。所以我无法规划整个行程。我必须进行分层规划。我必须想象,如果我想明天到巴黎,我首先得去机场赶飞机。现在我有一个子目标:去机场。我怎么去机场?我在纽约。所以我可以下楼到街上,打出租车。我怎么下楼?我得走到电梯或楼梯,按按钮,下楼,走出大楼。在此之前,我还有一个子目标:到达电梯或楼梯。我甚至怎么从椅子上站起来?你能用语言解释如何爬楼梯或从椅子上站起来吗?这是对现实世界的底层理解。在所有这些子目标中,到了某个点,你就能不假思索地完成任务,因为你习惯了从椅子上站起来。但想象你的行动后果(通过内部世界模型)然后规划一系列行动来完成任务的复杂性,是未来几年 AI 的重大挑战。我们还没到那一步。

I don't know how long it's going to take for a taxi to come by. I don't know if the light is going to be red or green. So I cannot plan my entire trip. I have to do hierarchical planning. I have to imagine that if I want to be in Paris tomorrow, I first have to go to the airport and catch a plane. Now I have a subgoal: going to the airport. How do I go to the airport? I'm in New York. So I can go down on the street, have a taxi. How do I go down the street? I have to walk to the elevator or the stairs, hit the button, go down, walk out the building. And before that, I have a subgoal of getting to the elevator or the stairs. How do I even stand up from my chair? Can you explain in words how you climb a stair or stand up from your chair? This is low-level understanding of the real world. At some point in all those subgoals, you get to a situation where you can just accomplish the task without really planning and thinking because you're used to standing up from your chair. But the complexity of this process of imagining what the consequences of your actions are going to be with your internal world model and then planning a sequence of actions to accomplish this task—this is the big challenge of AI for the next few years. We're not there yet.

Meta 的 Llama 开源策略 Meta's Open Source Strategy for Llama

Host

我一直想问一个问题:教授,这堂课很棒,我真的很感激。但我也想了解 Meta 当前对此的策略。Meta 决定,不管我们称之为开源、开放还是可用,Llama 都是一个巨大的工具。作为教育工作者,我很感激。我曾在纽约市立大学荣休,现在在石溪大学。正是因为 Llama,大学才能运行模型、从中学习并构建东西。我注意到,Meta 在 Llama 上的策略对行业大部分是破坏性的,但对学术和创业等开放发展是推动力。所以我想听您亲口说:您这样开放 Llama 背后的策略是什么?

One question I've been wanting to ask: this has been a great lesson, professor. I'm really grateful. But I also want to get to the current view of Meta's strategy on this. Meta has decided to go, whether we call it open source or open or available, but Llama is a tremendous tool. As an educator myself, I'm grateful. I was emeritus at CUNY but now I'm at Stony Brook. It's because of Llama that universities can run models and learn from them and build things. It struck me that Meta's strategy on Llama is a spoiler for much of the industry but an enabler for tremendous open development, whether academic or entrepreneurial. So I'd love to hear from the horse's mouth: what's the strategy behind opening up Llama in the way that you've done?

Yann LeCun

它正好是三家公司的破坏者。是的。它是成千上万公司的推动者。从纯粹的伦理角度看,这显然是正确的事情。Llama 2,以合格开源形式发布的 Llama 2,基本上完全启动了 AI 生态系统,不仅在工业和初创公司,也在学术界。学术界基本上没有能力训练与公司同等水平的自有基础模型。所以他们依赖这种开源平台来为 AI 研究做出贡献。这是 Meta 开源这些基础模型的主要原因之一:实现更快的创新。问题不在于这家或那家公司比另一家领先三个月——这确实是现状。问题是:我们目前的 AI 系统是否具备构建我们想要的产品的能力?答案是否定的。Meta 最终想要构建的产品是一个 AI 助手,或者一系列 AI 助手,时刻陪伴我们。也许它存在于我们的智能眼镜中,我们可以与之交谈。也许它在镜片上显示信息。要让这些东西发挥最大效用,它们需要具备人类水平的智能。迈向人类水平的智能不会是一个事件。不会有一天我们没有 AGI,第二天就有了 AGI。事情不会这样发生。如果真的发生了,我请你喝酒。但不会这样发生。真正的问题是:我们如何以最快的速度迈向人类水平的智能?由于这是我们面临的最大科学和技术挑战之一,我们需要来自世界各地的贡献。好的想法可以来自世界任何地方。我们最近看到了 DeepSeek 的例子,它让硅谷所有人都感到惊讶,但并没有让我们开源世界的许多人感到那么惊讶。这就是重点。它验证了整个开源理念。好的想法可以来自任何地方。没有人垄断好想法,除了那些有极度膨胀优越感的人。不是特指谁。但在美国某些地区,这种人高度集中。他们从传播“自己比别人强”的想法中获益。我认为这仍然是一个重大的科学挑战,我们需要每个人都做出贡献。在学术研究背景下,我们知道最好的方法是发表研究,尽可能开源代码,并让人们贡献。过去十几年 AI 的历史表明,进步很快是因为人们分享代码和科学信息。过去三年,一些参与者开始崛起,因为他们需要从技术中创收。在 Meta,我们不直接从技术本身创收。我们通过广告创收,而这些广告依赖于我们基于技术构建的产品质量、社交网络的网络效应,或者连接用户和人们的渠道。所以我们分发技术并不会在商业上伤害我们。事实上,它帮助我们。

It's a spoiler for exactly three companies. Yes. It's an enabler for thousands of companies. From a pure ethical point of view, it's obviously the right thing to do. Llama 2, the release of Llama 2 in qualified open source, has basically completely jumpstarted the AI ecosystem, not just in industry and startups but also in academia. Academia basically doesn't have the means to train their own foundation model at the same level as companies. So they rely on this kind of open source platforms to make contributions to AI research. That's one of the main reasons for Meta to release those foundation models in open source: to enable faster innovation. The question is not whether this or that company is three months ahead of the other, which is really the case right now. The question is: do we have the capabilities in the AI systems that we have at the moment to enable the products we want to build? The answer is no. The product that Meta wants to build ultimately is an AI assistant, or maybe a collection of AI assistants, that is with us at all times. Maybe it lives in our smart glasses, that we can talk to. Maybe it displays information in the lens. For those things to be maximally useful, they would need to have human-level intelligence. Moving towards human-level intelligence is not going to be an event. There's not going to be a day where we don't have AGI and a day after which we have AGI. It's just not going to happen that way. I'll buy you the drinks if that happens. But it's not going to happen this way. The question really is: how do we make the fastest possible progress towards human-level intelligence? Since it's one of the biggest scientific and technological challenges we've faced, we need contributions from anywhere in the world. Good ideas can come from anywhere in the world. We've seen an example with DeepSeek recently, which surprised everybody in Silicon Valley, but didn't surprise many of us in the open source world that much. That's the point. It's validation of the whole idea of open source. Good ideas can come from anywhere. Nobody has a monopoly on good ideas, except people who have an incredibly inflated superiority complex. Not that we're talking about anybody in particular. But there's a high concentration of those people in certain areas of the country. They have a vested interest in disseminating the idea that they are somehow better than everybody else. I think it's still a major scientific challenge and we need everybody to contribute. The best way we know how to do this in the context of academic research is to publish your research, publish code in open source as much as you can, and get people to contribute. The history of AI over the last dozen years really shows that progress has been fast because people were sharing code and scientific information. A few players in the space started climbing up over the last three years because they need to generate revenue from the technology. At Meta, we don't generate revenue from the technology itself. We generate revenue from ads, and those ads rely on the quality of products we built on top of the technology, the network effect of the social networks, or a conduit to the people and users. So the fact that we distribute our technology doesn't hurt us commercially. In fact, it helps us.

可穿戴设备与眼镜 Wearables and Glasses

Host

你提到了可穿戴设备和眼镜的话题,这当然总是引起我的注意。

You mentioned the topic of wearables and glasses, and that of course always sparks my attention.

智能眼镜与 AI 助手 Smart glasses and AI assistants

Host

去年十二月我有机会体验了谷歌的 Project Astra 眼镜,从那以后就一直印象深刻。它确实让我坚信,这是让世界情境化的绝佳下一步。我能在我们现在的状况和未来的可能性之间画出的那条线,不仅是那种体验给佩戴者带来的情境,而且对你、对 Meta、对那些创建这些系统的人来说,智能眼镜在现实世界中收集人类如何生活和运作的信息,可以成为你之前提到的知识的很好来源。我说得对吗,还是这只是拼图中的一小块?

I had the opportunity to check out Google's Project Astra glasses last December, and it stuck with me ever since. It really solidified my view of that being a wonderful next step for contextualizing the world. The line I've been able to draw between where we are now and where we're going is not only the context that experience gives the wearer, but for you, for Meta, and for those creating these systems, smart glasses out in the real world, taking in information on how humans live and operate in our physical world could be a really good source of knowledge to pull from for what you were talking about earlier. Am I on the right track, or is that just one small piece of the puzzle?

Yann LeCun

嗯,这是一块,重要的一块。但想法是你随时都有一个助手,它能看到你所见,听到你所闻,如果你允许的话。它是你的知己,甚至可能比人类助手更好地帮助你。这当然是一个重要的愿景。实际上,愿景是你不会只有一个助手;你会有一整队智能虚拟助手跟着你。就像我们都会成为老板。我的意思是,人们因为机器会比我们更聪明而感到威胁,但我们应该感到被赋能。它们会为我们工作。作为科学家或行业管理者,你能遇到的最好事情就是雇佣比你更聪明的人。那是理想情况,你不应该感到威胁,而应该感到被赋能。所以我认为那是我们应该设想的未来:一群聪明的助手,在日常生活中帮助你,也许比你更聪明,你给它们任务,它们完成得可能比你好,这很棒。

Well, it's a piece, an important piece. But the idea is that you have an assistant with you at all times that sees what you see, hears what you hear, if you let it. It's your confidant and can help you perhaps even better than a human assistant could. That's certainly an important vision. In fact, the vision is that you won't have a single assistant; you will have a whole staff of intelligent virtual assistants walking around with you. It's like all of us will be the boss. I mean, people feel threatened by the fact that machines would be smarter than us, but we should feel empowered by it. They're going to be working for us. As a scientist or a manager in industry, the best thing that can happen to you is to hire people who are smarter than you. That's the ideal situation, and you shouldn't feel threatened; you should feel empowered. So I think that's the future we should envision: a smart collection of assistants that help you in your daily lives, maybe smarter than you, you give them a task, they accomplish it perhaps better than you, and that's great.

Host

这联系到我之前想提的另一个点,关于开源。在那个未来,我们与数字世界的大部分互动都将由 AI 系统中介。这就是为什么谷歌现在有点慌乱,因为他们知道没人会再用搜索引擎了。你只会跟你的助手说话。所以他们正在谷歌内部尝试这个。这会通过眼镜实现。所以他们意识到他们可能必须制造这些眼镜;他们几年前就意识到了。所以我们有一点先发优势,但这确实会发生。我们将随时拥有那些 AI 系统,它们将中介我们所有的信息摄入。现在,如果你想想这个,如果你是世界任何地方的公民,你不会希望你的信息摄入来自美国西海岸或中国的少数几家公司构建的 AI 助手。你想要高度多样化的 AI 助手,首先能说你的语言,无论是生僻方言还是本地语言。其次,理解你的文化、你的价值体系、你的偏见,无论它们是什么。所以我们需要高度多样化的这类助手,原因和我们需要高度多样化的媒体一样。我意识到我在跟一位新闻学教授说话,但我说得对吗?

That connects to another point I wanted to make, related to the previous question, which is about open source. In that future, most of our interactions with the digital world will be mediated by AI systems. That's why Google is a little frantic right now because they know nobody is going to go to a search engine anymore. You're just going to talk to your assistant. So they're trying to experiment with this within Google. That's going to be through glasses. So they realize they probably have to build those; they realized this several years ago. So we have a bit of a head start, but that's really what's going to happen. We're going to have those AI systems with us at all times, and they're going to mediate all of our information diet. Now, if you think about this, if you are a citizen anywhere in the world, you do not want your information diet to come from AI assistants built by a handful of companies on the west coast of the US or China. You want a high diversity of AI assistants that first of all speak your own language, whether it's an obscure dialect or local language. Second of all, understand your culture, your value system, your biases, whatever they are. So we need a high diversity of such assistants for the same reason we need a high diversity of the press. And I realize I'm talking to a journalism professor here, but am I right?

Yann LeCun

阿门。事实上,我认为这正是我所赞颂的:互联网和下一代 AI 能做的事情是拆解大众媒体的结构,并在人类层面再次开放媒体。AI 让我们更人性化。我也希望如此。所以用当前技术实现这一点的唯一方法是,那些构建具有文化多样性等特性的助手的人能够访问强大的开源基础模型,因为他们没有资源训练自己的模型。我们需要能说世界上所有语言、理解所有价值体系、并拥有你能想象的所有文化、政治偏见等偏见的模型。所以会有成千上万个这样的模型供我们选择。它们将由世界各地的小团队构建,并且必须建立在像 Meta 这样的大公司或可能训练这些基础模型的国际联盟所训练的基础模型之上。我看到的市场演变图景类似于 90 年代末或 2000 年代初互联网软件基础设施的情况。在互联网早期,有 Sun、微软、惠普、IBM 等几家公司在推动提供互联网的硬件和软件基础设施,它们自己的 Unix 或 Windows NT 版本,以及它们自己机架上的网络服务器。所有这些都被 Linux 和通用硬件彻底淘汰了。被淘汰的原因是 Linux 是一个平台软件;它更便携、更可靠、更安全、更便宜,等等。所以谷歌是最早这样做的公司之一,在通用硬件和开源操作系统上构建基础设施。Meta 也做了完全相同的事情,现在每个人都在这样做,甚至微软。所以我认为市场会有类似的压力,使那些 AI 基础模型开放和免费,因为它是一种基础设施,就像互联网的基础设施一样。

Amen. In fact, I think that's what I celebrate: what the internet and next AI can do is to tear down the structure of mass media and open up media once again at a human level. AI lets us be more human. I hope so too. So the only way we can achieve this with current technology is if the people building those assistants with cultural diversity and everything have access to powerful open source foundation models, because they're not going to have the resources to train their own models. We need models that speak all the languages in the world, understand all the value systems, and have all the biases that you can imagine in terms of culture, political biases, whatever you want. So there's going to be thousands of those that we're going to have to choose from. They're going to be built by small shops everywhere around the world, and they're going to have to be built on top of foundation models trained by a large company like Meta or maybe an international consortium that trains those foundation models. The picture I see of the evolution of the market is similar to what happened with the software infrastructure of the internet in the late 90s or early 2000s. In the early days of the internet, you had Sun, Microsoft, HP, IBM, and a few others pushing to provide the hardware and software infrastructure of the internet, their own version of Unix or Windows NT, and their own web server on their own racks. All of this got completely wiped out by Linux and commodity hardware. The reason it got wiped out is because Linux is a platform software; it's more portable, more reliable, more secure, cheaper, everything. So Google was one of the first to build the infrastructure on commodity hardware and open source operating system. Meta did exactly the same thing, and everybody is doing it now, even Microsoft. So I think there's going to be a similar pressure from the market to make those AI foundation models open and free, because it's an infrastructure like the infrastructure of the internet.

学生与研究经费变化 Changes in students and research funding

Host

你教书多久了?

How long have you been teaching?

Yann LeCun

22 年了。

22 years.

Host

那么你看到现在的学生和他们的抱负在你的领域有什么不同?

So what differences do you see in students and their ambitions today in your field?

Yann LeCun

这很难说,因为过去十几年左右,我只教研究生。所以我没有看到博士生有什么显著变化,除了他们来自世界各地。我的意思是,现在美国正在发生一些绝对可怕的事情,研究经费被削减,还有不给外国学生签证的威胁。如果真按目前的方向实施,那将彻底摧毁美国的技术领导地位。

It's hard for me to tell because in the last dozen years or so, I've only taught graduate students. So I don't see any significant change in PhD students, other than the fact that they come from all over the world. I mean, there's something absolutely terrifying happening in the US right now, where funding for research is being cut, and threats of visas not being given to foreign students. It's completely going to destroy the technological leadership in the US if it's actually implemented the way it seems to be going.

结束语 Closing Remarks

Host

就像大多数 STEM 领域的博士生一样,你知道,科学、技术、工程、数学,都是外国人。在大多数工程学科的研究生阶段,这个比例甚至更高,主要是外国学生。大多数科技公司的创始人或 CEO 都是外国出生的。法国大学正在为美国研究人员提供去那里学习的机会。我还有一个问题要问你。你养猫吗?

Like most PhD students in STEM, you know, science, technology, engineering, mathematics, are foreign. And it's even higher in most engineering disciplines at the graduate level. They're mostly foreign students. Most founders or CEOs of tech companies are foreign-born. French universities are offering the opportunity for American researchers to go there. I've got one more question for you. Do you have a cat?

Yann LeCun

我没有,但我们小儿子养了一只猫,我们偶尔会照看它。

I don't, but our youngest son has a cat and we watch the cat occasionally.

Host

好的。我在想那是不是你的模型,对吧?好了,Yann,这真是太棒了。我知道我们比约定的时间稍微多留了你一会儿,耽误了你的日程。所以我们真的很感谢你抽出时间。是的,这真是太棒了,而且就像 Jeff 之前说的,能亲耳听到这些真是太好了,因为你在我们的谈话中经常被提到,我们真的很感谢你在 AI 领域的见解以及你多年来所做的一切。感谢你来到这里。这是我们的荣幸。感谢你为对话带来的理性。

Okay. I wonder if that was your model, right? Well, Yann, this has been wonderful. I know we've kept you just a slight bit longer than we had agreed to for your schedule. So, we really appreciate you carving out some time. Yeah, it's been really wonderful and it's wonderful to kind of hear some of this as Jeff said from the horse's mouth earlier because you come up in our conversations quite a lot and we really appreciate your perspective in the world of AI and all the work that you've done over the years. Thank you for being here with us. This has been an honor. Thank you for the sanity you bring to the conversation.

Yann LeCun

谢谢。是的。非常感谢你。和你谈话真的很愉快。

Thank you. Yes. Thank you so much. It's really been a pleasure talking with you.

Host

再次非常感谢 Yann LeCun 参加我们的 AI Inside 节目。希望将来能再请他回来,看看进展如何。当然,还要非常感谢我的联合主持人 Jeff Jarvis。他现在显然不在这里。JeffJarvis.com。如果你想看他所有精彩的书,关于这个节目的一切,这个播客都可以在我们的网站上找到。只要去那里,AI Inside。你会找到所有订阅方式,音频、视频,所有细节都在那里。如果你喜欢这个节目,你知道的,给我们一个评论。如果你特别喜欢这一集,给我们一个评论,留下你的意见。如果你在 YouTube 上观看,我们真的很想听听你对这次采访的看法。最后,如果你真的喜欢这个节目,你可以在 Patreon 上支持我们。那是 patreon.com/aiinsidow。你可以获得无广告节目、Discord 社区访问权限,如果你成为执行制作人,比如本周的执行制作人 Dr. Dub、Jeffrey Marachini、北卡罗来纳州阿什维尔的 WPVM 103.7、Dante St. James、Bono Day Rick、Jason Nefer 和 Jason Brady,你还会得到一件 AI Inside T 恤。你们太棒了。非常感谢你们每周的支持。感谢你们的收看。我是 Jason Howell。希望下周三在另一期 AI Inside 播客中见到大家。再见。

Huge thank you once again to Yann LeCun for joining us on AI Inside. Hope to have him back sometime down the line to check in on how things are going. And of course, a huge thank you to my co-host Jeff Jarvis. He's obviously not here right now. JeffJarvis.com. If you want to check out all of his wonderful books, everything you need to know about this show, this podcast can be found at our site. Just go there, AI Inside. You're going to find all the ways to subscribe and audio, video, all the details are there. And if you love this show, you know, give us a review. If you liked this episode, especially, give us a review, leave a comment. If you're watching this on YouTube, we really want to hear from you as far as what you think about this interview. And finally, if you really love this show, you can support us on Patreon. That's patreon.com/aiinsidow. You get ad-free shows, a Discord community access, you get an AI Inside t-shirt if you become an executive producer like this week's executive producers, Dr. Dub, Jeffrey Marachini, WPVM 103.7 in Asheville, North Carolina, Dante St. James, Bono Day Rick, Jason Nefer, and Jason Brady. Y'all are awesome. Thank you so much for your support each and every week. And thank you for being here. I'm Jason Howell. I hope to see you next Wednesday on another episode of the AI Inside podcast. Bye everybody.

互动版:逐字朗读 + 针对本期提问 →