Andrej Karpathy 谈自动驾驶、AGI 以及演示与产品之间的差距

Andrej Karpathy on Self-Driving Cars, AGI, and the Gap Between Demo and Product

安德烈·卡帕西 Andrej Karpathy · No Priors 播客 · 2024-09-05 · 约 44 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Andrej Karpathy 讨论自动驾驶现状,比较 Waymo 和特斯拉,并与 AGI 发展进行类比。

Andrej Karpathy discusses the state of self-driving cars, comparing Waymo and Tesla, and draws analogies to AGI development.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 26)

全文 · Full transcript(中英对照)

引言与自动驾驶汽车 Introduction and Self-Driving Cars

Host

听众朋友们好,欢迎回到 No Priors。今天我们请到了 Andrej Karpathy,他无需过多介绍。Andrej 是著名研究员、深受喜爱的 AI 教育者和程序员。他是 OpenAI 的早期成员、特斯拉 Autopilot 负责人,现在从事 AI 教育。我们将与他讨论研究现状、他的新公司以及我们对 AI 的期待。感谢你今天加入我们,很高兴你能来。

Hi listeners, welcome back to No Priors. Today we're hanging out with Andrej Karpathy, who needs no introduction. Andrej is a renowned researcher, beloved AI educator, and coder. An early team member from OpenAI, the lead for Autopilot at Tesla, and now working on AI for Education. We'll talk to him about the state of research, his new company, and what we can expect from AI. Thanks for joining us today, it's great to have you here.

Andrej

谢谢,我很高兴来到这里。

Thank you, I'm happy to be here.

Host

你曾领导特斯拉的 Autopilot,现在路上确实有了全自动驾驶的乘用车。你如何看待我们在能力集上的位置?我们应该多快看到能力提升或乘用车普及?

You led Autopilot at Tesla, and now we actually have fully self-driving cars, passenger vehicles on the road. How do you read that in terms of where we are in the capability set? How quickly should we see increased capability or pervasive passenger vehicles?

Andrej

是的,我在自动驾驶领域花了大约五年时间。我认为这是一个迷人的领域。现在这个领域正在发生的事情是……我从自动驾驶中汲取了很多与 AGI 的类比,也许只是因为我熟悉它。我有点觉得我们在自动驾驶方面已经达到了 AGI,因为今天有系统你可以作为付费客户乘坐。旧金山的 Waymo 非常普遍,你可能坐过 Waymo。我坐过很多次,它很棒。它可以带你去任何地方,你把它当作产品付费。Waymo 有趣的是,我第一次坐 Waymo 实际上是在十年前,差不多是 2014 年左右。我一个在那工作的朋友给我演示了一下,十年前它带我绕街区开了一圈,基本上完美。从演示到我可以付费的产品,花了十年时间,达到城市规模,并且还在扩展等等。

Yes, so I spent maybe five years in the self-driving space. I think it's a fascinating space. What's happening in the field right now is... I also draw a lot of analogies to AGI from self-driving, maybe just because I'm familiar with it. I kind of feel like we've reached AGI a little bit in self-driving, because there are systems today that you can take around as a paying customer. Waymo in San Francisco is very common, probably you've taken Waymo. I've taken it a bunch, and it's amazing. It can drive you all over the place, and you're paying for it as a product. What's interesting with Waymo is the first time I took Waymo was actually a decade ago, almost exactly 2014 or so. A friend of mine who worked there gave me a demo, and it drove me around the block 10 years ago, and it was basically perfect. It took 10 years to go from a demo to a product I can pay for, at a city scale, and it's expanding, etc.

Host

你认为其中有多少是监管因素,多少是技术因素?比如你认为技术什么时候准备好了?

How much of that do you think was regulatory versus technology? Like when do you think the technology was ready?

Andrej

我认为在单次 30 分钟的演示驾驶中,你不会遇到他们十年来必须处理的所有问题。所以演示和产品之间存在巨大差距。其中很多也是监管因素等等。但我确实认为我们在自动驾驶方面已经有点实现了 AI。然而真正迷人的是全球化根本没有发生。所以你有演示,可以在一个城市里乘坐它,但世界还没有改变,这需要很长时间。所以从演示到真正的全球化,我认为存在很大差距。这就是它与 AGI 的关系,因为我怀疑当我们得到 AGI 时,情况会类似。

I think in a single demo drive of 30 minutes, you're not running into all the stuff they had to deal with for a decade. So demo and product, there's a massive gap there. A lot of it also regulatory, etc. But I do think we've sort of achieved AI in the self-driving space in that sense, a little bit. Yet what's really fascinating is the globalization hasn't happened at all. So you have a demo and you can take it in a city, but the world hasn't changed yet, and that's going to take a long time. So going from a demo to actual globalization, I think there's a big gap there. That's how it relates to AGI, because I suspect it will look similar for AGI when we sort of get it.

Host

在自动驾驶领域停留一分钟,我认为人们认为 Waymo 领先于特斯拉。我个人认为特斯拉领先于 Waymo,我知道看起来不是这样,但我仍然非常看好特斯拉及其自动驾驶项目。我认为特斯拉有软件问题,Waymo 有硬件问题,这是我的看法。而且我认为软件问题容易得多。特斯拉在地球上大规模部署了所有这些汽车,而 Waymo 需要达到那个水平。所以一旦特斯拉达到能够实际部署并且真正有效的程度,我认为那将非常了不起。最新的版本,我昨天刚开过,它带我到处跑。他们最近取得了非常好的改进。我最近经常使用它,它实际上运行得相当好。昨天它为我做了一些神奇的驾驶,所以我对团队的工作印象深刻。所以我仍然认为特斯拉主要是软件问题,Waymo 主要是硬件问题。Waymo 现在看起来赢了,但当我们看 10 年后,谁真正达到规模,大部分收入来自哪里,我仍然认为他们在那个意义上领先。

Staying for a minute in the self-driving space, I think people think Waymo is ahead of Tesla. I think personally Tesla is ahead of Waymo, and I know it doesn't look like that, but I'm still very bullish on Tesla and its self-driving program. I think Tesla has a software problem and Waymo has a hardware problem, is the way I put it. And I think software problems are much easier. Tesla has deployment of all these cars on Earth at scale, and Waymo needs to get there. So the moment Tesla gets to the point where they can actually deploy this and it actually works, I think it's going to be really incredible. The latest builds, I just drove yesterday, it's just driving me all over the place. They've made really good improvements very recently. I've been using it a lot recently, and it actually works quite well. It did some miraculous driving for me yesterday, so I'm very impressed with what the team is doing. So I still think Tesla mostly has a software problem, Waymo mostly a hardware problem. Waymo looks like it's winning right now, but when we look in 10 years and who's actually at scale and where most of the revenue is coming from, I still think they're ahead in that sense.

Host

你认为我们离软件问题取得突破、达到某种等效性还有多远?因为显然如你所说,如果你看 Waymo 的汽车,它内置了很多昂贵的激光雷达和其他传感器,所以它能做到它所做的。这有助于支持软件系统。如果你能只用摄像头,也就是特斯拉的方法,那么你就能有效消除巨大的成本和复杂性,并且可以在许多不同类型的汽车上实现。你认为这种转变什么时候发生?

How far away do you think we are from the software problem turning the corner in terms of getting to some equivalency? Because obviously to your point, if you look at a Waymo car, it has a lot of expensive LiDAR and other sensors built into the car, so it can do what it does. It sort of helps support the software system. If you can just use cameras, which is the Tesla approach, then you effectively get rid of enormous cost and complexity, and you can do it in many different types of cars. When do you think that transition happens?

Andrej

我的意思是,在接下来的几年里,我希望如此。但实际上,真正有趣的是,我不确定人们是否意识到特斯拉实际上确实使用了很多昂贵的传感器,只是在训练时使用。所以有一批汽车带着激光雷达到处跑,它们做很多不可扩展的事情,有额外的传感器等等,它们进行地图绘制和所有那些东西。你在训练时做这些,然后将其蒸馏成一个测试时包,部署到汽车上,只使用视觉。这是一种传感器和成本的套利。我认为这实际上是一种巧妙的策略,我认为它没有被充分认识,而且我认为它会成功,因为像素包含信息。我认为网络将能够做到这一点。在训练时,这些传感器非常有用,但我不认为它们在测试时那么有用。

I mean, in the next few years, I'm hoping something like that. But actually, what's really interesting is that I'm not sure people appreciate that Tesla actually does use a lot of expensive sensors, they just do it at training time. So there are a bunch of cars that drive around with LiDARs, they do a bunch of stuff that doesn't scale, and they have extra sensors, etc., and they do mapping and all that stuff. You're doing it at training time, and then you're distilling that into a test-time package that is deployed to the cars and is vision-only. It's an arbitrage on sensors and expense. I think it's actually a kind of brilliant strategy that I don't think is fully appreciated, and I think it's going to work out well because the pixels have the information. I think the network will be capable of doing that. At training time, these sensors are really useful, but I don't think they're as useful at test time.

Host

似乎发生的另一件事或转变,基本上是从大量与边缘情况相关的启发式设计转向端到端深度学习。这是最近发生的另一个转变。你想谈谈这个吗?

It seems like the one other thing or transition that's happened is basically a move from a lot of sort of edge-case designed heuristics associated with it versus end-to-end deep learning. And that's another shift that's happened recently. Do you want to talk a little bit about that?

Andrej

是的,我认为这从一开始就是特斯拉的计划。我谈到过神经网络如何吞噬整个栈。当我加入时,有大量的 C++ 代码,而现在在汽车中运行的测试时包里 C++ 代码少了很多。后端仍然有很多东西,我们没在谈论。神经网络基本上贯穿整个系统。首先,它只在图像层面做检测,然后多个图像给出预测,然后多个图像随时间给出预测,你正在丢弃 C++ 代码。最终,你只给出转向命令。所以我认为特斯拉正在吞噬整个栈。我的理解是,目前的 Waymo 实际上并没有这样做;他们尝试过但最终没有这样做,但我不确定,因为他们不谈论这个。我基本上相信这种方法,而且我认为如果你想要,这是最后一块要落下的拼图。

Yeah, I think that was always the plan from the start at Tesla. I was talking about how the neural net can eat through the stack. When I joined, there was a ton of C++ code, and now there's much less C++ code in the test-time package that runs in the car. There's still a ton of stuff in the back end that we're not talking about. The neural net kind of takes through the system. First, it just does detection on the image level, then multiple images give you prediction, then multiple images over time give you a prediction, and you're discarding C++ code. Eventually, you're just giving steering commands. So I think Tesla is kind of eating through the stack. My understanding is that current Waymos are actually not doing that; they've tried but ended up not doing that, but I'm not sure because they don't talk about it. I do fundamentally believe in this approach, and I think that's the last piece to fall if you want.

端到端系统与渐进式开发 End-to-end systems and incremental development

Andrej

这样想的话,我确实怀疑特斯拉十年后的端到端系统就是一个神经网络。视频流进神经网络,命令输出。你必须逐步构建,一点一点来。即使是我们所做的所有中间预测和所有事情,我也不认为它们误导了开发。我认为它们是其中的一部分,因为有很多充分的理由。实际上,在驾驶中,当你只是模仿人类时,你只有很少的监督信号来训练一个巨大的神经网络,这些信号太少,无法训练数十亿的参数。所以这些中间表示等有助于你开发特征和检测器,然后让端到端部分的问题变得容易得多。所以我怀疑——虽然我不知道,因为我不是团队的一员——但有很多预训练在进行,这样你就可以为端到端做微调。所以基本上,我觉得有必要逐步推进,特斯拉就是这么做的。我认为这是正确的方法,而且看起来有效,所以我非常期待。如果你一开始就做端到端,你也不会有数据。

To think about it that way, and I do suspect that the end-to-end systems for Tesla in, say, 10 years, it is just a neural net. I mean, the videos stream into a neural net and commands come out. You have to sort of build up to it incrementally and do it piece by piece. And even all the intermediate predictions and all the things that we've done, I don't think they've actually misled the development. I think they're part of it, because there's a lot of solid reasons for this. So actually, in driving, when you're just imitating humans and so on, you have very few bits of supervision to train a massive neural net, and it's too few bits of signal to train so many billions of parameters. And so these intermediate representations and so on help you develop the features and the detectors for everything, and then it makes a much easier problem for the end-to-end part of it. And so I suspect, although I don't know because I'm not part of the team, but there's a ton of pre-training happening so that you can do the fine-tuning for end-to-end. And so basically, I feel like it was necessary to eat through it incrementally, and that's what Tesla has done. I think it's the right approach, and it looks like it's working, so I'm really looking forward. If you had started end-to-end, you wouldn't have had the data anyway.

Host

有道理。

That makes sense.

从汽车到人形机器人的迁移 Transfer from cars to humanoid robots

Host

所以你离开前在特斯拉人形机器人项目上工作过。我有很多问题,但其中一个就是,从汽车开始,什么可以迁移?基本上所有东西都可以迁移,我不认为人们意识到这一点。

So you worked on the Tesla humanoid robot before you left. I have so many questions, but one is like, starting here, what transfers? Basically everything transfers, and I don't think people appreciate it.

Andrej

好吧,这是个大胆的说法。实际上,汽车基本上就是机器人。汽车是机器人。而特斯拉,我不认为它是一家汽车公司。我认为这是误导。这是一家机器人公司,一家规模化机器人公司。因为我认为“规模化”本身就是一个独立的变量。他们不是在建造单个东西,而是在建造制造东西的机器,这完全是另一回事。所以我认为特斯拉是一家规模化机器人公司。而且我认为从汽车到人形机器人的迁移,根本不需要太多工作。事实上,早期版本的 Optimus,那个机器人,它以为自己是一辆车,因为它有完全相同的计算机,完全相同的摄像头。这很有趣,因为我们在机器人上运行了汽车网络,但它在办公室里走来走去,它试图识别可行驶空间,但现在那都是步行空间了。但它实际上有点泛化了,需要一些微调等等。但它以为自己是在驾驶,但实际上是在环境中移动。这是一种合理的思考方式,实际上它就是一个机器人。很多东西可以迁移,但缺少的比如驱动和动作数据。你确实缺少一些组件。另一部分我要说的是,迁移的东西太多了。Optimus 启动的速度,我觉得非常令人印象深刻,因为当 Elon 说我们要做这个时,人们就带着所有正确的工具出现了,所有东西都很快到位。所有这些 CAD 模型和供应链的东西,我感觉,哇,特斯拉内部有这么多构建机器人的专业知识。而且都是同样的工具,它们就像是从汽车重新配置而来,就像电影《变形金刚》一样,它们被重新配置和重组,但本质上是同样的东西。你需要所有相同的组件,你需要考虑所有相同的事情,无论是在硬件方面、规模化方面,还是在“大脑”方面。对于“大脑”,也有大量的迁移,不仅仅是特定的网络,还有所有的方法、标注团队、如何协调以及人们采用的方法。我认为有大量的迁移。

Okay, that's a big claim. Like, cars are basically robots when you actually look at it. Cars are robots. And Tesla, I don't think it's a car company. I think this is misleading. This is a robotics company, a robotics at scale company. Because I would say 'at scale' is also like a whole separate variable. They're not building a single thing; they're building the machine that builds the thing, which is a whole separate thing. And so I think robotics at scale company is what Tesla is. And I think in terms of the transfer from cars to humanoids, it was not that much work at all. In fact, the early versions of Optimus, the robot, it thought it was a car because it had the exact same computer, the exact same cameras. It was really funny because we were running the car networks on the robot, but it's walking around the office and so on, and it's trying to recognize drivable space, but it's all just walking space now, I suppose. But it actually kind of generalized a little bit, and there's some fine-tuning necessary and so on. But it thought it was driving, but it's actually like moving through an environment. It's a reasonable way to think of this as like actually it's a robot. Many things transfer, but you're just missing, for example, actuation and action data. You definitely miss some components. And the other part I would say is like so much transfers. The speed with which Optimus was started, I think to me was very impressive because the moment Elon said we're doing this, just people showed up with all the right tools and all the stuff just showed up so quickly. And all these CAD models and all the supply chain stuff, and I just felt like, wow, there's so much in-house expertise for building robotics at Tesla. And it's all the same tools, and they're just like okay, they're being reconfigured from a car, like Transformers the movie, they're just being reconfigured and reshuffled, but it's like the same thing. And you need all the same components, you need to think about all the same kinds of stuff, both on the hardware side, on the scale stuff, and also on the brains. And so for the brains, there was also a ton of transfer, not just of the specific networks, but also all of the approach and labeling team and how it all coordinates and the approaches people are taking. I just think there's a ton of transfer.

人形机器人的首批应用领域 First application areas for humanoid robots

Host

你认为人形机器人或人形形态的第一应用领域是什么?我认为很多人有这种愿景,比如帮你洗衣服等等。我认为那会来得晚。我不认为 B2C 应该是正确的起点,因为我不认为我们可以有一个机器人像“压碎奶奶”那样。法律风险太大了。我的意思是,它可能会摔倒之类的。你知道,这些东西还不完美,需要一些工作。所以我认为最好的客户首先是你自己。而且我认为特斯拉可能会这样做。我非常看好特斯拉,如果人们能看出来。第一个客户是你自己,你在工厂里孵化它等等,做很多物料搬运之类的事情。这样你就不需要与第三方签订合同。那真的很重,涉及律师等等。你孵化它,然后你再去。我认为第二步是 B2B,你去其他有大型仓库的公司。我们可以做物料搬运,我们会做所有的事情。合同起草,围栏设置,所有这些。然后一旦你在公司中孵化,我认为那时你才开始进入 B2C 应用。我确实认为我们也会看到 B2C 机器人。比如宇树科技等公司开始推出我真正想要的机器人。我买了一个,你也是?是的,好的,G1。所以我可能会买一个,可能也会有一个生态系统,人们在这些平台上构建。但就什么能大规模获胜而言,我会期待那种方法。但一开始,是大量的物料搬运,然后走向越来越具体的 HKC 事情。我特别兴奋的一个是 N Freedman 的吹叶机挑战。是的,我希望 Optimus 能沿着街道走,踮着脚尖,捡起每一片叶子,这样我们就不需要吹叶机了。我认为这会成功,这是一个了不起的任务。所以我希望那是第一批应用之一。即使是耙叶子,也应该可以,非常安静。

What do you think of the first application areas for humanoid robotics or human form stuff? I think a lot of people have this vision of it like doing your laundry, etc. I think that will come late. I don't think B2C should be the right start point because I don't think we can have a robot like crush grandma. It's like too much legal liability. I mean, it's just going to fall over or something like that. You know, these things are not perfect yet and they require some amount of work. So I think the best customer is yourself first. And I think probably Tesla's going to do this. I'm very bullish on Tesla, if people can tell. The first customer is yourself, and you incubate it in the factory and so on, doing maybe a lot of material handling, etc. This way you don't have to create contracts working with third parties. It's all really heavy, there's lawyers involved, etc. You incubate it, then you go. I think B2B second, and you go to other companies that have massive warehouses. We can do material handling, we're going to do all the stuff. Contracts get drafted up, fences get put around, all this kind of stuff. And then once you incubate in companies, I think that's when you start to go into the B2C applications. I do think we'll see B2C robots also. Like Unitree and so on are starting to come up with robots that I really want. I got one, you did? Yeah, okay, yeah the G1. Yeah, so I will probably buy one of those, and there's probably going to be an ecosystem of people building on those platforms too. But I think in terms of like what wins at scale, I would expect that kind of approach. But in the beginning, it's a lot of material handling, and then going towards more and more HKC things that are more specific. One that I'm really excited about is the N Freedman challenge of the leaf blower. Yeah, like I would love for an Optimus to walk down the street, like tiptoe down the street, and like pick up individual leaves so that we don't need leaf blowers. And I think this will work, and it's an amazing task. And so I would hope that that's one of the first applications. Even raking, that should work too, just very quietly.

人形机器人论点:统一平台 vs 专用 Humanoid thesis: one platform vs specialized

Host

我们能谈谈人形机器人论点吗?因为最简单的版本是,世界是为人类建造的,你建造一套硬件。正确的做法是构建一个模型,可以在这套硬件上完成越来越多的任务。我认为还有另一个阵营认为,人类在任何特定任务上都不是最优的,对吧?你可以让他们更强、更大或更小等等,为什么我们不应该做超人类的事情?你怎么看?

Can we talk about the humanoid thesis for a second? Because the simplest version of this is like the world is built for humans, and you build one set of hardware. The right thing to do is build a model that can do an increasing set of tasks in this set of hardware. I think there's another camp that believes like, well, humans are not optimal for any given task, right? You can make them stronger or bigger or smaller or whatever, and why shouldn't we do superhuman things? How do you think about this?

Andrej

我认为人们可能低估了任何单一平台所涉及的固定成本的复杂性。我认为任何单一平台你都要付出巨大的成本。

I think people are maybe underappreciating the complexity of any fixed cost that goes into any single platform. I think there's a large cost you're paying for any single platform.

统一人形平台的优势 Benefits of a unified humanoid platform

Andrej

我认为集中在一个单一平台上、让它能做所有事情是非常合理的。人形机器人这一点也很有吸引力,因为人们可以很容易地远程操作它,这对数据收集极其有帮助。我觉得这一点通常被忽视了。当然还有为人类设计的世界等方面,这也很重要。我认为人形机器人平台会有一些变体,但任何平台都有很大的固定成本。最后一点是,你能从不同任务之间的迁移学习中获益巨大。在 AI 中,你真正想要的是一个能多任务处理的单一神经元,做很多事情——这才是智能和能力的来源。这也是语言模型如此有趣的原因:你有一个单一的领域,比如文本,它同时处理各种不同的问题,所有问题之间共享知识,耦合在一个神经网络中。你想要那种平台。你希望为某个任务收集的所有数据都能惠及其他所有任务。如果你为某个特定任务构建专用设备,你就无法从其他任务之间的迁移中获益。

I think it makes a lot of sense to centralize that and have a single platform that can do all the things. The humanoid aspect is also very appealing because people can teleoperate it very easily, and that's extremely helpful for data collection. I think that's usually overlooked. There's also the aspect of world design for humans, which is important. I think we'll have some variations on the humanoid platform, but there is a large fixed cost to any platform. One last dimension is that you benefit a ton from transfer learning between different tasks. In AI, you really want a single neuron that is multitasking, doing lots of things—that's where you get all the intelligence and capability. That's also why language models are so interesting: you have a single regime, like text, multitasking all these different problems, and they all share knowledge between each other, coupled in a single neural net. You want that kind of platform. You want all the data you collect for one task to benefit all the other tasks. If you build a special-purpose thing for any one thing, you won't benefit from the transfer between all the other tasks.

Host

我觉得有一个论点:G1 大概要 3 万美元,对吧?似乎很难在某个成本以下造出非常能干的人形机器人。如果你想在轮子上装个手臂,它也能做事。也许一开始有更便宜的方法来构建通用平台。你觉得有道理吗?

I think there's one argument: the G1 is like 30 grand, right? It seems hard to build a very capable humanoid robot under a certain cost. If you wanted to put an arm on wheels, it could do things. Maybe there are cheaper approaches to a general platform at the beginning. Does that make sense to you?

Andrej

从硬件角度来说,更便宜的通用平台方法?嗯,我觉得有道理。你给它装上轮子而不是脚,等等。但我确实觉得这可能会让你陷入局部最优。我只是觉得,选择一个平台,把它做到完美,长期来看是相当不错的赌注。另一件事是,它会让人们感到熟悉,我认为人们会理解这一点。也许你想和它说话。心理层面也倾向于人形平台,除非人们害怕它,更喜欢更抽象的平台。但如果它只是一个做事的怪物,我不确定那是否更好。Unitree 的另一种形态是狗,它更友好、更熟悉。但人们看了《黑镜》之后,突然觉得狗会变成可怕的东西。这很难想清楚。我只是觉得,从心理上讲,人们很容易理解发生了什么。

Cheaper approaches to a general platform from a hardware perspective? Yeah, I think that makes sense. You put a wheel on it instead of feet, etc. I do feel like I wonder if it's taking you down a local minimum a little bit. I just feel like pick a platform, make it perfect is the long-term pretty good bet. The other thing is that it will be kind of familiar to people, and I think people will understand that. Maybe you want to talk to it. The psychological aspect also favors the human platform, unless people are scared of it and would prefer a more abstract platform. But if it's just a monster doing stuff, I don't know if that's better. The other form factor for the unitree is a dog, and it's almost more friendly and familiar. But then people watch Black Mirror and suddenly the dog flips to a scary thing. It's hard to think through. I just think psychologically it will be easy for people to understand what's happening.

机器人技术进步的技术里程碑 Technological milestones for robotics progress

Host

你认为在实现机器人、人形机器人或其他任何东西的未来方面,还缺少哪些技术里程碑?

What do you think is missing in terms of technological milestones for progress relative to substantiating this future for robotics, humanoid robots, or anything else?

Andrej

我不太清楚具体是什么。但我确实觉得有趣的是,对于人形形态,比如下半身,我不确定你是否想通过演示来做模仿学习,因为那主要是倒立摆控制。上半身则需要大量的远程操作、数据收集和端到端训练。从这个意义上说,一切都变得非常混合,我不知道这些系统如何交互。当我与从事这方面工作的人交谈时,他们关注的重点很多是驱动、操作和数字操作。我预计一开始会大量使用远程操作来启动,模仿它,并让系统达到 95%的可靠性。然后我们会讨论人与机器人的比例,逐步让人类成为机器人的监督者,而不是直接执行任务。所有这些都会随着时间的推移逐渐发生。我不认为存在任何我特别熟悉的单个障碍。我只是觉得这需要大量的苦力活。工具已经存在:Transformer 就像一块美丽的组织,可以处理任意任务。你只需要数据,把它整理成正确的形式,训练它,实验它,部署它,迭代它。这只是一大堆苦力活。我不认为有某个单一的技术因素在阻碍我们。

I don't know that I have a really good window into it. I do think it's interesting that for the human form factor, for example, for the lower body, I don't know that you want to do imitation learning from demonstration because it's a lot of inverted pendulum control. For the upper body, you need a lot of teleoperation and data collection and end-to-end training. Everything becomes very hybrid in that sense, and I don't know how those systems interact. When I talk to people working on it, a lot of what they focus on is actuation, manipulation, and digital manipulation. I do expect in the beginning it's a lot of teleoperation to get stuff off the ground, imitating it, and getting something that works 95% of the time. Then we talk about human-to-robot ratios and gradually having people who are supervisors of robots instead of doing the task directly. All this kind of stuff will happen over time, pretty gradually. I don't know that there are any individual impediments that I'm really familiar with. I just think it's a lot of grunt work. The tools are available: Transformers are this beautiful blob of tissue that can handle arbitrary tasks. You just need the data, put it in the right form, train it, experiment with it, deploy it, iterate on it. That's just a lot of grunt work. I don't know that there is a single individual thing holding us back technically.

大型 blob 研究与 Transformer 扩展现状 State of large blob research and Transformer scaling

Host

我们在大型“组织”研究方面处于什么阶段?

Where are we in the state of large blob research?

Andrej

我们处于一个非常好的状态。我不确定是否被充分认识到,但 Transformer 远比另一个神经网络更惊人。它是一个极其通用的神经网络。例如,当人们谈论神经网络中的缩放定律时,缩放定律在很大程度上其实是 Transformer 的一个属性。在 Transformer 之前,人们使用 LSTM 并堆叠它们,但你得不到清晰的缩放定律,而且训练效果也不好。Transformer 是第一个真正能扩展的东西,你得到了缩放定律,一切都有意义了。所以这个通用训练计算机——我把它看作一种计算机,但它是可微分的。你可以给它数十亿的输入和输出,用反向传播训练它。它实际上会自我组织成执行任务的东西。我认为这是我们在算法空间偶然发现的神奇东西。其中有几个单独的创新:残差连接、层归一化、注意力模块,以及没有像 tanh 这样饱和的非线性激活函数,因为它们会扼杀梯度信号。大概有四五个创新,它们都存在,然后被组合成这个 Transformer,这就是谷歌在论文中所做的。这个东西真的能训练,突然你就得到了缩放定律,你有了这块组织,它可以大规模训练。这是一个重大的突破。

We're in a really good state. I'm not sure if it's fully appreciated, but the Transformer is much more amazing than just another neural net. It's an extremely general neural net. For example, when people talk about scaling laws in neural networks, the scaling laws are actually to a large extent a property of the Transformer. Before the Transformer, people were playing with LSTMs and stacking them, but you don't actually get clean scaling laws, and it doesn't actually train well. The Transformer was the first thing that just kind of scales, and you get scaling laws, and everything makes sense. So this general-purpose training computer—I think of it as a kind of computer, but it's a differentiable computer. You can give it inputs and outputs, billions of them, and train it with backpropagation. It actually kind of arranges itself into a thing that does the task. I think it's a magical thing that we've stumbled on in the algorithm space. There are a few individual innovations that went into it: residual connections, layer normalization, the attention block, and the lack of saturating nonlinearities like tanh, which kill gradient signals. There are four or five innovations that all existed and were put together into this Transformer, and that's what Google did with their paper. This thing actually trains, and suddenly you get scaling laws, and you have this piece of tissue that just trains to a very large extent. It was a major unlock.

Host

你觉得我们还没有接近这个突破的极限,对吧?因为我认为有一个关于数据墙以及另一个有多昂贵的讨论……

You feel like we are not near the limit of that unlock, right? Because I think there is a discussion about the data wall and how expensive another...

扩展与瓶颈 Scaling and Bottlenecks

Host

关于规模生成,你怎么看?这就开始涉及——我认为神经网络架构已经不再是根本性的瓶颈了。它不再是短板,而在 Transformer 之前,它曾是短板,但现在不是了。所以现在我们更多地在讨论损失函数是什么、数据集是什么。这些几乎成了瓶颈。它不再是那种可以根据你的需求任意重构的通用组织。所以我认为很多活动已经转移到了那里。这就是为什么很多公司和应用这项技术的人不再考虑 Transformer,不再考虑架构。你看,Llama 发布后,Transformer 并没有太大变化。我们添加了 RoPE 位置编码和 RoPE 相对位置编码。那是主要的变化。其他东西都不太重要,只是在小部分上提升了 3%。但实际上,RoPE 是唯一插入的东西,这就是过去五年左右 Transformer 的变化。所以这方面创新不多。每个人都认为这是理所当然的,直接训练等等。然后大家主要都在数据集和损失函数的细节上创新。所以所有活动都集中到了那里,对吧?

Generation of scale would be like how do you think about that? That's where you start to get into like I don't think that the neural network architecture is like holding us back fundamentally anymore. It's like not the bottom leg, whereas I think in the previous before Transformer it was a bottom leg, but now it's not the bottom leg. So now we're talking a lot more about what is the loss function, what is the data set. We're talking a lot more about those and those have become the bottlenecks almost. It's not the general piece of tissue that reconfigures based on whatever you want it to be. And so that's where I think a lot of the activity has moved. And that's why a lot of the companies and someone who are applying this technology like they're not thinking about the Transformer, they're not thinking about the architecture. You know, the Llama release, the Transformer hasn't changed that much. We've added the RoPE positional and the RoPE relative position encodings. That's like the major change. Everything else doesn't really matter too much. It's like plus 3% on small few things. But really it's like RoPE is the only thing that's slotted in and that's the Transformer as it has changed since the last five years or something. So there hasn't been that much innovation on that. Everyone just takes it for granted, let's train it, etc. And then everyone's just innovating on the data set mostly and the loss function details. So that's where all the activity has gone to, right?

Andrej

但有人会争论说,在那个领域,当我们使用互联网数据时更容易,而现在互联网数据用完了?所以问题实际上围绕合成数据或更昂贵的数据收集。所以我认为这是个好观点。这就是现在 LLM 领域很多活动的方向。互联网数据并不是你真正想要给 Transformer 的数据。它就像一个最近邻,但令人惊讶地能带你走得很远。但互联网数据就是一堆网页,对吧?你真正想要的是你大脑的内部思维独白。是的,那是你大脑中的理想轨迹,你在解决问题时大脑中的轨迹。如果我们有十亿条那样的数据,AGI 基本上就实现了。我是说在很大程度上,但我们没有。所以我认为现在很多活动是,互联网数据实际上让你非常接近,因为互联网恰好包含了足够的推理痕迹和大量知识,而 Transformer 让它工作。所以我认为现在很多活动围绕将数据集重构为这些内部独白格式。我认为有大量的合成数据生成对此有帮助。所以有趣的是,当前模型在多大程度上帮助我们创建下一代模型。这有点像改进的阶梯。你认为合成数据有多重要,或者它能带我们走多远?因为正如你所说,每个模型都能帮助训练后续模型,至少创建工具、数据标注等,其中一部分就是合成数据。你认为合成数据这部分有多重要?

But what about the argument like in that domain that that was easier when we were taking internet data and we're out of internet data? And so the questions are really around like synthetic data or more expensive data collection. So I think that's a good point. So that's where a lot of the activity is now in LLMs. So the internet data is like not the data you want for your Transformer. It's like a nearest neighbor that actually gets you really far surprisingly. But the internet data is a bunch of internet web pages, right? It's just like what you want is the inner thought monologue of your brain. Yeah, that's the ideal trajectories in your brain, the trajectories in your brain as you're doing problem solving. If we had a billion of that, like AGI is here roughly speaking. I mean to a very large extent, and we just don't have that. So where a lot of activity is now I think is we the internet data that actually gets you like really close because it just so happens that internet has enough of reasoning traces in it and a bunch of knowledge and the Transformer just makes it work. Okay, so I think a lot of activity now is around refactoring the data set into these inner monologue formats. And I think there's a ton of synthetic data generation that's helpful for that. So what's interesting about that also is like the extent to which the current models are helping us create the next generation models. And so it's kind of like you know the staircase of improvement. How much do you think synthetic data is or how far does that get us right? Because to your point, on each data, each model helps you train the subsequent model better, at least create tools for it, data labeling, whatever may be part of it is synthetic data. How important do you think the synthetic data piece is?

合成数据与熵 Synthetic Data and Entropy

Host

是的,我认为这是我们取得进展的唯一途径。我们必须让它成功。我认为使用合成数据时必须小心,因为模型会无声崩溃,这是主要问题之一。如果你去问 ChatGPT 要一个笑话,你会发现它只知道大约三个笑话。大多数时候它只给你一个笑话,有时给你三个。这是因为模型崩溃了,而且是无声的。所以当你查看任何单个输出时,你只看到一个例子。但当你真正查看分布时,你会发现它不是一个非常多样化的分布。它无声地崩溃了。当你在做合成数据生成时,这是一个问题,因为你实际上非常需要那种熵。你需要数据集的多样性和丰富性。否则你会得到崩溃的数据集,而且你在查看任何单个例子时看不到,但分布已经失去了大量的熵和丰富性。所以它会无声地变差。这就是为什么你必须非常小心,确保在数据集中保持熵。有很多技术可以实现这一点。例如,有人发布了 Persona 数据集。Persona 数据集包含 10 亿个人格,像人类一样,有背景:“哦,我是一名教师”或“我是一名艺术家,我住在这里,我做这个等等。”它就像一小段虚构的人类背景。当你做合成数据生成时,不仅要“完成这个任务并以这种方式做”,还要“想象你在向这个人描述它”。你输入这些信息,现在你迫使它探索更多的空间,你得到了一些熵。所以我认为你必须非常小心地注入熵,保持分布,这是困难的部分,我认为总体上人们可能没有充分认识到这一点。所以我认为基本上合成数据绝对是未来。我的印象是我们不会用完数据。我只是认为你必须小心。

Yeah, I think it's the only way we can make progress. We have to make it work. I think with synthetic data you just have to be careful, because these models are silently collapsed is one of the major issues. So if you go to ChatGPT and you ask it to give you a joke, you'll notice that it only knows like three jokes. That's like the only it gives you like one joke I think most of the time, and sometimes it gives you like three jokes. And it's because the models are collapsed and it's silent. So when you're looking at any single individual output, you're just seeing a single example. But when you actually look at the distribution, you'll notice that it's not a very diverse distribution. It's silently collapsed. When you're doing synthetic data generation, this is a problem because you actually really want that entropy. You want the diversity and the richness in your data set. Otherwise you're getting collapsed data sets and you can't see it when you look at any individual example, but the distribution has lost a ton of entropy and richness. And so it silently gets worse. And so that's why you have to be very careful and you have to make sure that you maintain your entropy in your data set. And there's a ton of techniques for that. As an example, someone released this Persona dataset. The Persona dataset is a dataset of 1 billion personalities, like humans, like backgrounds: "Oh, I'm a teacher" or "I'm an artist, I live here, I do this, etc." And it's like little paragraphs of fictitious human background. And what you do when you do synthetic data generation is not only like "Oh, complete this task and do it in this way" but also "imagine you're describing it to this person." You put in this information and now you're forcing it to explore more of the space and you're getting some entropy. So I think you have to be just very careful to inject the entropy, maintain the distribution, and that's the hard part that I think maybe people aren't sufficiently appreciating as much in general. So I think basically synthetic data is absolutely the future. We're not going to run out of data is my impression. I just think you have to be careful.

了解人类认知 Learning About Human Cognition

Host

你认为我们从这项研究中学到了关于人类认知的什么?我不知道我们是否学到了。有人可能会说,弄清楚我们想要的推理轨迹的形状,例如,对于理解大脑如何工作是有启发性的。我会小心这些类比。但总的来说,我认为这是完全不同的事情。但我确实认为可以画一些类比。例如,我认为 Transformer 在很多方面实际上比人类大脑更好。我认为它们实际上是一个更高效的系统。它们不如人类大脑工作的原因主要是数据问题,粗略地说,我认为这是第一近似。实际上,举个例子,Transformer 记忆序列的能力比人类强得多。如果你给它一个序列,并对该序列进行一次前向-后向传播,那么如果你给它前几个元素,它就会完成序列的其余部分。它记住了那个序列,而且非常擅长。如果你给人类一次序列展示,他们不可能记住。所以 Transformer 实际上,我认为基于梯度的优化,我们一直用于训练神经网络的前向-后向更新,在某些方面实际上比大脑更高效。这些模型更好,只是还没有准备好大放异彩。但在许多认知方面,我认为它们……

What do you think we are learning now about human cognition from this research? I don't know if we're learning. One could argue that figuring out the shape of reasoning traces we want for example is instructive to actually understanding how the brain works. I would be careful with those analogies. But in general, I do think that it's a very different kind of thing. But I do think that there are some analogies you can draw. So as an example, I think Transformers are actually better than the human brain in a bunch of ways. I think they're actually a lot more efficient system. And the reason they don't work as good as the human brain is mostly data issue, roughly speaking, is the first story approximation I would say. And actually as an example, Transformer memorizing sequences is so much better than humans. If you give it a sequence and you do a single forward backward pass in that sequence, then if you give it the first few elements, it will complete the rest of the sequence. It memorized that sequence and it's so good at it. If you gave a human a single presentation of a sequence, there's no way that you can remember that. And so the Transformer is actually, I do think there's a good chance that the gradient-based optimization, the forward backward update that we do all the time for training neural nets, is actually more efficient than the brain in some ways. And these models are better, they're just not yet ready to shine. But in a bunch of cognitive sort of aspects, I think they...

人脑与 Transformer 的约束 Human brain vs. Transformer constraints

Host

也许在正确的输入下,它们会更好,这对各种应用来说普遍成立,对吧,就像你提到的记忆问题,没错。

might come out with the right inputs they will be better that that's generically true of computers for all sorts of applications right put memory to your point yeah exactly

Andrej

而且我认为人脑有很多限制,比如工作记忆非常小。我认为 Transformer 的工作记忆要大得多,而且这种情况会持续下去。它们的学习效率也高得多。人脑在各种约束下运作,它是否使用反向传播并不明显,也不清楚它是如何工作的。它是一个非常随机的动态系统,在环境条件等约束下运作。所以我确实认为我们拥有的东西实际上可能比大脑更好,只是还没达到那个水平。

and I think human brains just have a lot of constraints you know the working memory is very small I think Transformers have a lot lot bigger working memory and this will continue to be the case uh they're much more efficient Learners uh the human brains function under all kinds of constraints uh it's not obvious that the human brain is back propagation right it's not obvious how that would work it's uh very stochastic sort of dynamic system it has all these constraints it works under so ambient conditions Etc so I I do think that what we have is actually potentially better than the brain um and it's just not there yet

AI 增强人类 Human augmentation with AI

Host

你怎么看待随着时间推移,用不同 AI 系统增强人类?你认为这是可能的方向还是不太可能?

how do you think about um human augmentation with different AI systems over time do you think that's a likely direction do you think that's unlikely

Andrej

用 AI 模型增强人类?当然,但具体是哪种意义?我认为总体上绝对是的。有一个抽象版本,就是把它当作工具,这是外部版本;还有融合场景,很多人都在讨论。是的,我们已经在某种程度上融合了。问题在于输入输出瓶颈,但大多数情况下,如果你有这些模型,它们就在你的指尖。这有点不同,因为我认为人们已经争论了四五十年,技术工具只是人类能力的延伸。没错,计算机是人类思维的自行车。

augmentation augmentation of people with AI models oh of course I mean but in in what sense maybe I think in general absolutely because I mean there's an abstract version of it you're using as a tool that's the the external version there's the you know the merger scenario you know a lot of people end up talking about yeah yeah I mean we're already kind of merging the thing is like there's a you know there's the io bottleneck but for the most part you know at your fingertips if you have any of these models you're that's a little bit different because I mean people have been making that argument for I think 40 50 years where uh technological tools are just extension of human capabilities right yeah the computer is the bicycle for human mind exactly

Host

但是 AI 社区中有一部分人认为,我们化解与未来 AI 潜在冲突的方式,可能是通过某种形式的神经链接等。

so um but there's a subset of the AI community that thinks that for example the way that we subsume some potential conflict with future AI or something else would be through some form of yeah like the neuralink pitch Etc exactly

Andrej

嗯,我还不知道这种融合会是什么样子,但我确实看到你想减少工具使用的输入输出瓶颈。我认为这就像一种外皮层,建立在我们新皮层之上,只是下一层,结果它就在云端等等,但它是大脑的下一层。

um yeah I don't I don't know what this merger looks like uh yet but I can definitely see that you want to decrease the io to Tool use and I see this as kind of like an EXO cortex while building on top of our neocortex right and it's just the next layer and uh it just turns out to be in the cloud Etc but it is the next layer of the brain yeah

Host

21 世纪初的《加速》一书中有个版本,基本上所有东西都体现在一副计算连接到大脑的护目镜里,你戴着它,如果丢了,你会感觉失去了部分人格或记忆。我认为这很有可能。今天手机已经差不多是这样了,而且会更进一步。当你把科技产品拿开,你就像赤裸的人类,失去了部分智力,这很令人焦虑。一个很简单的例子就是地图。我注意到现在很多人不能很好地导航自己的城市,因为他们总是用逐向导航。如果我们有通用翻译器,我认为它不远了,如果你把设备拿开,你就无法和不说英语的人交流。我很乐意重新利用那部分大脑去做进一步研究。

accelerando uh book from the early 2000s has a version of this where basically everything is substantiated in a set of goggles that are computationally attached to your brain that you wear and then if you lose them you must feel like you're losing a part of your persona or memory I think that's very likely yeah and today the phone is already almost that and I think it's going to get worse when you put your techno stuff away from you you're just like naked human in nature as well you lose part of your intelligence it's very anxiety a very uh a very simple example of that is just Maps right so a lot of people now I've noticed can't actually navigate their City very well anymore because are always using turn by turn direction and if we have this for example like Universal translator which I don't think is too far away like you lose the ability to speak to people who don't speak English if you just put your stuff away I'm very comfortable repurposing that part of my brain to do further research

Andrej

我不知道你有没有看过一个视频,一个孩子拿着杂志,试图在上面滑动。我觉得有趣的是,这个孩子分不清什么是自然的东西,什么是建立在自然之上的技术,因为技术做得太透明了。我认为这可能会类似,人们会开始默认使用工具,当你拿走它们时,你才意识到人们不知道什么是技术,什么不是。如果你戴着这个总是为你翻译或做其他事情的东西,那么也许人们会失去基本的认知能力,这些能力可能就不存在了。我认为存在,因为本质上我们会专业化。你无法理解说西班牙语的人,这算怎么回事?或者像在迪士尼乐园,所有物体都是活的,我们可能会进入这样一个世界:为什么我不能和东西说话?就像今天你可以和 Alexa 说话,让它做事。我见过一些玩具公司,他们基本上试图在玩具中嵌入语言模型,让它能和孩子互动。你不觉得奇怪吗?当你走到一扇门前,你不能直接说“开门”,这算怎么回事?另一个我喜欢的例子,不知道你看过《超级战警》或《我,机器人》没有,人们嘲笑那种你不能和东西说话的想法,这算怎么回事?

I don't know if you saw the video of like the kid that's has a magazine and is trying to like swipe on the magazine what's fascinating to me about it is like this kid doesn't understand what comes with nature and what's technology technology on top of the nature because it's made so transparent and I think this might look similar where people will just start assuming the tools and then when you take them away you realize like I guess like people don't know what's technology and what's not if you're wearing this thing that's always translating one or like doing stuff like that for you then um maybe people like lose the basic cognitive abilities may not exist I think exist yeah like by Nature we're going to specialize you can't understand people who speak Spanish like what the hell or like if when you go to objects like in Disney uh all the objects are alive and I think we are going to potentially come to that kind of a world where why an I talk to things like already today you can talk to Alexa and you can ask her for things and so on yeah I've seen some toy companies like that where they're basically trying to embed an LM in a toy that can interact with a child yeah isn't it strange that when you go to a door you can't just say open like what the hell um another favorite example of that I don't know if you saw either demolition men or iRobot people make fun of the idea that you uh yeah like you can't just talk to things and what the hell

外脑与民主化 Exocortex and democratization

Host

如果我们谈论的是外皮层,那么民主化访问似乎是一个根本重要的事情。你认为当前 LLM 研究的市场结构如何?你知道只有少数大型实验室有机会进行下一代预训练,这如何转化为人们未来能访问的东西?

if we're talking about a um EXO cortex that feels like a pretty fundamentally um important thing to democratize access to how do you think like the current market structure of what's happening in llm research you know there's a small number of large Labs that actually have a shot at the Next Generation progressing training like how does that translate to what people have access to in the future

Andrej

你暗示的可能是生态系统的状态。我们有几个封闭平台的寡头垄断,然后有一个开放平台,比如 Meta 的 Llama 等,这有点像开源生态系统。我认为当我们开始把它视为外皮层时,加密货币中有句话:不是你的密钥,就不是你的代币。那么,如果不是你的权重,就不是你的大脑吗?这很有趣,因为一家公司实际上在控制你的外皮层,因此部分会感觉有点侵入性。如果这是我的外皮层,我认为人们会更关心所有权。是的,你意识到你在租用你的大脑,这似乎很奇怪。思想实验是:你愿意放弃所有权和控制权来租用一个更好的大脑吗?因为我是愿意的。所以我认为这是一种权衡。我们会看到它如何运作,但也许默认使用封闭版本,因为它们很棒,但在各种场景下有备用方案。我认为今天的情况已经如此,当一些封闭提供商的 API 宕机时,人们开始实施备用方案,比如完全控制的开放生态系统,他们感到被赋能。所以也许这只是大脑的延伸:如果发生什么事,你回退到开源的东西,但大多数时候你使用封闭的。所以开源继续进步非常重要,我 100%同意。这不是一个显而易见的观点,也不是人们现在都同意的。

so what you were kind of alluding to maybe is the state of the ecosystem right so we have kind of like an oligopoly of a few closed platforms and then we have an open platform that is kind of like behind so like metal Lama Etc and this is kind of like mirroring the open source uh kind of ecosystem I do think that when this stuff starts to when we start to think of it as like an exocortex uh so there's the there's a saying in crypto which is like not your keys not your not your tokens like is it the case that if it's like not your weights not your brain that's interesting because a company is effectively controlling your exocortex and therefore for part of it starts to feel kind of invasive if this is my exocortex I think people will care much more about ownership yes like you're yeah you're you realize you're renting your brain like it seems strange to rent your brain the thought experiment is like are you willing to give up ownership and control to rent a better brain because I am yeah yeah so I think that's the trade-off I think we'll see how that works but maybe it's possible to like by default use the closed versions because they're amazing but you have a fallback in various scenarios and I think that's kind of like the way things are shaping up today even right like um when apis go down on some of the closed source providers people start to implement fallbacks to like the open ecosystems for example that they fully control on they're in they feel empowered by that right so so maybe that's just the extension of what will look like for the brain is you fall back on the open source stuff um should anything happen but most of the time you actually so it's quite important that the open source stuff continues to progress I think so 100% and this is not like an obvious point or something that people maybe agree on right now but I

最小性能模型尺寸 Smallest performant model size

Host

我一直在想的一个问题是,你能得到的最小性能模型是什么,无论是参数规模还是其他衡量方式。我很好奇你的看法,因为你思考过很多关于蒸馏和小模型的问题。

I guess one thing I've been wondering about is what is the smallest performant model you can get, either in parameter size or however you want to think about it. I'm curious about your view because you've thought a lot about distillation and small models.

Andrej

我认为它可以小得惊人。我确实认为当前模型浪费了大量容量去记忆无关紧要的东西,比如 SHA 哈希或古老琐事,因为数据集没有很好地整理。我认为这种情况会消失,我们只需要达到认知核心,它可以非常小。它只是一个会思考的东西,如果需要查找信息,它知道如何使用工具。

I think it can be surprisingly small. I do think current models are wasting a ton of capacity remembering stuff that doesn't matter, like SHA hashes or ancient trivia, because the dataset isn't curated well. I think this will go away, and we just need to get to the cognitive core, which can be extremely small. It's just this thing that thinks, and if it needs to look up information, it knows how to use tools.

Host

那是 30 亿参数还是 200 亿?

Is that like 3 billion parameters or 20 billion?

Andrej

我认为甚至 10 亿就够了。我们可能会达到那个点,模型可以非常非常小。它们能非常小的根本原因是蒸馏出奇地有效。蒸馏就是用一个大模型或大量算力来监督一个小模型,你可以把很多能力塞进很小的规模里。

I think even a billion suffices. We'll probably get to that point, and the models can be very, very small. The reason they can be very small is fundamentally that distillation works surprisingly well. Distillation is where you get a really big model or a huge amount of compute to supervise a very small model, and you can stuff a lot of capability into a very small level.

Host

有没有某种数学或信息论的表述?因为感觉你现在应该能计算出来。

Is there some mathematical or information-theoretic formulation of that? Because it feels like you should be able to calculate that now.

Andrej

也许一种思考方式是,互联网数据集里大约 0.001%是认知,99.999%是信息,其中大部分对思考部分没有用。

Maybe one way to think about it is that the internet dataset is like 0.001% cognition and 99.999% information, most of which is not useful for the thinking part.

Host

也许另一种提问方式是:是否存在认知能力相对于模型大小的数学表示?如何根据你要完成的目标来捕捉认知的最小或最大值?也许没有好的表示方法。

Maybe another way to frame the question is: is there a mathematical representation of cognitive capability relative to model size? How do you capture cognition in terms of min or max relative to what you're trying to accomplish? Maybe there's no good way to represent that.

Andrej

我认为也许 10 亿参数能给你一个好的认知核心。我觉得甚至 10 亿都太多了。我不知道,我们拭目以待。考虑到边缘设备与云端的对比,以及使用模型的原始成本,这非常令人兴奋。

I think maybe a billion parameters gets you a good cognitive core. I think even 1 billion is too much. I don't know, we'll see. It's very exciting given the question of edge device versus cloud, and the raw cost of using the model.

Host

但低于 10 亿参数,我也可以在本地设备上拥有我的外部大脑。

But at less than a billion parameters, I have my Exocortex on a local device as well.

Andrej

而且可能不是一个单一模型。思考这将如何发展很有趣。你想要受益于并行化,而不是顺序过程。公司在某种程度上也是工作的并行化,但有层级结构,因为信息处理和缩减需要在组织内进行。所以我认为我们最终会看到模型公司。很可能你会拥有不同能力的模型,专门用于各种领域——也许一个程序员等等——它将在很大程度上开始类似于公司。你有程序员和项目经理,类似角色的 LLM 并行工作,为你编排计算。所以也许更像一个蜂群,一个生态系统,有专门的角色和生态位。你会根据问题的难度自动升级到蜂群的其他部分。CEO 是一个极其聪明的云端模型,但主力模型可以便宜得多,甚至可能是开源模型。我的成本函数和你的不同。所以这可能很有趣。

And probably it's not a single model. It's interesting to think about how this will play out. You want to benefit from parallelization, not a sequential process. Companies to some extent are also a parallelization of work, but with a hierarchy because information processing and reductions need to happen within an organization. So I think we'll end up with companies of models. It's not unlikely that you have models of different capabilities specialized to various domains—maybe a programmer, etc.—and it will start to resemble companies to a large extent. You have the programmer and the program manager, similar roles of LLMs working in parallel, orchestrating computation on your behalf. So maybe it's more like a swarm, an ecosystem, where you have specialized roles and niches. You'll have automatic escalation to other parts of the swarm depending on the difficulty of the problem. The CEO is a really brilliant cloud model, but the workhorse can be a lot cheaper, maybe even open-source models. My cost function is different from your cost function. So that could be interesting.

为何重视教育与赋能 Why education and empowering people

Host

你离开了 OpenAI,正在从事教育。你一直是个教育者。为什么这么做?

You left OpenAI, you're working on education. You've always been an educator. Why do this?

Andrej

我一直是个教育者。我热爱学习和教学,所以这是一个我长期充满热情的空间。另一件事是,我认为 AI 领域有很多活动,但大多数是为了取代或排挤人。但我总是对增强人的能力的事情更感兴趣。我站在人类这边。我对 AI 能做什么来增强人的能力感兴趣。我不希望未来人们被自动化边缘化;我希望他们处于被赋能的状态,甚至比今天更出色。另一个方面是,如果一个人在所有科目上都有完美的导师,他能走多远。我认为有了完美的课程,人们可以在任何事情上走得很远。我们看到富人有导师——他们确实走得很远。所以我认为我们可以用 AI 接近这一点,甚至超越它。80 年代有一对一辅导的明确文献:人们能提高一个标准差,还是两个?是布鲁姆的研究。

I've always been an educator. I love learning and teaching, so it's a space I've been passionate about for a long time. Another thing is that I think there's a lot of activity in AI, and most of it is to replace or displace people. But I'm always more interested in things that empower people. I'm on team human. I'm interested in what AI can do to empower people. I don't want a future where people are on the side of automation; I want them to be in an empowered state, even much more amazing than today. Another aspect is how far a person can go if they have the perfect tutor for all subjects. I think people could go really far with the perfect curriculum for anything. We see that with rich people who have tutors—they do go really far. So I think we can approach that with AI, or even surpass it. There's clear literature from the 80s on one-on-one tutoring: people get one standard deviation better, or is it two? It's the Bloom stuff.

Host

你如何看待通过 AI 实现这一点?首先会有哪些类型的产品真正有帮助?有像《钻石时代》这样的书,里面提到了年轻女士的图解入门书。

How do you view that substantiating through the lens of AI? What are the first types of products that will really help? There are books like The Diamond Age where they talk about the young lady's illustrated primer.

Andrej

我确实从中受到启发。在实践中,我正在尝试构建一门课程,就像你想学 AI 时会去上的那种课。问题在于,我已经教过像斯坦福的 231n 这样的课程,那是第一门深度学习课,相当成功。但问题是如何真正扩展这些课程,让你的目标受众可能是地球上的 80 亿人,说不同的语言,处于不同的能力水平。一个老师无法扩展到那样的受众。所以问题是如何利用 AI 来扩展一位优秀教师的能力。

I'm definitely inspired by aspects of it. In practice, what I'm doing is trying to build a single course that is like the course you would go to if you want to learn AI. The problem is that I've already taught courses like 231n at Stanford, which was the first deep learning class and was pretty successful. But the question is how to really scale these classes so that your target audience is maybe 8 billion people on Earth, speaking different languages, at different capability levels. A single teacher doesn't scale to that audience. So the question is how to use AI to do the scaling of a really good teacher.

AI 作为学生的前端 AI as front-end to students

Andrej

老师主要负责课程创作和课程设计,因为以目前的 AI 能力,我认为模型还不足以创作出好的课程。但我认为它们可以成为面向学生的前端,为学生解读课程。所以基本上,老师不再直接面对学生;老师不再是前端。老师在后端设计课程材料,而 AI 是前端。它能说各种不同的语言,并带你完成课程。

The teacher is doing a lot of the course creation and curriculum, because at current AI capability, I don't think the models are good enough to create a good course. But I think they're good to become the front end to the student and interpret the course to them. So basically, the teacher doesn't go to the people; the teacher is not the front end anymore. The teacher is on the back end designing the materials in the course, and the AI is the front end. It can speak all the different languages and takes you through the course.

Host

我应该把它理解为助教式的体验吗,还是说这个类比不太恰当?

Should I think of that as like the TA type experience, or is that not a good analogy?

Andrej

这是我考虑的一种方式:它是一个 AI 助教。我主要把它看作面向学生的前端,是实际与学生互动并带他们完成课程的东西。我认为这在今天是可行的,而且目前还不存在。我认为它可以做得非常好。然后随着时间的推移,随着能力的提升,你可能会以各种方式重构这个设置。我喜欢找到那些当前 AI 能力与对其良好理解相匹配的事情。我认为很多公司可能没有直观地理解当前的能力在哪里,最终构建的东西要么太超前于现有水平,要么不够有野心。所以我确实认为这是一个可能性的甜蜜点,也非常有趣和令人兴奋。

That is one way I'm thinking about it: it's an AI TA. I'm mostly thinking of it as this front end to the student, the thing that's actually interfacing with the student and taking them through the course. I think that's tractable today, and it just doesn't exist. I think it can be made really good. Then over time, as capability increases, you would potentially refactor the setup in various ways. I like to find things where the AI capability today and having a good model of it. I think a lot of companies that maybe don't quite understand intuitively where the capability is today end up building things that are too ahead of what's available or maybe not ambitious enough. So I do think this is a sweet spot of what's possible and also really interesting and exciting.

更好工具下的人类性能极限 Limits of human performance with better tooling

Host

我想回到你之前说的一句话,我觉得非常鼓舞人心,尤其是考虑到你的背景和对研究现状的理解,那就是我们不知道在拥有更好工具的情况下,人类从学习角度出发的表现极限在哪里。我认为有一个很简单的类比:一个月前我们刚举办了奥运会。一个跑步者的最佳英里成绩,或者任何运动项目,今天都比 10 年前好得多,抛开兴奋剂不谈,仅仅是因为你更早开始训练,有非常不同的训练计划,我们有更好的科学理解、技术和体育教育。你相信如果我们从工具和课程开始,人类可以走得更远,这太棒了。

I want to go back to something you said that I think is very inspiring, especially coming from your background and understanding of where exactly we are in research, which is essentially that we do not know what the limits of human performance from a learning perspective are given much better tooling. I think there's a very easy analogy: we just had the Olympics a month ago. A runner's best mile time or any sport today is much better than it was 10 years ago, putting aside performance-enhancing drugs, just because you start training earlier, you have a very different program, we have much better scientific understanding, technique, and PE. The fact that you believe we can get much further as humans if we're starting with the tooling and the curriculum is amazing.

Andrej

是的,我认为我们甚至还没有触及可能的皮毛。所以我认为基本上有两个维度。第一个是全球化的维度:我希望每个人都能获得真正好的教育。但另一个是单个个体能走多远。我认为这两者都非常有趣和令人兴奋。

Yeah, I think we haven't even scratched what's possible at all. So I think there are basically two dimensions to it. Number one is the globalization dimension: I want everyone to have really good education. But the other one is how far can a single person go. I think both of those are very interesting and exciting.

当前 AI 教育的适应性与覆盖范围 Adaptivity vs reach in AI education today

Host

通常当人们谈论一对一学习时,他们谈论的是适应性方面,即根据一个人的水平来挑战他们。你认为今天能用 AI 做到这一点吗,还是说这是未来的事情?今天更多的是关于覆盖范围、多语言和全球性的事情吗?

Usually when people talk about 1-on-1 learning, they talk about the adaptive aspect, where you challenge a person at the level they're at. Do you think you can do that with AI today, or is that something for the future? And is it more today about reach, multiple languages, and global things?

Andrej

例如,不同的语言是超级低垂的果实。我认为当前的模型在翻译方面确实很好,可以针对材料进行实时翻译。所以我认为很多事情都是低垂的果实。而对个人背景的适应性,我认为不是低垂的果实,但我觉得也不是太高或太远。但肯定有你需要的东西,因为不是每个人都有相同的背景。另外,非常有帮助的是,如果你过去熟悉其他学科,那么与你已知的事物进行类比非常有用,这在教育中极其强大。所以这绝对是你想要利用的一个维度。但我认为这开始变得不那么明显,需要一些工作。我认为简单的版本并不遥远:你可以想象直接提示模型,比如‘嘿,我懂物理’或‘我懂这个’,你可能会得到一些东西。但我想我指的是真正有效的东西,而不是那种可以演示但有时有效的东西。我的意思是它真正像人一样有效。

For example, different languages are super low-hanging fruit. I think the current models are actually really good at translation and can target the material and translate it on the spot. So I think a lot of things are low-hanging fruit. This adaptability to a person's background, I think, is not low-hanging fruit, but I don't think it's too high up or too far away. But there is something you definitely want, because not everyone comes in with the same background. Also, what's really helpful is if you're familiar with some other disciplines in the past, then it's really useful to make analogies to things you know, and that's extremely powerful in education. So that's definitely a dimension you want to take advantage of. But I think that starts to get to the point where it's not obvious and needs some work. I think the easy version of it is not too far: you can imagine just prompting the model, like 'Oh hey, I know physics' or 'I know this,' and you probably get something. But I guess what I'm talking about is something that actually works, not something that you can demo and works sometimes. I just mean it actually really works in the way a person would.

Host

是的,这就是我问适应性的原因,因为人们学习的速度不同,或者某些东西对一些人来说有挑战性而对另一些人没有,反之亦然。所以这有点像是如何根据那个背景建模。我想你可以随着时间的推移,将一个人擅长或不擅长的信息重新引入模型。这就是 AI 的问题:我觉得很多这些能力都只是提示一下就能得到,所以你总是能得到演示,但你真的能得到产品吗?所以从这个意义上说,我会说演示很近,但产品很远。

Yeah, and that's the reason I was asking about adaptability, because also people learn at different rates, or certain things they find challenging that others don't, or vice versa. So it's a little bit of how do you model relative to that context. And I guess you could have some reintroduction of what the person is good or bad at into the model over time. That's the thing with AI: I feel like a lot of these capabilities are just kind of prompt away, so you always get demos, but do you actually get a product? So in this sense, I would say the demo is near, but the product is far.

AI 教育的传承与传播 Lineage and propagation in AI education

Host

我们之前谈到的一件事,我觉得非常有趣,就是研究界中发生的某种传承,你来自某些实验室,每个人都在八卦彼此来自哪个实验室。我认为很高比例的诺贝尔奖得主实际上曾在另一位诺贝尔奖得主的实验室工作过。所以存在某种文化、知识、品牌等的传播。在以 AI 教育为中心的世界里,你如何保持传承,还是说这并不重要?你如何看待这些网络和知识的传播方面?

One thing we were talking about earlier which I think is really interesting is sort of lineages that happen in the research community, where you come from certain labs and everybody gossips about being from each other's labs. I think a very high proportion of Nobel laureates actually used to work in a former Nobel laureate's lab. So there's some propagation of culture, knowledge, branding, or what. In an AI education-centric world, how do you maintain lineage, or does it not matter? How do you think about those aspects of propagation of network and knowledge?

Andrej

我其实不想生活在一个传承太重要的世界里。所以我希望 AI 能帮助摧毁一点那种结构。这感觉像是某种有限稀缺资源的守门,比如‘哦,有这种传承的人数量有限。’所以我觉得有点那个意思。所以我希望它能摧毁它。

I don't actually want to live in a world where lineage matters too much. So I'm hoping that AI can help destroy that structure a little bit. It feels like kind of gatekeeping by some finite scarce resource, like 'Oh, there's a finite number of people who have this lineage.' So I feel like it's a little bit of that aspect. So I'm hoping it can destroy that.

Host

这肯定是一部分:实际学习,一部分是血统。是的,嗯,这也是集群效应的聚合。就像为什么所有 AI 社区都在湾区,或者为什么大多数金融科技社区在纽约?所以我认为很多也是将真正聪明、有共同兴趣和信念的人聚集在一起,然后他们从那个共同核心传播,并以有趣的方式分享知识。

It's definitely one piece: actual learning, one piece pedigree. Yeah, well, it's also the aggregation of a cluster effect. It's like why is all of the AI community in the Bay Area, or why is most of the fintech community in New York? So I think a lot of it is also just clustering really smart people with common interests and beliefs, and then they propagate from that common core and share knowledge in an interesting way.

Andrej

很多这种行为在某种程度上转移到了线上,尤其是对年轻人来说。我认为其中一个方面是教育方面:如果你今天是一个社区的一部分,你会得到大量的教育和学徒训练,这非常有帮助,让你在该领域达到一个 empowered 的状态。我认为另一部分是文化方面:你的动力是什么,你想做什么,文化重视什么,他们把什么奉为圭臬,以及他们基本上崇拜什么。

You got a lot of that behavior shifted online to some extent, particularly for younger people. I think one aspect of it is the educational aspect: if you're part of a community today, you're getting a ton of education and apprenticeship, which is extremely helpful and gets you to a point of empowered state in that area. I think the other piece of it is the cultural aspect: what you're motivated by, what you want to work on, what the culture prizes, what they put on a pedestal, and what they worship basically.

文化对动机与地位的影响 Cultural influences on motivation and status

Andrej

例如在学术界,大家都很在意 H 指数和发表论文的数量。我曾是这个群体的一员,看到了这一点。现在我去到了不同的地方,不同的群体有不同的偶像。我认为这对人们的动机、社会地位来源以及真正重要的事情有着巨大影响。我在斯洛伐克长大,那是一个非常不同的环境,在加拿大也是另一个非常不同的环境。那里什么重要?以冰球为例吧。在加拿大,我在多伦多大学,我不认为那是一个很有创业精神的环境。你甚至不会想到要创办公司。这不是人们在做的事情;你没有朋友在做这件事,你不知道你应该崇拜它。人们不会读关于创始人的书并谈论他们。这不是你渴望或关心的事情。大家谈论的都是你在哪里实习,之后去哪里工作。人们接受的是,有一组固定的公司供你选择,然后你把自己与其中一家绑定,这就是你所崇拜的。所以这些文化因素非常强大,甚至可能是主导变量。我几乎觉得今天教育方面反而更容易;大量资源已经可用。所以我认为主要是你所处的文化因素在起作用。

In the academic world, for example, the H-index and the number of papers you publish are what everyone cares about. I was part of that community and saw that. Now I've come to different places, and there are different idols in all the different communities. I think that has a massive impact on what people are motivated by, where they get their social status, and what actually matters to them. I was also part of different communities growing up in Slovakia, a very different environment, and being in Canada, also a very different environment. What mattered there? Hockey, I would say as an example. In Canada, I was at the University of Toronto, and I don't think it's a very entrepreneurial environment. It doesn't even occur to you that you should be starting companies. It's not something people are doing; you don't know friends who are doing it, you don't know that you're supposed to look up to it. People aren't reading books about founders and talking about them. It's just not something you aspire to or care about. What everyone is talking about is where you're getting your internship, where you're going to work afterwards. It's accepted that there's a fixed set of companies you're supposed to pick from and align yourself with one of them, and that's what you look up to. So these cultural aspects are extremely strong and maybe actually the dominant variable. I almost feel like today the education aspect is the easier one; a ton of stuff is already available. So I think mostly it's the cultural aspect you're part of.

学习 vs 娱乐与后 AGI 社会 Learning vs entertainment and post-AGI society

Host

关于这一点,几周前我们聊过,而且我想你也在网上发过相关内容:学习和娱乐是有区别的。学习实际上应该是困难的。我认为这与地位和偶像的问题有关。通过这样的系统,你认为你能在多大程度上改变人们的动机?如果这是一个阻碍因素,你是专注于给人们提供资源,让他们能根据自己的能力在序列中尽可能走得更远,比历史上任何时刻都更远,这本身就已经很鼓舞人心了?还是你实际上想改变有多少人愿意学习,或者至少让他们走上这条路?

On this point, one thing you and I were talking about a few weeks ago, and I think you also posted online about this, is there's a difference between learning and entertainment. Learning is actually supposed to be hard. I think it relates to this question of status and who the idol is. How much do you think you can change in terms of motivation through systems like this? If that's a blocking factor, are you focused on giving people the resources so they can get as far as possible in the sequence for their own capability, further than any other point in history already inspirational? Or do you actually want to change how many people want to learn, or at least bring themselves down the path?

Andrej

我想说,我想让学习变得更容易。那么也许人们可能不想学习。例如今天,人们出于实际原因想学习,比如找工作,这完全合理。所以在 AGI 之前的社会,教育是有用的,我认为人们会因此受到激励,因为他们正在经济上向上爬。在 AGI 之后的社会,我认为教育在很大程度上是娱乐,包括教育的成功成果,而不仅仅是让内容从你身上流过。

I would say I want to make it much easier to learn. Then maybe it is possible that people don't want to learn. Today, for example, people want to learn for practical reasons, like to get a job, which makes total sense. So in a pre-AGI society, education is useful, and I think people will be motivated by that because they're climbing up the ladder economically. In a post-AGI society, I think education is entertainment to a much larger extent, including successful outcomes of education, not just letting the content wash over you.

Host

成果指的是理解、学习、能够贡献新知识,或者无论你怎么定义。

Outcomes being understanding, learning, being able to contribute new knowledge, or however you define it.

Andrej

我认为这不是偶然的:如果你回到 200 或 300 年前,做科学的人是贵族或富人。我们都将成为和 Andrej 一起学习的贵族。

I think it's not an accident that if you go back 200 or 300 years, the people who were doing science were nobility or people of wealth. We will all be nobility learning with Andrej.

Host

我确实认为这非常类似于你之前的引述。我觉得学习就像去健身房,但针对的是大脑。就像去健身房一样。去健身房很有趣;人们喜欢举铁等等。有些人不去健身房。不,不,不,有些人去,但需要付出努力。是的,需要努力,但也挺有趣,而且你还会在各方面对自己感觉良好。我认为教育基本上等同于这个。所以当我说教育不应该有趣等等时,就是这个意思。它确实有点有趣,但我想是一种特定的有趣。

I do think that I see it very much equivalent to your quote earlier. I feel like learning something is kind of like going to the gym but for the brain. It feels like going to the gym. Going to the gym is fun; people like to lift, etc. Some people don't go to the gym. No, no, no, some people do, but it takes effort. Yeah, it takes effort, but it's also kind of fun, and you also have a payoff of how you feel about yourself in various ways. I think education is basically equivalent to that. So that's what I mean when I say education should not be fun, etc. It is kind of fun, but it's a specific kind of fun, I suppose.

Andrej

我确实认为,也许在 AGI 之后的世界,我希望发生的是人们经常去健身房,不仅是身体上的,还有精神上的,并且我们将其视为受过高等教育而崇拜。

I do think that maybe in a post-AGI world, what I would hope happens is people actually go to the gym a lot, not just physically but also mentally, and it's something we look up to as being highly educated.

Eureka 课程受众与时间线 Eureka course audience and timeline

Host

我能问最后一个关于 Eureka 的问题吗?因为我觉得人们会感兴趣。第一个课程的受众是谁?

Can I ask you one last question about Eureka, just because I think it would be interesting to people? Who is the audience for the first course?

Andrej

课程的受众,我主要把它看作一个本科水平的课程。所以如果你在技术领域读本科,我认为那是理想的受众。我确实认为我们现在看到的是,我们有一个过时的教育概念:你上学,然后毕业,去工作。显然这将会完全崩溃,尤其是在一个变化如此之快的社会中,随着技术快速变化,人们会更频繁地回到学校。所以它有点像本科水平,但我会说任何年龄达到那个水平的人都在范围内。我认为年龄会非常多样化。但我认为主要是那些真正想深入理解它的技术人员。

The audience for the course, I'm mostly thinking of this as an undergrad level course. So if you're doing an undergrad in a technical area, I think that would be the ideal audience. I do think that what we're seeing now is we have this antiquated concept of education where you go through school and then you graduate and go to work. Obviously this will totally break down, especially in a society that's turning over so quickly, that people are going to come back to school a lot more frequently as the technology changes very quickly. So it is kind of like undergrad level, but I would say anyone at that level at any age is kind of in scope. I think it will be very diverse in age as an example. But I think it is mostly technical people who actually want to understand it to a good amount.

Host

他们什么时候可以上这门课?

When can they take the course?

Andrej

我原本希望是今年年底。我确实有很多分心的事情在堆积,但我认为可能明年初是时间线。我正在努力让它变得非常好,这需要时间。

I was hoping it would be late this year. I do have a lot of distractions that are piling on, but I think probably early next year is kind of the timeline. I'm trying to make it very, very good, and it just takes time to get there.

给孩子的学习建议 Advice for kids on what to study

Host

我实际上还有一个问题,和这个有点相关。如果你今天有小孩,你认为他们应该学什么才能有一个有用的未来?

I have one last question actually that's pseudo-related to that. If you have little kids today, what do you think they should study in order to have a useful future?

Andrej

我心中有一个正确答案,正确答案主要是数学、物理、计算机科学这类学科。我这么说的原因是,我认为它有助于思维技能;在我看来,这是最好的思维技能核心。当然我有特定的背景,所以我会这么想,但这只是我的观点。我认为上物理课和其他这些课塑造了我的思维方式,而且我认为它对一般的问题解决非常有用。所以如果我们处于这样一个世界:在 AGI 之前这会有用,在 AGI 之后你仍然希望有赋权的人类能够在任何任意能力下运作。所以我只是认为这是人们应该做的正确答案。它要么有用,要么很好。我认为很多其他东西你可以稍后补上,但人们拥有大量时间和注意力的关键时期最好花在这些核心学科上。

There's a correct answer in my mind, and the correct answer is mostly math, physics, CS kind of disciplines. The reason I say that is because I think it helps for thinking skills; it's just the best thinking skill core, in my opinion. Of course I have a specific background, so I would think this, but that's just my view on it. I think taking physics classes and all these other classes just shaped the way I think, and I think it's very useful for problem solving in general. So if we're in this world where pre-AGI this is going to be useful, post-AGI you still want empowered humans who can function in any arbitrary capacity. So I just think this is the correct answer for people and what they should be doing. It's either useful or it's good. I think a lot of the other stuff you can tack on a bit later, but the critical period where people have a lot of time and attention is best spent on these core disciplines.

教育理念:聚焦操作密集型任务 Education Philosophy: Focus on Manipulation-Heavy Tasks

Andrej

我认为时间应该主要花在那些简单的、侧重操作的任务和工作上,而不是侧重记忆的任务和工作。我学过数学学位,当时感觉大脑里正在刻出一道新的沟回,而且这道沟回越晚刻越难。当然我也会加入其他很多东西——我不反对其他学科等等。我认为拥有多样化的东西其实很美,但我确实认为其中 80%应该是这类内容,因为与我们的工具相比,我们并不是高效的记忆者。

I think time should mostly be spent on doing these kinds of simple manipulation-heavy tasks and workloads, not memory-heavy tasks and workloads. I did a math degree and I felt like a new groove was being carved into my brain as I was doing that, and it's a harder groove to carve later. I would of course put in a bunch of other stuff as well—I'm not opposed to all the other disciplines, etc. I think it's actually beautiful to have a large diversity of things, but I do think 80% of it should be something like this, because we're not efficient memorizers compared to our tools.

结束语 Closing Remarks

Host

感谢你参与这次访谈。非常有趣。很高兴来到这里。在 Twitter 上关注我们 @no_prior_pod。如果你想看到我们的脸,请订阅我们的 YouTube 频道。在 Apple Podcasts、Spotify 或你收听的地方关注节目,这样你每周都能收到新一期。并在 no-pri.com 上注册邮件或查找每期节目的文字稿。

Thank you for doing this. So much fun. Great to be here. Find us on Twitter at @no_prior_pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-pri.com.

互动版:逐字朗读 + 针对本期提问 →