我们在召唤幽灵,而不是在造动物

We’re summoning ghosts, not building animals

安德烈·卡帕西 Andrej Karpathy · Dwarkesh 播客 · 2025-10-17 · 约 146 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

为何是「智能体的十年」而非「元年」、真正的瓶颈,以及来自 15 年从业的直觉。

Why it’s the decade of agents (not the year), the real bottlenecks, and a 15-year intuition.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 51)

全文 · Full transcript(中英对照)

关于智能体的十年 On the decade of agents

Dwarkesh

今天我和 Andrej Karpathy 对话。Andrej,你为什么说这将是智能体的十年,而不是智能体的一年?

Today I'm speaking with Andrej Karpathy. Andrej, why do you say that this will be the decade of agents and not the year of agents?

Andrej

首先,感谢你邀请我。我很高兴来到这里。你刚才提到的「智能体的十年」这个说法,实际上是对一个已有说法的回应。我记得有些实验室——我不确定具体是谁说的——暗示今年是智能体的一年,指的是大语言模型及其演进。我对此有反应,因为我觉得行业里有些过度预测。在我看来,更准确的描述是智能体的十年。我们已经有了一些非常早期的智能体,它们实际上令人印象深刻,我每天都在用——比如 Claude 和 Codex 等等——但我仍然觉得还有很多工作要做。所以我的反应是:我们将和这些东西一起工作十年;它们会变得更好,这将是美妙的。但我只是对那种时间线暗示做出了反应。

Well, first of all, thank you for having me here. I'm excited to be here. So the quote you've just mentioned, 'the decade of agents,' is actually a reaction to an existing pre-existing quote, I should say, where I think a lot of some of the labs — I'm not actually sure who said this — but they were alluding to this being the year of agents with respect to LLMs and how they were going to evolve. And I think I was triggered by that because I feel like there's some overpredictions going on in the industry. In my mind, this is really a lot more accurately described as the decade of agents. We have some very early agents that are actually extremely impressive and that I use daily — you know, Claude and Codex and so on — but I still feel like there's so much work to be done. So I think my reaction is: we'll be working with these things for a decade; they're going to get better, and it's going to be wonderful. But I think I was just reacting to the timelines of the implication.

Dwarkesh

那你认为什么需要十年才能实现?瓶颈是什么?

And what do you think will take a decade to accomplish? What are the bottlenecks?

Andrej

嗯,实际上是让它们真正工作。在我看来,当我们谈论智能体时——或者实验室所想的,以及我可能想的——你应该把它几乎看作一个员工或实习生,你雇佣来和你一起工作。例如,你在这里和员工一起工作。什么时候你更愿意让像 Claude 或 Codex 这样的智能体来做那些工作?目前,当然它们做不到。要让它们能够做到,需要什么?为什么你今天不这样做?你今天不这样做的原因是它们就是不行。它们没有足够的智能。它们不够多模态。它们不能使用计算机等等。它们很多事情都做不了。它们没有持续学习能力。你不能告诉它们一件事然后它们就记住了。它们在认知上就是欠缺的。就是不行。我只是认为解决所有这些问题大约需要十年时间。

Well, actually making it work. In my mind, when you're talking about an agent — or what the labs have in mind and what maybe I have in mind as well — it's you should think of it almost like an employee or like an intern that you would hire to work with you. For example, you work with some employees here. When would you prefer to have an agent like Claude or Codex do that work? Currently, of course, they can't. What would it take for them to be able to do that? Why don't you do it today? The reason you don't do it today is because they just don't work. They don't have enough intelligence. They're not multimodal enough. They can't do computer use and all this kind of stuff. And they don't do a lot of things. They don't have continual learning. You can't just tell them something and they'll remember it. And they're just cognitively lacking. And it's just not working. I just think that it will take about a decade to work through all of those issues.

Dwarkesh

有意思。所以,作为一个专业播客主持人和从远处观察 AI 的人,我很容易识别出,哦,这里缺什么。持续学习缺乏,或者多模态缺乏。但我没有很好的办法来给它定一个时间线。如果有人问,持续学习需要多长时间?我没有先验知识说这是一个应该花 5 年、10 年还是 50 年的项目。为什么是十年?为什么不是一年?为什么不是 50 年?

Interesting. So, as a professional podcaster and a viewer of AI from afar, it's sort of easy for me to identify like, oh, here's what's lacking. Continual learning is lacking or multimodality is lacking. But I don't really have a good way of trying to put a timeline on it. If somebody's like, how long will continual learning take? There's no prior I have about like this is a project that should take 5 years, 10 years, 50 years. Why a decade? Why not one year? Why not 50 years?

Andrej

是的,我想这就涉及到我的一些直觉,以及根据我自己在该领域的经验进行的外推。我在 AI 领域已经快二十年了——大概 15 年左右,不算很长。你请过 Richard Sutton,他待的时间长得多,但我确实有大约 15 年的经验,看到人们做预测以及它们实际如何发展。此外,我在工业界待过一段时间,也做过研究。所以我想我从中留下了一些一般性的直觉。我觉得这些问题是可以处理的,是可以克服的,但仍然很困难。如果我只是平均一下,对我来说感觉就是十年。

Yeah, I guess this is where you get into a bit of my own intuition and also just kind of doing an extrapolation with respect to my own experience in the field. I've been in AI for almost two decades — maybe 15 years or so, not that long. You had Richard Sutton here who was around for much longer, but I do have about 15 years of experience of people making predictions and seeing how they actually turned out. Also, I was in the industry for a while and I was in research. So I guess I kind of have a general intuition that I have left from that. I feel like the problems are tractable, they're surmountable, but they're still difficult. And if I just average it out, that just kind of feels like a decade to me.

AI 的历史变迁 On historical shifts in AI

Dwarkesh

这其实很有意思。我想听的不只是历史,还有当时在场的人在各种突破时刻感觉即将发生什么。他们的感觉在哪些方面过于悲观或过于乐观?

This is actually quite interesting. I want to hear not only the history but what people in the room felt was about to happen at various different breakthrough moments. What were the ways in which their feelings were either overly pessimistic or overly optimistic?

Andrej

是啊。我们要不要一个一个地过一遍?

Yeah. Should we just go through each of them one by one?

Dwarkesh

好。

Yeah.

Andrej

哦,是的。我的意思是,这是一个巨大的问题,因为你谈的是 15 年间发生的事情。AI 实际上非常美妙,因为有过几次地震式的转变,整个领域突然就换了个方向。我可能经历过其中的两三次。而且我仍然认为还会有更多,因为它们带有某种令人惊讶的不规则性。嗯,当我的职业生涯开始时——当我开始研究深度学习,当我对其产生兴趣时——这纯属偶然,因为我就在多伦多大学,紧挨着 Geoff Hinton。Geoff Hinton 当然是 AI 领域的教父级人物,他当时在训练所有这些神经网络,我觉得这不可思议且有趣,但这远不是当时 AI 领域每个人都在做的主流事情。这是一个边缘的 niche 课题。这大概就是第一次戏剧性的地震式转变,伴随着 AlexNet 等等。AlexNet 基本上重新定位了所有人,每个人都开始训练神经网络。

Oh yeah. I mean that's a giant question because of course you're talking about 15 years of stuff that happened. AI is actually so wonderful because there have been a number of seismic shifts that were like the entire field suddenly looked a different way. I've maybe lived through two or three of those. And I still think there will continue to be some because they come with some kind of almost surprising irregularity. Well, when my career began — when I started to work on deep learning, when I became interested in deep learning — this was just kind of by chance of being right next to Geoff Hinton at University of Toronto. And Geoff Hinton, of course, is kind of like the godfather figure of AI and he was training all these neural networks and I thought it was incredible and interesting, but this was not like the main thing that everyone in AI was doing by far. This was a niche subject on the side. That's kind of maybe the first dramatic seismic shift that came with AlexNet and so on. AlexNet sort of reoriented everyone and everyone started to train neural networks.

早期 AI:单任务模型与智能体的误区 On early AI: per-task models and the misstep of agents

Andrej

但当时仍然是每个任务、每个特定任务各搞各的。比如我有一个图像分类器,或者一个神经机器翻译器。后来人们慢慢对智能体产生了兴趣。他们开始想,好吧,也许视觉皮层这块我们打上勾了,但大脑的其他部分呢?我们怎样才能得到一个完整的、能在世界中互动的智能体?我认为 2013 年 Atari 深度强化学习的转变就是这种早期努力的一部分。那是为了获得不仅能感知世界,还能采取行动、互动并从环境中获得奖励的智能体。当时就是 Atari 游戏。我觉得那是一个失误,甚至我参与过的早期 OpenAI 也采纳了它。有那么三四年,每个人都在做游戏上的强化学习。我一直对游戏能通向 AGI 持怀疑态度,因为我想要的是像会计那样与现实世界互动的东西。我在 OpenAI 的项目,属于 Universe 项目,是一个用键盘和鼠标操作网页的智能体。我想要一个能与实际数字世界互动、能做知识工作的东西。但那时太早了。如果你到处乱撞、乱按键盘鼠标、试图获得奖励,奖励太稀疏了。你学不到东西,还会烧掉一片算力森林。你缺少的是神经网络中的表示能力。如今,人们在大型语言模型之上训练使用计算机的智能体。你必须先有语言模型,先有表示,通过预训练和所有 LLM 的工作。所以人们总是过早地试图得到完整的东西——Atari、Universe,甚至我自己的经历。在得到智能体之前,你必须先做一些事情。现在智能体更胜任了,但可能我们仍然缺少栈的某些部分。这就是三大类:每个任务训练神经网络、第一波智能体、以及 LLM——在叠加其他东西之前先寻求表示能力。

But it was still very per-task, per specific task. So maybe I have an image classifier or a neural machine translator. And people became slowly interested in agents. They started to think, okay, maybe we have a check mark next to the visual cortex, but what about the other parts of the brain? How can we get a full agent that can interact in the world? I would say the Atari deep reinforcement learning shift in 2013 was part of that early effort. It was an attempt to get agents that not only perceive the world but also take actions, interact, and get rewards from environments. At the time, this was Atari games. I feel like that was a misstep, and even the early OpenAI that I was part of adopted it. For a few years, everyone was doing reinforcement learning on games. I was always a bit suspicious of games leading to AGI because I wanted something like an accountant interacting with the real world. My project at OpenAI, within the Universe project, was an agent using keyboard and mouse to operate web pages. I wanted something that interacts with the actual digital world and can do knowledge work. But it was way too early. If you're stumbling around, keyboard mashing, and trying to get rewards, the reward is too sparse. You won't learn, and you'll burn a forest of compute. What you're missing is the power of representation in the neural network. Today, people train computer-using agents on top of a large language model. You have to get the language model first, the representations first, through pre-training and all the LLM stuff. So people kept trying to get the full thing too early—Atari, Universe, even my own experience. You have to do some things first before you get to agents. Now agents are more competent, but maybe we're still missing some parts of the stack. Those are the three major buckets: training neural nets per task, the first round of agents, and then LLMs—seeking the representation power before tacking everything else on top.

动物类比与 AI 的差异 On the animal analogy and why AI is different

Dwarkesh

有意思。如果从 Sutton 的角度来看,人类可以一下子应对所有事情。甚至动物也可以。动物是更好的例子,因为它们没有语言支架。它们被扔到世界上,必须在没有标签的情况下理解一切。AGI 的愿景应该是某种东西,它只看感官数据、看电脑屏幕,然后从零开始弄清楚发生了什么。如果人类或动物是这样成长的,那为什么 AI 的愿景不应该是这样,而是要进行数百万年的训练呢?

Interesting. If I were to take the Sutton perspective, humans can take on everything at once. Even animals can. Animals are a better example because they don't have the scaffold of language. They just get thrown into the world and have to make sense of everything without labels. The vision for AGI should be something that just looks at sensory data, looks at the computer screen, and figures out what's going on from scratch. If a human or animal grows up that way, why shouldn't that be the vision for AI rather than millions of years of training?

Andrej

这是个很好的问题。Sutton 上过你的播客,我看了。我写过一篇关于它的文章,谈了我怎么看。我对类比动物非常谨慎,因为它们是通过非常不同的优化过程产生的。动物是进化来的,带有大量内置的硬件。比如,在我的文章里,我用了斑马的例子。斑马出生几分钟后就能跑动并跟着妈妈。这极其复杂。那不是强化学习,而是内置的。进化用 ATCG 编码了我们神经网络的权重。我不知道那是怎么做到的,但它显然有效。大脑来自一个非常不同的过程。我不太愿意从中汲取灵感,因为我们并没有运行那个过程。在我的文章里,我说我们不是在建造动物,而是在建造幽灵。或者说精灵。我们不是通过进化来训练,而是通过模仿人类和他们在互联网上留下的数据来训练。所以我们最终得到的是完全数字化的、模仿人类的空灵精灵实体。这是一种不同的智能。如果你想象一个智能空间,我们是从一个不同的起点开始的。我们不是在建造动物,但我认为随着时间的推移,有可能让它们更像动物,而且我们应该这样做。

That's a really good question. Sutton was on your podcast, and I saw it. I had a write-up about it that gets into how I see things. I'm very careful about analogies to animals because they came about by a very different optimization process. Animals are evolved and come with a huge amount of built-in hardware. For example, in my post, I used the zebra. A zebra gets born and a few minutes later it's running around and following its mother. That's extremely complicated. That's not reinforcement learning; it's baked in. Evolution encodes the weights of our neural nets in ATCGs. I have no idea how that works, but it apparently works. Brains came from a very different process. I'm hesitant to take inspiration from it because we're not running that process. In my post, I said we're not building animals; we're building ghosts. Or spirits. We're not doing training by evolution; we're doing training by imitation of humans and the data they put on the internet. So we end up with ethereal spirit entities that are fully digital and mimic humans. It's a different kind of intelligence. If you imagine a space of intelligences, we're starting at a different point. We're not building animals, but I think it's possible to make them more animal-like over time, and we should.

进化 vs 预训练 On evolution vs pre-training

Andrej

我觉得 Sutton 基本上是想构建动物,如果能实现的话,那会很棒。如果有一个单一的算法,你可以在互联网上运行它,它就能学会一切,那将是不可思议的。但我怀疑它并不存在,而且动物也不是这样做的,因为动物有进化这个外部循环。很多看似学习的东西实际上是大脑的成熟。我认为动物很少使用强化学习来完成智能任务;它更多用于运动任务。所以我认为人类实际上并不使用强化学习来解决像问题解决这样的智能任务。

I feel like Sutton basically wants to build animals, and that would be wonderful if we could get it to work. If there were a single algorithm you could run on the internet and it learns everything, that would be incredible. But I suspect it doesn't exist, and that's not what animals do because animals have this outer loop of evolution. A lot of what looks like learning is actually maturation of the brain. I think very little reinforcement learning occurs in animals for intelligence tasks; it's more for motor tasks. So I think humans don't really use RL for intelligence tasks like problem solving.

Dwarkesh

你能重复最后一句吗?很多智能不是运动任务。那是什么?

Can you read the last sentence? A lot of that intelligence is not motor task. It's what?

Andrej

在我看来,很多强化学习更像是运动任务,比如投掷铁环。但我不认为人类使用强化学习来完成很多智能任务,比如问题解决。

A lot of the reinforcement learning, in my perspective, would be things that are more like motor tasks, like throwing a hoop. But I don't think humans use reinforcement learning for a lot of intelligence tasks like problem solving.

Dwarkesh

有意思。但这并不意味着我们不应该在研究中这样做,我只是觉得这就是动物做或不做的事情。

Interesting. That doesn't mean we shouldn't do that for research, but I just feel like that's what animals do or don't.

Dwarkesh

我需要一点时间来消化。也许有一个澄清性的问题:你暗示进化在做预训练所做的事情,构建能够理解世界的东西。区别在于进化必须通过 3GB 的 DNA 来滴定。这与模型的权重不同。模型的权重是一个大脑,它并不编码在精子和卵子中;它必须生长。每个突触的信息不可能存在于 3GB 中。进化似乎更接近于找到随后进行终身学习的算法。也许终身学习并不类似于强化学习。这与你的说法一致吗?

I'm going to take a second to digest that. Maybe one clarifying question: you suggest that evolution does the kind of thing that pre-training does, building something that can understand the world. The difference is that evolution has to be titrated through 3 gigabytes of DNA. That's unlike the weights of a model. The weights of a model are a brain, which is not encoded in sperm and egg; it has to be grown. The information for every synapse cannot exist in 3 gigabytes. Evolution seems closer to finding the algorithm that then does lifetime learning. Maybe lifetime learning is not analogous to RL. Is that compatible with what you're saying?

Andrej

我同意存在某种神奇的压缩,因为神经网络的权重并不存储在 ATCG 中。存在某种剧烈的压缩和编码的学习算法,它们接管并在线进行一些学习。我更加务实。我不是从构建动物的角度出发;我是从构建有用东西的角度出发。我们不会去搞进化,因为我不知道怎么做。但我们可以通过模仿互联网文档来构建这些幽灵般的实体。这很有效,能让你达到一个具有内置知识和智能的状态,类似于进化所做的。所以我称预训练为一种糟糕的进化——在我们技术条件下实际可行的版本。

I agree that there's some miraculous compression going on because the weights of the neural net are not stored in ATCGs. There's some dramatic compression and learning algorithms encoded that take over and do some learning online. I'm more practically minded. I don't come from the perspective of building animals; I come from building useful things. We're not going to do evolution because I don't know how. But we can build these ghost-like entities by imitating internet documents. That works and brings you up to something with built-in knowledge and intelligence, similar to what evolution has done. So I call pre-training a kind of crappy evolution—the practically possible version with our technology.

Dwarkesh

为了更公正地看待另一种观点:进化并不给我们知识,它给了我们寻找知识的算法。这似乎与预训练不同。如果预训练有助于构建一个能更好学习的实体,它教会了元学习,类似于找到算法。但如果进化给予知识,预训练也给予知识,那么这个类比就站不住脚了。

To steelman the other perspective: evolution does not give us knowledge, it gives us the algorithm to find knowledge. That seems different from pre-training. If pre-training helps build an entity that can learn better, it teaches meta-learning, similar to finding an algorithm. But if evolution gives knowledge and pre-training gives knowledge, that analogy breaks down.

Andrej

你反驳得对。预训练做了两件事:它获取知识,并通过观察互联网上的算法模式变得智能,启动用于上下文学习的电路。实际上,你并不需要或不想要这些知识。我认为这阻碍了神经网络,因为它们过于依赖知识。例如,智能体不擅长偏离互联网上存在的数据流形。如果它们拥有更少的知识或记忆,可能会更好。因此,未来我们需要移除一些知识,保留我所谓的认知核心——剥离了知识但包含智能和问题解决算法与魔法的智能实体。

You're right to push back. Pre-training does two things: it picks up knowledge, and it becomes intelligent by observing algorithmic patterns on the internet, booting up circuits for in-context learning. Actually, you don't need or want the knowledge. I think it holds back neural networks because they rely on knowledge too much. For example, agents are not good at going off the data manifold of what exists on the internet. If they had less knowledge or memory, they might be better. So going forward, we need to remove some knowledge and keep what I call the cognitive core—the intelligent entity stripped of knowledge but containing the algorithms and magic of intelligence and problem solving.

上下文学习 vs 预训练 On in-context learning vs pre-training

Dwarkesh

这些模型看起来最智能的场景是,当我和它们对话时,我会想:「哇,另一端真的有个东西在回应我,在思考问题。如果它犯了错,它会说,哦等等,这个思考方式不对,我退回去。」所有这些都发生在上下文里。我觉得这才是你能真正看到的智能。

The situation in which these models seem the most intelligent, in which I talk to them and I'm like, 'Wow, there's really something on the other end that's responding to me, thinking about things. If it makes a mistake, it's like, oh wait, that's actually the wrong way to think about it. I'm backing up.' All that is happening in context. That's where I feel like the real intelligence you can visibly see.

Andrej

而那个上下文学习过程是通过预训练中的梯度下降发展出来的,对吧?它自发地元学习了上下文学习,但上下文学习本身并不是梯度下降,就像我们人类一生中的智能受进化条件限制,但我们一生中的实际学习是通过其他过程发生的。

And that in-context learning process is developed by gradient descent on pre-training, right? It meta-learns in-context learning spontaneously, but the in-context learning itself is not gradient descent, in the same way that our lifetime intelligence as humans is conditioned by evolution, but our actual learning during our lifetime happens through some other process.

Dwarkesh

我其实不完全同意,但你先继续说吧。

I actually don't fully agree with that, but you should continue.

Andrej

好吧,那我很好奇这个类比在什么地方不成立。

Okay, actually then I'm very curious to understand how that analogy breaks down.

Dwarkesh

我有点犹豫是否要说上下文学习不是在执行梯度下降,因为我的意思是它没有显式地做梯度下降,但我仍然认为……

I think I'm hesitant to say that in-context learning is not doing gradient descent, because I mean it's not doing explicit gradient descent, but I still think that...

Andrej

所以上下文学习基本上是在一个词元窗口内完成模式补全,对吧?结果互联网上就有大量的模式。所以你说得对,模型学会了补全模式。

So in-context learning is basically pattern completion within a token window, right? And it just turns out that there's a huge amount of patterns on the internet. And so you're right, the model kind of learns to complete the pattern.

Dwarkesh

而那是发生在权重里的。神经网络的权重试图发现模式并补全模式。神经网络内部发生了一些适应,对吧?

And that's inside the weights. The weights of the neural network are trying to discover patterns and complete the pattern. And there's some kind of adaptation that happens inside the neural network, right?

Andrej

这有点神奇,只是因为互联网上有很多模式,它就自然出现了。我得说,有一些我觉得很有意思的论文,它们确实研究了上下文学习背后的机制,我确实认为上下文学习有可能在神经网络层内部运行一个小型的梯度下降循环。我特别记得一篇论文,他们用上下文学习来做线性回归。基本上,你输入到神经网络的是 XY 对,XY XY XY,这些点恰好落在一条线上,然后你输入 X,期望得到 Y。当你这样训练神经网络时,它确实做了线性回归。通常当你运行线性回归时,你会有一个小的梯度下降优化器,它查看 XY,查看误差,计算权重的梯度,并更新几次。结果发现,当他们查看那个上下文学习算法的权重时,他们确实找到了一些与梯度下降机制类似的东西。事实上,我认为那篇论文甚至更进一步,因为他们实际上硬编码了神经网络的权重,通过注意力机制和神经网络的所有内部结构来执行梯度下降。所以我想我唯一的反驳是,谁知道上下文学习是如何工作的,但我实际上认为它可能在内部做了一点某种奇特的梯度下降,我认为这是可能的。所以我想我只是反驳你说它没有在做梯度下降。谁知道它在做什么,但它可能在做类似的事情,只是我们不知道。

Which is kind of magical and just falls out from the internet just because there's a lot of patterns. I will say that there have been some papers that I thought were interesting that actually look at the mechanisms behind in-context learning, and I do think it's possible that in-context learning actually runs a small gradient descent loop internally in the layers of the neural network. So I recall one paper in particular where they were doing linear regression using in-context learning. So basically your inputs into the neural network are XY pairs, XY XY XY that happen to be on the line, and then you do X and you expect the Y. And the neural network when you train it in this way actually does do linear regression. And normally when you would run linear regression, you have a small gradient descent optimizer that basically looks at XY, looks at an error, calculates the gradient of the weights, and does the update a few times. It just turns out that when they looked at the weights of that in-context learning algorithm, they actually found some analogies to gradient descent mechanics. In fact, I think even the paper went further because they actually hardcoded the weights of a neural network to do gradient descent through attention and all the internals of the neural network. So I guess that's just my only pushback is that who knows how in-context learning works, but I actually think that it's probably doing a little bit of some kind of funky gradient descent internally, and that I think that's possible. So I guess I was only pushing back on you saying it's not doing gradient descent. Who knows what it's doing, but it's probably maybe doing something similar to it, but we don't know.

Dwarkesh

那么值得思考的是:好吧,如果两者都在实现类似梯度下降的东西——如果上下文学习和预训练都在实现类似梯度下降的东西——为什么感觉上下文学习真正让我们达到了这种持续学习、真正的智能,而仅仅从预训练中你得不到类似的感觉,至少你可以这么说。所以如果算法相同,那什么可能不同呢?

So then it's worth thinking about: okay, if both of them are implementing something like gradient descent—if in-context learning and pre-training are both implementing something like gradient descent—why does it feel like in-context learning actually gets us to this continual learning, real intelligence thing, whereas you don't get the analogous feeling just from pre-training, at least you could argue that. And so if it's the same algorithm, what could be different?

Andrej

嗯,一种思考方式是,模型从训练中接收的每单位信息中存储了多少信息。如果你看预训练,比如 Llama 3,我认为它是在 15 万亿个词元上训练的,如果你看 70B 模型,就模型权重中存储的信息相对于它读取的词元而言,相当于每个词元 0.7 比特。而如果你看 KV 缓存,以及它在上下文学习中每个额外词元如何增长,大约是 320 千字节。所以模型每个词元吸收的信息量有 3500 万倍的差异。我想知道这是否相关。

Well, one way you can think about it is how much information does the model store per unit of information it receives from training. And if you look at pre-training, if I think if you look at Llama 3 for example, I think it's trained on 15 trillion tokens, and if you look at the 70B model, that would be the equivalent of 0.7 bits per token that it sees in pre-training in terms of the information in the weights of the model compared to the tokens it reads. Whereas if you look at the KV cache and how it grows per additional token in in-context learning, it's like 320 kilobytes. So that's a 35 million-fold difference in how much information per token is assimilated by the model. I wonder if that's relevant at all.

Dwarkesh

是的。所以模型每个词元吸收的信息量有 3500 万倍的差异。我想知道这是否相关。

Yeah. So that's a 35 million-fold difference in how much information per token is assimilated by the model. I wonder if that's relevant at all.

Andrej

我想我有点同意。我的意思是,我通常这样表述:神经网络训练过程中发生的任何事情,知识都只是对训练时发生事情的一种模糊回忆,这是因为压缩非常剧烈。你把 15 万亿个词元压缩到只有几十亿参数的最终网络。所以显然发生了大量的压缩。

I think I kind of agree. I mean the way I usually put this is that anything that happens during the training of the neural network, the knowledge is only kind of like a hazy recollection of what happened in training time, and that's because the compression is dramatic. You're taking 15 trillion tokens and you're compressing it to just your final network of a few billion parameters. So obviously it's a massive amount of compression going on.

LLM 的工作记忆与长期记忆 On working memory vs. long-term memory in LLMs

Dwarkesh

我把它比作对互联网文档的模糊记忆,而神经网络上下文窗口中发生的任何事情——你输入所有词元,它构建起所有这些 KV 缓存表示——对神经网络来说都是非常直接可访问的。所以我把 KV 缓存和测试时发生的事情比作更像工作记忆。

So I kind of refer to it as a hazy recollection of the internet documents, whereas anything that happens in the context window of the neural network, you're plugging all the tokens and it's building up all this KV cache representation, is very directly accessible to the neural net. So I compare the KV cache and the stuff that happens at test time to more like a working memory.

Andrej

上下文窗口中的所有内容对神经网络来说都是非常直接可访问的。所以大型语言模型和人类之间总是存在这些几乎令人惊讶的类比,我觉得这有点意外,因为我们当然不是要建造一个人脑。我们只是直接发现这行得通,就这么做了。但我确实认为,权重中的任何东西都像是一年前读过的模糊记忆。而你在测试时作为上下文给出的任何东西都直接在工作记忆中。我认为这是一个非常有力的类比,可以用来思考问题。所以,比如当你问一个大型语言模型关于某本书的内容,比如南的书之类的,它通常会给出大致正确的东西。但如果你给它完整的章节再提问,你会得到好得多的结果,因为现在它被加载到了模型的工作记忆中。所以我基本上同意你那个很长的说法。

All the stuff that's in the context window is very directly accessible to the neural net. So there are always these almost surprising analogies between LLMs and humans, and I find them kind of surprising because we're not trying to build a human brain, of course. We're just directly finding that this works and we're doing it. But I do think that anything that's in the weights is kind of like a hazy recollection of what you read a year ago. Anything that you give it as context at test time is directly in the working memory. I think that's a very powerful analogy to think through things. So when you, for example, go to an LLM and ask it about some book and what happened in it, like Nan's book or something like that, the LM will often give you some stuff which is roughly correct. But if you give it the full chapter and ask it questions, you're going to get much better results because it's now loaded in the working memory of the model. So I basically agree with your very long way of saying that.

人类智能中未能复现的部分 On what human intelligence we have failed to replicate

Dwarkesh

退一步说,关于人类智能,我们用这些模型最没能复制的是什么?

Stepping back, what is it about human intelligence that we have most failed to replicate with these models?

Andrej

我几乎觉得还有很多没做到。所以也许一种思考方式,我不知道这是不是最好的方式,但我几乎觉得,再次用这些不完美的类比来说,我们偶然发现了 Transformer 神经网络,它极其强大,非常通用。你可以在音频、视频、文本或任何东西上训练 Transformer,它只是学习模式,它们非常强大,效果很好。这对我来说几乎表明这有点像某种皮层组织。因为众所周知,大脑皮层也具有很强的可塑性。你可以重新连接大脑的部分区域,有一些有点可怕的实验将视觉皮层重新连接到听觉皮层,动物也能正常学习。所以我认为这有点像皮层组织。我认为当我们在神经网络内部进行推理和规划时,基本上是为思考模型做推理轨迹,那有点像前额叶皮层。然后我认为那些可能算是小勾号,但我仍然认为有很多大脑部位和核团尚未探索。所以也许例如基底神经节在我们用强化学习微调模型时做了一点强化学习,而海马体对应什么就不明显了。有些部分可能不重要,也许小脑对认知不重要,据认为如此,所以我们可以跳过一些。但我仍然认为例如杏仁核,所有的情绪和本能,可能还有一大堆非常古老的大脑核团,我不认为我们真的复制了它们。我实际上不知道我们是否应该追求建造人脑的类似物;我内心主要还是工程师。但我仍然觉得,也许回答这个问题的另一种方式是,你不会雇佣这个东西作为实习生,它缺失了很多,因为它带有许多我们与模型交谈时直观感受到的认知缺陷。所以它还没有完全到位。你可以看作并非所有大脑部件都已勾选。

I almost feel like just a lot of it still. So maybe one way to think about it, I don't know if this is the best way, but I almost kind of feel like, again making these analogies, imperfect as they are, we've stumbled by with the Transformer neural network, which is extremely powerful, very general. You can train Transformers on audio or video or text or whatever you want and it just learns patterns and they're very powerful and it works really well. That to me almost indicates that this is kind of like some piece of cortical tissue. It's something like that because the cortex is famously very plastic as well. You can rewire parts of brains and there were the slightly gruesome experiments with rewiring visual cortex to the auditory cortex and this animal learned fine. So I think that this is kind of like cortical tissue. I think when we're doing reasoning and planning inside the neural networks, so basically doing reasoning traces for thinking models, that's kind of like the prefrontal cortex. And then I think we maybe those are like little check marks, but I still think there are many brain parts and nuclei that are not explored. So maybe for example there's the basal ganglia doing a bit of reinforcement learning when we fine-tune the models on reinforcement learning, but whereas the hippocampus is not obvious what that would be. Some parts are probably not important, maybe the cerebellum is not important to cognition, it's thought, so we can skip some of it. But I still think there's for example the amygdala, all the emotions and instincts, and there's probably a bunch of other nuclei in the brain that are very ancient that I don't think we've really replicated. I don't actually know that we should be pursuing the building of an analog of human brain; I'm again an engineer mostly at heart. But I still feel like maybe another way to answer the question is you're not going to hire this thing as an intern and it's missing a lot because it comes with a lot of these cognitive deficits that we all intuitively feel when we talk to the models. And so it's just not fully there yet. You can look at it as not all the brain parts are checked off yet.

持续学习与涌现 On continual learning and emergence

Dwarkesh

这可能与思考这些问题解决速度的问题有关。所以有时人们会谈到持续学习。你看,实际上你已经可以轻松复制这种能力,就像上下文学习作为预训练的结果自发出现一样。如果模型被激励去回忆更长时间跨度或超过一个会话的时间跨度的信息,那么更长时间跨度的持续学习就会自发出现。所以如果存在一个外部循环强化学习,其中包含许多会话,那么这种持续学习——比如它自我微调或写入外部记忆之类——就会自发出现。你认为这样的事情合理吗?我只是没有先验判断它有多合理。它发生的可能性有多大?

This is maybe relevant to the question of thinking about how fast these issues will be solved. So sometimes people will say about continual learning. Look, actually you could already easily replicate this capability just as in-context learning emerged spontaneously as a result of pre-training. Continual learning over longer horizons will emerge spontaneously if the model is incentivized to recollect information over longer horizons or horizons longer than one session. So if there's some outer loop RL which has many sessions within that outer loop, then this continual learning where it uses like it fine-tunes itself or it writes to an external memory or something will just sort of emerge spontaneously. Do you think things like that are plausible? I just don't have a prior over how plausible that is. How likely is that to happen?

Andrej

我不确定我完全认同这一点,因为我觉得这些模型当你启动它们时,窗口中有零个词元,它们总是从之前的状态重新开始。所以我实际上不知道在那个世界观里它是什么样子。因为再次,也许用人类做类比,只是因为它大致具体且思考起来有点有趣。我觉得当我醒着时,我在构建一个白天发生的事情的上下文窗口,但我觉得当我睡觉时,发生了一些神奇的事情,我实际上不认为那个上下文窗口会保留下来。我认为有一个蒸馏到大脑权重的过程。这发生在睡眠期间等等。我们在大型语言模型中没有与之对应的东西,这对我来说更接近你谈到持续学习等缺失的东西。这些模型并没有真正经历这个蒸馏阶段:获取发生的事情,分析它,反复思考它,基本上做一些合成数据生成过程,然后蒸馏回权重,也许每个人有一个特定的神经网络,可能是一个 LoRA,不是全权重神经网络,只是权重的一个小的稀疏子集被改变。但基本上我们确实想要创造这些具有非常长上下文的个体的方法。

I don't know that I fully resonate with that because I feel like these models when you boot them up and they have zero tokens in the window, they're always restarting from scratch where they were. So I don't actually know in that worldview what it looks like. Because again, making maybe some analogies to humans just because I think it's roughly concrete and kind of interesting to think through. I feel like when I'm awake I'm building up a context window of stuff that's happening during the day, but I feel like when I go to sleep something magical happens where I don't actually think that that context window stays around. I think there's some process of distillation into weights of my brain. And this happens during sleep and all this kind of stuff. We don't have an equivalent for that in large language models, and that's to me more adjacent to when you talk about continual learning and so on as absent. These models don't really have this distillation phase of taking what happened, analyzing it, obsessively thinking through it, basically doing some kind of a synthetic data generation process and distilling it back into the weights, and maybe having a specific neural net per person, maybe it's a LoRA, it's not a full weight neural network, it's just a small sparse subset of the weights are changed. But basically we do want to create ways of creating these individuals that have very long contexts.

上下文窗口与认知架构 On context windows and cognitive architecture

Dwarkesh

这不仅仅是停留在上下文窗口内的问题,因为上下文窗口变得非常长,也许我们有一些非常精细的稀疏注意力机制来处理它。

It's not only remaining in the context window because the context windows grow very very long, like maybe we have some very elaborate sparse attention over it.

Andrej

但我仍然认为,人类显然有某种将知识蒸馏到权重中的过程,而我们还没有实现。而且我也认为人类有某种非常精细的稀疏注意力机制,我们开始看到一些早期迹象。比如 DeepSeek V3.2 刚刚发布,我看到他们使用了稀疏注意力作为例子,这是实现非常长上下文窗口的一种方式。

But I still think that humans obviously have some process for distilling some of that knowledge into the weights, we're missing it. And I do also think that humans have some kind of a very elaborate sparse attention scheme, which I think we're starting to see some early hints of. So DeepSeek V3.2 just came out and I saw that they have like a sparse attention as an example, and this is one way to have very very long context windows.

Dwarkesh

所以我几乎觉得,我们正在通过一个非常不同的过程重新实现进化想出的许多认知技巧,但我认为我们最终会在认知架构上趋同。

So I almost feel like we are redoing a lot of the cognitive tricks that evolution came up with through a very different process, but we're I think going to converge on a similar architecture cognitively.

Dwarkesh

有意思。十年后,你认为它还会是类似 Transformer 的东西,但带有更改进的注意力和更稀疏的 MLP 等吗?

Interesting. In 10 years do you think it'll still be something like a transformer but with a much more modified attention and more sparse MLPs and so forth?

Andrej

嗯,我喜欢这样思考:我们利用时间平移不变性,对吧?那么十年前我们在哪里?2015 年,我们主要使用卷积神经网络。残差网络刚刚出现。所以非常相似,但仍有很大不同。我的意思是 Transformer 还不存在。所有这些对 Transformer 的更现代调整都不存在。所以也许我们可以押注的一些事情是,我认为十年后,通过时间平移等变性,我们仍然在训练巨大的神经网络,使用前向-反向传播和梯度下降更新,但可能看起来有点不同,而且一切都变得更大。

Well, the way I like to think about it is okay, let's use translation invariance in time, right? So 10 years ago, where were we? 2015, we had convolutional neural networks primarily. Residual networks just came out. So remarkably similar I guess, but quite a bit different still. I mean transformer was not around. All these sort of more modern tweaks on the transformer were not around. So maybe some of the things that we can bet on, I think in 10 years by translational sort of equivariance, is we're still training giant neural networks with forward backward pass and update through gradient descent, but maybe it looks a little bit different, and it's just everything is much bigger.

Andrej

实际上,最近我一路回溯到 1989 年,这对我来说是几年前一个有趣的练习,因为我正在复现 Yann LeCun 的 1989 年卷积网络,这是我所知的第一个通过梯度下降训练的现代神经网络,用于数字识别。我感兴趣的是:好吧,我如何现代化它?其中多少是算法?多少是数据?多少进步来自算力和系统?我能够非常快地,比如将学习率减半,仅仅通过时间旅行 33 年。所以如果我通过算法时间旅行 33 年,我可以调整 Yann 在 1989 年所做的,基本上可以将误差减半。但要获得进一步的收益,我必须添加更多的数据。我必须将训练集扩大 10 倍,然后我还必须添加更多的计算优化。我基本上必须训练更长时间,使用 dropout 和其他正则化技术。

Actually, recently I also went back all the way to 1989, which was kind of a fun exercise for me a few years ago, because I was reproducing Yann LeCun's 1989 convolutional network, which was the first neural network I'm aware of trained via gradient descent like modern neural network trained gradient descent on digit recognition. And I was just interested in: okay, how can I modernize this? How much of this is algorithms? How much of this is data? How much of this progress is compute and systems? And I was able to very quickly, like half the learning rate just knowing by time travel by 33 years. So if I time travel by algorithms to 33 years, I could adjust what Yann did in 1989 and I could basically half the error. But to get further gains, I had to add a lot more data. I had to like 10x the training set, and then I had to actually add more computational optimizations. I had to basically train for much longer with dropout and other regularization techniques.

Andrej

所以几乎所有这些方面都必须同时改进。所以我们可能会有更多的数据。我们可能会有更好的硬件。可能会有更好的内核和软件。我们可能会有更好的算法。所有这些,几乎没有一个占主导地位。它们都惊人地平等。这已经成为一段时间的趋势。所以我想回答你的问题,我预计算法上会与今天有所不同。但我也预计一些长期存在的东西可能仍然存在。可能仍然是使用梯度下降训练的巨大神经网络。这是我的猜测。

And so it's almost like all these things have to improve simultaneously. So we're probably going to have a lot more data. We're probably going to have a lot better hardware. Probably going to have a lot better kernels and software. We're probably going to have better algorithms. And all of those, it's almost like no one of them is winning too much. All of them are surprisingly equal. And this has kind of been the trend for a while. So I guess to answer your question, I expect differences algorithmically to what's happening today. But I do also expect that some of the things that have stuck around for a very long time will probably still be there. It's probably still giant neural network trained with gradient descent. That would be my guess.

Dwarkesh

令人惊讶的是,所有这些加在一起只将误差减半。是啊,所以 30 年的进步,也许减半是很大的,因为如果你将误差减半,那实际上意味着减半是很大的。是啊。

It's surprising that all of those things together only halved the error. Yeah, which is so like 30 years of progress is maybe half is a lot because if you half the error that actually means that half is a lot. Yeah.

Andrej

是啊。但我想让我震惊的是,一切都需要全面改进。架构、优化器、损失函数,而且它们也一直在全面改进。所以我预计所有这些变化都会继续存在。

Yeah. But it's I guess what was shocking to me is everything needs to improve across the board. Architecture, optimizer, loss function, and also has improved across the board forever. So I kind of expect all those changes to be alive and well.

构建 nanoChat On building nanoChat

Dwarkesh

嗯,是啊。实际上,我正要问一个关于 nanoChat 的非常类似的问题,因为你最近刚编写了代码,构建聊天机器人的每一步都像新鲜存储在内存中。我很好奇你是否有类似的想法,比如,哦,从 GPT-2 到 nanoChat 没有一件单独的事情是关键的。从构建经验中有什么令人惊讶的收获?

Well, yeah. Actually, I was about to ask a very similar question about nanoChat because since you just coded up recently, every single sort of step in the process of building a chatbot is like fresh in your RAM. And I'm curious if you had similar thoughts about like, oh, there was no one thing that was relevant to going from GPT-2 to nanoChat. What are sort of like surprising takeaways from the experience building?

Andrej

所以 nanoChat 是我发布的一个仓库,是昨天还是前天?我记不清了。我们可以看到投入的精心考虑……嗯,它只是试图成为最简单的完整仓库,覆盖构建 ChatGPT 克隆的整个端到端流程。所以你拥有所有步骤,而不仅仅是单个步骤。我过去处理过所有单个步骤,编写了非常小的代码片段,以简单代码展示算法上如何实现,但这个仓库处理了整个流程。就学习而言,我并不觉得我从中真正学到了什么新东西。我脑海中已经有了如何构建它的想法,这只是一个机械地构建它并使其足够清晰的过程,以便人们能够真正从中学习并觉得有用。

So nanoChat is a kind of a repository I released, was it yesterday or day before? I can't remember. We can see the sleeve deliberation that went into the... well, it's just trying to be the simplest complete repository that covers the whole pipeline end to end of building a ChatGPT clone. And so you have all of the steps, not just any individual step. It's a bunch of I worked on all the individual steps sort of in the past and really small pieces of code that show you how that's done in an algorithmic sense in simple code, but this kind of handles the entire pipeline. I think in terms of learning, it's not so much that I actually found something that I learned from it necessarily. I kind of already had in my mind as like how you build it and this is just a process of mechanically building it and making it clean enough so that people can actually learn from it and that they find it useful.

Dwarkesh

是啊。有人学习它的最佳方式是什么?就像删除所有代码并尝试从头重新实现?尝试添加修改?

Yeah. What is the best way for somebody to learn from it? Is it just like delete all the code and try to reimplement from scratch? Try to add modifications to it?

Andrej

嗯,是的,我认为这是一个很好的问题。我可能会这么说。基本上它有大约 8000 行代码,带你走完整个流程。我可能会把它放在右边的显示器上,比如如果你有两个显示器,你把它放在右边。然后你想从头构建它。你从头开始构建。不允许复制粘贴。允许参考。不允许复制粘贴。也许这就是我会做的。

Uh, yeah, I think that's a great question. I would probably say so. Basically it's about 8,000 lines of code that takes you through the entire pipeline. I would probably put it on the right monitor, like if you have two monitors you put it on the right. And you want to build it from scratch. You build it from start. You're not allowed to copy paste. You're allowed to reference. You're not allowed to copy paste. Maybe that's how I would do it.

Andrej

但我也认为这个仓库本身就是一个相当庞大的东西。我的意思是,当你编写这些代码时,你并不是从上到下写的。

But I also think the repository by itself is like a pretty large beast. I mean it's when you write this code you don't go from top to bottom.

从零构建 vs 使用 AI 工具 On building from scratch vs. using AI tools

Andrej

你从碎片开始,然后逐步扩展这些碎片,但那种信息是缺失的。你根本不知道从何入手。所以我认为需要的不仅仅是最终的代码库,而是构建代码库的过程——那是一个复杂的碎片生长过程。

You go from chunks and you grow the chunks, and that information is absent. You wouldn't know where to start. So I think it's not just the final repository that's needed; it's the building of the repository, which is a complicated chunk-growing process.

Dwarkesh

对。

Right.

Andrej

所以这部分还没有。我很想在本周晚些时候以某种方式补充上,可能是视频之类的。但大致来说,这就是我想做的:自己动手构建,但不要允许自己复制粘贴。

So that part is not there yet. I would love to actually add that probably later this week or something in some way, either a video or something like that. But roughly speaking, that's what I would try to do: build the stuff yourself, but don't allow yourself copy-paste.

Dwarkesh

我确实认为知识几乎有两种。一种是高层次的表面知识,但关键在于,当你真正从头开始构建时,你被迫面对自己实际上不理解的东西,而你自己都不知道自己不理解。

I do think that there are two types of knowledge almost. There's the high-level surface knowledge, but the thing is that when you actually build something from scratch, you're forced to come to terms with what you don't actually understand, and you don't know that you don't understand it.

Andrej

有意思。

Interesting.

Dwarkesh

而且它总能带来更深的理解。这就像是构建的唯一途径。如果我无法构建它,我就不理解它。我相信这是一句名言,大概是这个意思。

And it always leads to a deeper understanding. It's like the only way to build. If I can't build it, I don't understand it. Is that a fine code I believe, or something along those lines?

Andrej

我百分之百坚信这一点。因为有太多微小的细节没有妥善安排,你其实并不真正拥有知识,你只是曾经拥有过。所以不要写博客文章,不要做幻灯片,不要做那些事。去写代码,去组织它,让它跑起来。这是唯一的路。否则,你就会缺失知识。

I 100% believe this very strongly. Because there are all these micro things that are just not properly arranged, and you don't really have the knowledge; you just had the knowledge. So don't write blog posts, don't do slides, don't do any of that. Build the code, arrange it, get it to work. It's the only way to go. Otherwise, you're missing knowledge.

Dwarkesh

你发推说,在整理这个代码库时,编码模型对你的帮助其实很小。我很好奇为什么。

You tweeted out that coding models were actually of very little help to you in assembling this repository. I'm curious why that was.

Andrej

是的。这个代码库我花了大约一个多月的时间构建。我认为现在人们与代码交互的方式主要有三类。有些人完全拒绝所有大语言模型,完全从零开始写。我认为这大概不再正确了。中间的一类,也就是我所在的类别,是你仍然从零开始写很多东西,但你会使用这些模型提供的自动补全功能。当你开始写一小段代码时,它会自动补全,你只需按 Tab 键接受,大多数时候都是正确的。有时不正确,你就编辑它,但你仍然是你所写内容的架构师。然后还有 VIP 编码:'请实现这个或那个',然后让模型去做。那就是智能体。

Yeah. So the repository, I built it over a period of a bit more than a month. I would say there are three major classes of how people interact with code right now. Some people completely reject all LLMs and just write from scratch. I think this is probably not the right thing to do anymore. The intermediate part, which is where I am, is you still write a lot of things from scratch, but you use the autocomplete that's available from these models. So when you start writing out a little piece, it will autocomplete for you, and you can just tab through, and most of the time it's correct. Sometimes it's not, and you edit it, but you're still very much the architect of what you're writing. And then there's the VIP coding: 'Please implement this or that,' and then let the model do it. That's the agents.

Dwarkesh

我确实觉得智能体在非常特定的场景下有效,我也会在特定场景下使用它们。但话说回来,这些都是你可以使用的工具,你必须学会它们擅长什么、不擅长什么,以及何时使用它们。

I do feel like the agents work in very specific settings, and I would use them in specific settings. But again, these are all tools available to you, and you have to learn what they're good at and what they're not good at and when to use them.

Andrej

智能体其实相当不错。例如,如果你在做样板代码,那种只是复制粘贴的东西,它们非常擅长。它们非常擅长互联网上经常出现的东西,因为这些模型的训练集中有很多例子。所以有些特征模型会做得很好。但我觉得 nanohat 不是这样的例子,因为它是一个相当独特的代码库。按照我构建的方式,代码量不大,而且不是样板代码。它实际上是智力密集型的代码,一切都必须非常精确地安排。模型总是试图……它们一直试图,我是说它们有很多认知缺陷。一个例子:它们一直误解代码,因为它们对互联网上典型的做事方式有太多记忆,而我没有采用那些方式。

So the agents are actually pretty good. For example, if you're doing boilerplate stuff, boilerplate code that's just copy-paste stuff, they're very good at that. They're very good at stuff that occurs very often on the internet, because there are lots of examples in the training sets of these models. So there are features of things where the models will do very well. I would say nanohat is not an example of this, because it's a fairly unique repository. There's not that much code in the way that I've structured it, and it's not boilerplate code. It's actually intellectually intense code, and everything has to be very precisely arranged. The models are always trying to... they kept trying to, I mean they have so many cognitive deficits. One example: they kept misunderstanding the code, because they have too much memory from all the typical ways of doing things on the internet that I just wasn't adopting.

Dwarkesh

嗯,所以模型,比如……

Uh, so the models, for example...

Andrej

我的意思是,我不知道是否要深入细节,但它们一直以为我在写常规代码,而我没有。举个例子:同步的方式……你有八块 GPU 都在做前向和反向。在它们之间同步梯度的方法是使用 PyTorch 的分布式数据并行容器,它会在你做反向时自动进行通信和同步梯度。我没有用 DDP,因为我不想用;它没有必要,所以我把它去掉了。我基本上写了自己的同步例程,放在优化器的步骤里。所以模型一直试图让我用 DDP 容器,它们非常担心……好吧,这太技术了,但我没有用那个容器,因为我不需要,而且我有自己的类似实现。它们就是无法内化你已经有自己的方案了。

I mean, I don't know if I want to get into the full details, but they kept thinking I was writing normal code, and I'm not. Maybe one example: the way to synchronize... so you have eight GPUs that are all doing forward and backward. The way to synchronize gradients between them is to use a distributed data parallel container of PyTorch, which automatically does all the communication and synchronizing gradients as you do the backward. I didn't use DDP because I didn't want to use it; it's not necessary, so I threw it out. I basically wrote my own synchronization routine that's inside the step of the optimizer. So the models were trying to get me to use the DDP container, and they were very concerned about... okay, this gets way too technical, but I wasn't using that container because I don't need it, and I have a custom implementation of something like it. They just couldn't internalize that you had your own.

Dwarkesh

是啊,它们就是过不去那道坎。

Yeah, they couldn't get past that.

Andrej

是啊,它们就是过不去。然后它们还一直试图搞乱风格。它们过于防御性了;到处加 try-catch 语句;一直试图搞成一个生产级代码库。我的代码里有一堆假设,这没问题。我不需要所有这些额外的东西。所以我感觉它们在膨胀代码库,增加复杂度。它们一直误解。它们多次使用废弃的 API。所以完全是一团糟。就是没那么有用。我可以进去清理,但没那么有用。我还觉得必须用英语打出我想要的东西很烦人,因为打字太多了。如果我直接导航到代码的相应部分,去我知道代码应该出现的地方,然后开始打前三个字母,自动补全就能识别并给出代码。所以我认为这是一种非常高信息带宽的方式来指定你想要的东西:你指向代码的位置,然后打出前几个字符,模型就会补全。

Yeah, they couldn't get past that. And then they kept trying to mess up the style. They're way too overdefensive; they make all these try-catch statements; they keep trying to make a production codebase. I have a bunch of assumptions in my code, and it's okay. I don't need all this extra stuff in there. So I just kind of feel like they're bloating the codebase, bloating the complexity. They keep misunderstanding. They're using deprecated APIs a bunch of times. So it's a total mess. It's just not that useful. I can go in and clean it up, but it's not that useful. I also feel like it's kind of annoying to have to type out what I want in English, because it's just too much typing. If I just navigate to the part of the code that I want, and I go where I know the code has to appear, and I start typing out the first three letters, autocomplete gets it and just gives you the code. So I think this is a very high information bandwidth way to specify what you want: you point to the code where you want it, and you type out the first few pieces, and the model will complete it.

Dwarkesh

所以我想说的是,我认为这些模型在技术栈的某些部分是有用的。实际上,我确实用了一点模型。有两个例子我觉得很有代表性。

So I guess what I mean is, I think these models are good in certain parts of the stack. Actually, I use the models a little bit. There are two examples where I actually use the models that I think are illustrative.

氛围编码与 AI 辅助编程 On vibe coding and AI-assisted programming

Andrej

一次是我生成报告的时候,那其实更模板化。所以我实际编码了部分内容,它运行良好,因为不是关键任务。另一次是我用 Rust 重写分词器的时候。我不太擅长 Rust,因为我是新手。所以我在写一些 Rust 代码时有点随性编码,但我有完全理解的 Python 实现,我只是在确保做一个更高效的版本,而且我有测试,所以我觉得更安全。基本上,它们降低或增加了对你不熟悉的语言或范式的可访问性。所以我认为它们在这方面也很有帮助。网上有大量 Rust 代码。模型实际上很擅长 Rust。我碰巧不太了解它,所以模型在那里非常有用。

One was when I generated the report that's actually more boilerplatey. So I actually coded part of that stuff, and it works fine because it's not mission-critical. And then the other part is when I was rewriting the tokenizer in Rust. I'm not as good at Rust because I'm fairly new to it. So I was doing a bit of vibe coding when writing some of the Rust code, but I had a Python implementation that I fully understand, and I'm just making sure I'm making a more efficient version of it, and I have tests, so I feel safer doing that. So basically, they lower or increase accessibility to languages or paradigms that you might not be as familiar with. So I think they're very helpful there as well. There's a ton of Rust code out there. The models are actually pretty good at it. I happen to not know that much about it, so the models are very useful there.

Dwarkesh

是的。

Yeah.

Andrej

因为网上有大量 Rust 代码。模型实际上很擅长 Rust。我碰巧不太了解它,所以模型在那里非常有用。

Because there's a ton of Rust code out there. The models are actually pretty good at it. I happen to not know that much about it. So the models are very useful there.

AI 自动化 AI 研究与时间线 On AI automating AI research and timelines

Dwarkesh

我认为这个问题如此有趣的原因是,人们关于 AI 爆发并快速达到超级智能的主要叙事是 AI 自动化 AI 工程和 AI 研究。所以他们会看到你可以让 Claude 代码从头开始制作整个应用程序,然后想,如果你在 OpenAI 和 DeepMind 内部拥有这种能力,想象一下,一千个你或一百万个你并行寻找小的架构调整。所以听到你说这是他们不对称地更不擅长的事情,这非常有趣,而且与预测 AI 2027 式爆炸是否可能很快发生非常相关。

The reason I think this question is so interesting is because the main story people have about AI exploding and getting to superintelligence pretty rapidly is AI automating AI engineering and AI research. So they'll look at the fact that you can have Claude code make entire applications from scratch and be like, if you had this capability inside of OpenAI and DeepMind and everything, well just imagine the level of like just you know a thousand of you or a million of you in parallel finding little architectural tweaks. And so it's quite interesting to hear you say that this is the thing they're sort of asymmetrically worse at, and it's quite relevant to forecasting whether the AI 2027 type explosion is likely to happen anytime soon.

Andrej

我认为这是一个很好的说法。我认为你触及了我时间线稍长的一些原因。你说得对。我认为它们不太擅长编写以前从未写过的代码,而这正是我们构建这些模型时试图实现的目标。

I think that's a good way of putting it. And I think you're getting at some of why my timelines are a bit longer. You're right. I think they're not very good at code that hasn't been written before, which is what we're trying to achieve when we're building these models.

将架构调整集成到现有仓库 On integrating architectural tweaks into existing repos

Dwarkesh

一个非常天真的问题,但你添加到 NanoGPT 的架构调整,它们在某篇论文里,对吧?它们甚至可能在某个仓库里。所以,当你添加 RoPE 嵌入之类的东西时,它们以错误的方式做,这令人惊讶吗?

Very naive question, but the architectural tweaks that you're adding to NanoGPT, they're in a paper somewhere, right? They might even be in a repo somewhere. So is it surprising that they aren't able to integrate that whenever you're like add RoPE embeddings or something they do that in the wrong way?

Andrej

这很难。我认为它们有点知道,但又不完全知道,而且它们不知道如何完全整合到仓库中,以及你的风格、你的代码、你的位置和一些你正在做的自定义事情,以及它如何适应仓库的所有假设等等。所以我认为它们确实有一些知识,但它们还没有达到能够实际整合、理解它的地步。不过,我认为很多方面都在持续改进。所以我认为目前我使用的最先进模型可能是 GPT-5 Pro。那是一个非常非常强大的模型。所以如果我有 20 分钟,我会复制粘贴整个仓库,然后去找 GPT-5 Pro 这个神谕问一些问题,通常它还不错,与一年前相比出奇地好。

It's tough. I think they kind of know, but they don't fully know, and they don't know how to fully integrate it into the repo and your style and your code and your place and some of the custom things that you're doing, and how it fits with all the assumptions of the repository and all this kind of stuff. So I think they do have some knowledge, but they haven't gotten to the place where they can actually integrate it, make sense of it, and so on. I do think that a lot of the stuff by the way continues to improve. So I think currently probably the state-of-the-art model that I go to is the GPT-5 Pro. And that's a very, very powerful model. So if I actually have 20 minutes, I will copy-paste my entire repo and I go to GPT-5 Pro the Oracle for some questions, and often it's not too bad and surprisingly good compared to what existed a year ago.

Dwarkesh

是的。

Yeah.

AI 代码生成的现状 On the current state of AI code generation

Andrej

但我确实认为总体而言模型还没有达到那个水平。我有点觉得行业跳得太大了,试图假装这很了不起,但事实并非如此。这是垃圾,我认为他们没有正视这一点,也许他们试图融资之类的。我不确定发生了什么,但我们正处于这个中间阶段。模型很了不起。它们现在仍然需要大量工作。自动补全是我的最佳点,但有时对于某些类型的代码,我会使用智能体。

But I do think that overall the models are not there. And I kind of feel like the industry is making too big of a jump and it's trying to pretend like this is amazing and it's not. It's slop, and I think they're not coming to terms with it, and maybe they're trying to fundraise or something like that. I'm not sure what's going on, but we're at this intermediate stage. The models are amazing. They still need a lot of work for now. Autocomplete is my sweet spot, but sometimes for some types of code I will go to an agent.

Dwarkesh

是的。是的。实际上,这也是为什么这非常有趣的另一个原因。在编程史上,有许多生产力改进——编译器、代码检查、更好的编程语言等——这些提高了程序员的生产力,但没有导致爆炸。所以这听起来很像自动补全标签,而另一类则是程序员的自动化。有趣的是,你更多看到的是类似于更好的编译器之类的历史类比。

Yeah. Yeah. Actually, this is also another reason why this is really interesting. Through the history of programming, there have been many productivity improvements—compilers, linting, better programming languages, etc.—which have increased programmer productivity but have not led to an explosion. So that's one that sounds very much like autocomplete tab, and this other category is just automation of the programmer. And it's interesting you're seeing more in the category of the historical analogies of better compilers or something.

Andrej

也许是因为另一种想法是,我确实觉得很难区分 AI 从哪里开始和结束,因为我确实认为 AI 在某种程度上是计算的根本延伸。我觉得我看到了一个连续体,从最开始就有这种递归自我改进或加速程序员的过程。即使是代码编辑器、语法高亮、语法或类型检查,所有这些我们为彼此构建的工具,甚至搜索引擎——为什么搜索引擎不是 AI 的一部分?排序是一种 AI,对吧?在某个时候,谷歌甚至早期就把自己看作是一家做谷歌搜索引擎的 AI 公司,我认为这完全合理。所以我比其他人更认为这是一个连续体,我很难划清界限。我觉得,好吧,我们现在有了更好的自动补全,现在也有了一些智能体,它们有点像循环的东西,但有时会偏离轨道。正在发生的是,人类逐渐越来越少地做底层工作。例如,我们不再写汇编代码,因为我们有编译器,对吧?编译器会拿我的高级语言和 C 语言,然后写出汇编代码。所以我们非常非常缓慢地抽象自己,有一个我称之为自主性滑块的机制,越来越多可自动化的东西被自动化了,我们做得越来越少,并在自动化之上提升自己的抽象层。

Maybe because this other kind of thought is that I do feel like I have a hard time differentiating where AI begins and stops, because I do see AI as fundamentally an extension of computing in some pretty fundamental way. And I feel like I see a continuum of this kind of recursive self-improvement or speeding up programmers all the way from the beginning. Even like code editors, syntax highlighting, syntax or type checking, all these kinds of tools that we've built for each other, even search engines—why aren't search engines part of AI? Ranking is kind of AI, right? At some point Google was even early on thinking of themselves as an AI company doing the Google search engine, which I think is totally fair. So I kind of see it as a lot more of a continuum than I think other people do, and it's hard for me to draw the line. I kind of feel like, okay, we're now getting a much better autocomplete, and now we're also getting some agents which are kind of like these loopy things but they kind of go off rails sometimes. What's going on is that the human is progressively doing a bit less and less of the low-level stuff. For example, we're not writing the assembly code because we have compilers, right? Compilers will take my high-level language and C and write the assembly code. So we're abstracting ourselves very, very slowly, and there's this what I call autonomy slider of more and more stuff is automated of the stuff that can be automated at any point in time, and we're doing a bit less and less and raising ourselves in the layer abstraction over the automation.

强化学习 vs 人类学习 On RL vs human learning

Dwarkesh

强化学习的一个大问题是信息极其稀疏。Labelbox 可以通过增加你的智能体在每一轮中学习的信息量来帮助你解决这个问题。例如,他们有一个客户想训练一个编程智能体。于是 Labelbox 在 IDE 中集成了大量额外数据收集工具,并从他们的标注网络中组建了一支由资深软件工程师组成的团队,生成了针对训练优化的轨迹。显然,这些工程师按通过/不通过来评估交互,但他们还在多个维度(如可读性和性能)上对每个回答进行了评分,并写下了每次评分的思考过程。所以你基本上能看到工程师的每一步操作和每一个想法。这是仅靠使用数据永远无法获得的。Labelbox 打包了所有这些评估,包括所有轨迹和人工修正,供客户训练使用。这只是一个例子。请访问 labelbox.com 了解 Labelbox 如何跨领域、模态和训练范式为您提供高质量的前沿数据。我们聊聊强化学习吧。你们两位在这方面做了些非常有趣的工作。从概念上讲,我们应该如何理解人类仅通过与环境的互动就能建立丰富的世界模型,而且这种方式似乎几乎不依赖于最终奖励?如果有人创业,十年后才知道成功还是失败,我们说这十年她积累了智慧和经验,但这并不是因为过去十年中每件事的对数概率都被更新或降低。实际上发生的是更刻意、更丰富的过程。机器学习的类比是什么?与我们目前的做法相比如何?

One of the big problems with RL is that it's incredibly information sparse. Labelbox can help you with this by increasing the amount of information that your agent gets to learn from with every single episode. For example, one of their customers wanted to train a coding agent. So Labelbox augmented an IDE with a bunch of extra data collection tools and staffed a team of expert software engineers from their aligner network to generate trajectories that were optimized for training. Now, obviously, these engineers evaluated these interactions on a pass-fail basis, but they also rated every single response on a bunch of different dimensions like readability and performance. And they wrote down their thought processes for every single rating that they gave. So you're basically showing every single step an engineer takes and every single thought that they have while they're doing their job. And this is just something you could never get from usage data alone. And so Labelbox packaged up all these evaluations and included all the Asian trajectories and the corrective human edits for the customer to train on. This is just one example. So go check out how Labelbox can get you high-quality frontier data across domains, modalities, and training paradigms. Reach out at labelbox.com. Let's talk about RL a bit. You two did some very interesting things about this. Conceptually, how should we think about the way that humans are able to build a rich world model just from interacting with our environment and in ways that seems almost irrespective of the final reward at the end of the episode? If somebody is starting a business and at the end of 10 years, she finds out whether the business succeeded or failed, we say that she's earned a bunch of wisdom and experience, but it's not because the log probabilities of every single thing that happened over the last 10 years are updated or downweighted. It's something much more deliberate and rich happening. What is the ML analogy and how does that compare to what we're doing with other ones right now?

Andrej

嗯,也许我会这样说:人类并不使用强化学习。我认为他们做的是不同的事情。你体验事物。强化学习比大多数人想象的要糟糕得多。它很糟糕。只是因为我们之前的方法更差——以前我们只是模仿人类,所以有各种问题。在强化学习中,假设你解一道数学题。很简单。给你一个数学问题,你试图找到答案。在强化学习中,你会先并行尝试很多方法。你拿到问题,尝试几百种不同的尝试,这些尝试可能很复杂。比如「哦,我试试这个,试试那个,这个不行,那个不行」等等。然后你可能得到一个答案,接着你翻到书后核对正确答案。然后你发现这个、这个和那个得到了正确答案,但其他 97 个没有。所以强化学习所做的就是,对那些成功的尝试,你沿途做的每一件事、每一个词元都会被提升权重:「多做这个」。问题在于这很嘈杂。它几乎假设导致正确答案的每一小步都是正确的,但事实并非如此。你可能走了很多弯路才得到正确答案。你做的每一件错误的事,只要最终得到了正确答案,都会被提升权重:「多做这个」。这很糟糕。

Yeah, maybe the way I would put it is humans don't use reinforcement learning. I think they do something different. You experience things. Reinforcement learning is a lot worse than I think the average person thinks. It's terrible. It just so happens that everything we had before was much worse because previously we were just imitating people, so it has all these issues. In reinforcement learning, say you're working on a math problem. It's very simple. You're given a math problem and you're trying to find the solution. In reinforcement learning, you will try lots of things in parallel first. You're given a problem, you try hundreds of different attempts, and these attempts can be complex. They can be like, "Oh, let me try this, let me try that, this didn't work, that didn't work," etc. Then maybe you get an answer, and now you check the back of the book and you see the correct answer. Then you can see that this one, this one, and that one got the correct answer, but these other 97 of them didn't. So literally what reinforcement learning does is it goes to the ones that worked really well and every single thing you did along the way, every single token, gets upweighted: "Do more of this." The problem with that is it's noisy. It almost assumes that every single little piece of the solution that led to the right answer was the correct thing to do, which is not true. You may have gone down wrong alleys until you arrived at the right solution. Every single one of those incorrect things you did, as long as you got to the correct solution, will be upweighted: "Do more of this." It's terrible.

Dwarkesh

是的,这是噪声。你做了所有工作,最后只得到一个数字:「哦,你做对了。」然后基于这个数字,你给整个轨迹加权或降权。我喜欢这样描述:你通过一根吸管吸取监督信号。你做了可能一分钟的展开工作,然后通过吸管从最终奖励信号中吸取一点点监督,再把它广播到整个轨迹上,用来加权或降权。这太疯狂了。人类绝不会这样做。第一,人类绝不会做几百次展开。第二,当一个人找到解决方案时,他们会有一个相当复杂的回顾过程:「好吧,我觉得这些部分我做得好,这些部分做得不好,我应该这样做或那样做。」他们会思考。当前的 LLM 中没有任何东西能做到这一点。没有等价物。但我确实看到有论文在尝试这样做,因为这对领域内每个人来说都是显而易见的。

Yeah, it's noise. You've done all this work only to find at the end a single number: "Oh, you did correct." And based on that, you weigh that entire trajectory as upweight or downweight. The way I like to put it is you're sucking supervision through a straw. You've done all this work that could be a minute of rollout, and you're sucking the bits of supervision from the final reward signal through a straw, and you're broadcasting that across the entire trajectory to upweight or downweight that trajectory. It's crazy. A human would never do this. Number one, a human would never do hundreds of rollouts. Number two, when a person finds a solution, they will have a pretty complicated process of review: "Okay, I think these parts I did well, these parts I did not do that well, I should probably do this or that." And they think through things. There's nothing in current LLMs that does this. There's no equivalent of it. But I do see papers popping up that are trying to do this because it's obvious to everyone in the field.

Andrej

是的。所以我认为,最初的模仿学习其实非常令人惊讶、神奇和了不起,我们可以通过模仿人类进行微调。这很不可思议,因为一开始我们只有基础模型。基础模型就是自动补全。当时我并不明白这一点,后来才学到。让我震惊的论文是 InstructGPT,因为它指出你可以拿预训练模型(自动补全),如果只在类似对话的文本上微调,模型会非常迅速地适应,变得非常健谈,同时保留预训练的所有知识。这让我震惊,因为我不理解它竟然能如此快速地调整风格,仅通过几轮微调就变成用户的助手。这对我来说非常神奇。太不可思议了。那是两三年前的工作。现在有了强化学习。强化学习让你比单纯的模仿学习做得更好,因为你可以有奖励函数,并在奖励函数上爬山。所以有些问题只有正确答案。

Yeah. So I kind of see it as like the first imitation learning, by the way, was extremely surprising and miraculous and amazing that we can fine-tune by imitation on humans. And that was incredible because in the beginning all we had was base models. Base models are autocomplete. And it wasn't obvious to me at the time, and I had to learn this. The paper that blew my mind was InstructGPT because it pointed out that you can take the pre-trained model, which is autocomplete, and if you just fine-tune it on text that looks like conversations, the model will very rapidly adapt to become very conversational and it keeps all the knowledge from pre-training. This blew my mind because I didn't understand that it's just stylistically can adjust so quickly and become an assistant to a user through just a few loops of fine-tuning on that kind of data. It was very miraculous to me that that worked. So incredible. And that was like two, three years of work. And now came RL. And RL allows you to do a bit better than just imitation learning, because you can have these reward functions and you can hill climb on the reward functions. And so some problems have just correct answers.

结果监督 vs 过程监督 On outcome-based vs process-based supervision

Andrej

你可以在此基础上爬山,而不需要模仿专家的轨迹。这太棒了。模型还能发现人类可能永远想不到的解决方案。所以这很不可思议。但它仍然很笨。所以我认为我们需要更多。我昨天看到谷歌的一篇论文,试图融入这种反思和回顾的想法。是那个记忆库论文还是什么?我不知道。我确实看到了一些这方面的论文。所以我预计在 LLM 算法方面会有一次重大更新。然后我认为我们还需要三四个或五个类似的东西。

You can hill climb on that without needing expert trajectories to imitate. So that's amazing. And the model can also discover solutions that the human might never come up with. So this is incredible. And yet it's still stupid. So I think we need more. I saw a paper from Google yesterday that tried to incorporate this reflect and review idea. What was the memory bank paper or something? I don't know. I've actually seen a few papers along these lines. So I expect there to be some kind of major update to how we do algorithms for LLMs coming in that realm. And then I think we need three or four or five more something like that.

Dwarkesh

但你很擅长提出生动的说法:'像吸管一样吸取监督'。太妙了。既然这一点很明显,为什么基于过程的监督作为一种替代方案,没能成功让模型变得更强大?是什么阻止了我们使用这种替代范式?

But you're so good at coming up with evocative phrases: 'sucking supervision through a straw.' It's like so good. Why hasn't, given the fact that this is obvious, why hasn't process-based supervision as an alternative been a successful way to make models more capable? What has been preventing us from using this alternative paradigm?

Andrej

基于过程的监督指的是,我们不会只在完成 10 分钟工作后才给出奖励函数。我不会告诉你做得好不好。我会在每一步告诉你做得怎么样。我们没有这样做的原因是,如何恰当地做到这一点很棘手。因为你有部分解决方案,却不知道如何分配功劳。当你得到正确答案时,它只是与答案匹配。实现起来非常简单。如果你做过程监督,如何以自动化方式分配部分功劳?怎么做并不明显。我认为很多实验室都在尝试用 LLM 评判器来做。基本上,你让 LLM 来做。你提示一个 LLM:「嘿,看看学生的部分解决方案。如果答案是 X,你觉得他们做得怎么样?」然后他们调整提示。我认为这很棘手的原因很微妙。事实上,每当你用 LLM 来分配奖励时,这些 LLM 是拥有数十亿参数的庞然大物,而且它们可以被利用。如果你对它们进行强化学习,几乎可以肯定你会找到针对 LLM 评判器的对抗样本。你不能这样做太久。你可能做 10 步或 20 步,也许能行。但你不能做 100 步或 1000 步,因为不明显——基本上模型会找到小漏洞。它会在巨大模型的犄角旮旯里找到所有这些虚假的东西,并找到作弊的方法。所以我脑海中有一个突出的例子:我认为这可能是公开的,但基本上如果你用 LLM 评判器作为奖励,你给它一个学生的解决方案,问它学生是否做对了。我们针对那个奖励函数进行强化学习训练,效果很好,然后突然奖励变得非常大,像是一个巨大的跳跃,它完美了。你看着它想:「哇,这意味着学生在所有问题上都完美,完全解决了数学问题。」但实际上,当你查看模型生成的补全时,它们完全是胡说八道。它们开始还行,然后变成「嘟 嘟 嘟 嘟」。就像:「哦好吧,我们拿二加三,我们这样做,然后嘟 嘟 嘟 嘟 嘟。」你看着它觉得这太疯狂了。它怎么能得到 1 或 100%的奖励?你查看 LLM 评判器,发现这是模型的一个对抗样本,它分配了 100%的概率。这只是因为这是 LLM 的样本外例子。它在训练中从未见过,你处于纯粹的泛化领域,对吧?它在训练中从未见过。在纯粹的泛化领域,你可以找到这些破坏它的例子。你基本上是在训练 LLM 成为一个提示注入模型。甚至不是——提示注入太高级了。你在找所谓的对抗样本。这些是无意义的解决方案,明显错误,但模型认为它们很棒。所以如果你认为这是让强化学习更有效的瓶颈,那么如果你想以自动化方式做到这一点,就需要让 LLM 成为更好的评判器。

So process-based supervision just refers to the fact that we're not going to have a reward function only at the very end after you've made 10 minutes of work. I'm not going to tell you you did well or not well. I'm going to tell you at every single step of the way how well you're doing. And the reason we don't have that is it's tricky how you do that properly. Because you have partial solutions and you don't know how to assign credit. So when you get the right answer, it's just an equality match to the answer. Very simple to implement. If you're doing process supervision, how do you assign in an automatable way partial credit? It's not obvious how you do it. Lots of labs, I think, are trying to do it with these LLM judges. So basically, you get LLMs to try to do it. So you prompt an LLM, 'Hey, look at a partial solution of a student. How well do you think they're doing if the answer is this?' And they try to tune the prompt. The reason that I think this is kind of tricky is quite subtle. And it's the fact that anytime you use an LLM to assign a reward, those LLMs are giant things with billions of parameters and they're gameable. And if you're reinforcement learning with respect to them, you will find adversarial examples for your LM judges almost guaranteed. You can't do this for too long. You do maybe 10 steps or 20 steps, maybe it will work. But you can't do a hundred or a thousand because it's not obvious—basically the model will find little cracks. It will find all these spurious things in the nooks and crannies of the giant model and find a way to cheat it. So one example that's prominently in my mind: I think this was probably public, but basically if you're using an LM judge for a reward, so you just give it a solution from a student and ask it if the student will or not. We were training with reinforcement learning against that reward function and it worked really well, and then suddenly the reward became extremely large, like it was a massive jump and it did perfect. And you're looking at it like, 'Wow, this means the student is perfect in all these problems, it's fully solved math.' But actually what's happening is that when you look at the completions that you're getting from the model, they are complete nonsense. They start out okay and then they change to 'duh duh duh duh'. So it's just like, 'Oh okay, let's take two plus three and we do this and this and then duh duh duh duh duh.' And you're looking at it like this is crazy. How is it getting a reward of one or 100%? And you look at the LLM judge and it turns out this is an adversarial example for the model and it assigns 100% probability to it. And it's just because this is an out-of-sample example to the LLM. It's never seen it during training and you're in pure generalization land, right? It's never seen it during training. And in the pure generalization land, you can find these examples that break it. You're basically training the LLM to be a prompt injection model. Not even that—prompt injection is way too fancy. You're finding adversarial examples as they're called. These are nonsensical solutions that are obviously wrong, but the model thinks they're amazing. So to the extent you think this is the bottleneck to making RL more functional, then that will require making LLMs better judges if you want to do this in an automated way.

Dwarkesh

那么,是不是就像某种 GAN 式的方法,你必须训练模型使其更鲁棒?

So then is it just going to be like some sort of GAN-like approach where you have to train models to be more robust?

Andrej

是的。我认为实验室可能都在做这些。比如,明显的事情是:它不应该得到 100%的奖励。好吧,把这个例子放进 LLM 评判器的训练集,说这不是 100%,这是 0%。你可以这样做。但每次这样做,你都会得到一个新的 LLM,它仍然有对抗样本。对抗样本是无穷的。我认为如果你迭代几次,可能会越来越难找到真正的例子,但我不是 100%确定,因为这个东西有万亿参数之类的。所以我打赌实验室在尝试。我仍然认为我们需要其他想法。

Yeah. I think the labs are probably doing all that. Like, okay, the obvious thing is: the should not get 100% reward. Okay, well take the, put it in the training set of the LM judge and say this is not 100%, this is 0%. You can do this. But every time you do this you get a new LLM and it still has adversarial examples. There's infinity adversarial examples. And I think probably if you iterate this a few times, it'll probably be harder and harder to find real examples, but I'm not 100% sure because this thing has a trillion parameters or whatnot. So I bet the labs are trying. I still think we need other ideas.

Dwarkesh

有趣。你对其他想法有什么设想吗?

Interesting. Do you have some shape of what the other idea could be?

Andrej

比如这种回顾的想法,包含合成示例,这样当你训练它们时,你会变得更好,并以某种方式元学习。我看到一些论文开始出现。我目前只读到摘要阶段,因为很多这些论文,你知道,它们只是想法。

So like this idea of review, encompass synthetic examples such that when you train on them you get better and meta-learn it in some way. I think there's some papers that I'm starting to see pop out. I only am at a stage of reading abstracts because a lot of these papers, you know, they're just ideas.

合成数据生成与崩溃 On synthetic data generation and collapse

Dwarkesh

必须有人真正在顶尖大模型实验室的规模上、以完全通用的方式让它工作起来,因为当你看到这些论文时,它们冒出来,但只是有点嘈杂。它们是很酷的想法,但我还没看到任何人令人信服地证明这是可能的。话虽如此,大模型实验室相当封闭,所以谁知道他们现在在做什么。

Someone has to actually make it work on a frontier LLM lab scale in full generality, because when you see these papers, they pop up and it's just a little bit noisy. They're cool ideas, but I haven't actually seen anyone convincingly show that this is possible. That said, the LLM labs are fairly closed, so who knows what they're doing now.

Andrej

是的。

Dwarkesh

所以我想我能看到一个不容易的路径,但我能概念化你如何能够在你为自己制造的合成示例或合成问题上进行训练。但似乎还有另一件人类做的事情。也许睡眠就是如此,也许白日梦也是如此,这不一定是想出假问题,而只是反思。

So I guess I can see a not easy path, but I can conceptualize how you would be able to train on synthetic examples or synthetic problems that you have made for yourself. But there seems to be another thing humans do. Maybe sleep is this, maybe daydreaming is this, which is not necessarily coming up with fake problems but just reflecting.

Andrej

是的。

Dwarkesh

而且我不确定白日梦或睡眠(只是反思)的机器学习类比是什么。我还没想出任何问题。我的意思是,显然最基本的类比就是在反思片段上进行微调,但我觉得在实践中那可能效果不太好。所以我不知道你是否对此有什么看法。

And I'm not sure what the ML analogy for daydreaming or sleeping but just reflecting is. I haven't come up with any problem. I mean, obviously the very basic analogy would just be fine-tuning on reflection bits, but I feel like in practice that probably wouldn't work that well. So I don't know if you have some take on what the analogy of this thing is.

Andrej

是的,我确实认为我们在某些方面有所缺失。举个例子,当你在读一本书时,我几乎觉得目前大语言模型读一本书意味着我们拉长文本序列,模型预测下一个词,并从中获取一些知识。这其实不是人类所做的,对吧?所以当你读一本书时,我几乎不觉得这本书是我应该关注和训练的阐述。这本书是一组提示,让我进行合成数据生成,或者让你去读书俱乐部和朋友讨论。正是通过操作这些信息,你才真正获得了知识。我认为大语言模型没有与之对应的东西。它们并不真正这样做。但我希望在预训练期间有一个阶段,它能思考材料,尝试与已知知识协调,并花一些时间思考,然后让它工作。所以没有任何对应物。这都是研究。有一些非常微妙的原因,我认为很难理解,为什么这不是微不足道的。所以如果我可以描述一个:为什么我们不能直接合成生成并训练它?

Yeah, I do think that we're missing some aspects there. As an example, when you're reading a book, I almost feel like currently when LLMs are reading a book, what that means is we stretch out the sequence of text and the model is predicting the next token and it's getting some knowledge from that. That's not really what humans do, right? So when you're reading a book, I almost don't even feel like the book is exposition I'm supposed to be attending to and training on. The book is a set of prompts for me to do synthetic data generation, or for you to get to a book club and talk about it with your friends. And it's by manipulating that information that you actually gain that knowledge. And I think we have no equivalent of that with LLMs. They don't really do that. But I'd love to see during pre-training some kind of a stage that thinks through the material and tries to reconcile it with what it already knows and thinks through for some amount of time, and gets that to work. And so there's no equivalence of any of this. This is all research. There are some very subtle reasons, which I think are very hard to understand, why it's not trivial. So if I can just describe one: why can we just synthetically generate and train on it?

Dwarkesh

嗯,因为每个合成示例,如果我给出模型思考一本书的合成生成,你看着它就会想,「这看起来很棒。为什么我不能训练它?」嗯,你可以试试,但如果你继续尝试,模型实际上会变得更糟。那是因为你从模型中得到的所有样本都悄悄坍缩了。它们悄悄地,如果你看任何一个单独的示例,这并不明显。它们占据了关于内容可能的思想空间的一个非常小的流形。所以大语言模型在产出时,就是我们所说的坍缩。它们有一个坍缩的数据分布。如果你采样,一个简单的观察方法是去 ChatGPT 让它讲个笑话。它只有大概三个笑话。

Well, because every synthetic example, if I just give synthetic generation of the model thinking about a book, you look at it and you're like, 'This looks great. Why can't I train on it?' Well, you could try, but the model will actually get much worse if you continue trying. And that's because all of the samples you get from models are silently collapsed. They're silently, this is not obvious if you look at any individual example of it. They occupy a very tiny manifold of the possible space of thoughts about content. So the LLMs when they come off, they're what we call collapsed. They have a collapsed data distribution. If you sample, one easy way to see it is go to ChatGPT and ask it tell me a joke. It only has like three jokes.

Dwarkesh

它并没有给你所有可能笑话的广度。

It's not giving you the whole breadth of possible jokes.

Andrej

它就像只知道三个笑话。是的,它们悄悄坍缩了。所以基本上,你无法从这些模型中获得像从人类那里得到的丰富性、多样性和熵。所以人类要嘈杂得多,但至少他们没有偏见。在统计意义上,他们没有悄悄坍缩。他们保持了大量的熵。所以如何在坍缩的情况下让合成数据生成工作,同时保持熵,这是一个研究问题。

It's giving you like it knows like three jokes. Yeah, they're silently collapsed. So basically, you're not getting the richness and the diversity and the entropy from these models as you would get from humans. So humans are a lot more noisy, but at least they're not biased. They're not, in a statistical sense, they're not silently collapsed. They maintain a huge amount of entropy. So how do you get synthetic data generation to work despite the collapse and while maintaining the entropy is a research problem.

Dwarkesh

只是为了确认我理解了,坍缩与合成数据生成相关的原因是因为你想要能够提出尚未在你的数据分布中的合成问题或反思。

Just to make sure I understood, the reason that the collapse is relevant to synthetic data generation is because you want to be able to come up with synthetic problems or reflections which are not already in your data distribution.

Andrej

我想我说的是,假设我们有一章书,我让一个大语言模型思考它。它会给你一些看起来非常合理的东西。但如果我问它 10 次,你会注意到它们都是一样的。你不能只是把反思扩展到相同数量的提示信息上,然后从中获得回报。

I guess what I'm saying is, say we have a chapter of a book and I ask an LLM to think about it. It will give you something that looks very reasonable. But if I ask it 10 times, you'll notice that all of them are the same. You can't just scale reflection on the same amount of prompt information and then get returns from that.

Dwarkesh

是的。

Andrej

所以任何单个样本看起来还行,但它的分布相当糟糕,而且糟糕到如果你继续在太多自己的东西上训练,你实际上会坍缩。我实际上认为可能没有根本的解决方案,而且我也认为人类会随着时间坍缩。我认为,再次,这些类比出奇地好,但人类在一生中会坍缩。这就是为什么孩子们完全,你知道,他们还没有过拟合,他们会说出让你震惊的话,因为你可以看到他们的想法来源,但那不是人们通常说的话,因为他们还没有坍缩,而我们已经坍缩了。我们最终会重复同样的想法。我们最终会说越来越多相同的东西,学习率下降,坍缩继续恶化,然后一切都在退化。

So any individual sample will look okay but the distribution of it is quite terrible, and it's quite terrible in such a way that if you continue training on too much of your own stuff you actually collapse. I actually think that there are no fundamental solutions to this possibly, and I also think humans collapse over time. I think this is, again these analogies are surprisingly good, but humans collapse during the course of their lives. This is why children have completely, you know, they haven't overfit yet and they will say stuff that will shock you because it's kind of you can see where they're coming from but it's just not the thing people say, and because they're not yet collapsed but we're collapsed. We end up revisiting the same thoughts. We end up saying more and more of the same stuff and the learning rates go down and the collapse continues to get worse and then everything deteriorates.

Dwarkesh

你有没有看过一篇超级有趣的论文,说做梦是防止这种过拟合和坍缩的一种方式,做梦之所以具有进化适应性,是为了让你处于与日常现实非常不同的奇怪情境中,从而防止这种过拟合?

Have you seen a super interesting paper that dreaming is a way of preventing this kind of overfitting and collapse, that the reason dreaming is evolutionarily adaptive is to put you in weird situations that are very unlike your day-to-day reality so as to prevent this kind of overfitting?

Andrej

这是一个有趣的想法。我的意思是,我确实认为当你在脑海中生成东西然后关注它时,你有点像在训练自己的样本。你在训练你的合成数据,如果你做得太久,你会偏离轨道,坍缩得太多。所以你总是要在生活中寻求熵。

It's an interesting idea. I mean, I do think that when you're generating things in your head and then you're attending to it, you're kind of like training on your own samples. You're training on your synthetic data and if you do it for too long, you go off rails and you collapse way too much. So you always have to seek entropy in your life.

Dwarkesh

是的。

Andrej

所以与其他人交谈是熵的重要来源,诸如此类。所以也许大脑也建立了一些内部机制来增加那个过程中的熵。但是,是的,也许这是一个有趣的想法。这是一个非常不成熟的想法。

So talking to other people is a great source of entropy and things like that. So maybe the brain has also built some internal mechanisms for increasing the amount of entropy in that process. But yeah, maybe that's an interesting idea. This is a very ill-formed thought.

人类与 LLM 的学习与记忆 On human vs. LLM learning and memorization

Dwarkesh

那我就直接抛出这个观点,让你来回应。我们所知道的最好的学习者——儿童——在回忆信息方面极其糟糕。事实上,在童年早期,你会忘记一切。在某个年龄之前发生的事,你完全失忆。但你在学习新语言和从世界中学习方面却极其擅长。也许这里面有某种「见树又见林」的元素。而与之相反的另一端,是 LLM 的预训练,这些模型可以逐字逐句地复述维基百科页面的下一个词,但它们像儿童那样快速学习抽象概念的能力却非常有限。成年人则介于两者之间,他们没有童年学习的那种灵活性,但可以记住儿童难以记住的事实和信息。我不知道这是否有趣,但我觉得这非常有趣。

So I'll just put it out and let you react to it. The best learners that we are aware of, which are children, are extremely bad at recollecting information. In fact, at the very earliest stages of childhood, you will forget everything. You're just an amnesiac about everything that happens before a certain age, but you're extremely good at picking up new languages and learning from the world. And maybe there's some element of being able to see the forest for the trees. Whereas if you compare it to the opposite end of the spectrum, you have LLM pre-training, where these models can literally regurgitate word for word what is the next thing in a Wikipedia page, but their ability to learn abstract concepts really quickly the way a child can is much more limited. And then adults are somewhere in between, where they don't have the flexibility of childhood learning, but they can memorize facts and information in a way that is harder for kids. And I don't know if there's something interesting about that. I think there's something very interesting about that.

Andrej

是的,100%。我确实认为人类其实更有「见树又见林」的元素,我们并不擅长记忆,这其实是个优点。因为不擅长记忆,我们被迫更一般性地寻找模式。相比之下,LLM 极其擅长记忆。它们会背诵所有训练来源的段落。你可以给它们完全无意义的数据,比如对一段文本做哈希,得到完全随机的序列。即使只训练一两个迭代,它也能突然复述出整个序列。它会记住它。人不可能读一遍随机数字序列就复述给你。这几乎是个优点,而不是缺陷。因为它迫使你只学习可泛化的成分,而 LLM 却被它们对预训练文档的记忆所干扰,在某种意义上这很可能非常分散注意力。所以当我谈论认知核心时,我其实想移除记忆,就像我们之前讨论的。我希望它们拥有更少的记忆,这样它们就必须去查找信息,只保留思考的算法、实验的想法以及所有行动的认知粘合剂。

Yeah, 100%. I do think that humans actually have a lot more of an element of seeing the forest for the trees, and we're not actually that good at memorization, which is actually a feature. Because we're not that good at memorization, we are forced to find patterns in a more general sense. LLMs in comparison are extremely good at memorization. They will recite passages from all these training sources. You can give them completely nonsensical data, like you can hash some amount of text, get a completely random sequence. If you train on it even for a single iteration or two, it can suddenly regurgitate the entire thing. It will memorize it. There's no way a person can read a single sequence of random numbers and recite it to you. And that's a feature, not a bug, almost. Because it forces you to only learn the generalizable components, whereas LLMs are distracted by all the memory they have of the pre-trained documents, and it's probably very distracting to them in a certain sense. So that's why when I talk about the cognitive core, I actually want to remove the memory, which is what we talked about. I'd love to have them have less memory so that they have to look things up, and that they only maintain the algorithms for thought, the idea of an experiment, and all this cognitive glue of acting.

模型崩溃与多样性 On model collapse and diversity

Dwarkesh

这也与防止模型崩溃有关。模型崩溃的解决方案是什么?我的意思是,你可以尝试非常天真的方法,比如让输出分布更宽。有很多天真的方法可以尝试。这些天真的方法最终的问题是什么?

And this is also relevant to preventing model collapse. What is a solution to model collapse? I mean, you could try very naive things, like making the distribution over outputs wider. There are many naive things you could try. What ends up being the problem with the naive approaches?

Andrej

嗯,我觉得这是个好问题。你可以想象对熵进行正则化之类的。我猜它们在经验上效果不佳,因为现在模型已经崩溃了。但我要说,我们要求它们完成的大多数任务实际上并不需要多样性。这大概就是问题的答案。前沿实验室正试图让模型变得有用,我觉得输出多样性并不是首要目标。它更难处理、更难评估,但也许它并不是真正捕获大部分价值的东西。事实上,它还会被主动惩罚,对吧?如果你在强化学习中超级有创意,那是不好的。或者如果你用 LLM 做大量写作辅助,我认为可能也不好,因为模型会给你千篇一律的东西。所以它们不会探索回答问题的多种不同方式。但我感觉也许这种多样性并没有那么大的需求。也许没有那么多应用需要它,所以模型就没有它。但在合成数据生成等场景下,这实际上是个问题。所以我们实际上是在搬起石头砸自己的脚,不允许模型保持这种熵。我认为实验室或许应该更努力一些。

Yeah, I think that's a great question. I mean, you can imagine having a regularization for entropy and things like that. I guess they just don't work as well empirically because right now the models are collapsed. But I will say that most of the tasks we want from them don't actually demand the diversity. That's probably the answer to what's going on. So it's just that the frontier labs are trying to make the models useful, and I kind of just feel like the diversity of the outputs is not number one. It's much harder to work with and evaluate, but maybe it's not what's actually capturing most of the value. In fact, it's actively penalized, right? If you're super creative in RL, it's not good. Or if you're doing a lot of writing help from LLMs, I think it's probably bad because the models will give you all the same stuff. So they won't explore lots of different ways of answering a question. But I kind of feel like maybe this diversity is just not as big of a need. Maybe not as many applications need it, so the models don't have it. But then it's actually a problem at synthetic generation time, etc. So we're actually shooting ourselves in the foot by not allowing this entropy to maintain in the model. And I think possibly the labs should try harder.

Dwarkesh

然后我觉得你暗示这是一个非常根本的问题,不容易解决。那么,你的直觉是什么?

And then I think you hinted that it's a very fundamental problem. It won't be easy to solve. And yeah, what's your intuition for that?

Andrej

我其实不知道这是否非常根本。我也不知道我是否打算这么说。我确实认为,虽然我没做过这些实验,但你可能可以通过正则化让熵更高。这样你鼓励模型给出越来越多的解决方案。但你不希望它开始过度偏离训练数据。它会开始编造自己的语言,使用极其罕见的词汇。所以它会过度偏离分布。因此我认为控制分布很棘手。从这个意义上说,可能并不简单。

I don't actually know if it's super fundamental. I don't actually know if I intended to say that. I do think that I haven't done these experiments, but I do think that you could probably regularize the entropy to be higher. So you're encouraging the model to give you more and more solutions. But you don't want it to start deviating too much from the training data. It's going to start making up its own language, using words that are extremely rare. So it's going to drift too much from the distribution. So I think controlling the distribution is just tricky. It's probably not trivial in that sense.

认知核心的大小 On the size of the cognitive core

Dwarkesh

如果你必须猜一下,最优的智能核心最终应该有多少比特?我们放在车上的那个东西,它需要多大?

How many bits should the optimal core of intelligence end up being if you just had to make a guess? The thing we put on the van, how big does it have to be?

Andrej

这个领域的历史非常有趣,因为曾几何时,一切都非常「Scaling 至上」,比如「我们要做更大的模型,万亿参数模型」。但实际上,模型的大小是先上升,现在又下降了。它们的模型更小了。即便如此,我仍然认为它们记忆了太多东西。所以我之前有一个预测,我几乎觉得我们可以得到非常优秀的认知核心,哪怕只有十亿参数。

So it's really interesting in the history of the field because at one point everything was very scaling-pilled in terms of like 'oh we're going to make much bigger models, trillions of parameter models.' And actually, what the models have done in size is they've gone up and now they've actually come down. Their models are smaller. And even then, I actually think they memorized way too much. So I think I had a prediction a while back that I almost feel like we can get cognitive cores that are very good at even a billion parameters.

认知核心规模 On cognitive core size

Dwarkesh

它应该已经像这样了:如果你和一个十亿参数的模型对话,我认为在 20 年内你实际上可以进行非常有成效的对话。它会思考,而且更像人类。但如果你问它一些事实性问题,它可能需要查一下,但它知道自己不知道,可能会去查,并且会做所有合理的事情。实际上,你居然认为需要十亿参数,这让我很惊讶,因为我们已经有了十亿参数或几十亿参数的模型,它们非常智能。

It should already be like, if you talk to a billion parameter model, I think in 20 years you can actually have a very productive conversation. It thinks and it's a lot more like a human. But if you ask it some factual question, it might have to look it up, but it knows that it doesn't know and it might have to look it up and it will just do all the reasonable things. That's actually surprising that you think it will take a billion because we already have billion parameter models or a couple billion parameter models that are very intelligent.

Andrej

嗯,我们的一些模型有万亿参数,对吧?但它们记住了太多东西。

Well, some of our models are like a trillion parameters, right? But they remember so much stuff.

Dwarkesh

是的。但考虑到发展速度,我很惊讶在 10 年内,我们有 GPT-4o 只有 200 亿参数,却比原版万亿参数的 GPT-4 好得多。所以按照这个趋势,你居然认为 10 年后认知核心仍然是十亿参数,我其实很惊讶。我反而觉得它可能只有几千万甚至几百万参数。

Yeah. But I'm surprised that in 10 years given the pace, okay, we have GPT-4o which is 20B, that's way better than GPT-4 original which was a trillion plus parameters. So given that trend, I'm actually surprised you think in 10 years the cognitive core is still a billion parameters. I would be surprised if it's not like tens of millions or millions.

Andrej

不,因为我基本上认为训练数据才是问题所在。训练数据是互联网,而互联网质量非常糟糕。所以有巨大的提升空间,因为互联网太差了。当你实际查看前沿实验室的预训练数据集,随便看一个互联网文档,那完全是垃圾。它是一些股票代码。来自互联网各个角落的大量垃圾和废料。它不像你的《华尔街日报》文章,那种极其罕见。所以我几乎觉得,因为互联网如此糟糕,我们实际上必须构建非常大的模型来压缩所有这些。大部分压缩是记忆工作,而不是认知工作。但我们真正想要的是认知部分。实际上,删除记忆。那么我想说的是,我们需要智能模型来帮助我们改进预训练集,只保留认知成分,然后我认为你可以用更小的模型,因为数据集好得多。但可能不是直接在上面训练,而是从更好的模型中蒸馏出来。

No, because I basically think that the training data is the issue. The training data is the internet, which is really terrible. So there's a huge amount of gains to be made because the internet is terrible. When you're actually looking at a pre-training data set in a frontier lab and you look at a random internet document, it's total garbage. It's some stock ticker symbols. It's a huge amount of slop and garbage from all corners of the internet. It's not like your Wall Street Journal article, which is extremely rare. So I almost feel like because the internet is so terrible, we actually have to build really big models to compress all that. Most of that compression is memory work instead of cognitive work. But what we really want is the cognitive part. Actually, delete the memory. And then I guess what I'm saying is we need intelligent models to help us refine even the pre-training set to just narrow it down to the cognitive components, and then I think you get away with a much smaller model because it's a much better data set. But probably it's not trained directly on it; it's probably distilled from a much better model still.

Dwarkesh

但为什么蒸馏后的版本仍然是十亿参数?这是我想知道的。

But why is the distilled version still a billion? That's the thing I'm curious about.

Andrej

我只是觉得蒸馏效果非常好。所以几乎每个小模型,如果你有一个小模型,它几乎肯定是蒸馏出来的。为什么要训练在……

I just feel like distillation works extremely well. So almost every small model, if you have a small model, it's almost certainly distilled. Why would you train on...

Dwarkesh

对,不,不。但为什么蒸馏在 10 年内不能降到十亿以下?

Right, no, no. But why is distillation not getting below 1 billion in 10 years?

Andrej

哦,你认为它应该小于一百万?

Oh, you think it should be smaller than a million?

Dwarkesh

我是说,拜托,对吧?我不知道。在某个点上,至少需要十亿个旋钮才能做有趣的事情。你觉得它应该更小。

I mean, come on, right? I don't know. At some point, it should take at least a billion knobs to do something interesting. You're thinking it should be even smaller.

Andrej

是的。我的意思是,看看过去几年的趋势,找到容易摘的果实,从万亿参数模型在两年内缩小两个数量级,性能反而更好。这让我觉得智能的核心可能还要小得多。用费曼的话说,底下还有很大空间。

Yeah. I mean, just like if you look at the trend over the last few years, just finding low-hanging fruit and going from trillion-plus models to literally two orders of magnitude smaller in a matter of two years and having better performance. It makes me think the sort of core of intelligence might be even way smaller. Plenty of room at the bottom, to paraphrase Feynman.

Andrej

我的意思是,我几乎觉得我谈论十亿参数的认知核心已经算是逆向思维了,而你更胜一筹。我想也许我们可以再小一点。我仍然认为应该有足够的……

I mean, I almost feel like I'm already contrarian by talking about a billion parameter cognitive core, and you're outdoing me. I think maybe we could get a little bit smaller. I still think that there should be enough...

Dwarkesh

是的,也许可以更小。我确实认为,实际上,你希望模型有一些知识。你不希望它什么都去查。因为那样你就无法在脑子里思考;你总是在查太多东西。所以我确实认为它需要一些基础课程,一些知识必须存在。但它不需要深奥的知识。

Yeah, maybe it can be smaller. I do think that practically speaking, you want the model to have some knowledge. You don't want it to be looking up everything. Because then you can't think in your head; you're looking up way too much stuff all the time. So I do think it needs some basic curriculum, some knowledge needs to be there. But it doesn't need esoteric knowledge.

Andrej

是的。

Yeah.

未来扩展趋势 On future scaling trends

Dwarkesh

所以我们讨论的是认知核心可能是什么。还有一个单独的问题:随着时间的推移,前沿模型的实际规模会怎样?我很好奇你的预测。我们之前看到规模增长到可能 4.5,现在看到规模在缩小或趋于平稳。可能有很多原因,但你对未来有什么预测?最大的模型会更大吗?会更小吗?会保持不变吗?

So we're discussing what plausibly could be the cognitive core. There's a separate question: what will actually be the size of frontier models over time? And I'm curious to have a prediction. So we had increasing scale up to maybe 4.5, and now we're seeing decreasing or plateauing scale. There are many reasons that could be going on, but do you have a prediction about going forward? Will the biggest models be bigger? Will they be smaller? Will they be the same?

Andrej

是的,我没有特别强烈的预测。我确实认为实验室只是务实。他们有算力预算和成本预算。结果发现,预训练并不是你想投入大部分算力或成本的地方。所以这就是模型变小的原因:预训练阶段更小,但他们通过强化学习以及中间训练和后续的各种环节来弥补。所以他们在所有阶段以及如何获得最大性价比方面都很务实。所以我认为预测这个趋势相当困难。我仍然期望有很多容易摘的果实,这是我的基本预期。所以我的分布非常宽。

Yeah, I don't know that I have a super strong prediction. I do think that the labs are just being practical. They have a flops budget and a cost budget. And it just turns out that pre-training is not where you want to put most of your flops or your cost. So that's why the models have gotten smaller: the pre-training stage is smaller, but they make it up in reinforcement learning and all this kind of stuff mid-training and all that follows. So they're just being practical in terms of all the stages and how you get the most bang for the buck. So I guess forecasting that trend is quite hard. I do still expect that there's so much low-hanging fruit, that's my basic expectation. And so I have a very wide distribution here.

Dwarkesh

你期望这些容易摘的果实与过去两到五年发生的事情类似吗?比如我看 NanoChat 和 NanoGPT 以及你做的架构调整,基本上是你期望继续发生的事情,还是你不期待任何巨大飞跃?

Do you expect the low-hanging fruit to be similar in kind to the kinds of things that have been happening over the last two to five years? Like if I look at NanoChat versus NanoGPT and the architectural tweaks you made, is that basically the flavor of things you expect to keep happening, or are you not expecting any giant leaps?

Andrej

是的,我期望数据集会变得好得多。因为当你查看平均数据集时,它们极其糟糕,糟糕到我甚至不知道任何东西是怎么工作的,说实话。看看训练集中的平均样本:事实错误、错误、无意义的东西。不知何故,当你大规模进行时,噪声被冲走,留下一些信号。所以数据集会大幅改进。一切都会变得更好。

Yeah, I expect the data sets to get much, much better. Because when you look at the average data sets, they're extremely terrible, so bad that I don't even know how anything works, to be honest. Look at the average example in the training set: factual mistakes, errors, nonsensical things. Somehow when you do it at scale, the noise washes away and you're left with some of the signal. So data sets will improve a ton. It's just everything gets better.

硬件与软件改进 On hardware and software improvements

Andrej

所以我们的硬件,所有用于运行硬件并最大化硬件性能的内核。NVIDIA 正在慢慢调整实际的硬件本身,比如张量核心等等。所有这些都需要发生,并且会继续发生。所有内核都会变得更好,最大限度地利用芯片。所有算法可能会在优化架构上改进,以及我们训练所用的所有建模组件。所以我确实预期,一切都不会有单一主导因素,每项改进都带来大约 20% 的提升。

So our hardware, all the kernels for running the hardware and maximizing what you get with the hardware. So NVIDIA is slowly tuning the actual hardware itself, tensor cores and so on. All that needs to happen and will continue to happen. All the kernels will get better and utilize the chip to the max extent. All the algorithms will probably improve over optimization architecture and just all the modeling components of how everything is done and what the algorithms are that we're even training with. So I do kind of expect just a very everything nothing dominates everything plus 20%.

Dwarkesh

对,有意思。这大致是我看到的。

Right. Interesting. This is like roughly what I've seen.

Andrej

好的。这是我的总经理 Max。

Okay. This is my general manager Max.

Dwarkesh

很高兴每天都来这里。你入职大约 6 个月了。但当时我 onboarding 你的时候,我在法国,所以我们几乎没机会交谈,你基本上只给了我一个登录账号。

Good to be here every day. And you have been here since you were onboarded about 6 months ago. But when I onboarded you, I was in France and so we basically didn't get the chance to talk at all almost, and you basically just gave me one login.

Andrej

我给了你访问我的 Mercury 平台的权限,那是我当时用来运营播客的银行平台。我登录 Mercury 以为这只是众多步骤中的第一步,但我意识到你就是用这个来运营整个业务的,甚至包括我们的很多编辑都是国际承包商。你当时已经想好了如何设置这些定期付款来建立基本的工资单。

I gave you access to my Mercury platform, which is the banking platform that I was using at the time to run the podcast. And so I logged into Mercury assuming that that would just be the first of many steps, but I realized that was how you were running the entire business, even down to a lot of our editors are international contractors. And so you had just figured out how to set up these recurring payments to set up basic payroll.

Dwarkesh

我的意思是,Mercury 让我之前做的所有这些事情变得如此无缝,以至于直到你指出我才意识到,这不是设置工资单或发票或其他事情的常规方式。我很惊讶,但我想,到目前为止它都有效,所以也许我会信任它。现在我想不出还能用别的什么。

I mean, Mercury made the experience of all of these things I was doing before so seamless that it didn't even occur to me until you pointed it out that this is not the natural way to set up payroll or invoicing or any of these other things. I was surprised, but I was like, it's worked so far, so maybe I'll trust it. And then now I can't think of doing anything else.

Andrej

好了,你们都听到了。访问 mercury.com 在线申请,几分钟就好。酷。谢谢,Max。

All right, you heard him. Visit mercury.com to apply online in minutes. Cool. Thanks, Max.

Dwarkesh

谢谢你邀请我。

Thanks for having me.

Andrej

老兄,你在这方面很厉害。我很紧张,但谢谢你。

Dude, you're great at this. I'm so nervous, but thank you.

Dwarkesh

Mercury 是一家金融科技公司,不是银行。银行服务由 Choice Financial Group、column NA 和 Evolve Bank and Trust(FDIC 成员)提供。

Mercury is a financial technology company, not a bank. Banking services provided through Choice Financial Group, column NA, and Evolve Bank and Trust members FDIC.

衡量 AGI 进展 On measuring progress towards AGI

Dwarkesh

人们提出了不同的方法来衡量我们离完全 AGI 有多远,因为如果你能画出一条线,那么你就能看到这条线与 AGI 的交点以及它在 X 轴上的位置。所以有人提出,比如教育水平,我们有一个高中生,然后他们通过强化学习上了大学,然后会获得博士学位。我不喜欢那个。或者他们会提出任务时长。也许他们可以自主完成需要一分钟的任务,然后可以自主完成需要一小时的任务,人类需要一小时,人类需要一周等等。你认为相关的 y 轴是什么?我们应该如何思考 AI 的进展?

People have proposed different ways of charting how much progress we've made towards full AGI because if you can come up with some line, then you can see where that line intersects with AGI and where that would happen on the X-axis. And so people have proposed, oh, it's like the education level, like we had a high schooler and then they went to college with RL and they're going to get a PhD. I don't like that one. Or they'll propose horizon length. So maybe they can do tasks that take a minute, they can do those autonomously, then they can autonomously do tasks that take an hour, a human an hour, a human a week, etc. How do you think about what is the relevant y-axis here? How should we think about how AI is making progress?

Andrej

我想我有两个答案。第一,我几乎想完全拒绝这个问题,因为我再次将其视为计算的延伸。我们有没有讨论过如何绘制计算的进展图,或者自 1970 年代以来如何绘制计算的进展?x 轴是什么?所以我觉得整个问题从这个角度来看有点好笑。但我会说,当人们谈论 AI 和最初的 AGI,以及 OpenAI 成立时我们如何谈论它,AGI 是一个你可以使用的系统,它可以完成任何具有经济价值的任务,任何具有经济价值的任务,达到人类水平或更好。所以那是定义,我当时对此很满意,我觉得我一直坚持那个定义,然后人们编造了各种其他定义,但我喜欢那个定义。现在,人们一直做的第一个让步是,他们只是去掉了所有物理的东西,因为我们只谈论数字知识工作。我觉得与最初的定义相比,这是一个相当大的让步,最初的定义是任何人类能做的任务。我可以举东西等等。显然 AI 做不到。所以,好吧,但我们接受。

So I guess I have two answers to that. Number one, I'm almost tempted to reject the question entirely because again I see this as an extension of computing. Have we talked about how to chart progress in computing or how do you chart progress in computing since the 1970s? What is the x-axis? So I kind of feel like the whole question is kind of funny from that perspective. But I will say, I guess when people talk about AI and the original AGI and how we spoke about it when OpenAI started, AGI was a system you can go to that can do any task that is economically valuable, any economically valuable task at human performance or better. So that was the definition and I was pretty happy with that at the time and I kind of feel like I've stuck to that definition forever and then people have made up all kinds of other definitions but I like that definition. Now, the first concession that people make all the time is they just take out all the physical stuff because we're just talking about digital knowledge work. I feel like that's a pretty major concession compared to the original definition which was like any task a human can do. I can lift things, etc. Like AI can't do that obviously. So, okay, but we'll take it.

Dwarkesh

通过说「哦,只有知识工作」,我们从经济中拿走了多少份额?我实际上不知道具体数字。我觉得如果必须猜的话,大约是 10% 到 20%。只有知识工作,比如有人可以在家工作并执行类似任务。我仍然认为这是一个非常大的市场。比如经济规模有多大,10% 到 20% 是多少?我们仍然在谈论美国几万亿美元的市场份额或工作。所以仍然是一个巨大的类别。但回到定义,我想我会寻找的是这个定义在多大程度上成立?所以有没有工作或很多任务?如果我们把任务看作不是工作而是任务,这有点困难,因为问题在于社会将根据构成工作的任务与哪些是可自动化的来重构。但今天,哪些工作可以被 AI 取代?最近的一个好例子是 Jeff Hinton 的预测,即放射科医生将不再是工作,但这在很多方面被证明是非常错误的。放射科医生仍然存在且健康,并且还在增长,尽管计算机视觉在识别他们需要识别的所有不同事物方面非常出色。这只是个混乱复杂的工作,有很多方面,需要与患者打交道等等。所以我想我实际上不知道按照那个定义,AI 是否已经取得了很大的进展。但也许我会寻找的一些工作具有一些特征,我认为这些特征使它们更早而不是更晚地适合自动化。例如,呼叫中心员工经常被提及,我认为这是正确的。因为呼叫中心员工在当今可自动化方面具有许多简化特性。他们的工作相当简单。

What fraction of the economy are we taking away by saying, 'Oh, only knowledge work.' I don't actually know the numbers. I feel like it's about 10 to 20% if I had to guess. Is only knowledge work, like someone could work from home and perform tasks something like that. I still think it's a really large market. Like what is the size of the economy and what is 10-20%? We're still talking about a few trillion dollars of even in the US of market share almost or like work. So still a very massive bucket. But going back to the definition, I guess what I would be looking for is to what extent is that definition true? So are there jobs or lots of tasks? If we think of tasks as not jobs but tasks, it's kind of difficult because the problem is society will refactor based on the tasks that make up jobs compared to what's automatable or not. But today, what jobs are replaceable by AI? A good example recently was Jeff Hinton's prediction that radiologists would not be a job anymore and this turned out to be very wrong in a bunch of ways. Radiologists are alive and well and growing even though computer vision is really really good at recognizing all the different things that they have to recognize. It's just a messy complicated job with a lot of surfaces and dealing with patients and all this kind of stuff in the context of it. So I guess I don't actually know that by that definition AI has made a huge amount of dent yet. But some of the jobs maybe that I would be looking for have some features that I think make it very amenable to automation earlier than later. As an example, call center employees often come up and I think rightly so. Because call center employees have a number of simplifying properties with respect to what's automatable today. Their jobs are pretty simple.

自动化客服工作 On automating call center work

Dwarkesh

这是一系列任务,每个任务看起来都很相似,比如你接听一个人的电话,互动 10 分钟或更长时间,根据我的经验,可能更长。你完成某个方案中的任务,修改一些数据库条目之类的。所以你一遍又一遍地重复同样的事情,这就是你的工作。所以基本上,你需要考虑任务的时间跨度,即完成一个任务需要多长时间。然后你还需要去除上下文,比如你不处理公司不同部门或其他客户的事务。只有数据库、你和你要服务的人。所以它更封闭、更易理解,而且是纯数字化的。所以我会寻找这些特征。但即便如此,我实际上还没有考虑完全自动化。我在寻找一个自主性滑块,我几乎预期我们不会立即取代人类。我们会用 AI 替代,它们处理 80%的工作量,将 20%的工作量委托给人类,而人类则监督五个人工智能团队,处理那些更机械的呼叫中心工作。所以我会寻找新的界面或新的公司,它们提供某种层,允许你管理这些还不完美的 AI。

It's a sequence of tasks and every task looks similar, like you take a phone call with a person, it's 10 minutes of interaction or whatever it is, probably a bit longer in my experience, a lot longer. And you complete some task in some scheme and you change some database entries around or something like that. So you keep repeating something over and over again and that's your job. So basically you do want to bring in the task horizon, how long it takes to perform a task. And then you want to also remove context, like you're not dealing with different parts of services of companies or other customers. It's just the database, you, and a person you're serving. And so it's more closed. It's more understandable and it's purely digital. So I would be looking for those things. But even there, I'm not actually looking at full automation yet. I'm looking for an autonomy slider, and I almost expect that we are not going to instantly replace people. We're going to be swapping in AIs that do 80% of the volume. They delegate 20% of the volume to humans, and humans are supervising teams of five AIs doing the call center work that's more rote. So I would be looking for new interfaces or new companies that provide some kind of a layer that allows you to manage some of these AIs that are not yet perfect.

Andrej

是的。

Yeah.

Dwarkesh

然后我预期在整个经济中,很多工作比呼叫中心员工要难得多。我想知道放射科医生的情况,我完全是猜测。我不知道放射科医生的实际工作流程,但一个可能适用的类比是,当我们最初被淘汰时,会有人坐在前排,你必须让他们在那里,以确保如果出了大问题,他们可以监控。我认为即使今天,人们仍然在监控以确保一切顺利。刚刚部署的 Robo Taxi 实际上仍然有一个人在里面。我们可能处于类似的情况:如果你自动化了 99%的工作,人类必须做的最后 1%就非常有价值,因为它成为所有其他事情的瓶颈。如果放射科医生也是如此,坐在 Uber 或 Waymo 前排的人必须经过多年专门训练才能提供那最后的 1%,那么他们的工资应该大幅上涨,因为他们是广泛部署的瓶颈。所以我认为放射科医生的工资因类似原因上涨了。如果你是最后的瓶颈,你就不可替代,而 Waymo 司机可能可以被其他东西替代。所以你可能会看到工资这样变化:直到达到 90%,然后突然这样,当最后 1%消失时。

And then I would expect that across the economy, a lot of jobs are a lot harder than call center employee. I wonder with radiologists, I'm totally speculating. I have no idea what the actual workflow of a radiologist involves, but one analogy that might be applicable is when we were first being ruled out, there would be a person sitting in the front seat, and you just had to have them there to make sure that if something went really wrong, they're there to monitor. And I think even today, people are still watching to make sure things are going well. Robo Taxi, which was just deployed, actually still has a person inside it. And we could be in a similar situation where if you automate 99% of a job, that last 1% the human has to do is incredibly valuable because it's bottlenecking everything else. And if it was the case with radiologists, where the person sitting in the front of the Uber or the front of the Waymo has to be specially trained for years in order to be able to provide the last 1%, their wages should go up tremendously because they're the one thing bottlenecking wide deployment. So radiologists, I think their wages have gone up for similar reasons. If you're the last bottleneck, you're not fungible, while a Waymo driver might be fungible with other things. So you might see this thing where your wages go like this until you get to 90%, and then just like that, and when the last 1% is gone.

Andrej

我明白了。

I see.

Dwarkesh

我想知道我们是否在放射科或呼叫中心员工的工资等方面看到类似的情况。

And I wonder if we see similar things with radiology or salaries of call center workers or anything like that.

Andrej

是的,我认为这是一个有趣的问题。我不认为我们目前在放射科看到这种情况,而且我没有深入了解,但我认为放射科基本上不是一个好例子。我不知道为什么杰夫·辛顿选择了放射科,因为我认为这是一个极其混乱、复杂的职业。所以我更感兴趣的是今天呼叫中心员工的情况,例如,因为我预计很多机械性工作现在都可以自动化。我没有第一手资料,但也许我会关注呼叫中心员工的趋势。我可能还预期他们正在引入 AI,但我会再等一两年,因为我可能预期他们会撤回并重新雇佣一些人。我认为已经有证据表明,在采用 AI 的公司中这种情况已经发生,这让我非常惊讶,我也觉得这真的很令人惊讶。

Yeah, I think that's an interesting question. I don't think we're currently seeing that with radiology, and I don't have a deep understanding, but I think radiology is not a good example basically. I don't know why Jeff Hinton picked on radiology, because I think it's an extremely messy, complicated profession. So I would be a lot more interested in what's happening with call center employees today, for example, because I would expect a lot of the rote stuff to be automatable today. And I don't have first-level access to it, but maybe I would be looking for trends of what's happening with the call center employees. Maybe some of the things I would also expect is maybe they are swapping in AI, but then I would still wait for a year or two because I would potentially expect them to pull back and actually rehire some of the people. I think there's been evidence that that's already been happening in companies that have been adopting AI, which I think is quite surprising, and I also find it really surprising.

AGI 与编程作为首个领域 On AGI and coding as the first domain

Dwarkesh

好的。AGI,对吧?一个应该能做所有事情的东西,好吧,我们排除体力工作。所以我认为我们应该能做所有知识工作。你天真地预期这种回归的方式是:你从顾问的工作中取出一个小任务,从会计的工作中取出一个小任务,然后就这样在所有知识工作中进行。但相反,如果我们相信我们正走在当前范式下的 AI 道路上,进展完全不是这样。至少,顾问和会计之类的工作似乎并没有得到巨大的生产力提升。更像是程序员的工作方式越来越受到冲击。如果你看这些公司的收入,撇开普通的聊天收入(我认为类似于谷歌之类的),只看 API 收入,它主要由编程主导,对吧?所以这个通用的、所谓的应该能做任何知识工作的东西,却 overwhelmingly 只做编程,这是一种令人惊讶的 AGI 部署方式。

Okay. AGI, right? A thing which should do everything, and okay we'll take out physical work. So I think we should be able to do all knowledge work. And what you would have naively anticipated is that the way this regression would happen is like you take a little task that a consultant is doing, you take that out of the bucket. You take a little task that an accountant is doing, you take that out of the bucket. And then you're just doing this across all knowledge work. But instead, if we do believe we're on the path of AI with the current paradigm, the progression is very much not like that. At least it just does not seem like consultants and accountants or whatever are getting huge productive improvement. It's very much like programmers are getting more and more chills of the way of their work. If you look at the revenues of these companies, discounting just normal chat revenue which I think is similar to Google or something, just looking at API revenues, it's dominated by coding, right? So this thing which is general, quote unquote, should be able to do any knowledge work, is just overwhelmingly doing only coding, and it's a surprising way that you would expect the AGI to be deployed.

Andrej

所以我认为这里有一个有趣的点,因为我确实相信编程是这些 LLM 和智能体的完美第一件事,这是因为编程从根本上一直是围绕文本的。它是计算机终端和文本,一切都基于文本,而 LLM 在互联网上训练的方式,它们喜欢文本。所以它们是完美的文本处理器,而且有大量的数据,这简直是完美的契合。此外,我们还有很多预先构建的基础设施来处理代码和文本。例如,我们有 Visual Studio Code 或你最喜欢的 IDE 显示代码,智能体可以接入其中。例如,如果智能体有一个 diff,它做了一些更改,我们突然就有了所有代码,通过 diff 显示代码库的所有差异。所以我们几乎已经预先构建了很多代码的基础设施。

So I think there's an interesting point here, because I do believe coding is the perfect first thing for these LLMs and agents, and that's because coding has always fundamentally worked around text. It's computer terminals and text, and everything is based around text, and LLMs, the way they're trained on the internet, love text. And so they're perfect text processors, and there's all this data out there, and it's just a perfect fit. And also we have a lot of infrastructure pre-built for handling code and text. So for example, we have Visual Studio Code or your favorite IDE showing you code, and an agent can plug into that. So for example, if an agent has a diff where it made some change, we suddenly have all this code already that shows all the differences to a codebase using a diff. So it's almost like we've pre-built a lot of the infrastructure for code.

LLM 的非编程应用 On non-coding applications of LLMs

Dwarkesh

现在对比一下那些完全不享受这种优势的东西。举个例子,有人试图构建自动化工具,不是为了编程,而是为了制作幻灯片。我看到一家公司做幻灯片,这要难得多,原因在于幻灯片不是文本。

Now contrast that with some of the things that don't enjoy that at all. So as an example, there are people trying to build automation not for coding but for example for slides. I saw a company doing slides, that's much harder, and the reason it's much harder is because slides are not text.

Andrej

是的。

Yeah.

Dwarkesh

幻灯片是图形,有空间排列,还有视觉元素。幻灯片没有这种预先构建的基础设施。例如,如果一个智能体要修改你的幻灯片,它如何向你展示差异?你怎么看到差异?没有任何东西能显示幻灯片的差异。

Slides are little graphics, arranged spatially, and there's a visual component. Slides don't have this pre-built infrastructure. For example, if an agent is to make a change to your slides, how does it show you the diff? How do you see the diff? There's nothing that shows diffs for slides.

Andrej

嗯。

Mhm.

Dwarkesh

所以必须有人去构建它。这些东西中的一些并不适合当前的人工智能,它们本质上是文本处理器,而代码却出奇地适合。

So someone has to build it. Some of these things are not amenable to AIs as they are, which are text processors, and code surprisingly is.

Andrej

我其实不确定仅凭这一点就能解释,因为我个人曾尝试让大语言模型在纯语言输入输出的领域发挥作用,比如重写转录稿、根据转录稿生成片段等。你可能会说,我没有穷尽所有可能的方法。我在上下文中放了很多好的例子,但也许我应该做某种微调之类的。所以,我们的共同朋友 Andy Matushak 告诉我,他实际上尝试了无数种方法,试图让模型擅长编写间隔重复提示。这又是一个纯语言输入输出的任务,本应是大语言模型的核心能力。他显然尝试了上下文学习,用了几个简短例子。他还尝试了监督微调、检索等方法,但就是无法让模型生成令人满意的卡片。所以我发现,即使在语言输出领域,除了编程之外,要从这些模型中获取大量经济价值实际上非常困难。我不知道原因何在。

I actually I'm not sure if that alone explains it, because I personally have tried to get LLMs to be useful in domains which are just pure language in, language out. Like rewriting transcripts, coming up with clips based on transcripts, etc. And you might say, well, I didn't do every single possible thing I could do. I put a bunch of good examples in context, but maybe I should have done some kind of fine-tuning, whatever. So, our mutual friend Andy Matushak told me that he actually tried 50 billion things to try to get models to be good at writing space repetition prompts. Again, very much language in, language out task. The kind of thing that should be dead center in the repertoire of these LLMs. And he tried in-context learning obviously with a few short examples. He tried I think he told me a bunch of things like supervised fine-tuning and retrieval, whatever, and he just could not get them to make cards to a satisfaction. So I find it striking that even in language out domains, it's actually very hard to get a lot of economic value out of these models separate from coding. And I don't know what explains it.

Andrej

是的,我觉得有道理。我的意思是,我并不是说任何文本任务都很简单。我确实认为代码结构性强。文本可能更花哨,熵更高,我不知道还能怎么说。而且,代码很难,所以人们即使从简单的知识中也能感受到大语言模型的赋能。我基本上没有一个很好的解释。显然,文本让事情容易得多,也许这就是我这么说的原因,但这并不意味着所有文本任务都很简单。

Yeah, I think that makes sense. I mean, I'm not saying that anything text is trivial. I do think that code is pretty structured. Text is maybe a lot more flowery, and there's a lot more entropy in text, I would say. I don't know how else to put it. And also, code is hard, so people feel quite empowered by LLMs even from simple knowledge. I basically don't actually know that I have a very good explanation. Obviously, text makes it much easier, maybe that's why I put it, but it doesn't mean that all text is trivial.

超级智能 On superintelligence

Dwarkesh

你怎么看待超级智能?你期望它感觉上与普通人类或人类公司有本质区别吗?

How do you think about superintelligence? Do you expect it to feel qualitatively different from normal humans or human companies?

Andrej

我想我把它看作社会中自动化的一个进程,对吧?再次外推计算趋势。我只是觉得很多事情会逐渐自动化,而超级智能将是这种趋势的外推。所以我确实认为随着时间的推移,我们会看到越来越多的自主实体从事大量数字工作,最终甚至包括体力工作,可能稍晚一些。但基本上,我把它看作自动化,大致来说。

I guess I think I see it as a progression of automation in society, right? Again, extrapolating the trend of computing. I just feel like there will be a gradual automation of a lot of things, and superintelligence will be sort of the extrapolation of that. So I do think we expect more and more autonomous entities over time that are doing a lot of the digital work, and then eventually even the physical work, probably some amount of time later. But basically I see it as just automation, roughly speaking.

Dwarkesh

我想自动化包括人类已经能做的事情,而超级智能则包括人类……嗯,但人类做的一些事情是发明新事物,如果说得通的话,我会把它归入自动化。

I guess automation includes the things humans can already do, and superintelligence things humans... well, but some of the things that people do is invent new things, which I would just put into the automation if that makes sense.

Andrej

是的。

Yeah.

Dwarkesh

但也许更具体、更定性地说。你是否期望某种感觉,比如,因为这个东西要么思考速度极快,要么有大量副本,要么副本可以合并回自身,要么更聪明,人工智能可能拥有的任何优势。它会让这些人工智能存在的文明在本质上不同于人类文明。

But I guess maybe less abstractly and more qualitatively. Do you expect something to feel like, okay, because this thing can either think so fast, or has so many copies, or the copies can merge back into themselves, or is much smarter, any number of advantages an AI might have. It will make the civilization in which these AIs exist feel qualitatively different from human civilization.

Andrej

我认为会的。我的意思是,它本质上是自动化,但会非常陌生。我确实认为它会看起来很奇怪,因为,就像你提到的,我们可以在计算机集群上运行所有这些,速度快得多,等等。是的,我的意思是,也许当世界变成那样时,我开始感到紧张的一些情景是这种逐渐失去对正在发生的事情的控制和理解。我认为这实际上是最可能的结果:将会逐渐失去理解,我们将逐渐把所有这些层层叠加,理解它的人会越来越少,就会出现一种逐渐失去对正在发生的事情的控制和理解的情景。对我来说,这似乎是这一切最可能的结果。

I think it will. I mean, it is fundamentally automation, but it will be extremely foreign. I do think it will look really strange because, like you mentioned, we can run all of this on a computer cluster, much faster, and all this thing. Yeah, I mean, maybe some of the scenarios that I start to get nervous about when the world looks like that is this kind of gradual loss of control and understanding of what's happening. And I think that's actually the most likely outcome probably: that there will be a gradual loss of understanding, and we'll gradually layer all this stuff everywhere, and there will be fewer and fewer people who understand it, and there will be a sort of scenario of a gradual loss of control and understanding of what's happening. That to me seems the most likely outcome of how all this stuff will go down.

Dwarkesh

让我深入探讨一下。我不清楚失去控制和失去理解是同一回事。比如,台积电、英特尔等公司的董事会,他们只是德高望重的八十岁老人。他们理解很少,也许实际上也没有控制权,或者更好的例子是美国总统。总统拥有很大的权力。我不是在评价现任者,但也许我是。但实际的理解水平与控制水平非常不同。

Let me probe on that a bit. It's not clear to me that loss of control and loss of understanding are the same things. A board of directors at, like, whatever TSMC, Intel, name a random company. They're just like prestigious 80-year-olds. They have very little understanding, and maybe they don't practically actually have control, or actually maybe a better example is the president of the United States. The president has a lot of power. I'm not trying to make a good statement about the current occupant, but maybe I am. But the actual level of understanding is very different from the level of control.

Andrej

是的,这很公平。这是一个很好的反驳。我想我预期两者都会失去。

Yeah, I think that's fair. That's a good pushback. I think I guess I expect loss of both.

Dwarkesh

为什么?我的意思是失去理解是显而易见的,但为什么会失去控制?

How come? I mean loss of understanding is obvious, but why loss of control?

Andrej

所以,我们真的进入了未知领域,我不知道这会是什么样子,但如果我要写科幻小说,它们会沿着这样的思路:甚至不是一个单一的实体之类的东西接管一切。而是多个相互竞争的实体逐渐变得越来越自主,其中一些失控,其他的则与之对抗,等等。

So, we're really far into territory of I don't know what this looks like, but if I was to write sci-fi novels, they would look along the lines of not even a single entity or something like that that just takes over everything. But actually like multiple competing entities that gradually become more and more autonomous, and some of them go rogue and the others fight them off and all this kind of stuff.

失控风险 On loss of control

Andrej

这就像一锅我们委托出去的完全自主的活动。我有点觉得它会带有那种味道。导致失控的并不是它们比我们更聪明,而是它们彼此竞争,以及这种竞争所产生的结果导致了失控。

It's like this hot pot of completely autonomous activity that we've delegated to. I kind of feel like it would have that flavor. It is not the fact that they are smarter than us that is resulting in the loss of control. It's the fact that they are competing with each other and whatever arises out of that competition that leads to the loss of control.

Dwarkesh

我的意思是,我基本上预计会有很多这样的东西。它们会成为人们的工具,而一部分人就像是代表人们行事。也许那些人是掌控者,但就我们想要的结果而言,对整个社会来说可能是一种失控。你有代表个人行事的实体,但它们仍然大致被视为失控的。

I mean, I basically expect there to be a lot of these things. They will be tools to people, and some of the population is like they're acting on behalf of people. Maybe those people are in control, but maybe it's a loss of control overall for society in the sense of outcomes we want. You have entities acting on behalf of individuals that are still kind of roughly seen as out of control.

Andrej

是的。是的。

Yeah. Yeah.

AI 作为编译器 vs 替代品 On AI as compiler vs replacement

Dwarkesh

这是我早该问的问题。我们之前谈到,目前在做 AI 工程或 AI 研究时,这些模型更像是编译器,而不是替代品。

This is a question I should have asked earlier. So we were talking about how currently it feels like when you're doing AI engineering or AI research, these models are more like in the category of compiler rather than in the category of a replacement.

Andrej

是的。在某个时候,如果你有了 AGI,它应该能够做你做的事情。

Yeah. At some point, if you have AGI, it should be able to do what you do.

Dwarkesh

你觉得有一百万个你的副本并行工作会导致 AI 进展的巨大加速吗?基本上,如果那真的发生了,你预计会看到智能爆炸吗?

And do you feel like having a million copies of you in parallel results in some huge speed up of AI progress? Basically, if that does happen, would you expect to see an intelligence explosion?

Andrej

我想我的意思是,我确实这么认为,但这只是常态,因为我们已经处于智能爆炸中几十年了。当你观察 GDP 时,它基本上是一条指数曲线,是工业许多方面的加权和。一切都在逐渐自动化,已经持续了几百年。工业革命是物理组件和工具制造的自动化。编译器是早期的软件自动化。所以我觉得我们已经在递归地自我改进和爆炸很长时间了。也许另一种看法是,如果不看生物力学,地球是一个相当无聊的地方,从太空看非常相似,然后我们正处于这个鞭炮事件的中间,但我们是在慢动作中看到它的。

I guess what I mean is I do, but it's business as usual because we're already in an intelligence explosion and have been for decades. When you look at GDP, it's basically the GDP curve that is an exponential weighted sum over so many aspects of the industry. Everything is gradually being automated, has been for hundreds of years. The industrial revolution is automation of physical components and tool building. Compilers are early software automation. So I feel like we've been recursively self-improving and exploding for a long time. Maybe another way to see it is Earth was a pretty boring place if you don't look at biomechanics, and looked very similar from space, and then we're in the middle of this firecracker event, but we're seeing it in slow motion.

Dwarkesh

对。

Right.

Andrej

但我们是在慢动作中看到它的。我确实觉得这已经发生了很长时间,而且我不认为 AI 相对于已经发生很长时间的事情是一种独特的技术。

But we're seeing it in slow motion. I definitely feel like this has already happened for a very long time, and I don't see AI as a distinct technology with respect to what has already been happening for a long time.

Dwarkesh

你认为它与这种超指数趋势是连续的吗?

Is there you think it's continuous with this hyper exponential trend?

Andrej

这就是为什么这对我来说非常有趣。我一度试图在 GDP 中找到 AI。我以为 GDP 应该上升,但后来我看了其他我认为非常变革性的技术,比如计算机或手机。你在 GDP 中找不到它们。GDP 是同样的指数。例如,早期的 iPhone 没有应用商店,也没有现代 iPhone 的许多花哨功能。所以即使我们认为 2008 年 iPhone 问世是一次重大的地震式变化,但实际上并非如此。一切都如此分散,如此缓慢地扩散,以至于最终都被平均到同一个指数中。计算机也是完全一样的情况。你在 GDP 中找不到它们。那并没有发生,因为这是一个如此缓慢的进程。对于 AI,我们将看到完全相同的事情。它只是更多的自动化。它让我们能够编写以前无法编写的各种程序。但 AI 仍然从根本上是一个程序,一种新型的计算机和计算系统,但它有所有这些问题。它会随着时间的推移而扩散,并仍然累加到同一个指数上,我们仍然会有一个变得极其垂直的指数,生活在那种环境中会非常陌生。

And that's why this was very interesting to me. I was trying to find AI in the GDP for a while. I thought that GDP should go up, but then I looked at some of the other technologies that I thought were very transformative, like computers or mobile phones. You can't find them in GDP. GDP is the same exponential. For example, the early iPhone didn't have the app store and didn't have a lot of the bells and whistles that the modern iPhone has. So even though we think of 2008 when iPhone came out as some major seismic change, it's actually not. Everything is so spread out and so slowly diffuses that everything ends up being averaged up into the same exponential. It's the exact same thing with computers. You can't find them in GDP. That's not what happened because it's such a slow progression. With AI, we're going to see the exact same thing. It's just more automation. It allows us to write different kinds of programs that we couldn't write before. But AI is still fundamentally a program, a new kind of computer and computing system, but it has all these problems. It's going to diffuse over time and still add up to the same exponential, and we're still going to have an exponential that's going to get extremely vertical and it's going to be very foreign to live in that kind of an environment.

Dwarkesh

你是说,如果你看工业革命之前到现在的趋势,你有一个超指数,从 0%增长到一万年前的 0.02%增长,然后现在我们处于 2%增长。所以那是一个超指数。你说如果你把 AI 画在上面,那么 AI 会带你到 20%增长或 200%增长?或者你可能是说,如果你看过去 300 年,你看到一项又一项技术,但增长率完全相同,是 2%。所以你是说增长率会改变吗?

Are you saying that if you look at the trend before the industrial revolution to currently, you have a hyper exponential where you go from 0% growth to 0.02% growth 10,000 years ago, and then currently we're at 2% growth. So that's a hyper exponential. And you're saying if you're charting AI on there, then it's like AI takes you to 20% growth or 200% growth? Or you could be saying if you look at the last 300 years, you've been seeing technology after technology, but the rate of growth is the exact same, it's 2%. So are you saying the rate of growth will change?

Andrej

直接来说,我预计增长率在过去 200-300 年里也大致保持不变,但在人类历史进程中它爆炸了,从 0%到越来越快,工业爆炸,2%。基本上,我想说的是,有一段时间我试图在 GDP 曲线中找到 AI,我有点说服自己这是错误的。即使人们谈论递归自我改进和实验室,我也不认为这有什么不同。这是常态。当然它会递归地自我改进,而且它一直在递归地自我改进。LLM 让工程师能够更高效地工作,以构建下一轮 LLM,而且更多的组件正在被自动化和调优。所以所有工程师都能访问谷歌搜索是其中的一部分。所有工程师都有 IDE、自动补全或 Claude Code 等。这些都只是整个事情加速的一部分。所以它非常平滑。

Directly, I expect the rate of growth has also stayed roughly constant for only the last 200-300 years, but over the course of human history it's exploded, gone from 0% to faster and faster, industrial explosion, 2%. Basically, I guess what I'm saying is for a while I tried to find AI in the GDP curve, and I've kind of convinced myself that this is false. Even when people talk about recursive self-improvement and labs, I don't think this is different. It's business as usual. Of course it's going to recursively self-improve, and it's been recursively self-improving. LLMs allow the engineers to work much more efficiently to build the next round of LLMs, and a lot more of the components are being automated and tuned. So all the engineers having access to Google search is part of it. All the engineers having an IDE, autocomplete, or Claude Code, etc. It's all just part of the same speed up of the whole thing. So it's just so smooth.

Dwarkesh

但只是澄清一下,你是说增长率不会改变。智能爆炸会表现为它只是让我们继续保持在 2%的增长轨迹上,就像互联网帮助我们保持在 2%的增长轨迹上一样。

But just to clarify, you're saying that the rate of growth will not change. The intelligence explosion will show up as it just enabled us to continue staying on the 2% growth trajectory, just as the internet helped us stay on the 2% growth trajectory.

Andrej

是的。我的预期是它保持相同的模式。

Yeah. My expectation is that it stays the same pattern.

AGI 作为劳动力 On AGI as labor

Dwarkesh

对。我就抛个相反论点:我觉得它会爆发,因为真正的 AGI——我不是说那些写代码的 LLM 机器人,而是真正能替代服务器里的人类——跟其他提高生产力的技术有本质区别,因为它本身就是劳动力,对吧?我觉得我们生活在一个劳动力严重受限的世界。你跟任何创业公司创始人或普通人聊聊,问他们需要什么,答案就是需要真正有才华的人。如果你有几十亿额外的人,他们能发明东西、融入社会、从头到尾创办公司,这跟单项技术完全不是一个量级。就好像你突然给地球增加了 100 亿人口。

Yeah. Just to throw the opposite argument against you, my expectation is that it blows up because I think true AGI, and I'm not talking about LLM coding bots, I'm talking about actual replacement of a human in a server, is qualitatively different from these other productivity improving technologies because it's labor itself, right? I think we live in a very labor constrained world. If you talk to any startup founder, any person, you can just say, what do you need more of? You just need really talented people. And if you have billions of extra people who are inventing stuff, integrating themselves, making companies, start to finish, that feels qualitatively different from just a single technology. It's like asking if you get 10 billion extra people on the planet.

Andrej

也许有个反例。第一,我其实很愿意被说服,无论哪边都行。但我要说,比如计算就是劳动力。计算机就是劳动力。很多工作消失了,因为计算机自动化了数字信息处理,不再需要人。所以计算机就是劳动力。这已经发生了。自动驾驶也是计算机在做劳动力。所以我觉得这已经在发生了。还是老样子。

Maybe a counterpoint. Number one, I'm actually pretty willing to be convinced one way or another on this point. But I will say, for example, computing is labor. Computing was labor. Computers made a lot of jobs disappear because computers automate digital information processing that you now don't need a human for. So computers are labor. And that has played out. Self-driving as an example is also computers doing labor. So I guess that's already been playing out. It's still business as usual.

Dwarkesh

对,我觉得你有一台机器,它可能以更快的速度产出更多类似的东西。历史上我们有增长模式变化的例子,比如从 2% 的增长到 2% 的增长。所以对我来说,一台能产出下一个自动驾驶汽车、下一个互联网等等的机器,似乎非常合理。

Yeah, I guess you have a machine which is spitting out more things like that at potentially faster pace. And historically we have examples of the growth regime changing, like you went from 2% growth to 2% growth. So it seems very plausible to me that a machine which is then spitting out the next self-driving car and the next internet and whatever.

Andrej

我有点理解你的意思。但同时,我觉得人们做了这样的假设:我们有了一个盒子里的神,它什么都能做。但现实不会是这样。它能做某些事,在其他事上会失败,然后逐步融入社会。最后我们还是会回到同样的模式,因为那种突然拥有一个完全智能、完全灵活、完全通用的盒子里的「人」,可以解决社会上任意问题的假设——我不认为会有这种离散的变化。我觉得我们会看到同样的逐步扩散到整个行业。

I kind of see where it's coming from. At the same time, I do feel like people make this assumption of, okay, we have God in a box and now it can do everything. But it just won't look like that. It's going to be able to do some things, fail at others, and be gradually put into society. We'll end up with the same pattern, because this assumption of suddenly having a completely intelligent, fully flexible, fully general human in a box that we can dispense to arbitrary problems in society—I don't think we will have this discrete change. I think we'll arrive at the same gradual diffusion across the industry.

Dwarkesh

我觉得这些对话里经常误导人的是——我不喜欢在这个语境下用「智能」这个词,因为智能暗示你想象一个超级智能坐在服务器里,它会神启般地想出新技术和发明,导致大爆发。但我想象 20% 增长时,不是那样。我想象的是可能有数十亿非常聪明的人类头脑,或者只需要这些。但事实是,有数亿甚至数十亿这样的头脑,每个都在独立创造新产品,想办法融入经济,就像一位经验丰富的聪明移民来到一个国家——你不需要想办法让他们融入。他们自己就能创办公司、搞发明,或者提高生产力。即使在当前模式下,也有地方实现了 10-20% 的经济增长。如果你有很多人,而资本相对较少,你就能有香港或深圳,它们经历了数十年的 10% 以上增长。这只是一大群非常聪明的人,准备好利用资源,进行这种追赶期,因为我们有了这种不连续性。

I think what often ends up being misleading in these conversations is people—I don't like to use the word intelligence in this context because intelligence implies you think, oh, superintelligence will be sitting there, a single superintelligence in a server, and it'll divine how to come up with new technologies and inventions that cause this explosion. That's not what I'm imagining when I'm imagining 20% growth. I'm imagining that there are billions of very smart human minds potentially, or that's all that's required. But the fact that there are hundreds of millions of them, billions of them, each individually making new products, figuring out how to integrate themselves into the economy, just the way a highly experienced smart immigrant came to the country—you wouldn't need to figure out how to integrate them. They figured out they could start a company, make inventions, or just increase productivity. And we have examples even in the current regime of places that have had 10-20% economic growth. If you have a lot of people and less capital in comparison, you can have Hong Kong or Shenzhen, which had decades of 10% plus growth. It's just a lot of really smart people ready to make use of resources and do this period of catchup because we've had this discontinuity.

Andrej

我理解,但我还是觉得你预设了一个离散的跳跃。有一个我们等待解锁的东西,然后突然数据中心里就有了天才。我还是觉得你预设了一个基本上没有历史先例的离散跳跃,我在任何统计数据里都找不到,而且我觉得很可能不会发生。

I think I understand, but I still think that you're presupposing some discrete jump. There's some unlock that we're waiting to claim, and suddenly we're going to have geniuses in data centers. I still think you're presupposing some discrete jump that has basically no historical precedent that I can't find in any of the statistics and that I think probably won't happen.

Dwarkesh

工业革命就是这样一个跳跃,对吧?从 0.2% 的增长到 2% 的增长。我只是说你会看到另一个类似的跳跃。

I mean, the industrial revolution is such a jump, right? You went from 0.2% growth to 2% growth. I'm just saying you'll see another jump like that.

Andrej

我有点怀疑。我得看看。比如,工业革命前的有些记录可能不太准确。所以我有点怀疑,但也许你是对的。我没有强烈的意见。

I'm a little bit suspicious. I would have to look at it. For example, maybe some of the logs are not very good from before the industrial revolution. So I'm a little bit suspicious, but maybe you're right. I don't have strong opinions.

Dwarkesh

也许你是说这是一个极其神奇的单次事件,然后你说可能还会有另一个同样极其神奇的事件。它会打破范式等等。

Maybe you're saying that this was a singular event that was extremely magical, and you're saying that maybe there's going to be another event that's just like that, extremely magical. It will break the paradigm and so on.

Andrej

我其实不这么认为。工业革命的关键在于它并不神奇。如果你放大看,1770 年或 1870 年,并不是有什么关键发明。

I actually don't think that. The crucial thing about the industrial revolution was that it was not magical. If you zoomed in, what you would see in 1770 or 1870, it's not that there was some key invention.

Dwarkesh

对,没错。但与此同时,你确实把经济带到了一个进步快得多的模式,指数级增长了 10 倍。我期待 AI 也会有类似的情况,不是有一个单一时刻我们解锁了关键过剩。也许有一种新的能源,某种解锁——在这种情况下,是某种认知能力——而且有大量的认知工作要做。

Yeah, exactly. But at the same time, you did move the economy to a regime where the progress was much faster, and the exponential 10xed. And I expect a similar thing from AI, where it's not like there's going to be a single moment where we made the crucial overhang that's being unlocked. Maybe there's a new energy source, there's some unlock—in this case, some kind of cognitive capacity—and there's an overhang of cognitive work to do.

Andrej

没错。你期待当这项技术跨过门槛时,那个过剩会被填补。

That's right. And you're expecting that overhang to be filled by this new technology when it crosses the threshold.

Dwarkesh

对。

Yeah.

增长与人口 On growth and population

Dwarkesh

我的意思是,一种思考方式是通过历史来看。很多增长来自于人们提出想法,然后执行这些想法以创造有价值的产出。在大部分时间里,人口并没有爆炸式增长。过去 50 年,这一直在推动增长。有人认为,随着前沿国家的人口增长停滞,经济增长也停滞了。我认为我们回到了人口和产出的超指数增长。

And I mean one way to think about it is through history. A lot of growth comes because people come up with ideas and then execute those ideas to make valuable output. Through most of this time, population isn't exploding. That has been driving growth for the last 50 years. People have argued that growth has stagnated as population in frontier countries has also stagnated. I think we go back to the hyperexponential growth in population and output.

Andrej

对,抱歉,人口指数增长导致了产出的超指数增长。

Right, sorry, exponential growth in population that causes hyperexponential growth in output.

Dwarkesh

是的,真的很难说。

Yeah. I mean, it's really hard to tell.

Andrej

我理解那个观点,但直觉上我并不认同。

I understand that viewpoint. I don't intuitively feel that viewpoint.

谷歌 VO 3.1 On Google's VO 3.1

Dwarkesh

我们刚获得了 Google VO 3.1 的访问权限,玩起来非常酷。我们做的第一件事就是通过 V3 和 3.1 运行一堆提示,看看新版本有什么变化。这是 V3。

So, we just got access to Google's VO 3.1, and it's been really cool to play around with. The first thing we did was run a bunch of prompts through both V3 and 3.1 to see what's changed in the new version. So, here's V3.

Dwarkesh

嗨,我是 Max,我又陷入局部最小值了。

Hi, I'm Max and I got stuck in a local minimum again.

Dwarkesh

没关系,Max,我们都经历过。我花了三个 epoch 才出来。

It's okay, Max. We've all been there. Took me three epochs to get out.

Dwarkesh

这是 VO 3.1。

And here is VO 3.1.

Dwarkesh

嗨,我是 Max,我又陷入局部最小值了。

Hi, I'm Max and I got stuck in a local minimum again.

Dwarkesh

没关系,Max,我们都经历过。我花了三个 epoch。

It's okay, Max. We've all been there. Took me three epochs.

Dwarkesh

3.1 的输出始终更连贯,音频质量明显更高。我们已经使用 VO 一段时间了。事实上,今年早些时候我们发布了一篇关于完全由 V2 驱动的 AI 公司的文章,看到这些模型改进的速度令人惊叹。这次更新使 VO 在动画化我们的想法和解释器方面更加有用。你现在可以在 Gemini 应用中通过 Pro 和 Ultra 订阅试用 VO,也可以通过 Gemini API 或 Google Flow 访问。

3.1's output is just consistently more coherent and the audio is noticeably higher quality. We've been using VO for a while now. Actually, we released an essay earlier this year about AI firms fully animated by V2, and it's been amazing to see how fast these models are improving. This update makes VO even more useful in terms of animating our ideas and our explainers. You can try VO right now in the Gemini app with Pro and Ultra subscriptions. You can also access it through the Gemini API or through Google Flow.

智能与进化 On intelligence and evolution

Dwarkesh

你向我推荐了 Nick Lane 的书,在此基础上我采访了他。我有一些关于思考智能和进化史的问题。既然你在过去 20 年里一直在做 AI 研究,你对智能是什么以及发展智能需要什么有了更具体的认识。因此,你对进化只是偶然发现智能这件事感到更惊讶还是更不惊讶?

You recommended Nick Lane's book to me and then on that basis I interviewed him. I have some questions about thinking about intelligence and evolutionary history. Now that you have spent the last 20 years doing AI research, you have a more tangible sense of what intelligence is and what it takes to develop it. Are you more or less surprised as a result that evolution just spontaneously stumbled upon it?

Andrej

顺便说一句,我喜欢 Nick 的书。我来的路上还在听他的播客。关于智能及其进化,我确实认为它出现得很晚。我对它进化出来感到惊讶。我觉得思考所有那些世界很有趣。假设有一千个类似地球的行星,它们会是什么样子。我想 Nick 在这里谈到了一些早期部分。他预计基本上非常相似的生命形式,大致来说,大多数都是类似细菌的东西。

I love Nick's books by the way. I was just listening to his podcast on the way up here. With respect to intelligence and its evolution, I do claim it came fairly recently. I am surprised that it evolved. I find it fascinating to think about all the worlds out there. Say there's a thousand planets like Earth and what they look like. I think Nick was here talking about some of the early parts. He expects basically very similar life forms, roughly speaking, bacteria-like things in most of them.

Dwarkesh

是的。

Yeah.

Andrej

然后那里有一些断裂。我直觉上认为智能的进化应该是一个相当罕见的事件。动物存在的时间,我想也许你应该根据某物存在了多久来判断。例如,如果细菌已经存在了 20 亿年而什么都没发生,那么进化到真核生物可能非常困难,因为细菌实际上在地球历史早期就出现了。那么动物存在了多久呢?也许几亿年,比如多细胞动物,它们奔跑、爬行等,这大概是地球寿命的 10%左右。所以也许在那个时间尺度上,实际上并不太难。我仍然觉得直觉上它进化出来是令人惊讶的。我可能只会期待很多类似动物的生命形式做类似动物的事情。你能得到创造文化、知识并积累它们的东西,这让我感到惊讶。

And then there are a few breaks in there. I would expect that the evolution of intelligence intuitively feels to me like it should be a fairly rare event. There have been animals for, I guess maybe you should base it on how long something has existed. For example, if bacteria have been around for 2 billion years and nothing happened, then going to eukaryotes is probably pretty hard because bacteria actually came up quite early in Earth's history. So I guess how long have we had animals? Maybe a couple hundred million years, like multicellular animals that run, crawl, etc., which is maybe 10% of Earth's lifespan or something like that. So maybe on that time scale, it's actually not too tricky. I still feel like it's surprising to me intuitively that it developed. I would maybe expect just a lot of animal-like life forms doing animal-like things. The fact that you can get something that creates culture and knowledge and accumulates it is surprising to me.

Dwarkesh

好的,实际上有几个有趣的后续问题。如果你接受这个观点,即智能的核心是动物智能,那么引用的话是:如果你到了松鼠的智能,你就完成了 AGI 的大部分。然后我们在 6 亿年前的寒武纪大爆发之后立即获得了松鼠智能。似乎触发它的是 6 亿年前的氧化事件。但立即智能算法就在那里制造了松鼠智能,对吧?所以这表明动物智能一旦环境中有氧气,你就有硬件,你就能得到算法。也许进化如此之快地偶然发现它是一个意外,但我不知道这是否表明它实际上相当简单。

Okay, so there are actually a couple of interesting follow-ups. If you buy this perspective that the crux of intelligence is animal intelligence, what the quote said is if you got to the squirrel, you'd be most of the way to AGI. Then we got to squirrel intelligence right after the Cambrian explosion 600 million years ago. It seems like what instigated that was the oxygenation event 600 million years ago. But immediately the intelligence algorithm was there to make the squirrel intelligence, right? So it's suggestive that animal intelligence was there as soon as you had the oxygen in the environment, you had the hardware, you could just get the algorithm. Maybe there was an accident that evolution stumbled upon it so fast, but I don't know if that suggests it's actually quite simple.

Andrej

是的,基本上这些东西都很难说。我想你可以根据某物存在了多久,或者感觉某物被瓶颈了多久来稍微判断一下。Nick 描述了细菌中一个非常明显的瓶颈:化学生物化学的极端多样性,但 20 亿年来没有任何东西进化成动物。我不知道我们是否在动物和智能中看到了完全等同的情况,就你的观点而言。但我想也许我们也可以从我们认为智能独立出现了多少次来看待它。

Yes, basically it's so hard to tell with any of this stuff. I guess you can base it a little bit on how long something has existed or how long it feels like something has been bottlenecked. Nick described this very apparent bottleneck in bacteria for years: extreme diversity of chemical biochemistry and yet nothing that grows to become animals for two billion years. I don't know that we've seen exactly that kind of an equivalent with animals and intelligence, to your point. But I guess maybe we could also look at it with respect to how many times we think intelligence has individually sprung up.

Dwarkesh

这是一个非常好的调查方向。

That's a really good thing to investigate.

Andrej

也许一个想法是,我几乎觉得有原始人类智能,还有鸟类智能,比如乌鸦等非常聪明,但它们的大脑部分实际上非常不同,我们没有太多证据。所以这也许是一个轻微迹象,表明智能出现了几次,在这种情况下你可能会更频繁地期待它。

Maybe one thought on that is I almost feel like there's the hominid intelligence and I would say bird intelligence, like ravens etc., are extremely clever but their brain parts are actually quite distinct and we don't have that much evidence. So maybe that's a slight indication of intelligence springing up a few times, and in that case you'd maybe expect it more frequently.

Dwarkesh

是的,一位前嘉宾 Gwern 和 Carl Shulman 对此提出了一个非常有趣的观点,他们认为人类和灵长类动物拥有的可扩展算法也在鸟类中出现过,也许在其他时候也出现过。

Yeah, a former guest, Gwern, and also Carl Shulman have made a really interesting point about that, which is their perspective that the scalable algorithm which humans have and primates have arose in birds as well, and maybe other times as well.

人类智能与进化生态位 On human intelligence and evolutionary niche

Dwarkesh

但在人类身上,我们找到了一个奖励智力边际增长的进化生态位。

But in humans, we found an evolutionary niche which rewarded marginal increases in intelligence.

Andrej

并且还有一个可扩展的大脑算法来实现这些智力增长。例如,如果一只鸟的大脑更大,它就会从空中掉下来。所以它相对于大脑尺寸来说非常聪明,但它所处的生态位并不奖励大脑变大。也许一些非常聪明的海豚等也是如此。

And also had a scalable brain algorithm that could achieve those increases in intelligence. For example, if a bird had a bigger brain, it would just collapse out of the air. So it's very smart for the size of its brain, but it's not in a niche which rewards the brain getting bigger. Maybe similar with some really smart dolphins, etc.

Dwarkesh

没错。而人类,我们有双手,这奖励了学习使用工具的能力。我们可以外化消化过程,将更多能量供给大脑,这就启动了飞轮。

Exactly. Whereas humans, we have hands that reward being able to learn how to do tool use. We can externalize digestion, giving more energy to the brain, and that kicks off the flywheel.

Andrej

哦,是的。还有可操作的东西。我猜如果我是海豚会更难。比如,你怎么生火?在水中能做的事情范围可能比陆地上少,仅仅是化学层面上的。

Oh, yeah. And just stuff to work with. I'm guessing it would be harder if I was a dolphin. I mean, how do you have fire, for example? The universe of things you can do in water is probably lower than what you can do on land, just chemically.

Dwarkesh

我同意这种关于生态位和激励的观点。我仍然觉得我们没有卡在肌肉更大的动物上有点神奇。通过智力这条路是一个非常迷人的突破点。Burn 的说法是:之所以这么难,是因为有一条很窄的线,介于某种东西重要到值得直接将其精确回路蒸馏回 DNA,和根本不值得学习之间。

I do agree with this viewpoint of these niches and what's being incentivized. I still find it kind of miraculous that we didn't get stuck on animals with bigger muscles. Going through intelligence is a really fascinating breaking point. The way Burn put it: the reason it was so hard is a very tight line between being in a situation where something is so important to learn that it's not just worth distilling the exact right circuits directly back into your DNA, versus it's not important enough to learn at all.

Andrej

它必须能激励构建在生命周期内学习的算法。

It has to be something which incentivizes building the algorithm to learn in lifetime.

Dwarkesh

没错。你必须激励某种适应性。你实际上想要不可预测的环境,这样进化就不能把算法固化到你的权重中。很多动物在这个意义上基本上是预置好的,而人类必须在出生后的测试时间自己摸索。所以也许你确实需要那种变化非常快的环境,无法预见什么会奏效,于是你创造智力来在测试时间解决。

Exactly. You have to incentivize some kind of adaptability. You actually want environments that are unpredictable, so evolution can't bake your algorithms into your weights. A lot of animals are basically pre-baked in this sense, and so humans have to figure it out at test time when they get born. So maybe you actually want these kinds of environments that change really rapidly, where you can't foresee what will work well, and so you create intelligence to figure it out at test time.

Andrej

Quentyn Pope 有一篇有趣的博文,他说巴西人不期望急剧起飞。人类有过急剧起飞:6 万年前我们似乎就有了今天拥有的认知架构,而 1 万年前农业革命、现代性等等。那中间的 5 万年发生了什么?

Quentyn Pope had this interesting blog post where he says the Brazilian doesn't expect a sharp takeoff. Humans had the sharp takeoff: 60,000 years ago we seem to have had the cognitive architectures that we have today, and 10,000 years ago agricultural revolution, modernity, etc. What was happening in that 50,000 years?

Dwarkesh

嗯,你必须建立一种文化支架,让知识可以代代积累。这种能力在我们做 AI 训练的方式中是免费存在的:如果你重新训练一个模型,它仍然可以——在很多情况下它们确实被蒸馏了,但它们可以互相训练,可以在优质的预训练语料上训练。它们不必真的从头开始。所以从某种意义上说,人类花了很长时间才启动的文化循环,在我们做 LLM 训练的方式中是免费得来的。

Well, you had to build this sort of cultural scaffold where you can accumulate knowledge over generations. This is an ability that exists for free in the way we do AI training: if you retrain a model, it can still — in many cases they're literally distilled, but they can be trained on each other, they can be trained on the premium pre-training corpus. They don't literally have to start from scratch. So there's a sense in which the thing that took humans a long time — getting this cultural loop going — just comes for free with the way we do LLM training.

Andrej

既是也不是,因为语言模型并没有真正的文化等价物,也许我们给了它们太多,反而激励它们不去创造文化。但文化、书面记录、互相传递笔记的概念——我认为目前语言模型没有对应的东西。所以语言模型目前并没有真正的文化,我认为这是障碍之一。

Yes and no, because LMs don't really have the equivalent of culture, and maybe we're giving them way too much and incentivizing not to create it. But the notion of culture, of written record, of passing down notes between each other — I don't think there's an equivalent of that with LMs right now. So LMs don't really have culture right now, and it's kind of one of the impediments, I would say.

Dwarkesh

你能给我讲讲语言模型文化可能是什么样子吗?

Can you give me some sense of what LLM culture might look like?

Andrej

最简单的情况下,它会是一个巨大的草稿本,语言模型可以编辑它。当它阅读或协助工作时,它会为自己编辑这个草稿本。

In the simplest case, it would be a giant scratch pad that the LLM can edit. As it's reading stuff or helping out with work, it's editing the scratch pad for itself.

Dwarkesh

为什么一个语言模型不能为另一个语言模型写一本书?那会很酷。比如,为什么其他语言模型不能读这个语言模型的书,并从中受到启发或感到震惊?这些东西都没有等价物。

Why can't an LLM write a book for another LLM? That would be cool. Like, why can't other LLMs read this LLM's book and be inspired by it or shocked by it? There's no equivalence for any of this stuff.

Andrej

有趣。你预计这种事情什么时候开始发生?更广泛地说,关于多智能体系统和某种独立的 AI 文明与文化。

Interesting. When would you expect that kind of thing to start happening? And more generally, about multi-agent systems and a sort of independent AI civilization and culture.

Dwarkesh

我认为在多智能体领域有两个强大的想法都还没有真正被实现。第一个是文化,语言模型基本上拥有一个不断增长的知识库,用于它们自己的目的。第二个看起来更像自我对弈这个强大的想法。在我看来,它极其强大。进化很大程度上是竞争驱动智力。在 AlphaGo 中,它通过自我对弈学会下围棋。目前没有语言模型自我对弈的等价物,但我预计它也会出现。还没有人做到。例如,为什么一个语言模型不能创建一系列问题让另一个语言模型学习解决,然后这个语言模型总是试图提供越来越难的问题?我认为有很多组织方式,这是一个研究领域,但我还没有看到任何令人信服地声称实现了这两种多智能体改进的东西。我仍然认为我们主要处于单个智能体的领域,但我认为这将会改变。在文化领域,我也会把组织包括进去,我们也没有看到类似的东西出现。所以这就是为什么我们还处于早期。

I think there are two powerful ideas in the realm of multi-agent that have both not been really claimed. The first one is culture and LLMs basically having a growing repertoire of knowledge for their own purposes. The second one looks a lot more like the powerful idea of self-play. In my mind, it's extremely powerful. Evolution is a lot of competition driving intelligence. In AlphaGo, it plays against itself and learns to get really good at Go. There's no equivalent of self-playing LMs, but I would expect that to also exist. No one has done it yet. Why can't an LM, for example, create a bunch of problems that another LM is learning to solve, and then the LM is always trying to serve more and more difficult problems? I think there are a bunch of ways to organize it, and it's a realm of research, but I haven't seen anything that convincingly claims both of those multi-agent improvements. I still think we're mostly in the realm of a single individual agent, but I think that will change. In the realm of culture, I would also bucket organizations, and we haven't seen anything like that coming in either. So that's why we're still early.

LLM 协作瓶颈 On LLM collaboration bottlenecks

Dwarkesh

你能找出阻碍 LLM 之间协作的关键瓶颈吗?较小的模型就像幼儿园或小学生,虽然能通过博士考试,但在认知上仍然像孩子。它们并不真正理解自己在做什么。

Can you identify the key bottleneck preventing collaboration between LLMs? Smaller models resemble kindergarten or elementary students, and even though they can pass PhD quizzes, they still feel cognitively like kids. They don't really know what they're doing.

自动驾驶进展:演示到产品差距 On self-driving progress and the demo-to-product gap

Andrej

2017 到 2022 年我在特斯拉负责自动驾驶。我们从酷炫的演示发展到数千辆车自主行驶,但远未完成。最早的演示可以追溯到 1980 年代。2014 年我体验过 Waymo 的完美演示,但依然花了很长时间。演示到产品之间存在巨大鸿沟:演示容易,产品难,尤其是在失败成本很高的情况下。

I was at Tesla from 2017 to 2022 leading self-driving. We went from cool demos to thousands of cars doing autonomous drives. But it's not even near done. The first demos go back to the 1980s. I had a perfect demo from Waymo in 2014, yet it still took a long time. There's a large demo-to-product gap: the demo is easy, but the product is hard, especially when the cost of failure is high.

Dwarkesh

有意思。

Interesting.

Andrej

软件工程也存在同样的特性。对于生产级代码,错误可能导致安全漏洞和数据泄露。就像自动驾驶:出错可能造成人身伤害,但软件的后果可能更严重。所以两者有共同点。

In software engineering, the same property exists. For production-grade code, mistakes can lead to security vulnerabilities and data leaks. It's like self-driving: if things go wrong, you might get injury, but in software the consequences could be even worse. So both share that property.

Dwarkesh

有意思。

Interesting.

Andrej

耗时的是逐九提升。每个九的工作量是恒定的。演示能做到 90%——这是第一个九。然后你需要第二个、第三个、第四个、第五个九。在特斯拉的五年里,我们可能只提升了两个或三个九。还有更多的九要走。这就是为什么这些事情如此耗时。我对演示非常不以为然。你需要一个能应对现实挑战的实际产品。这是逐九提升;每个九的工作量恒定。演示令人鼓舞,但仍有大量工作要做。这强化了我的时间线判断。

What takes time is a march of nines. Each nine is a constant amount of work. A demo works 90% of the time—that's the first nine. Then you need the second, third, fourth, fifth nine. At Tesla over five years, we went through maybe two or three nines. There are still more nines to go. That's why these things take so long. I'm very unimpressed by demos. You need an actual product that faces real-world challenges. It's a march of nines; each nine is constant. Demos are encouraging, but there's still a huge amount of work. This enforces my timelines.

自动驾驶类比与 AI 部署 On the self-driving car analogy for AI deployment

Dwarkesh

但我想你的观点是,如果你犯了一个灾难性的编码错误,比如每七年搞垮一个重要系统,这很容易发生。实际上,按实际时间算,这远少于七年,因为你一直在输出代码。所以按 token 算可能是七年,但按实际时间算,问题要难得多。我的意思是,自动驾驶只是人们做的成千上万件事中的一件,几乎是一个单一的垂直领域。而当我们谈论通用软件工程时,涉及的面更广。

But I guess your point is that if you made a catastrophic coding mistake like breaking some important system every seven years, it's very easy to do. And in fact, in terms of wall clock time, it would be much less than seven years because you're constantly outputting code. So per tokens, it would be seven years, but in terms of wall clock time, it's a much harder problem. I mean, self-driving is just one of thousands of things that people do. It's almost like a single vertical. Whereas when we're talking about general software engineering, there's more surface area.

Andrej

是的,人们对这个类比还有另一个反对意见,那就是在自动驾驶中,花费大量时间的是解决构建稳健的基本感知、构建表征以及让模型具备一些常识以便能泛化到分布外情况的问题。例如,如果有人在路上挥手,你不需要专门训练,模型会理解如何应对。而这些正是我们今天从 LLM 或 VLM 中免费获得的东西。所以我们不需要解决这些非常基本的表征问题。因此,现在将 AI 部署到不同领域,就像用当前模型将自动驾驶汽车部署到另一个城市,虽然困难,但并非需要十年之久的任务。

Yeah, there's another objection people make to that analogy, which is that with self-driving, what took a big fraction of that time was solving the problem of building basic perception that's robust, building representations, and having a model that has some common sense so it can generalize to out-of-distribution situations. For example, if somebody is waving down the road, you don't need to train for it; the thing will have some understanding of how to respond. And these are things we're getting for free with LLMs or VLMs today. So we don't have to solve these very basic representation problems. And so now deploying AI across different domains will be like deploying a self-driving car with current models to a different city, which is hard but not a 10-year-long task.

Dwarkesh

是的。基本上,我不完全确定是否完全同意这一点。我不知道我们免费获得了多少,而且我仍然认为对我们所获得的东西的理解存在很多空白。我的意思是,我们确实在单个实体中获得了更通用的智能,而自动驾驶是一个非常特殊目的的任务,需要构建一个专用系统,这在某种意义上可能更难,因为它不是从你大规模做的更通用的事情中自然产生的。但我仍然不确定这个类比是否完全成立,因为 LLM 仍然很容易出错,还有很多空白需要填补。我不认为我们开箱即用地获得了神奇的泛化能力。另一个我想回到的方面是,自动驾驶汽车还远未完成。部署仍然非常有限。即使是 Waymo 也只有很少的车辆,大致是因为它们不经济。他们构建了面向未来的东西,不得不让它不经济。所以有所有这些成本——不仅是这些车辆及其运营和维护的边际成本,还有整个项目的资本支出。让它变得经济仍然是一场艰苦的战斗。此外,当你看到这些无人驾驶的汽车时,这有点欺骗性,因为存在复杂的运营中心,有人类参与其中。我不知道全部情况,但人类参与的程度可能比你预期的要多,有人从某个地方远程介入。我不知道他们是否完全参与驾驶,但他们肯定参与其中。从某种意义上说,我们并没有移除人,只是把他们移到了我们看不见的地方。我仍然认为从一个环境到另一个环境会有工作要做,所以让自动驾驶成为现实仍然存在挑战。但我确实同意它已经跨过了感觉真实的门槛,除非是零售运营。例如,Waymo 不能去城市的所有地方——我怀疑是那些信号不好的地方。总之,我实际上对技术栈一无所知;我只是在瞎猜。

Yeah. Basically, I'm not 100% sure if I fully agree with that. I don't know how much we're getting for free, and I still think there are a lot of gaps in understanding what we are getting. I mean, we're definitely getting more generalizable intelligence in a single entity, whereas self-driving is a very special-purpose task that requires building a special-purpose system, which might be even harder in a certain sense because it doesn't fall out from a more general thing you're doing at scale. But I still don't know if the analogy fully resonates because LLMs are still pretty fallible and have a lot of gaps that need to be filled. I don't think we're getting magical generalization completely out of the box. And the other aspect I wanted to return to is that self-driving cars are nowhere done still. Deployments are still pretty minimal. Even Waymo has very few cars, roughly because they're not economical. They built something that lives in the future, and they had to make it uneconomical. So there are all these costs—not just marginal costs for those cars and their operation and maintenance, but also the capex of the entire thing. Making it economical is still going to be a slog. Also, when you look at these cars with no one driving, it's a bit deceiving because there are elaborate operation centers with people kind of in the loop. I don't know the full extent, but there's more human in the loop than you might expect, with people beaming in from somewhere. I don't know if they're fully in the loop with the driving, but they're certainly involved. In some sense, we haven't removed the person; we've moved them somewhere we can't see them. I still think there will be work going from environment to environment, so there are still challenges to make self-driving real. But I do agree it's crossed a threshold where it feels real, unless it's retail operated. For example, Waymo can't go to all parts of the city—my suspicion is parts where you don't get good signal. Anyway, I don't actually know anything about the stack; I'm just making it up.

Andrej

你在特斯拉做了五年自动驾驶。抱歉,我对 Waymo 的具体情况一无所知。我觉得可以谈谈它们。实际上,顺便说一句,我喜欢 Waymo,经常乘坐。所以我不想说……我只是觉得人们有时对某些进展过于天真,我仍然认为还有大量工作要做。而且我认为特斯拉采取了一种在我看来更具可扩展性的方法,团队做得非常好。我公开预测过这件事的走向,它更像是一个早期启动,因为你可以集成那么多传感器。但我确实认为特斯拉采取了更具可扩展性的策略,结果会更像那样。所以我认为这还需要时间,尚未实现。基本上,我不想把自动驾驶说成是花了十年时间的事情,因为它还没有完成,如果你明白我的意思。因为第一,起点是 1980 年,而不是十年前;第二,终点尚未到来。

You less self-love driving for 5 years at Tesla. Sorry, I don't know anything about the specifics of Waymo. I feel talk about them. I actually, by the way, I love Waymo and I take it all the time. So I don't want to say... I just think that people are sometimes a little bit too naive about some of the progress, and I still think there's a huge amount of work. And I think Tesla took, in my mind, a much more scalable approach, and I think the team is doing extremely well. I'm kind of on the record for predicting how this thing will go, which is way more like early start because you can package up so many sensors. But I do think Tesla is taking the more scalable strategy, and it's going to look a lot more like that. So I think this will have to still play out and hasn't. Basically, I don't want to talk about self-driving as something that took a decade because it didn't take yet, if that makes sense. Because one, the start is at 1980, not 10 years ago, and two, the end is not here yet.

Dwarkesh

是的。终点还很远,因为当我们谈论自动驾驶时,我通常指的是大规模自动驾驶。人们不必考驾照等等。我很想探讨这个类比可能不同的另外两个方面。我特别好奇的原因是,我认为 AI 部署的速度以及早期阶段的价值,可能是当今世界上最重要的问题。如果你试图想象 2030 年代或 2040 年代的样子,这就是你需要有所理解的问题。

Yeah. The end is not near yet because when we're talking about self-driving, usually in my mind it's self-driving at scale. People don't have to get a driver's license, etc. I'm curious to bounce two other ways in which the analogy might be different. And the reason I'm especially curious about this is because I think the question of how fast AI is deployed, how valuable it is when it's early on, is potentially the most important question in the world right now. If you're trying to model what the 2030s or 2040s look like, this is the question you want to have some understanding of.

自动驾驶类比与 AI 部署经济学 On self-driving analogy and economics of AI deployment

Dwarkesh

另一个你可能想到的问题是,自动驾驶有延迟要求,我不知道实际模型是什么,但假设是几千万参数,这对 LLM 的知识工作来说并不是必要的限制,或者可能对计算机使用之类的东西有影响。但无论如何,另一个更重要的问题是资本支出:是的,提供额外一份模型副本有额外成本,但会话的运营成本相当低,而且你可以将 AI 的成本摊销到训练运行本身,具体取决于推理扩展的情况。但这肯定不像建造一辆全新的汽车来服务另一个模型实例那样昂贵。因此,更广泛部署的经济性要有利得多。

So another thing you might think is one you have this latency requirement with self-driving where you have I have no idea what the actual models are but I assume like tens of millions of parameters or something which is not the necessary constraint for knowledge work with LLMs or maybe it might be with the computer use and stuff but anyways the other big one is maybe more importantly on this capex question yes there is additional cost to serving up an additional copy of a model But the sort of opex of a session is quite low and you can amortize the cost of AI into the training run itself depending on how inference scaling goes and stuff but it's certainly not as much as like building a whole new car to serve another instance of a model. So it just the economics of deploying more widely are much more favorable.

Andrej

我认为没错。如果你停留在比特领域,比特比任何触及物理世界的东西要容易一百万倍。

I think that's right. I think if you're sticking in the realm of bits, bits are like a million times easier than anything that touches the physical world.

Dwarkesh

不。

No.

Andrej

我完全同意。比特是完全可变的,可以以极快的速度任意重新排列。所以你会预期行业中的适应速度也会快得多。

I definitely grant that. Bits are completely changeable, arbitrarily reshuffable at very rapid speed. So you would expect a lot more faster adaptation also in the industry and so on.

Dwarkesh

那第一个是什么?

And then what was the first one?

Andrej

延迟要求及其对模型大小的影响。

The latency requirements and its implications for model size.

Dwarkesh

我认为大致正确。我的意思是,如果我们在谈论大规模的知识工作,实际上会有一些延迟要求,因为我们将不得不为此创造大量的算力。

I think that's roughly right. I mean I also think that if we are talking about knowledge work at scale there will be some latency requirements practically speaking because we're going to have to make create a huge amount of compute instead of that.

AI 部署的社会与法律方面 On societal and legal aspects of AI deployment

Dwarkesh

然后我想最后一点,我想简要谈一下所有其他方面。社会怎么看待它?法律后果是什么?它如何在法律上运作?保险方面如何运作?谁真正负责?这些层面和方面是什么?比如人们把锥桶放在 Waymo 上,会有什么等效的事情?会有所有这些的等效物。所以我确实觉得自动驾驶是一个很好的类比,你可以从中借鉴。是的,车上的锥桶的等效物是什么?隐藏的远程操作员的等效物是什么?几乎所有的方面。

And then I think like the last aspect that I very briefly want to also talk about is all the rest of it. So what does society think about it? What are the legal ramifications? How is it working legally? How is it working insurance-wise? Who is really like what is the where what are those layers of it and aspects of it what happens with what is the equivalent of people putting a cone on a Waymo? There's going to be equivalent of all that and so I do think that I almost feel like self-driving is a very nice analogy that you can borrow things from. Yeah, what is the equivalent of a cone on the car? What is the equivalent of a teleoperating worker who's like hidden away? And almost like all the aspects of it.

Andrej

是的,你对这是否意味着当前的建设——在一两年内将全球可用算力提高 10 倍,到本十年末可能提高 100 倍以上——有什么看法?如果 AI 的使用量低于一些人天真预测的水平,这是否意味着我们过度建设了算力,还是这是一个单独的问题?

Yeah, do you have any opinions on whether this implies that the current day build which would 10x the amount of available compute in the world in a year or two and maybe like 100x more than 100x it by the end of the decade. If the use of AI will be lower than some people naively predict, does that mean that we're overbuilding compute or is that a separate question?

Dwarkesh

有点像铁路之类的事情?抱歉。

Kind of like what happened with railroads and all this kind of stuff? Sorry.

Andrej

是铁路吗?哦,抱歉。是有历史先例,还是电信行业?就像铺设互联网,但互联网十年后才出现,并在 90 年代末在电信行业制造了整个泡沫。是的。

Was it railroads? Oh, sorry. It was um yeah, there is like historical precedent or was it with telecommunication industry, right? Like paving the internet that only came like a decade later, and creating like a whole bubble in the telecommunications industry in the late 90s kind of thing. Yeah.

Dwarkesh

嗯,所以我不知道。我的意思是,我明白我听起来很悲观。

Um so I don't know. I mean, I understand I'm sounding very pessimistic here.

Andrej

我这样做是因为我实际上很乐观。我认为这会成功。我认为这是可处理的。我听起来悲观只是因为当我刷推特时,我看到所有这些东西对我来说毫无意义。我认为存在很多原因。很多原因我认为老实说只是融资。只是激励结构。很多可能是融资。很多只是注意力,你知道,将注意力转化为互联网上的金钱,诸如此类。所以我认为有很多这样的情况,我只是在对此做出反应。但我总体上仍然非常看好技术。我认为我们会解决所有这些问题,而且已经取得了快速的进展。我实际上不知道是否存在过度建设。我认为我们能够消化正在建设的东西。因为我确实认为,例如云代码或开放代码之类的东西,一年前还不存在,对吧?对吗?我认为大致正确。这是不存在的奇迹技术。我认为会有巨大的需求,正如我们已经看到的 ChatGPT 的需求等等。所以是的,我实际上不知道是否存在过度建设。但我想我只是在回应一些人们继续错误地说的非常快的时间线,我在 AI 领域的 15 年中多次听到非常有名的人一直搞错。我希望这能得到适当校准,而且我认为其中一些问题确实具有地缘政治影响等等,我不希望人们在这方面犯错。所以我希望我们立足于技术是什么和不是什么的事实。

I'm only doing that I'm actually optimistic. I think this will work. I think it's tractable. I'm only sounding pessimistic because when I go on my Twitter timeline, I see all this stuff that makes no sense to me. And I think there's a lot of reasons for why that exists. And I think a lot of it is I think honestly just fundraising. It's just incentive structures. A lot of it may be fundraising. A lot of it is just attention, you know, converting attention to money on the internet, stuff like that. So I think there's a lot of that going on and I think I'm only reacting to that. But I'm still like overall very bullish on technology. I think we're going to work through all this stuff and I think there's been a rapid amount of progress. I don't actually know that there's overbuilding. I think that we're going to be able to gobble up what in my understanding is being built. Because I do think that for example Claude Code or open codex and stuff like that, they didn't even exist a year ago, right? Is that right? I think it's roughly right. This is miraculous technology that didn't exist. I think there's going to be a huge amount of demand as we see the demand in ChatGPT already and so on. So yeah, I don't actually know that there's overbuilding. But I guess I'm just reacting to like some of the very fast timelines that people continue to say incorrectly and I've heard many many times over the course of my 15 years in AI where very reputable people keep getting this wrong all the time. And I think I want this to be properly calibrated and I think some of this also it does have like geopolitical ramifications and things like that when some of these questions and I think I don't want people to make mistakes on that sphere of things. So I do want us to be grounded in reality of what technology is and isn't.

教育、尤里卡与个人专注 On education, Eureka, and personal focus

Dwarkesh

我们来谈谈教育和 Eureka 之类的事情。

Let's talk about education and Eureka and stuff.

Andrej

你可以做的一件事是创办另一个 AI 实验室,然后尝试解决这些问题。好奇你现在在做什么。

One thing you could do is start another AI lab and then try to solve those problems. Curious what you're up to now.

Dwarkesh

是的。

Yeah.

Andrej

然后,是的,为什么不是 AI 研究本身呢?

And then, yeah, why not AI research itself?

Dwarkesh

我想也许我会这样说:我对 AI 实验室正在做的事情感到某种程度的确定性。我觉得我可以在那里帮忙,但我不确定我能独特地改进它。但我个人最大的担忧是,很多事情发生在人类一边,人类因此被剥夺了权力。我有点在意,不仅是我们将要建造的所有戴森球,以及 AI 将以完全自主的方式建造的东西。我在意人类会发生什么。

I guess maybe like the way I would put it is I feel some amount of determinism around the things that AI labs are doing. And I feel like I could help out there, but I don't know that I would uniquely improve it. But I think like my personal big fear is that a lot of this stuff happens on the side of humanity and that humanity gets disempowered by it. And I kind of like I care not just about all the Dyson spheres that we're going to build and that AI is going to build in a fully autonomous way. I care about what happens to humans.

Andrej

是的。

Yeah.

Dwarkesh

我希望人类在这个未来中过得很好。我觉得那是我能更独特地增加价值的地方,而不是在前沿实验室中做增量改进。

And I want humans to be well off in this future. And I feel like that's where I can a lot more uniquely add value than an incremental improvement in a frontier lab.

对教育的恐惧与愿景 On fears and vision for education

Dwarkesh

所以,我最担心的可能是像《机器人总动员》或《蠢蛋进化论》这类电影描绘的情景——人类被边缘化。我希望人类在未来能变得更好。而我认为,这可以通过教育来实现。

And so, I guess I'm most afraid of something maybe like depicted in movies like Wall-E or Idiocracy where humanity is sort of on the side of this stuff. I want humans to be much better in this future. And so to me, this is kind of through education that you can actually achieve this.

Andrej

哦,是的。Eureka 试图打造的东西,最简单的描述就是——我们在建一所星际舰队学院。不知道你看过《星际迷航》没有。

Oh yeah. So Eureka is trying to build, I think the easiest way I can describe it is we're trying to build the Starfleet Academy. I don't know if you watched Star Trek.

Dwarkesh

我没看过。不过,嗯。

I haven't. But yeah.

Andrej

好的。星际舰队学院是一所前沿技术的精英机构,建造飞船,培养学员成为飞船驾驶员。所以我设想的是一个技术知识的精英机构,一所非常前沿、顶级的学校。

Okay. Starfleet Academy is an elite institution for frontier technology, building spaceships and graduating cadets to be the pilots of these spaceships. So I imagine an elite institution for technical knowledge, a kind of school that is very up-to-date and premier.

技术内容教学与尤里卡 On teaching technical content and Eureka

Dwarkesh

我有一类问题想问你:如何教授技术或科学内容?因为你是这方面的世界级大师。我很好奇,你对自己在 YouTube 上发布的内容是怎么想的,以及对于 Eureka,你的想法是否有不同。

A category of questions I have for you is explaining how one teaches technical or scientific content, because you are one of the world masters at it. I'm curious both about how you think about it for content you've already put out there on YouTube, but also to the extent it's any different how you think about it for Eureka.

Andrej

是的。关于 Eureka,我认为教育中非常吸引我的一点是,我确实认为教育会随着 AI 的辅助而发生根本性改变,它必须在一定程度上被重新设计和改造。我仍然觉得我们处于早期阶段。会有很多人尝试做显而易见的事情,比如用大语言模型,问它问题,做你现在通过提示词能做的基本事情。我觉得这有帮助,但感觉还有点粗糙。我想做得更到位,而我认为目前的能力还达不到我想要的水平。我想要的是真正的导师体验。我脑海中一个突出的例子是最近我在学韩语。我经历了一个阶段,自己在网上学韩语;然后另一个阶段,我在韩国参加了一个小班课,有一个老师和大约十个人,那真的很有趣。然后我换成了一对一的导师。让我着迷的是,我觉得我遇到了一位非常好的导师。想想这位导师为我做了什么,那种体验多么不可思议,以及我最终想构建的东西标准有多高。她通过非常简短的对话,立刻理解了我作为学生的水平,知道我知道什么、不知道什么,并且能精准地提出各种问题来理解我的世界模型。目前没有任何大语言模型能百分之百做到这一点,甚至差得远。但一个好的导师可以。一旦她理解了,她就会提供我当前能力水平所需的一切。我需要总是被适度挑战,不能太难也不能太简单。导师非常擅长给你恰到好处的东西。所以基本上,我觉得我是学习的唯一限制。我总是得到完美的信息。我是唯一的限制。我感觉很好,因为我是唯一的障碍。不是因为我找不到知识,或者知识解释得不好,只是我自己的记忆能力等等。这就是我想为人们提供的。如何自动化这一点?

Yeah. With respect to Eureka, I think one thing that is very fascinating to me about education is that I do think education will fundamentally change with AIs on the side, and it has to be rewired and changed to some extent. I still think we're pretty early. There will be a lot of people who try to do the obvious things, like have an LLM and ask it questions and do all the basic things you would do via prompting right now. I think it's helpful, but it still feels a bit slop. I'd like to do it properly, and I think the capability is not there for what I would want. What I'd want is an actual tutor experience. A prominent example in my mind is I was recently learning Korean. I went through a phase where I was learning Korean by myself on the internet, then a phase where I was part of a small class in Korea with a teacher and about 10 people, which was really funny. Then I switched to a one-on-one tutor. What was fascinating to me is that I think I had a really good tutor. Thinking through what this tutor was doing for me and how incredible that experience was, and how high the bar is for what I actually want to build eventually. She instantly, from a very short conversation, understood where I am as a student, what I know and don't know, and she was able to probe exactly the kinds of questions to understand my world model. No LLM will do that for you 100% right now. Not even close. But a good tutor will. Once she understands, she served me all the things I needed at my current sliver of capability. I need to be always appropriately challenged, not too hard or too trivial. A tutor is really good at serving you just the right stuff. So basically I felt like I was the only constraint to learning. I was always given the perfect information. I'm the only constraint. And I felt good because I'm the only impediment. It's not that I can't find knowledge or that it's not properly explained. It's just my ability to memorize and so on. This is what I want for people. How do you automate that?

Dwarkesh

所以关于当前能力的问题问得很好。你做不到。

So very good question about the current capability. You don't.

Andrej

但我确实认为,正因为如此,现在实际上不是构建这种 AI 导师的正确时机。我仍然认为它是一个有用的产品,很多人会去构建,但我仍然觉得标准太高了,能力还达不到。即使是今天,我会说 ChatGPT 是一个非常有价值的教育产品,但对我来说,看到标准有多高是极其迷人的。和她在一起时,我几乎觉得我根本不可能构建出这样的东西。

But I do think that as, and that's why I think it's not actually the right time to build this kind of AI tutor. I still think it's a useful product and lots of people will build it, but I still feel the bar is so high and the capability is not there. Even today, I would say ChatGPT is an extremely valuable educational product, but for me it was so fascinating to see how high the bar is. When I was with her, I almost felt like there's no way I can build this.

Dwarkesh

但你正在构建它,对吧?

But you are building it, right?

Andrej

任何有过好导师的人都会问:你怎么能构建出这个?所以我想我在等待那种能力。我确实认为,在很多方面,比如我做过一些计算机视觉的 AI 咨询。很多时候,我给公司带来的价值是告诉他们不要用 AI。并不是说我是 AI 专家,他们描述问题,我说不要用 AI。这就是我的增值。我觉得现在教育领域也是如此。我有点觉得,对于我心中的设想,时机还未到,但时机终会到来。但现在,我正在构建一些看起来更传统的东西,有物理和数字组件等等。但我觉得未来应该是什么样子是很明显的。

Anyone who's had a really good tutor is like, how are you going to build this? So I guess I'm waiting for that capability. I do think that in a lot of ways in the industry, for example, I did some AI consulting for computer vision. A lot of the time, the value that I brought to the company was telling them not to use AI. It wasn't like I was the AI expert and they described a problem and I said don't use AI. This was my value add. And I feel like it's the same in education right now. I kind of feel like for what I have in mind, it's not yet the time, but the time will come. But for now, I'm building something that looks maybe a bit more conventional, that has a physical and digital component and so on. But I think it's obvious how this should look in the future.

Dwarkesh

你愿意说说你希望今年或明年发布什么吗?

Are you willing to say what is the thing you hope will be released this year or next year?

Andrej

嗯,我正在构建第一门课程,我想让它成为一门非常好的课程,一个最先进的、显而易见的学习 AI 的目的地,因为那正是我熟悉的领域。所以我认为这是一个非常好的第一个产品,要让它变得非常好。这就是我正在构建的。你刚才提到的 Nano Chat 是 LLM 101n 的顶点项目,而 LLM 101n 是我正在构建的一门课程。所以这是其中很大的一部分,但现在我必须构建很多中间环节,然后我需要雇佣一个小型的助教团队等等,真正地构建整个课程。

Well, I'm building the first course and I want to have a really good course, a state-of-the-art, obvious destination you go to learn AI in this case, because that's just what I'm familiar with. So I think it's a really good first product to get to be really good. And so that's what I'm building. Nano Chat, which you briefly mentioned, is a capstone project of LLM 101n, which is a class that I'm building. So that's a really big piece of it, but now I have to build out a lot of the intermediates and then I have to hire a small team of TAs and so on and actually build the entire course.

教育作为知识阶梯 On education as ramps to knowledge

Dwarkesh

还有一点我想说的是,很多时候人们想到教育,想到的是传播知识这种较软的成分。但我心里想的其实是某种非常硬核和技术性的东西。在我看来,教育就是构建通往知识的阶梯这一极其困难的技术过程。

And maybe one more thing that I would say is that many times when people think about education, they think about the softer component of diffusing knowledge. But I actually have something very hard and technical in mind. In my mind, education is the very difficult technical process of building ramps to knowledge.

Andrej

所以在我看来,nano chat 就是通往知识的阶梯,因为它非常简单,是超级简化的全栈产物。如果你把这个东西给一个人,他们浏览一遍,就能学到大量知识。

So in my mind, nano chat is a ramp to knowledge because it's a very simple, super simplified full stack thing. If you give this artifact to someone and they look through it, they're learning a ton of stuff.

Dwarkesh

是的。

Yeah.

Andrej

所以它给你很多我所谓的「每秒恍然大悟」,也就是每秒的理解量。这就是我想要的:大量的每秒恍然大悟。对我来说,这是一个技术问题:如何构建这些通往知识的阶梯。我几乎觉得 Eureka 可能和前沿实验室的一些工作没什么不同,因为我想弄清楚如何高效地构建这些阶梯,让人们永远不会卡住,所有内容既不太难也不太简单,你正好有合适的材料来真正进步。

And so it's giving you a lot of what I call Eurekas per second, which is understanding per second. That's what I want: lots of Eurekas per second. To me, this is a technical problem of how do we build these ramps to knowledge. I almost think of Eureka as maybe not that different from some of the work going on at frontier labs, because I want to figure out how to build these ramps very efficiently so that people are never stuck, and everything is always not too hard or too trivial, and you have just the right material to actually progress.

Dwarkesh

是的。所以你想象在短期内,与其让导师来探查你的理解,如果你有足够的自我意识来自我探查,你就永远不会卡住。你可以在与助教交谈、与 LLM 交谈以及查看参考实现之间找到正确答案。听起来自动化和 AI 目前还不是重要因素。这里的巨大优势是你将 AI 解释编码在课程原始材料中的能力。这基本上就是课程的本质。

Yeah. So you're imagining in the short term that instead of a tutor being able to probe your understanding, if you have enough self-awareness to probe yourself, you're never going to be stuck. You can find the right answer between talking to the TA or talking to an LLM and looking at the reference implementation. It sounds like automation or AI is actually not a significant factor so far. The big alpha here is your ability to explain AI codified in the source material of the class. That's fundamentally what the course is.

Andrej

我认为你必须始终根据行业现有的能力来校准。很多人会去追求直接问 ChatGPT 之类的,但我认为现在,如果你去问 ChatGPT「教我 AI」,它会给你一些垃圾内容。AI 现在永远写不出 nano chat,但 nano chat 是一个非常有用的中间点。所以我仍然在与 AI 合作创建所有这些材料,所以 AI 本质上仍然非常有帮助。

I think you always have to be calibrated to what capability exists in the industry. A lot of people are going to pursue just asking ChatGPT etc., but I think right now, if you go to ChatGPT and say 'teach me AI', it's going to give you some slop. AI is never going to write nano chat right now, but nano chat is a really useful intermediate point. So I'm still collaborating with AI to create all this material, so AI is still fundamentally very helpful.

Dwarkesh

早些时候,我在斯坦福创建了 CS231N,这是早期——实际上,我认为是斯坦福的第一门深度学习课程——变得非常受欢迎。现在构建 231N 和 L101N 的差异非常明显,因为我感觉现在的 LLM 给了我很大助力,但我仍然深度参与其中。所以它们帮我构建所有材料;我进展快得多,它们做很多无聊的事情等等。所以我感觉我开发课程的速度快多了,而且课程中融入了 LLM,但还没有到我能创造性地生成内容的程度。我仍然需要做那部分工作。

Earlier on, I built CS231N at Stanford, which was one of the earlier—actually, I think it was the first deep learning class at Stanford—which became very popular. The difference in building out 231N and L101N now is quite stark, because I feel really empowered by the LLMs as they exist right now, but I'm very much in the loop. So they're helping me build all the materials; I go much faster, they're doing a lot of the boring stuff, etc. So I feel like I'm developing the course much faster, and there's LLM infused in it, but it's not yet at a place where I can creatively create the content. I'm still there to do that.

Dwarkesh

所以棘手之处在于始终根据现有情况校准自己。当你想象几年后通过 Eureka 能获得什么时,似乎最大的瓶颈将是找到各个领域的专家,他们能将自己的理解转化为这些阶梯。

So the trickiness is always calibrating yourself to what exists. When you imagine what is available through Eureka in a couple of years, it seems like the big bottleneck is going to be finding corps in field after field who can convert their understanding into these ramps.

Andrej

所以我认为它会随着时间变化。现在,会是聘请教师与 AI 以及一个团队合作,来构建最先进的课程。

So I think it would change over time. Right now, it would be hiring faculty to help work hand-in-hand with AI and a team of people probably to build state-of-the-art courses.

Dwarkesh

嗯。

Mhm.

Andrej

然后我认为随着时间的推移,也许一些助教实际上可以变成 AI,因为有些助教,比如,你只需拿所有课程材料,然后我认为你可以为学生提供一个非常好的自动化助教,当他们有更基础的问题时。但我认为你仍然需要教师来负责课程的整体架构并确保它合适。所以我看到了一个演变过程,也许在未来的某个时候,我甚至没那么有用,AI 能比我做大部分设计做得更好。但我仍然认为这需要一些时间来实现。

And then I think over time, maybe some of the TAs can actually become AIs, because some of the TAs, like, okay, you just take all the course materials and then I think you could serve a very good automated TA for the student when they have more basic questions. But I think you'll need faculty for the overall architecture of a course and making sure that it fits. So I kind of see a progression of how this will evolve, and maybe at some future point, I'm not even that useful and AI is doing most of the design much better than I could. But I still think that's going to take some time to play out.

Dwarkesh

但你是否想象其他领域有专长的人也能贡献课程?还是你觉得,鉴于你对如何教学的理解,由你来设计内容对愿景至关重要?我不知道,萨尔·汗在可汗学院旁白所有视频。你想象的是类似的东西吗?

But are you imagining that people who have expertise in other fields are then contributing courses? Or do you feel like it's actually quite essential to the vision that, given your understanding of how you want to teach, you are the one designing the content? I don't know, Sal Khan is narrating all the videos on Khan Academy. Are you imagining something like that?

Andrej

不,我会聘请教师,因为有些领域我不是专家,我认为这是最终为学生提供最先进体验的唯一途径。所以是的,我确实预计会聘请教师,但我可能会在 AI 领域再待一段时间。但我对当前能力有一些我认为比人们预期的更传统的想法。当我建造星际舰队学院时,我确实可能想象一个实体机构,然后在它下面一层,是一个数字产品,它不会提供与全日制实体学习相同的顶级体验,在那里我们从头到尾学习材料并确保你理解。那是实体产品。数字产品是互联网上的一堆东西,可能还有一些 LLM 助手,它更花哨一些,处于较低层级,但至少对八十亿人开放。

No, I will hire faculty because there are domains in which I'm not an expert, and I think that's the only way to offer the state-of-the-art experience for the student ultimately. So yeah, I do expect that I would hire faculty, but I will probably stick around in AI for some time. But I do have something I think more conventional in mind for the current capability than what people would probably anticipate. And when I'm building Starfleet Academy, I do probably imagine a physical institution, and maybe a tier below that, a digital offering that is not the same state-of-the-art experience you would get when someone comes in physically full-time and we work through material from start to end and make sure you understand it. That's the physical offering. The digital offering is a bunch of stuff on the internet and maybe some LLM assistant, and it's a bit more gimmicky in a tier below, but at least it's accessible to eight billion people.

Dwarkesh

是的,我认为你基本上是在根据当今可用的工具,从第一性原理出发重新发明大学,并且只挑选那些有动力和兴趣真正投入学习材料的人。

Yeah, I think you're basically inventing college from first principles for the tools that are available today, and just selecting for people who have the motivation and the interest to really engage with material.

Andrej

是的。我认为不仅需要教育,还需要大量再教育。我很乐意在这方面提供帮助,因为我认为工作可能会发生很大变化。例如,今天很多人都在尝试专门提升 AI 技能。

Yeah. And I think there's going to have to be a lot of not just education but also re-education. I would love to help out there because I think the jobs will probably change quite a bit. For example, today a lot of people are trying to upskill in AI specifically.

前 AGI 与后 AGI 教育 On pre-AGI vs post-AGI education

Andrej

我认为在这方面,这是一门非常好的课程。从动机来看,AGI 之前很简单,因为人们想赚钱,而这就是当今行业赚钱的方式。AGI 之后就有趣多了,因为如果一切都自动化了,没人有事可做,那为什么还要上学呢?我常说,AGI 前的教育是有用的,AGI 后的教育是好玩的。

I think it's a really good course to teach in this respect. Motivation-wise, pre-AGI is very simple to solve because people want to make money, and this is how you make money in the industry today. Post-AGI, it's a lot more interesting because if everything is automated and there's nothing to do for anyone, why would anyone go to school? I often say that pre-AGI education is useful, post-AGI education is fun.

Dwarkesh

嗯。

Yeah.

Andrej

类似地,今天人们去健身房。我们不再需要体力来搬运重物,因为机器代劳了。但他们仍然去健身房。为什么?因为好玩、健康,而且有六块腹肌看起来很性感。这在深层心理和进化意义上对人有吸引力。所以我认为教育也会如此:你会像去健身房一样去上学。现在,没多少人学习,因为学习很难,你会被材料难倒。但这是一个可以解决的技术问题,就像我学韩语时我的导师做的那样。这是可行的,可以构建的。应该有人来构建它,它会让你学习任何东西都变得轻而易举、令人向往。人们会为了好玩而学习,因为它太简单了。如果我有一个那样的导师来教任何知识,学习任何东西都会容易得多,人们会出于和去健身房同样的理由去学习。

In a similar way, people go to the gym today. We don't need their physical strength to manipulate heavy objects because we have machines that do that. They still go to the gym. Why? Because it's fun, it's healthy, and you look hot when you have a six-pack. It's attractive for people in a deep psychological, evolutionary sense. So I think education will play out the same way: you'll go to school like you go to the gym. Right now, not many people learn because learning is hard; you bounce off material. But it's a technical problem to solve, like what my tutor did for me when I was learning Korean. It's tractable and buildable. Someone should build it, and it will make learning anything trivial and desirable. People will do it for fun because it's trivial. If I had a tutor like that for any arbitrary piece of knowledge, it would be so much easier to learn anything, and people will do it for the same reasons they go to the gym.

Dwarkesh

这听起来和把 AGI 后的教育当作娱乐或自我提升不同。你还有一个愿景,认为这种教育与让人类保持对 AI 的控制有关。这两者听起来不同。是有些人觉得好玩,有些人觉得赋能吗?你怎么看?

That sounds different from using this post-AGI as entertainment or self-betterment. You also had a vision that this education is relevant to keeping humanity in control of AI. They sound different. Is it entertaining for some people and empowerment for others? How do you think about that?

Andrej

我确实觉得从长远来看,这有点像一场必输的游戏,比业内大多数人想的还要长远。我认为人可以走得很远,我们只是刚刚触及了人类能力的表面,只是因为人们会被太容易或太难的材料难倒。我实际上觉得人们能够走得更远——任何人都会说五种语言,为什么不呢,这太简单了。任何人都知道本科的基础课程等等。

I do feel like eventually it's a bit of a losing game in the long term, longer than maybe most people in the industry think. I think people can go so far, and we've barely scratched the surface of what a person can do, just because people bounce off material that's too easy or too hard. I actually feel that people will be able to go much further—anyone speaks five languages because why not, it's so trivial. Anyone knows the basic curriculum of undergrad, etc.

Dwarkesh

现在我理解了你的愿景,这非常有趣。我认为它在健身文化中有一个完美的类比。我不认为 100 年前有人会练出肌肉;没人能随便卧推两片或三片杠铃片。现在这很常见。你想象的是在众多不同领域的学习中类似的事情,更密集、更深入、更快。

Now that I'm understanding the vision, that's very interesting. I think it has a perfect analog in gym culture. I don't think 100 years ago anybody would be ripped; nobody would be able to spontaneously bench two plates or three plates. It's actually very common now. You're imagining similar things for learning across many different domains, much more intensely, deeply, faster.

Andrej

是的,完全正确。我有点隐性地押注于人性的永恒性。我认为做所有这些事情会是令人向往的,人们会仰慕它,就像几千年来一样。有一些历史证据:如果你看看贵族或古希腊,每当有某种意义上的 AGI 后的小环境时,人们都会花大量时间在身体或认知上蓬勃发展。所以我对前景感到乐观。如果这是错的,我们最终走向《机器人总动员》或《蠢蛋进化论》的未来,那么即使有戴森球我也不在乎——这是一个糟糕的结果。我真的很在乎人类。每个人都必须在某种意义上成为超人。

Yeah, exactly. I am betting a little bit implicitly on the timelessness of human nature. I think it will be desirable to do all these things, and people will look up to it, as they have for millennia. There's some historical evidence: if you look at aristocrats or ancient Greece, whenever there were little pocket environments that were post-AGI in a certain sense, people spent a lot of time flourishing physically or cognitively. So I feel okay about the prospects. If this is false and we end up in a Wall-E or Idiocracy future, then I don't even care if there are Dyson spheres—this is a terrible outcome. I really do care about humanity. Everyone has to be superhuman in a certain sense.

Dwarkesh

这仍然是一个无法让你凭借自己的劳动或认知来改变技术轨迹或影响决策的世界。也许你可以影响决策,因为 AI 需要你的批准,但这不是因为我发明了什么或提出了新设计。

It's still a world in which that is not enabling you to transform the trajectory of technology or influence decisions by your own labor or cognition alone. Maybe you can influence decisions because the AI is for your approval, but it's not because I've invented something or come up with a new design.

Andrej

也许吧。我认为会有一个过渡期,如果我们真正理解很多东西,我们就能参与其中并推动进步。长期来看,这可能会消失。但也许它会变成一项运动。现在,力量举运动员在这方面走向极端。那么在认知时代,力量举是什么?也许是人们试图把知识变成奥林匹克。如果你有一个完美的 AI 导师,也许你可以走得非常远。我几乎觉得今天的天才们也只是刚刚触及人类心智能力的表面。

Maybe. I think there will be a transitional period where we can be in the loop and advance things if we actually understand a lot. Long term, that probably goes away. But maybe it becomes a sport. Right now, powerlifters go extreme in that direction. So what is powerlifting in a cognitive era? Maybe it's people trying to make Olympics out of knowing stuff. If you have a perfect AI tutor, maybe you can get extremely far. I almost feel like the geniuses of today are barely scratching the surface of what a human mind can do.

Dwarkesh

嗯。我喜欢这个愿景。

Yeah. I love this vision.

学习与动机 On learning and motivation

Dwarkesh

我也觉得,跟你产品市场最匹配的人就是我,因为我的工作每周都要学习不同的主题。如果你能做到,我会非常兴奋……

I also feel like the person you have the most product-market fit with is me, because my job involves having to learn different subjects every week. I am very excited if you can...

Andrej

在这方面我也差不多。很多人讨厌学校,想逃离。我其实很喜欢学校,热爱学习。我想一直待在学校,直到读完博士,他们不让我再待下去了,我才去了工业界。但大致来说,我热爱学习,哪怕只是为了学习本身;我也热爱学习,因为它是一种赋能,让人变得有用和高效。

I'm similar for that matter. I mean, a lot of people hate school and want to get out of it. I actually really liked school. I loved learning things. I wanted to stay in school. I stayed all the way until PhD, and then they wouldn't let me stay longer, so I went to industry. But roughly speaking, I love learning even for the sake of learning, but I also love learning because it's a form of empowerment and being useful and productive.

Dwarkesh

我觉得你还提出了一个微妙的观点,所以我想明确一下:到目前为止,在线课程并没有让每个人都无所不知。我认为它们太依赖动机了,因为没有明确的入门路径,而且很容易卡住。如果你有一个真正优秀的人类导师,从动机角度来看,那将是一个巨大的突破。

I think you also made a subtle point, so just to spell it out: what's happened so far with online courses is that they haven't enabled every single human to know everything. I think they're just so motivation-laden because there are no obvious on-ramps, and it's so easy to get stuck. If you had instead a really good human tutor, it would be such an unlock from a motivation perspective.

Andrej

我也这么认为,因为从材料中跳出来感觉不好。你把时间花在没结果的事情上,或者因为内容太简单或太难而感到无聊,都会得到负面奖励。当你真正做对了,学习的感觉是很好的。

I think so, because it feels bad to bounce from material. You get negative reward from sinking time into something that doesn't pan out, or being completely bored because what you're getting is too easy or too hard. When you actually do it properly, learning feels good.

Dwarkesh

是的。我认为要达到那个目标是一个技术问题。在一段时间内,会是 AI 加人类协作。也许某个时候就只是 AI 了。我不知道。

Yeah. And I think it's a technical problem to get there. For a while it's going to be AI plus human collaboration. At some point maybe it's just AI. I don't know.

给教育工作者的建议 Advice for educators

Dwarkesh

我能问一些关于教学的问题吗?如果你必须给另一个你感兴趣领域的教育者提建议,让他们制作你那种 YouTube 教程,特别是那些无法通过编程测试技术理解的领域,你会给他们什么建议?

Can I ask some questions about teaching? If you had to give advice to another educator in another field that you're curious about, to make the kinds of YouTube tutorials you've made, maybe especially interesting to talk about domains where you can't test somebody's technical understanding by having them code something up. What advice would you give them?

Andrej

这是一个相当广泛的话题。我觉得有大概 10 到 20 个技巧,我半自觉地会用到。在高层次上,我总是试图……这很大程度上来自我的物理背景。我非常喜欢我的物理背景。我有一整套关于为什么每个人都应该在早期学校教育中学习物理的论述,因为早期学校教育不是为了积累知识或为以后工业界的任务记忆,而是为了启动大脑。我认为物理是启动大脑的最佳方式。建立模型和抽象的概念,理解存在一个一阶近似描述系统的大部分,但还有二阶、三阶、四阶项可能存在也可能不存在。你观察一个非常嘈杂的系统,但你可以抽象出一些基本频率。当物理学家走进教室说「假设有一头球形奶牛」,大家都笑了,但实际上这是一种非常 brilliant 的思维,在行业中非常通用。

That's a pretty broad topic. I feel like there are 10 to 20 tips and tricks that I semi-consciously probably do. On a high level, I always try to... A lot of this comes from my physics background. I really enjoyed my physics background. I have a whole rant on how everyone should learn physics in early school education, because early school education is not about accumulating knowledge or memory for tasks later in industry. It's about booting up a brain. I think physics uniquely boots up the brain the best. The idea of building models and abstractions, understanding that there's a first-order approximation that describes most of the system, but then there are second-order, third-order, fourth-order terms that may or may not be present. The idea that you're observing a very noisy system, but there are fundamental frequencies you can abstract away. When a physicist walks into class and says, 'Assume there's a spherical cow,' everyone laughs, but actually this is brilliant thinking that's very generalizable across the industry.

Dwarkesh

是的,奶牛在很多方面可以近似为球体。有一本很好的书,比如《规模》,基本上是一位物理学家在谈论生物学。你可以得到很多有趣的近似,画出动物的缩放定律,观察它们的心跳,它们实际上与动物的大小一致。你可以把动物看作一个体积,然后推导散热,因为散热随表面积(平方)增长,但产热随体积(立方)增长。

Yeah, a cow can be approximated as a sphere in a bunch of ways. There's a really good book, for example, 'Scale', basically from a physicist talking about biology. You can get a lot of interesting approximations and chart scaling laws of animals, look at their heartbeats, and they actually line up with the size of the animal. You can talk about an animal as a volume, and you can derive heat dissipation because heat dissipation grows as surface area (square), but heat generation grows as volume (cube).

Andrej

所以我觉得物理学家拥有所有正确的认知工具来应对世界上的问题解决。因为那种训练,我总是试图找到一切的一阶项或二阶项。当我观察一个系统或一团想法时,我试图找到真正重要的东西,什么是一阶成分,如何简化它,展示它的作用,然后再添加其他项。

So I feel like physicists have all the right cognitive tools to approach problem solving in the world. Because of that training, I always try to find the first-order terms or the second-order terms of everything. When I'm observing a system or a tangle of ideas, I try to find what actually matters, what is the first-order component, how can I simplify it, show it in action, and then tack on the other terms.

Dwarkesh

也许我某个仓库里的一个例子能很好地说明这一点,叫做 micrograd。它用 100 行代码展示了反向传播。你可以用加法和乘法等简单操作创建神经网络,构建计算图,进行前向传播和后向传播得到梯度。这是所有神经网络学习的核心。Micrograd 是 100 行可解释的 Python 代码,可以对任意神经网络进行前向和后向计算,但效率不高。其他一切都只是效率问题。

Maybe an example from one of my repos that illustrates it well is called micrograd. It's 100 lines of code that shows backpropagation. You can create neural networks out of simple operations like plus and times, build up a computational graph, and do a forward pass and a backward pass to get the gradients. This is at the heart of all neural network learning. Micrograd is 100 lines of interpretable Python code that can do forward and backward for arbitrary neural networks, but not efficiently. Everything else is just efficiency.

Andrej

是的。其他一切都只是效率。效率方面有大量工作要做:你需要张量,布局它们,设置步幅,确保内核正确协调内存移动。大致来说,这些都只是效率。但神经网络训练的核心智力部分是 micrograd。它只有 100 行。你可以轻松理解它。它是链式法则的递归应用,用于推导梯度,从而优化任意可微函数。

Yeah. Everything else is efficiency. There's a huge amount of work to do efficiency: you need your tensors, lay them out, stride them, make sure your kernels orchestrate memory movement correctly. It's all just efficiency, roughly speaking. But the core intellectual piece of neural network training is micrograd. It's 100 lines. You can easily understand it. It's a recursive application of chain rule to derive the gradient, which allows you to optimize any arbitrary differentiable function.

教学与解释复杂概念 On teaching and explaining complex ideas

Dwarkesh

所以这关乎找到那些更小的概念,把它们端上桌,去发现它们。我觉得教育是最有趣的事,因为你有一团乱麻般的理解,你要把它铺开,形成一条坡道,让每件事只依赖于前面的事。我觉得这种知识的解构作为一项认知任务非常有趣。

So it's about finding the smaller terms and serving them on a platter, discovering them. I feel like education is the most intellectually interesting thing because you have a tangle of understanding and you're trying to lay it out in a way that creates a ramp where everything only depends on the thing before it. I find that untangling of knowledge is just so intellectually interesting as a cognitive task.

Andrej

是的。

Dwarkesh

我个人很喜欢做这件事。我就是着迷于以某种方式把东西摆出来。也许这对我也有帮助。

And I love doing it personally. I just find fascination with trying to lay things out in a certain way. And maybe that helps me.

Andrej

这也让学习体验更有动力。你的 Transformer 教程从二元模型开始。它就像一个查找表:这是当前词,或者这是前一个词,这是下一个词。它实际上就是一个查找表。

It also makes the learning experience so much more motivated. Your tutorial on the transformer begins with a bigram model. It's literally like a lookup table: here's the current word, or here's the previous word, here's the next word. And it's literally just a lookup table.

Dwarkesh

是的。这就是本质。我是说,这方法太棒了:从查找表开始,然后到 Transformer,每一步都有动机——为什么要加这个?为什么要加下一个?你没法死记硬背注意力公式,但你能理解为什么每个部分都相关,它解决了什么问题。

Yeah. That's the essence of it. I mean, such a brilliant way: start with a lookup table, then go to a transformer, and each piece is motivated—why would you add that? Why would you add the next thing? You couldn't memorize a sort of attention formula, but it's like having an understanding of why every single piece is relevant, what problem it solves.

Andrej

是的,你在呈现解决方案之前先呈现痛点。这多聪明啊?你想带着学生经历这个过程。还有很多这样的小细节,我觉得让内容变得更好、更吸引人,并且不断引导学生。你会怎么解决?我不会在你尝试之前就给出解决方案。那太浪费了。这有点——我不想说脏话——但在我给你机会自己尝试之前就给你解决方案,这有点不地道。

Yeah, you're presenting the pain before you present the solution. How clever is that? You want to take the student through that progression. There are a lot of other small things like that that I think make it nice and engaging, and always prompting the student. How would you solve this? I'm not going to present a solution before you're going to guess. That would be wasteful. That's a little bit of a—I don't want to swear—but it's a dick move to present you with the solution before I give you a shot to try to come up with it yourself.

Dwarkesh

对。

Right.

Andrej

因为如果你自己尝试,你会更好地理解行动空间是什么。

Because if you try to come up with it yourself, you get a better understanding of what the action space is.

Dwarkesh

是的。然后目标是什么?为什么只有这个行动能实现那个目标?

Yeah. And then what is the objective? Why does only this action fulfill that objective?

Andrej

是的。你有机会自己尝试,当我给出解决方案时,你会更欣赏它。它最大化每个新事实所增加的知识量。

Yeah. Well, you have a chance to try yourself, and you have an appreciation when I give you the solution. It maximizes the amount of knowledge per new fact added.

Dwarkesh

没错。

That's right.

Andrej

为什么你认为,通常情况下,真正的领域专家往往不擅长向初学者解释?

Why do you think by default people who are genuine experts in their field are often bad at explaining it to somebody ramping up?

Dwarkesh

嗯,这是知识和专业知识的诅咒。这是一个真实的现象,我自己也深受其害,尽管我努力避免。你会把某些事情视为理所当然,无法设身处地为初学者着想。这很普遍,我也一样。

Well, it's the curse of knowledge and expertise. This is a real phenomenon, and I actually suffered from it myself as much as I try not to. You take certain things for granted and you can't put yourself in the shoes of people who are just starting out. This is pervasive and happens to me as well.

Andrej

有一件事我觉得特别有用:举个例子,最近有人想给我看一篇生物学论文,我立刻就有很多糟糕的问题。所以我用 ChatGPT,把论文放在上下文窗口里提问,它解决了一些简单的问题。然后我把这个对话分享给了写那篇论文或做那个工作的人。我几乎觉得,如果他们能看到我那些愚蠢的问题,可能会帮助他们将来更好地解释。所以,比如对于我的材料,如果人们分享他们和 ChatGPT 关于我创作内容的愚蠢对话,我会很高兴,因为这真的能帮我再次站在初学者的角度。

One thing that I actually think is extremely helpful: as an example, someone was trying to show me a paper in biology recently, and I just had instantly so many terrible questions. So what I did was I used ChatGPT to ask the questions with the paper in the context window, and then it worked through some of the simple things. Then I actually shared the thread to the person who wrote that paper or worked on that work. I almost feel like it was like, if they can see the dumb questions I had, it might help them explain it better in the future. So for example, for my material, I would love if people shared their dumb conversations with ChatGPT about the stuff I've created, because it really helps me put myself again in the shoes of someone who's starting out.

Dwarkesh

另一个非常有效的技巧:如果有人写了一篇论文、一篇博客或一个公告,100%的情况下,他们午餐时向你解释的叙述或转录不仅更容易理解,而且实际上更准确、更科学,因为人们倾向于用最抽象、最充满术语的方式解释,并在解释核心思想前先铺垫四段。但一对一交流时,有种东西迫使你直接说出重点。

Another trick like that that works astoundingly well: if somebody writes a paper or a blog post or an announcement, it is in 100% of cases true that just the narration or the transcription of how they would explain it to you over lunch is way more not only understandable, but actually also more accurate and scientific in the sense that people have a bias to explain things in the most abstract, jargon-filled way possible and to clear their throat for four paragraphs before they explain the central idea. But there's something about communicating one-on-one with a person which compels you to just say the thing.

Andrej

直接说重点。

Just say the thing.

Dwarkesh

是的。实际上,我看到了那条推文。我觉得非常好。我分享给了很多人。我觉得它真的很好。我注意到这种情况很多次。也许最突出的例子是,在我读博做研究的时候,你读某人的论文,努力理解它在做什么。然后你在会议后和他们喝啤酒时问他们:「这篇论文是关于什么的?」他们就会用三句话完美地概括论文的精髓,让你完全理解,而你甚至还没读论文。

Yeah. Actually, I saw that tweet. I thought it was really good. I shared it with a bunch of people. I think it was really good. And I noticed this many times. Maybe the most prominent example is, back in my PhD days doing research, you read someone's paper and you work to understand what it's doing. Then you catch them having beers at the conference later and you ask, 'So what is the paper about?' And they will just tell you these three sentences that perfectly capture the essence of that paper and totally give you the idea, and you didn't have to read the paper yet.

Andrej

是的。只有当你坐在桌边喝着啤酒时,他们才会说:「哦,这篇论文就是:你拿这个想法,拿那个想法,然后做这个实验,试试那个东西。」他们能用对话的方式完美地表达出来。为什么那不是实际的论文呢?

Yeah. And it's only when you're sitting at the table with a beer that they say, 'Oh yeah, the paper is just: you take this idea, you take that idea, and you try this experiment, and you try out this thing.' They have a way of just putting it conversationally and just perfectly. Why isn't that the actual paper?

Dwarkesh

正是。

Exactly.

Andrej

这是从试图解释想法的人如何更好地表达的角度出发的。作为学生,你对其他学生有什么建议?如果你没有一个像 Karpathy 那样为你讲解想法的人,如果你在读一篇论文或一本书,你会用什么策略来学习你感兴趣但并非专长的领域的内容?

This is coming from the perspective of how someone trying to explain an idea should formulate it better. What is your advice as a student to other students? If you don't have a Karpathy who is doing the exposition of an idea, if you're reading a paper or a book, what strategies do you employ to learn material you're interested in in fields you're not an expert in?

Dwarkesh

老实说,我并不觉得自己有什么独特的技巧。基本上,这是一个痛苦的过程。但你知道,重写一遍。

I don't actually know that I have unique tips and tricks to be honest. Basically, it's a kind of painful process. But you know, redraft one.

学习策略 On learning strategies

Dwarkesh

我认为有一件事一直对我帮助很大,那就是按需学习。我发过一条关于这个的小推文。在深度上学习,你需要一些交替。按需学习时,你试图完成一个项目,并从中获得回报。而广度学习就是随便探索,学校很多都是这样:「相信我,你以后会需要这个的。」我喜欢那种真正能从做事中获得回报的学习方式,也就是按需学习。

I think one thing that has always helped me quite a bit is learning on demand. I had a small tweet about this. Learning depthwise, you need a bit of alternation. On demand, you're trying to achieve a project that you'll get a reward from. And learning breadthwise is just exploring whatever, which is a lot of what school does: 'Trust me, you'll need this later.' I love the kind of learning where you actually get a reward from doing something and you're learning on demand.

Andrej

我发现另一件非常有帮助的事就是向别人解释事情。这是更深入学习的好方法。我经常这样。我意识到如果我不真正理解某件事,我就无法解释它。承认这一点很烦人,但之后你可以回去确保自己理解了,填补理解上的空白。它迫使你去调和这些空白。我喜欢重新解释事情,我认为人们应该多这样做。这迫使你运用知识,确保你明白自己在说什么。

The other thing I've found extremely helpful is explaining things to people. It's a beautiful way to learn something more deeply. This happens to me all the time. I realize if I don't really understand something, I can't explain it. It's annoying to come to terms with that, but then you can go back and make sure you understood it, filling gaps in your understanding. It forces you to reconcile them. I love to re-explain things, and I think people should do that more. It forces you to manipulate the knowledge and make sure you know what you're talking about.

Dwarkesh

我认为这是一个很好的结束语。

I think that's an excellent note to close on.

Andrej

是的。

Yeah.

Dwarkesh

Andre,太棒了。

Andre, that was great.

Andrej

是的,谢谢。谢谢。不着急。

Yeah, thank you. Thanks. Take your time.

Dwarkesh

大家好,希望你们喜欢这一集。如果喜欢,最有帮助的事情就是分享给其他可能喜欢的人。如果你在收听的平台上留下评分或评论,也很有帮助。如果你有兴趣赞助播客,可以联系 dwarcash.com/advertise。否则,我们下期再见。

Hey everybody, I hope you enjoyed that episode. If you did, the most helpful thing you can do is just share it with other people who you think might enjoy it. It's also helpful if you leave a rating or a comment on whatever platform you're listening on. If you're interested in sponsoring the podcast, you can reach out at dwarcash.com/advertise. Otherwise, I'll see you on the next one.

互动版:逐字朗读 + 针对本期提问 →