Jeremy Howard: The Art of Understanding AI
打开互动全文版(中英对照 + 朗读 + 问答)→深度学习先驱 Jeremy Howard 探讨 AI 编程中的控制幻觉、交互式学习的重要性,以及他 20 年来致力于阻止不人道工作方式的使命。
Deep learning pioneer Jeremy Howard discusses the illusion of control in AI coding, the importance of interactive learning, and his 20-year mission to stop inhumane work practices.
这真的让我感到厌恶。我真的认为这很不人道。我的使命和 20 年来一样,就是阻止人们这样工作。
It literally disgusts me. Like I literally think it's inhumane. My mission remains the same as it has been for like 20 years, which is to stop people working like this.
杰瑞米·霍华德,深度学习先驱,Kaggle 大师。他大力倡导通过交互循环、笔记本、涟漪效应——不断试探问题直到它给出反馈——来真正理解我们构建的东西。他认为这才是真正洞察发生的地方。有趣的是,他们俩都是对的。LLM 在角色扮演理解事物。它们假装理解。实际上没有人能比以前多创造 50 倍的高质量软件。嗯,我们刚做了一项研究,人们实际交付的东西只有微小的增长。基于 AI 的编程就像老虎机,你有一种控制的错觉。你知道,你可以精心设计你的提示词、MCP 列表、技能等等,但最后你拉下杠杆,对吧?这是一段没人理解的代码。
Jeremy Howard, a deep learning pioneer, a Kaggle grandmaster. He is a huge advocate for actually understanding what we are building through an interactive loop, a notebook, a ripple, the act of poking at a problem until it pushes back. He argues this is where the real insight happens. And the funny thing is they're both right. LLMs cosplay understanding things. They pretend to understand things. No one's actually creating 50 times more high-quality software than they were before. Um, so we've actually just done a study of this and there's a tiny uptick in what people are actually shipping. The thing about AI based coding is that it's like a slot machine in that you have an illusion of control. You know, you could get to craft your prompt and your list of MCPs and your skills and whatever but then in the end you pull the lever, right? Here's a piece of code that no one understands.
是的。
Yeah.
我要把公司产品押注在上面吗?答案是不知道,因为没人遇到过这种情况。它们真的很不擅长软件工程。而且我认为这可能永远都是事实。人类如果能实时操作计算机内的对象、研究它们、移动它们、组合它们,就能用计算机做更多事情。无论你听谁说,费曼还是其他人,你总会听到伟大的科学家们如何通过构建心智模型来建立更深层的直觉,这些模型是通过与所学事物互动逐渐获得的。机器可以通过深度学习模型查看海量文本语料的统计相关性,来构建关于世界是什么以及如何运作的有效抽象层次结构。这就是我的前提。
And am I going to bet my company's product on it? And the answer is I don't know because no one's been in this situation. They're really bad at software engineering. And I think that's possibly always going to be true. The idea that a human can do a lot more with a computer when the human can manipulate the objects inside that computer in real time and study them and move them around and combine them together. Whoever you listen to, whether it be Feynman or whatever, you always hear from the great scientists how they build deeper intuition by building mental models which they get over time by interacting with the things that they're learning about. A machine could kind of build an effective hierarchy of abstractions about what the world is and how it works entirely through looking at the statistical correlations of a huge corpus of text using a deep learning model. That was my premise.
本期视频由 Nvidia GTC 赞助。大会将于 3 月 16 日至 19 日在圣何塞举行,并免费在线直播。今年的关键主题包括智能体式 AI 与推理、高性能推理与训练、开放模型,以及物理 AI 与机器人。我对 DJX Spark 非常兴奋。我已经排队等了一年多。这是一台个人超级计算机,大小与 Mac Mini 相当。顺便说一句,它是 MacBook Pro 的完美搭配。你可以用这东西微调一个 700 亿参数的语言模型。我要免费送出一台。你只需注册会议并通过描述中的链接参加一场会议即可。至于会议,我有兴趣参加 Ammon Sang 的演讲。他是 Cursor 的 CTO,他的会议主题是“带上下文的代码:构建真正理解你代码库的智能体式 IDE”。当然,Jensen 的主题演讲在 3 月 16 日。他说他将推出一款震惊世界的新芯片。他们的下一代架构 Vera Rubin 已经全面投产。有猜测我们甚至可能提前一睹他们的新 Feynman 架构。所以别忘了,链接在描述中。如果你虚拟参会,完全免费。不要错过。
This video is brought to you by Nvidia GTC. It's running March 16th until the 19th in San Jose and streaming free online. The key topics this year are agentic AI and reasoning, high performance inference and training, open models, and physical AI and robotics. I'm so excited about the DJX Spark. I've been on the waiting list for over a year now. It's a personal supercomputer that is about the size of a Mac Mini. It's the perfect adornment to a MacBook Pro, by the way. And you can fine-tune a 70 billion parameter language model with one of these things. And I'm giving one away for free. All you have to do is sign up to the conference and attend one of the sessions using the link in the description. As for the sessions, I'm interested in attending Ammon Sang's talk. So, he's the CTO of Cursor and his session is code with context. Build an agentic IDE that truly understands your codebase. Now, obviously, Jensen's keynote is on March 16th. He said he's going to unveil a new chip that will surprise the world. Their next generation architecture, Vera Rubin, is already in full production. And there's speculation we might even get an early glimpse of their new Feynman architecture. So don't forget folks, the link is in the description. If you're attending virtually, it's completely free. Don't miss it.
杰瑞米·霍华德,欢迎来到 MLST。
Jeremy Howard, welcome to MLST.
嗯,欢迎来到我家。谢谢你来。
I mean, welcome to my home. Thanks for coming.
是的。那么,我们现在在哪里?
Yeah. Well, where are we now?
我们在昆士兰东南部美丽的莫顿湾。我们在海边,在我家后院。
We are in beautiful Morton Bay in southeast Queensland. We are by the sea in my backyard.
天气没让人失望。
The weather didn't disappoint.
确实没有。通常不会,但如果你昨天在这里,那就大不一样了。
It certainly didn't. It doesn't often, but if you were here yesterday, it would have been very different.
嗯,我不知道从何说起。我大概从 2017、2018 年就是你的超级粉丝了。当然,你有著名的 ULMFiT 论文。我在微软时,记得做过一个关于它的演示,因为实际上——我的意思是现在我们理所当然地认为我们在文本语料上微调语言模型,然后继续训练并专门化它们。但显然这并非公认的智慧。
Well, I don't know where to start. So, I've been a huge fan probably since about 2017, 2018. Of course, you had the famous ULMFiT paper. And when I was at Microsoft, I remember doing a presentation about that because it was actually—I mean now we take it for granted that we fine-tune language models on a corpus of text and then we kind of continue to train them and specialize them. But apparently this was not received wisdom.
不,这是第一次发生。是的,算是第一或第二次。所以,Quoc Le 和 Andrew Dai 几年前做过一些工作,但他们错过了关键点:你预训练的东西必须是一个通用语料。
No, this was the first time it happened. Yeah, kind of the first or second. So, Quoc Le and Andrew Dai had done something a few years ago, but they had missed the key point, which is the thing you pre-train on has to be a general purpose corpus.
所以,没人完全意识到这个关键点。也许我有点幸运,我的背景是哲学和认知科学。所以我花了几十年思考这个问题。
So, no one quite realized this key thing. And maybe I had a bit of fortune here that my background was in philosophy and cognitive science. And so, I'd spent some decades thinking about this.
ULMFiT 的技术架构。简单描述一下。
The technical architecture of ULMFiT. Just sketch that out.
我是正则化的超级粉丝。我热衷于采用一个极其灵活的模型,然后通过添加正则化而不是减小架构规模来使其更受约束。即使在当时,这也极具争议。但这绝非我们独有的见解。Stephen Merity 所做的是,他采用了 LSTM 的极端灵活性——一种经典的带状态循环神经网络,如今事物正逐渐回归——并添加了五种不同类型的正则化。他添加了你所能想象到的每一种正则化。然后那成了我的起点:好吧,我现在有一个极其灵活的深度学习模型,它可以像我想要的那样强大,也可以像我需要的那样受约束。然后我需要一个非常大的文本语料。有趣的是,这也是 Stephen。他曾在 Common Crawl 工作,我认为他帮助或制作了维基百科数据集。然后我意识到维基百科数据集做了很多假设。它有很多像 UNK 这样的未知词标记,因为它假设了经典的 NLP 方法。所以我重做了整个事情,创建了一个新的维基百科数据集,那就是我的通用语料。然后我使用 AWD-LSTM 进行训练。实际上只用了一夜。在游戏 GPU 上训练了八小时,你知道,因为我在旧金山大学,我们没有大量资源。我猜大概是 2080 Ti 之类的。然后第二天早上我醒来时,我使用了我们今天使用的相同三阶段架构:预训练、中训练、后训练。所以我想,好吧,我现在训练了一个预测维基百科下一个词的东西,它一定对世界了解很多。然后我想,如果我随后在一个特定语料上微调它——我们现在可以称之为监督微调数据集——在这个案例中是电影评论数据集,它会特别擅长预测这些评论的下一个词。所以它会学到很多关于电影的知识。
I'm a huge fan of regularization. And I'm a huge fan of taking a model that's incredibly flexible and then making it more constrained not by decreasing the size of the architecture but by adding regularization. So even that at the time was extremely controversial. But that was by no means a unique insight of ours. So what Stephen Merity had done is he took the extreme flexibility of an LSTM—a kind of classic stateful recurrent neural net towards which things are kind of gradually heading back nowadays—and added five different types of regularization. He added every type of regularization you can imagine. And then that was my starting point: to say okay, I now have a massively flexible deep learning model that can be as powerful as I want it to be and it can also be as constrained as I need it to be. And then I needed a really big corpus of text. Funnily enough, this is also Stephen. He had been at Common Crawl and I think he helped or made the Wikipedia dataset. And then I realized actually the Wikipedia dataset made lots of assumptions. It had all these like UNK for unknown words because it all assumed classic NLP approaches. So I redid the whole thing, created a new Wikipedia dataset and that was my general corpus. And then I used an AWD-LSTM and trained it. So it was actually overnight. So for eight hours on a gaming GPU, you know, because I was at the University of San Francisco, we didn't have heaps of resources. Probably like a 2080 Ti or something, I suspect. And then the next morning when I woke up, I then used the same three-stage architecture that we do today: pre-training, mid-training, post-training. So then I figured, okay, I've now trained something to predict the next word of Wikipedia, it must know a lot about the world. I then figured if I then fine-tune it on a corpus specific—so what we could now call supervised fine-tuning dataset—which in this case was a dataset of movie reviews, it would become especially good at predicting the next word of those. So it would learn a lot about movies.
我花了大约一个小时做预训练,然后花了几分钟微调下游分类器。那是一个经典的学术数据集,被认为是最难的:5000 词的电影评论,判断正面或负面情感。今天这很容易,但当时只有高度专业化的模型——人们整个博士论文的成果——才能做好。我五分钟后就击败了他们所有的结果,通过微调那个模型。太神奇了。
I did that for about an hour, then a few minutes of fine-tuning the downstream classifier. It was a classic academic dataset, considered the hardest one: 5,000-word movie reviews to classify positive or negative sentiment. Today that's easy, but back then only highly specialized models—people's entire PhDs—did it well. I beat all their results five minutes later by fine-tuning that model. It was amazing.
另一个有趣的点是你做微调的方法论。
And the other interesting thing is the methodology around how you do the fine-tuning.
是的。我们做微调的方法是在 fast.ai 开发的,那时我们刚起步。我们做的一个极具争议的事情是专注于微调现有模型,因为我们觉得这很重要。其他人也在做类似工作——Jason Yosinski 在他的博士期间做了很棒的研究,关于如何微调模型。在计算机视觉领域,我们是第一批。我们觉得对整个模型使用单一学习率毫无意义,因为不同层的行为不同。我们提出了先只训练最后一层,然后最后两层,再最后三层的想法,并使用判别式学习率——不同层用不同的学习率。另一个关键洞察是,你必须微调每个批归一化层,因为它们会整体上移或下移或改变尺度。这一点我们告诉了所有人,但多年没人意识到。对于 ULMFiT,我们最终解冻了所有层,但只需要最后两层就能接近最先进水平。这只需要几秒钟。
Yeah. The way we did fine-tuning was developed at fast.ai, in our very early days. One extremely controversial thing we did was focus on fine-tuning existing models because we thought it was important. Others were working on it too—Jason Yosinski did great research during his PhD on how to fine-tune models. In computer vision, we were among the first. We felt using a single learning rate for the whole model made no sense because different layers behave differently. We developed the idea of training only the last layer first, then the last two, then the last three, using discriminative learning rates—different learning rates for different layers. Another critical insight, which no one realized for years even though we told everyone, is that you must fine-tune every batch normalization layer, because they shift the whole thing up or down or change its scale. With ULMFiT, we ended up unfreezing all layers, but only the last two were really needed to get close to state-of-the-art. It took seconds.
判别式学习率很有意思。当时普遍的看法是,微调时如果学习率太高,会破坏表征。所以我想共识是使用非常低的学习率,以免破坏表征。
The discriminative learning rate thing is interesting. At the time, the received wisdom was that if the learning rate is too high during fine-tuning, you blow out the representations. So I guess the wisdom was to use a very low learning rate to avoid destroying representations.
当时没有普遍看法,因为没人讨论。没人在乎。迁移学习根本不是大家考虑的事情。Rachel 和我认为它比什么都重要——只需要一个人训练一次大模型,其他人就可以微调它。所以我们花了很多时间尝试各种方法。最终,直觉很简单:直觉上应该有效的方法几乎总是有效。这与今天人们做 ML 研究的方式大不相同——他们认为一切都得靠消融实验,不能做假设。但我发现几乎所有我预期有效的东西第一次就能成功,因为我建立了关于梯度行为的直觉。
There was no received wisdom because nobody talked about it. No one cared. Transfer learning was just not something anybody thought about. Rachel and I felt it matters more than anything—only one person has to train a really big model once, and the rest of us can all fine-tune it. So we spent a lot of time trying lots of things. In the end, the intuition was straightforward: what intuitively seemed like it ought to work almost always did. That's a big difference from how people still do ML research today—they think it's all about ablations and you can't make assumptions. But I find nearly everything I expect to work works first time, because I build up intuitions about how gradients behave.
持续学习(保持通用性)和针对特定任务微调之间存在二分法。想法是你可以让模型变得专门化,但会失去通用性并降低表征质量。请谈谈这个。
There's a dichotomy between continual learning—keeping generality—and fine-tuning for a specific task. The idea is that you can make a model specific, but you lose generality and degrade representations. Tell me about that.
有一定道理,但没你想的那么严重。大问题是人们不看激活值和梯度。在我们的 fast.ai 软件中,我们内置了能一眼看到整个网络的能力。练习几次后,花几个小时就能学会看出某层是否过拟合、欠拟合或出错。例如,如果神经元趋向无穷大,就会变成死神经元,梯度为零。你总能修复它。这远没有人们想的那么糟糕。一个在持续学习上训练良好的模型,如果小心处理,也可以很好地微调用于特定任务。从某种意义上说,你确实希望神经元死亡,以弯曲行为并引入隐式约束——没有约束就没有创造力或推理。所以你想说,‘别那么做,做点别的。’
There's some truth, but not as much as you think. The big problem is people don't look at their activations or gradients. In our fast.ai software, we built in the ability to see your entire network at a glance. After a few times, it takes a couple of hours to learn to see if something is overtrained, undertrained, or wrong at a layer. For example, dead neurons with zero gradient often happen if they head off to infinity. You can always fix that. It's not as bad as people think. Something that trains well for continual learning can also be fine-tuned for a particular task if you're careful. In a sense, you do want neurons to die to bend behavior and introduce implicit constraints—without constraints, there's no creativity or reasoning. So you want to say, 'Don't do that, do something else.'
我不这么看。我发现用人类来类比 AI 很有帮助。它们的行为相似之处多于不同。对人类来说,学习新东西不是要忘记旧东西。当我让模型学习两个相似的任务时,它们几乎总是在两个任务上都比只学一个任务的模型表现更好。
I don't think of it that way. I find thinking about humans helpful for AI. They behave more similarly than differently. With a human, learning something new isn't about unlearning something else. When I got models to learn two somewhat similar tasks, they almost always got better at both than one that only learned one.
是的。半监督和自监督学习是一个非常被低估的领域。Yann LeCun 也是在这方面工作的人之一。
Yeah. Semi-supervised and self-supervised learning was such an unappreciated area. Yann LeCun was one of the guys working on it too.
我实际上发过一篇帖子,因为我很恼火几乎没人在乎半监督学习。
I actually did a post because I was so annoyed at how few people cared about semi-supervised learning.
我几年前写过一篇完整的文章。闫卢坤也帮我看了看,还推荐了几篇我遗漏的其他工作。我有点惊讶于提出一个预文本任务竟然如此有用。在视觉领域,我们在 ULMFiT 之前就做过这个。比如在医学影像中,你拿一张组织切片,遮住几个方块,然后预测那里原来是什么。我在 USF 的一些学生就在做这个。这基本上是在借鉴我们和其他人在视觉领域已经做过的工作。
I did a whole post about it years ago. Yan Lukun looked at it for me as well and suggested a few other pieces of work that I had missed. I was kind of surprised at how incredibly useful it is to basically come up with a pretext task. In vision, we did this before ULMFiT. For example, in medical imaging, you take a histology slide, mask out a few squares, and predict what used to be there. Some of my students at USF were doing stuff with that. It was basically taking stuff that we and others had already done in vision.
是的。
Yeah.
所以这个遮住方块的想法,不是我们发明的。遮住单词是显而易见的事。逐渐解冻层的想法我们之前在计算机视觉中做过。从通用预训练模型开始的想法在计算机视觉中已经存在。2015 年左右有一篇经典的计算机视觉论文,完全是经验性的,说当我们拿一个预训练的 ImageNet 模型来预测这个雕塑是由哪位雕塑家创作的,或者预测这是什么建筑风格,结果在每个任务上都达到了最先进的水平。这让我很惊讶。人们没有看到这一点并想,我打赌这应该也适用于其他所有领域,无论是基因组序列还是语言等等。但人们有点缺乏想象力。他们倾向于认为事情只在一个特定领域有效。这真的很真实。
So this idea of masking out squares, we didn't invent it. Masking out words was the obvious thing. The idea of gradually unfreezing layers we had done before in computer vision. The whole idea of starting with a pre-trained model that was general purpose had been in computer vision. There was a classic paper in computer vision around 2015 that was entirely empirical, saying look what happens when we take a pre-trained ImageNet model predicting what sculptor created this sculpture or predicting what architecture style this is. In every task, it got the state-of-the-art result. It really surprised me. People didn't look at that and think, I bet that ought to work in every other area as well, whether it be genome sequences or language or whatever. But people have a bit of a lack of imagination. They tend to assume things only work in one particular field. That's really true.
是的。
Yeah.
我的意思是,我想那里有两件事。首先,我们有点在暗示古德哈特定律或捷径规则的概念,即你得到你优化的东西,但以其他一切为代价。但情况似乎并非如此,因为我们可以优化语言模型中的困惑度,正如你所说,似乎发生的是我们有点进入了分布假设。所以你知道词以类聚。当我们有大量的关联数据时,它可能是主自动预测或任何这些东西。模型似乎构建了某种我们可以称之为理解的东西。
I mean, I guess there's two things there. First of all, we were kind of hinting at this notion of Goodhart's law or the shortcut rule that you get exactly what you optimize for at the cost of everything else. But that doesn't seem to be the case because we can optimize for perplexity in the case of language models and as you say, what seems to happen is we're getting into the distributional hypothesis here a little bit. So you know the word by the company it keeps. When we have an incredible amount of associative data, it might be master auto prediction or any of these things. The model seems to build something that we might call an understanding.
我一直认为它是一个抽象层次结构。如果要预测下一个词,它需要知道一些关于国际象棋记谱法或至少开局的知识。如果是像'这被 1956 年的美国总统否决了',你不仅需要知道总统是谁,还需要知道有总统、有领导人、有等级制度的人群、有人、有物体。如果不了解所有这些,你就无法很好地预测一个句子的下一个词。我创建 ULMFiT 的假设是,为了尽可能好地压缩这些知识,它必须在模型深处创建这些抽象、这些抽象层次结构。否则,它怎么可能做好预测下一个词的工作?因为深度学习模型是通用学习机器,我们有一种通用的训练方法,我想如果我们数据正确、硬件足够好,那么理论上我们应该能够构建那个下一个词预测机器,它应该会隐式地构建一个对文本所描述事物的层次结构理解。
I have always thought of it as a hierarchy of abstractions. If it's going to predict the next word, it needs to know something about chess notation or at least openings. If it's like 'this was vetoed by the 1956 US president', you need to know not just who the president was but the idea that there are presidents, leaders, groups of people with hierarchies, people, objects. You can't predict the next word of a sentence well without knowing all of these things. My hypothesis for why I created ULMFiT is that to compress that knowledge as well as possible, it would have to create these abstractions, these hierarchies of abstractions somewhere deep inside its model. Otherwise, how could it possibly do a good job of predicting the next word? Because deep learning models are universal learning machines, and we had a universal way to train them, I figured if we get the data right and the hardware is good enough, then in theory, we ought to be able to build that next word predicting machine, which ought to implicitly build a hierarchical structural understanding of the things described by the text it is learning to predict.
我认为它们可以以一种相当肤浅的方式知道。有大量的表面统计关系,它们泛化得非常好。这很神奇。
I think that they can know in quite a superficial way. There's a myriad of surface statistical relationships and they generalize extraordinarily well. It's miraculous.
是的。
It is.
但问题是,我想把这个与你关于创造力的其他评论进行对比。我认为知识是关于约束的,而创造力是知识的演化,尊重这些约束。因此,AI 没有创造力。你也说过同样的话。你说过 AI 没有创造力。那么一方面,你怎么能说它们知道,却不认为它们可以有创造力呢?
But the thing is I want to contrast this with other comments you've made about creativity. I think knowledge is about constraints and I think creativity is the evolution of knowledge, respecting those constraints. Therefore, AI is not creative. And you've said the same thing. You've said AI isn't creative. So on the one hand, how can you say that they know and not think that they can be creative?
我的意思是,我不认为我用过那个确切的表达。我记得和彼得·诺维格在镜头前聊天,我们俩都说,嗯,实际上,它们有点创造力。我们只是需要小心用词。我非常尊敬的彼得·沃兹尼亚克重新发现了间隔重复,建立了 SuperMemo 系统,是现代记忆大师。他一生致力于记忆的全部原因是他相信创造力来自于记住很多东西,也就是说,以有趣的方式把你记住的东西组合起来是创造力的好方法。LLM 实际上很擅长这个,但有一种创造力它们完全不擅长,那就是走出分布。
I mean, I don't think I've used that exact expression. I remember chatting with Peter Norvig on camera and both of us said, well, actually, they kind of are creative. We just got to be a bit careful about our choices of words. Peter Wozniak, who I really respect, rediscovered spaced repetition, built the SuperMemo system, and is the modern-day guru of memory. The entire reason he's based his life around remembering things is because he believes that creativity comes from having a lot of stuff remembered, which is to say putting together stuff you've remembered in interesting ways is a great way to be creative. LLMs are actually quite good at that, but there's a kind of creativity they're not at all good at, which is moving outside the distribution.
我想这正是你问题的方向。
Which I think is where you're heading with your question.
我这样表述是为了说明你必须对这件事非常细致入微。如果你说它们没有创造力,可能会给你错误的想法,因为它们能做出看起来很有创造力的事情。但如果是说,它们真的能在训练分布之外进行外推吗?答案是否定的,它们不能。但训练分布如此之大,在它们之间进行插值的方式如此之多,我们还不真正知道其局限性。我每天都看到这一点,因为我的工作是研发。我经常处于训练数据的边缘或之外。我在做以前从未做过的事情。有一种奇怪的现象我每天都会看到多次,即 LM 从极其聪明变得比愚蠢还糟糕,不理解世界运作的最基本前提。
I'm framing it this way to say you have to be so nuanced about this stuff. If you say they're not creative, it can give you the wrong idea because they can do very creative-seeming things. But if it's like, can they really extrapolate outside the training distribution? The answer is no, they can't. But the training distribution is so big and the number of ways to interpolate between them is so vast, we don't really know yet what the limitations of that are. I see it every day because my work is R&D. I'm constantly on the edge of and outside the training data. I'm doing things that haven't been done before. There's this weird thing I see multiple times every day where the LM goes from being incredibly clever to worse than stupid, not understanding the most basic fundamental premises about how the world works.
是的。
Yeah.
然后就像,哦,糟糕。我落到了训练数据分布之外。它变笨了。再继续讨论就没有意义了。
And it's like, oh, whoops. I fell outside the training data distribution. It's gone dumb. There's no point having that discussion any further.
是的。我的意思是,我喜欢玛格丽特·鲍登,她提出了这种创造力的层次结构。
Yes. I mean I love Margaret Bowden, she had this kind of hierarchy of creativity.
所以有组合性、探索性和变革性的创造力。模型当然可以做组合性创造力,但对我来说,一切都关乎约束。这是 Boden 说的,甚至达芬奇也说创造力全在于约束。你谈到过对话工程,但实际情况是,当我们与语言模型对话时,这是一个规范获取问题。我们来回交流。当我们认为智能的过程是在脑海中构建这个想象中的乐高积木并尊重各种约束时,当你尊重这些约束并持续演进,这些东西就被认为是创造性的。语言模型在通过监督、批评者或验证器添加约束时,是有创造性的。我们见过很多例子。但错觉在于,它们自己,没有约束时,只有我们讨论的行为塑造。它们没有硬约束,这就是为什么它们无法超出分布。我认为它们无法超出分布,因为那种数学模型就是做不好。它能做,但做不好。
So there's combinatorial, exploratory, and transformative creativity. Models can certainly do combinatorial creativity, but for me it's all about constraints. This is what Boden said, and even Leonardo da Vinci said creativity is all about constraints. You've spoken about dialogue engineering, but what happens is when we talk with language models, it's a specification acquisition problem. We go back and forth. When we think the process of intelligence is about building this imaginary Lego block in our mind and respecting various constraints, when you respect those constraints and continue to evolve, those things are said to be creative. Language models, when you add constraints via supervision, critics, or verifiers, are creative. We've seen many examples of this. But the illusion is on their own, sans constraints, they have this behavioral shaping we're talking about. They don't have hard constraints, and that's why they can't go outside their distribution. I think they can't go outside their distribution because it's just something that type of mathematical model can't do well. It can do it, but it won't do it well.
当你看到将曲线拟合到数据的二维情况时,一旦超出数据覆盖的区域,曲线就会以疯狂的方向消失到空间中。我们做的就是这些,只不过是在多维空间中。我认为 Boden 可能会对组合性创造力在组合整个人类知识库时能走多远感到震惊。这就是人们经常困惑的地方。例如,我昨天和 Chris Lattner 谈到 Anthropic 让 Claude 写了一个 C 编译器。他们说:‘哦,这是一个洁净室 C 编译器。你可以看出它是洁净室的,因为它用 Rust 创建。’Chris 创建了当今最广泛使用的 C/C++ 编译器,基于 LLVM。他们说:‘Chris 没用 Rust。我们没有给它任何编译器源代码,所以是洁净室实现。’但这误解了 LLM 的工作原理。Chris 的所有工作都在训练数据中出现了很多次。LLVM 被广泛使用,很多东西都建立在它之上,包括许多 C 和 C++ 编译器。将其转换为 Rust 是训练数据各部分之间的插值。这是一个风格迁移问题。最多算是组合性创造力,如果你能称之为创造力的话。当你查看它创建的仓库时,它复制了 LLVM 代码的部分。Chris 说:‘哦,我犯了个错误。我不该那样做。没别人那样做。’看,他们是唯一另一个那样做的。这不会偶然发生。这是因为你实际上并没有创造。你只是在 Rust 相关和构建编译器相关之间找到了训练数据中的非线性平均点。
When you look at the 2D case of fitting a curve to data, once you go outside the area the data covers, the curves disappear off into space in wild directions. That's all we're doing, but in multiple dimensions. I think Boden might be pretty shocked at how far compositional creativity can go when you can compose the entirety of the human knowledge corpus. This is where people often get confused. For example, I was talking to Chris Lattner yesterday about how Claude, Anthropic had Claude write a C compiler. They said, 'Oh, this is a clean room C compiler. You can tell it's clean room because it was created in Rust.' Chris created the most widely used C/C++ compiler nowadays, built on top of LLVM. They said, 'Chris didn't use Rust. We didn't give it access to any compiler source code, so it's a clean room implementation.' But that misunderstands how LLMs work. All of Chris's work was in the training data many times. LLVM is used widely, and lots of things are built on it, including many C and C++ compilers. Converting it to Rust is an interpolation between parts of the training data. It's a style transfer problem. It's definitely compositional creativity at most, if you can call it creative at all. When you look at the repo it created, it copied parts of the LLVM code. Chris says, 'Oh, I made a mistake. I shouldn't have done it that way. Nobody else does it that way.' Look, they're the only other one that did it that way. That doesn't happen accidentally. That happens because you're not actually being creative. You're just finding the nonlinear average point in your training data between Rust things and building compiler things.
这一切都是真的。首先,我认为我们不应该低估这种组合性创造力有多大。代码在互联网上,但他们还有一大堆测试作为脚手架。每次提交代码时,他们都可以运行测试。他们基本上有一个批评者,可以进行这种自主反馈循环。从某种意义上说,这非常类似于 OpenAI 和 Gemini 最近的研究,你试图解决一个数学问题,并且已经有了一个评估函数。ARC 奖也是如此。你有一个评估函数,而人们忽略的是,即使知道评估函数是什么,也是对问题的部分了解。然后你可以暴力搜索。你可以使用统计模式匹配,将验证器作为约束,实际上你可以……
All of that is true. First, I think we shouldn't underestimate how big this combinatorial creativity is. The code is on the internet, but they also had a whole bunch of tests which were scaffolded. Every time some code was committed, they could run the test. They basically had a critic and could do this autonomous feedback loop. In a sense, it's very similar to recent research by OpenAI and Gemini, where you're trying to solve a math problem and you already have an evaluation function. Same on the ARC prize. You have an evaluation function, and what people discount is even knowledge of what the evaluation function is is partial knowledge of the problem. You can then brute force search. You can use statistical pattern matching, use the verifier as a constraint, and you can actually...
而且他们甚至不需要那样做。他们实际上已经知道如何通过这些测试,因为有很多软件已经做到了。
And they don't even need to do that. They literally already know how to pass those tests because there's lots of software that already does it.
所以它只是利用这些,并将它们翻译成 Rust。它所做的就是这些,这令人印象深刻。
So it just uses that and translates them to Rust. That's all it did, which is impressive.
是的。
Yeah.
我对数学远不如对计算机科学熟悉,但通过与数学家交谈,他们告诉我 Erdos 问题等也是如此。其中一些被新解决了。
I'm much less familiar with math than computer science, but from talking to mathematicians, they tell me that's also what's happening with Erdos problems and stuff. Some of them are newly solved.
是的。
Yeah.
但它们不是灵感的火花。它们解决的是那些可以通过将人类已经弄清楚的密切相关的东西拼接在一起就能解决的问题。
But they are not sparks of insight. They're solving ones that you can solve by meshing together very closely related things that humans have already figured out.
关于 Claude Code,我知道你广泛谈论过 vibe coding。Rachel 有一些有趣的研究。她引用了 meter 研究,该研究表明当人们进行 vibe coding 时,生产力实际上下降了,但我认为……
On the subject of Claude Code, I know you've spoken extensively about vibe coding. Rachel had some interesting work out. She quoted the meter study which showed that productivity actually went down when people were vibe coding, but I think...
而且他们以为生产力提高了,这是最有趣的。
And they thought that they went up, which is the most interesting.
然后还有 Anthropic 的研究。也许我们应该往回退一点。Dario 前几天发了一篇文章,我想标题是《技术的青春期》之类的。他基本上是在说,看,我们在 Anthropic 有所有这些了不起的软件工程师,他们非常高效。他正在推断到普通软件工程师,所以会有大规模失业,因为很快我们就能用 AI 自动化所有这些。
And then there was the Anthropic study. Maybe we should rewind a little. Dario had this essay out the other day, I think it was called 'The Adolescence of Technology' or something like that. He was basically saying, look, we have all these amazing software engineers at Anthropic and they are just so productive. He was extrapolating to the average software engineer, so there's going to be mass unemployment because soon we're going to be able to automate all of this with AI.
我的意思是,这毫无道理。Elon Musk 几天前也说了类似的话,他说:‘哦,LLM 会直接吐出机器码。我们不需要库、编程语言了。’
I mean, it doesn't make any sense. Elon Musk said something a bit similar a few days ago, saying, 'Oh, LLMs will just spit out the machine code directly. We won't need libraries, programming languages.'
是的。
Yeah.
听着,问题是这些人最近都没有做过软件工程师。我不确定 Dario 是否曾经是软件工程师。软件工程是一门不寻常的学科,很多人误以为它和往 IDE 里输入代码是一回事。编码是另一个风格迁移问题。你拿到要解决问题的规范,然后利用你的组合性创造力,找到训练数据中在它们之间插值解决该问题的部分,再将其与目标语言的语法插值,你就得到了代码。
Look, the thing is none of these guys have been software engineers recently. I'm not sure Dario's ever been a software engineer at all. Software engineering is an unusual discipline, and a lot of people mistake it for being the same as typing code into an IDE. Coding is another one of these style transfer problems. You take a specification of the problem to solve, and you can use your compositional creativity to find the parts of the training data which interpolated between them solve that problem and interpolate that with syntax of the target language, and you get code.
弗雷德·布鲁克斯几十年前写过一篇非常著名的文章《没有银弹》。听起来几乎就像是在谈论今天。他当时专门回应了一种类似的观点:那时候大家都在说,‘哦,那些第四代语言什么的,我们再也不需要程序员了,不需要软件工程师了,因为软件现在太好写了,谁都能写。’而他说,他估计最多能提高 30%。他特别指出是未来十年提高 30%,但我认为他没必要限制得那么死,因为软件工程中的绝大部分工作并不是敲代码。
There's a very famous essay by Fred Brooks written many decades ago, 'No Silver Bullet.' It almost sounded like he was talking about today. He was specifically responding to something very similar: in those days, it was all like, 'Oh, what about all these new fourth-generation languages? We're not going to need any coders anymore, any software engineers anymore, because software is now so easy to write, anybody can write it.' And he said that he guessed you could get at maximum a 30% improvement. He specifically said a 30% improvement in the next decade, but I don't think he needed to limit it that much because the vast majority of work in software engineering isn't typing in the code.
是的。所以从某种意义上说,达里奥说的有一部分是对的。现在对很多人来说,他们的大部分代码都是由语言模型敲出来的。对我来说也是如此,大概 90%吧。但这并没有让我变得高效那么多,因为那从来都不是瓶颈。它也在研究方面帮了我很多,比如找出哪些文件会被改动。但每当我试图让 LLM 设计一个以前没有被大量设计过的解决方案时,结果都很糟糕。因为它每次给我的都是一个表面上看起来有点相似的设计。而这往往是一场彻底的灾难,因为那些表面上相似的东西——而我恰恰是在试图创造新东西来摆脱相似的东西——这非常具有误导性。
Yeah. So in some sense, parts of what Dario said were right. For quite a few people now, most of their code is being typed by a language model. That's true for me. Maybe like 90%. But it hasn't made me that much more productive, because that was never the slow bit. It's also helped me with research a lot, figuring out which files are going to be touched. But anytime I've made any attempt at getting an LLM to design a solution to something that hasn't been designed lots of times before, it's horrible. Because what it gives me every time is the design of something that looks on its surface a bit similar. And often that's going to be an absolute disaster, because things that look on their surface a bit similar—and I'm literally trying to create something new to get away from the similar thing—it's very misleading.
首先,我对科技界那种误解认知科学和哲学的倾向感到恼火。我们在 MLST 上采访过很多非常有趣的人,比如塞萨尔·伊达尔戈,他写了《知识的法则》这本书,还有神经科学哲学家马尔瓦·奇拉马。她谈到知识是视角性的。我认为知识是视角性的;我不认为知识可以是那种抽象的、无视角的东西,可以存在于维基百科上。我还认为知识是具身的、活的。它是存在于我们体内的东西,而组织的目的是保存和演化知识。所以当你开始将认知任务委托给语言模型时,实际上会产生一种奇怪的悖论效应:你侵蚀了组织内部的知识。
First of all, I'm exasperated by what I see as the tech bro predilection to misunderstand cognitive science and philosophy. We've spoken to so many really interesting people on MLST, like César Hidalgo, who wrote the book 'The Laws of Knowledge,' and even Marwa Chirama, a philosopher of neuroscience. She was talking about how knowledge is perspectival. I think that knowledge is perspectival; I don't think that knowledge can be this abstract, perspective-free thing that can exist on Wikipedia. I also think that knowledge is embodied and alive. It's something that exists in us, and the purpose of an organization is to preserve and evolve knowledge. So when you start delegating cognitive tasks to language models, you actually have this weird paradoxical effect that you erode the knowledge inside the organization.
嗯,这确实是真的,而且很可怕。网上经常有这样的争论:一边有人说‘LLM 什么都不懂,它们只是假装理解’,另一边有人说‘别荒谬了,看看这个 LLM 为我做了什么’。有趣的是,他们两边都对。LLM 是在角色扮演理解事物,它们假装理解。早期认知科学中丹尼尔·丹尼特的工作就很有意思。中文屋实验基本上就是这个意思,对吧?有一个人在一个房间里,完全不会中文,但他看起来确实会,因为你可以输入问题,他给出答案,但他实际上只是在一大堆书或机器中查找。假装智能和真正智能之间的区别,只要在你假装有效的范围内,就完全不重要。所以对于很多任务来说,LLM 只是假装智能其实没问题,因为从所有实际目的来看,这无关紧要,直到它无法再假装下去。然后你才意识到,‘天哪,这东西太蠢了。’
Well, that's true and that's terrifying. There's often these arguments online between people who are like, 'LLMs don't understand anything. They're just pretending to understand.' And then other people are like, 'Don't be ridiculous. Look what this LLM just did for me.' The funny thing is they're both right. LLMs cosplay understanding things. They pretend to understand things. This was the interesting thing about early cognitive science work with Daniel Dennett. That's basically what the Chinese room experiment is, right? You've got a guy in a room who can't speak Chinese at all, but he sure looks like he does because you can feed in questions and he gives you back answers, but all he's actually doing is looking up things in a huge array of books or machines. The difference between pretending to be intelligent and actually being intelligent is entirely unimportant as long as you're in the region in which the pretense is actually effective. So it's actually fine for a great many tasks that LLMs only pretend to be intelligent, because for all intents and purposes, it doesn't matter until you get to the point where it can't pretend anymore. And then you realize, 'Oh my god, this thing's so stupid.'
顺便说一句,我是塞尔的支持者。他说理解是因果可还原的,但本体论上不可还原,他还说理解有一个现象学成分。但你甚至不需要走到那一步。知识是视角性的这个观点有趣之处在于,世界是复杂的。我们没有人能完全理解它。就像盲人摸象,我们都有不同的视角。它非常复杂。所以我们都在做这种建模。但有趣的是,语言模型有时似乎理解了,它们理解是因为监督者把它们放在了一个框架里。在那个框架内,当你有了大象的视角时,它们实际上出奇地连贯,但我们忽略了监督者把模型放在那个框架里的事实。
I'm a fan of Searle, by the way. He said that understanding is causally reducible but ontologically irreducible, and he said there was a phenomenal component to understanding. But you don't even need to go there. The interesting thing about knowledge being perspectival is this idea that the world is a complex place. None of us understand it. It's like the blind men and the elephant. We all have different perspectives. It's very complex. And so we all do this kind of modeling. But the interesting thing is that the language models sometimes seem to understand, and they understand because the supervisor places them in a frame. So inside that frame, when you have that perspective of the elephant, they're actually surprisingly coherent, but we discount the supervisor placing the models in that frame.
是的。所以塞尔对丹尼特,或者说塞尔和丹尼特,是我本科读哲学时大家都在讨论的。我记得《意识的解释》大概是那时候出版的,中文屋可能更早一点。有趣的是,当时的讨论和我们现在进行的讨论是一样的,但它们从抽象讨论变成了现实讨论。如果人们回顾那些抽象讨论会很有帮助,因为它能帮你摆脱被那些如此擅长角色扮演智能的东西所分心,回到根本问题上来。所以,我只是想提一下,我们现在处于一个有趣的情况:很容易对 AI 能做什么产生错误的认识,特别是当你不理解编码和软件工程之间的区别时。
Yeah. So that's Searle versus Dennett, or Searle and Dennett, was what everybody was talking about when I was doing my undergrad in philosophy. I think 'Consciousness Explained' came out about then, probably the Chinese room a little bit before. It's interesting because the discussions were the same discussions we're having now, but they've gone from being abstract discussions to being real discussions. It's helpful if people go back to the abstract discussions, because it helps you get out of the distraction of looking at something that's cosplaying intelligence so well and go back to the fundamental question. So anyway, I just wanted to mention that it's this interesting situation we're now in where it's very easy to really get the wrong idea about what AI can do, particularly when you don't understand the difference between coding and software engineering.
是的。这就引出了你的观点,即这对组织的影响。很多组织基本上是在赌一个投机的前提:AI 将能比人类做得更好,或者至少编码方面比人类做得更好。我非常担心这一点,既为组织也为人类。对人类来说,当你不积极使用你的设计、工程和编码肌肉时,你就不会成长。你甚至可能萎缩,但至少不会成长。而作为一家研发初创公司的 CEO,如果我的员工不成长,那我们就会失败。我们不能让这种情况发生。
Yeah. Which then takes me to your point about the implications of that for organizations. A lot of organizations are basically betting their futures on a speculative premise: that AI is going to be able to do everything better than humans, or at least everything in coding better than humans. I worry about this a lot, both for the organizations and for the humans. For the humans, when you're not actively using your design and engineering and coding muscles, you don't grow. You might even wither, but you at least don't grow. And speaking of the CEO of an R&D startup, if my staff aren't growing, then we're going to fail. We can't let that happen.
而提高特定的提示技巧,无论当前一代 AI CLI 框架的细节如何,都不是成长。你知道,这就像在不了解互联网工作原理的情况下学习某个 AWS API 的细节一样。这不是可重复使用的知识;它是短暂的知识。所以如果你愿意,你实际上可以把它用作学习超能力,但它也可能产生相反的效果。它自然会随着时间的推移削弱你的信心。
And getting better at the particular prompting skills, whatever details of the current generation of AI CLI frameworks, isn't growing. You know, that's like as helpful as learning about the details of some AWS API when you don't actually understand how the internet works. It's not reusable knowledge; it's ephemeral knowledge. So if you wanted to, you can actually use it as a learning superpower, but also it can do the opposite. The natural thing it's going to do is remove your confidence over time.
我同意这是自然而然的事情。所以这对你尤其相关,因为你的职业生涯基本上就是教育人们获得技术和 AI 素养。所以默认行为非常类似于自动驾驶汽车,存在一个临界点,你不再参与其中。你不再注意,你授权了能力,你产生了理解债务。这是默认情况。所以几周前 Anthropic 的这项研究完全反驳了 Dario,因为它甚至说,研究中确实有少数人在问概念性问题,实际上是在保持对事物的掌控,他们有一个学习梯度,但大多数人没有。我对此的假设是,生成式 AI 编码的理想情况是像我们这样,我们已经编写了几十年的软件。我们已经有了这种抽象理解。我们在我们熟悉的领域使用它,我们可以指定,我们可以消除大量歧义。我们可以跟踪,我们可以来回交流,我们可以保持与过程的联系。但发生的情况是,默认的吸引子是让人们进入这种自动驾驶模式,他们完全不知道发生了什么,这实际上让他们变得更笨。
I agree that that's the natural thing. So this is especially pertinent for you because your career has been around basically educating people to get technology and AI literacy. So the default behavior is very similar to a self-driving car where there's this tipping point at which you're not engaged anymore. You're not paying attention, and you get this delegation of competence and you get understanding debt. That's the default thing. So this study from Anthropic a couple of weeks ago contradicted Dario completely because it even said that yeah, there were a few people in the study that were asking conceptual questions that are actually kind of keeping on top of things and they had a gradient of learning, but most people didn't. And my hypothesis about that is the ideal situation for Gen AI coding is that like us, we've been writing software for decades. We already have this abstract understanding. We're using it in domains that we know well and we can specify, we can remove loads of ambiguity. We can track and we can go back and forth and we can stay in touch with the process. But what happens is that the default attractor is for people to just go into this autopilot mode and they've got no idea what's happening and it's actually making them dumber.
我在 2014 年创建了第一家深度学习医学公司,名为 Enlitic。我们最初的重点是放射学,很多人担心这会导致放射科医生在放射学方面变得不那么有效。
I created the first deep learning for medicine company called Enlitic back in 2014. Our initial focus was on radiology, and a lot of people were worried that this would cause radiologists to become less effective at radiology.
是的。
Yeah.
而我强烈感觉相反,我对此做了相当多的研究,比如飞机上的电传操纵或汽车上的防抱死刹车等。如果你能成功自动化那些真正可自动化的任务部分,你就可以让专家专注于他们需要关注的事情。我们看到了这种情况发生。所以在放射学中,我们发现如果我们能自动识别肺部 CT 扫描中可能的结节,我们实际上很擅长这一点,确实如此。然后放射科医生可以专注于查看结节并判断它们是否是恶性的或如何处理。所以这又是这些微妙的事情之一。所以如果有事情你可以有效地完全自动化,从而减轻人类的认知负担,让他们专注于需要关注的事情,那可能是好的。我不知道我们在软件开发中处于什么位置,因为我已经编程大约 40 年了。所以我写了很多代码,我可以扫一眼代码屏幕,除非是相当奇怪或复杂的东西,我马上就能告诉你它做什么以及是否有效。我凭直觉就能看出可以改进的地方,需要注意的可能问题。我不确定如果我没有写过大量代码,我是否能达到那个水平。所以我发现现在真正能从 AI 中受益的人要么是完全不会编码的初级人员,他们现在可以编写他们脑海中的一些应用,只要这些应用在当前 AI 能力下能相当快地工作,他们就满意了;要么是像我或 Chris Lattner 这样非常有经验的人,因为我们基本上可以让 AI 帮我们做一些打字和研究工作。中间的人,也就是大多数时候的大多数人,这真的让我担心,因为你怎么从 A 点到达 B 点?
And I strongly felt the opposite, which is, and I did quite a bit of research into this of like what happens when there's fly by wire in airplanes or anti-lock brakes in cars or whatever. If you can successfully automate parts of a task that really are automatable, you can allow the expert to focus on the things that they need to focus on. And we saw this happen. So in radiology, we found if we could automate identifying the possible nodules in a lung CT scan, we were actually good at it, which we were. And then the radiologist can focus on looking at the nodules and trying to decide if they're malignant or what to do about it. So again, it's one of these subtle things. So if there are things which you can fully automate effectively in a way that you can remove that cognitive burden from a human so that they can focus on things that they need to focus on, that can be good. I don't know where we sit in software development because I've been coding for 40ish years. So I've written a lot of code and I can glance at a screen of code and, unless it's something quite weird or sophisticated, I can immediately tell you what it does and whether it works. I can kind of see intuitively things that could be improved, possible things to be careful of. I'm not sure I could have got to that point if I hadn't have written a lot of code. So the people I'm finding who can really benefit from AI right now are either really junior people who can't code at all who can now write some apps that they have in their head and as long as they work reasonably quickly with the current AI capabilities then they're happy, and then really experienced people like me or like Chris Lattner because we can basically have it do some of our typing for us, and some of our research for us. People in the middle, which is most people most of the time, it really worries me because how do you get from point A to point B?
是的。
Yeah.
不通过打字代码,也许有可能,但我们没有这方面的经验。这可能吗?你会怎么做?这有点像回到学校,在小学我们不让孩子们使用计算器,这样他们就能锻炼数字能力?我们是否需要在作为开发者的头五年这样做?你必须自己写所有代码。我不知道。但如果我是一个有 2 到 20 年经验的开发者,我会经常问自己这个问题,否则你可能正在让自己变得过时。
Without typing code, it might be possible, but we have no experience of that. Is it possible? How would you do it? Is it kind of like going back to school where at primary school we don't let kids use calculators so that they develop their number muscle? Do we need to do that for the first five years as a developer? You have to write all the code yourself. I don't know. But if I was a developer with between two and 20 years of experience, I would be asking that question of myself a lot because otherwise you might be in the process of making yourself obsolete.
是的。嗯,这是关于知识的另一件事,César Hidalgo 说过。他说知识是不可替代的,这意味着它不能被交换。所以他意思是学习过程在某种重要意义上不可简化,对吧?所以你必须拥有经验,而经验必须有摩擦。当我们构建世界模型时,我们实际上在学习。有句话叫‘现实会反击’,所以我们犯很多错误,更新我们的模型,我们在模型中放置这些一致性约束,这就是我们学习的方式。所以你使用 Claude 代码,过程中摩擦非常小。这正是 Anthropic 那项研究所说的。它说摩擦太小了,他们什么也没学到。
Yeah. Well, this is another thing about knowledge that César Hidalgo said. He said that knowledge is non-fungible, which means it can't be exchanged. So what he means by that is the process of learning is in some important sense not reducible, right? So you have to have the experience and the experience has to have friction. And when we build models of the world, we actually learn. There's this phrase 'reality pushes back', so we make lots of mistakes and we update our models and we're placing these coherence constraints in our model and that's how we come to learn. So you use Claude code and there's so little friction in the process. That's exactly what this study from Anthropic said. It said there was so little friction they didn't learn anything.
对。是的。不,完全正确。理想难度是教育中出现的概念。但即使追溯到 19 世纪最初提出重复间隔学习的艾宾浩斯,以及更近的彼得·沃兹尼亚克,我们发现了同样的事情。我们知道记忆不会形成,除非形成它们很费力。所以这就是你得到这个有点令人惊讶的结果的地方,即复习太频繁是个坏主意,因为它会太快地被想起。所以对于像 Anki 和 SuperMemo 这样的重复间隔学习,算法试图在你即将忘记的那一刻之前安排闪卡。所以这很费力。我学了 10 年中文,试图自己了解学习。我真的注意到这一点,我使用 Anki,因为它总是在我即将忘记之前安排我的卡片,所以总是非常费力。
Right. Yeah. No, exactly. Desirable difficulty is the concept that kind of comes up in education. But even going back to the work of Ebbinghaus, who was the original repetitive spaced learning guy in the 19th century, and then Piotr Woźniak more recently, we find the same. We know that memories don't get formed unless it is hard work to form them. So that's where you get this somewhat surprising result that says revising too often is a bad idea because it comes to mind too quickly. And so with repetitive spaced learning with stuff like Anki and SuperMemo, the algorithm tries to schedule the flashcards just before the moment you're about to forget. So then it's hard work. I studied Chinese for 10 years in order to try to learn about learning myself. And I really noticed this that I used Anki and because it was always scheduling my cards just before I was about to forget them, it was always incredibly hard work.
是的。
Yeah.
做复习,因为几乎所有的卡片都是我即将忘记的。这绝对令人筋疲力尽。但天哪,效果很好。我现在就在这里。我已经 15 年以上没有学习过了,但我仍然记得我的中文。
To do reviews because almost all the cards were ones I was on the verge of forgetting. It was absolutely exhausting. But my god, it worked well. Here I am. I haven't done any study for 15 plus years and I still remember my Chinese.
回到你那个放射学的例子,人们常举的一个例子是呼叫中心。我们有一种观念,认为组织里有高智力角色和低智力角色。对我来说,智力就是知识的适应性获取和综合。所以我们假设低智力角色,比如呼叫中心的工作,是不适应的,这意味着组织做的某些事情是不变的,所以我们可以自动化它们,不需要更新知识。我认为这其实低估了——以放射学为例——拥有这种整体知识的重要性:在呼叫中心,有那么多奇怪的边缘情况出现,那么多奇怪的事情发生,这些会向上过滤,我们随时间适应。所以当你开始自动化时,你实际上失去了创造最初那个流程的能力,也失去了组织中知识的可进化性。你其实是在自断双腿。
Well, coming back to your radiology example, one example people give is call centers. We have this notion that in an organization there are high intelligence roles and low intelligence roles. For me, intelligence is just the adaptive acquisition and synthesis of knowledge. So we assume that low intelligence roles, like call center work, don't adapt, meaning there are certain things an organization does that do not change, so we could automate them and we don't need to update our knowledge. I think that discounts that, actually, with the radiology example, having this holistic knowledge—in a call center there are so many weird edge cases that come in, so many weird things happen, and that filters up in the organization and we adapt over time. So when you start to automate things, you actually lose the competence to create the process which created the thing in the first place, and you lose the evolvability of that knowledge in the organization. You're actually kind of cutting your legs off.
是的,完全同意。在我的公司,我经常告诉员工:我几乎唯一关心的是你们个人能力增长了多少。我真的不在乎你们做了多少个 PR,多少个功能。有个不错的例子,TCL 的 John Oster 最近发布了他的一些斯坦福周五讲座,其中有一个叫‘一点斜率抵得上很多截距’。大意是,在你的人生中,如果你能专注于做那些让你成长更快的事情,那远比专注于你已经擅长的事情要好得多,后者只是高截距。
Yeah, absolutely. In my company, I tell our staff all the time: almost the only thing I care about is how much your personal human capabilities are growing. I don't actually care how many PRs you're doing, how many features you're doing. There's that nice John Oster, the TCL guy, recently released some of his Stanford Friday takeaway lectures, and he has this nice one called 'A little bit of slope makes up for a lot of intercept.' Basically the idea that in your life, if you can focus on doing things that cause you to grow faster, it's way better than focusing on the things you're already good at, which have that high intercept.
对。
Yeah.
所以我真正关心的,也是我认为对公司唯一重要的事情,就是我的团队专注于他们的斜率。
So the only thing I really care about, and I think is the only thing that matters for my company, is that my team is focusing on their slope.
对。如果你只专注于在当前 AI 能力极限下产出结果,你关心的只是截距。所以我认为这基本上是公司和员工走向过时的道路。我很惊讶现在有多少大公司高管在推动这个,因为如果他们错了——很可能就是——而且他们无从判断,因为这是他们完全不熟悉的领域,MBA 里也没学过。他们基本上是在把自己的公司推向毁灭。
Yeah. If you focus on just driving out results at the limit of whatever AI can do right now, you're only caring about the intercept. So I think it's basically a path to obsolescence for a company and the people in it. I'm really surprised how many executives of big companies are pushing this now, because it feels like if they're wrong—which they probably are—and they have nowhere to tell if they are, because this is an area they're not at all familiar with, they never learned it in their MBAs. They're basically setting up their companies to be destroyed.
对。
Yeah.
而且我很惊讶股东们会允许他们这么做——进行如此投机性的行动。感觉很多公司会因为积累的技术债务而失败,导致无法再维护或构建产品。有很多人,比如 Frans Lanting,真正理解这一点。他明白,这始终是关于领域认知模型的模因共享,以及我们如何共同完善它。这是生成式 AI 编码的另一个大 Scaling(规模扩张)问题。理想情况下:我做过这个,我非常了解一个领域,可以极其详细地指定它,然后告诉 Claude Code 去做这件事,我脑子里的模型不重要。然后你进入一个组织,现在我需要把我的知识分享给所有其他人。我肯定你的公司也有这个问题。这种知识获取瓶颈是组织中一个非常严重的问题。所以当只有我一个人时,我觉得使用 Claude Code 我的效率可能提高了 50 倍。这绝对是魔法,我理解为什么人们如此兴奋。但人们似乎没有理解这个瓶颈,以及它如何无法真正应用到许多现实世界的组织中。
And I'm really surprised that shareholders would let them do that—set up such an incredibly speculative action. It feels like a lot of companies are going to fail as a result of the amassed tech debt that causes them to not be able to maintain or build their products anymore. There are loads of folks out there like Frans Lanting who really gets it. He understands this, and he's always said it's about this kind of memetic sharing of cognitive models about the domain and how we refine it together. This is another big scaling problem with Gen AI coding. The ideal case: I've done this, I know a domain really well, and I can specify it with exquisite detail, and I tell Claude Code to go and do this thing, and the models in my mind don't matter. Then you go into an organization, and now I need to share my knowledge with all the other people. I'm sure you have this in your company as well. This knowledge acquisition bottleneck is a real serious problem in organizations. So when it's just me, I think I'm probably 50 times more productive using Claude Code. It's absolutely magic, and I can see why people are so excited about it. But people don't seem to understand the bottleneck and how that doesn't really translate to many real world organizations.
实际上没有人比之前多创造了 50 倍的高质量软件。我们刚刚做了一项研究,人们实际交付的东西只有微小的增长。这是事实。显然,我是 AI 及其能力的爱好者,但我的妻子 Rachel 最近在一篇文章中指出:让赌博上瘾的所有要素都存在于暗流中。
No one's actually creating 50 times more high-quality software than they were before. We've actually just done a study of this, and there's a tiny uptick in what people are actually shipping. That's the facts. Obviously, I'm an enthusiast of AI and what it can do, but also my wife Rachel recently pointed out in an article: all of the pieces that make gambling addictive are present in dark flow.
对。我正想提这个。你得跟我们说说编码的事。
Yeah. I was going to bring that up. You have to tell us about coding.
是的,这是一个非常尴尬的情况:我认识的几乎所有最近几个月对 AI 驱动编码非常热衷的人,当他们最终回头看看那些热情高涨的日子里构建的东西——我今天在用吗?我的客户今天在用吗?我今天从中赚钱了吗?——他们都完全改变了看法。几乎所有的钱都被网红或生产 token 的公司赚走了。AI 编码的问题在于,它就像一台老虎机,你有一种控制的错觉。你可以精心设计你的提示词、你的 MCP 列表、你的技能等等。但最终,你拉下杠杆——输入提示词,然后得到一些东西,就像樱桃樱桃。你会想,‘哦,下次我改一下提示词,加一点上下文。’再拉一次杠杆。这是随机的。你偶尔会赢一次,比如‘哦,我赢了,我得到了一个功能。’所以它具备了所有伪装成胜利的失败特征:有点随机、控制感——所有游戏公司试图在游戏室中设计的东西。这并不意味着 AI 没有用,但天哪,真的很难说。
Yeah, it's this really awkward situation where almost everybody I know who got very enthusiastic about AI-powered coding in recent months have totally changed their mind about it when they finally went back and looked at how much stuff that I built during those days of great enthusiasm am I using today? Are my customers using today? Am I making money from today? Almost all the money is being made by influencers, or by the companies that produce the tokens. The thing about AI-based coding is that it's like a slot machine in that you have an illusion of control. You can get to craft your prompt and your list of MCPs and your skills and whatever. But then in the end, you pull the lever—you put in the prompt and something comes back, and it's like cherry cherry. It's like, 'Oh, next time I'll change my prompt a bit. I'll add a bit more context.' Pull the lever again. It's the stochastic thing. You get the occasional win that's like, 'Oh, I won. I got a feature.' So it's got all these hallmarks of loss disguised as a win, somewhat stochastic, feeling of control—all the stuff that gaming companies try to engineer into their gaming rooms. Now, none of that means that AI is not useful, but gosh, it's hard to tell.
我知道。而且 Rachel,明确一下,她还说赌博的一个特点是你会自欺欺人地认为自己了解情况,但实际上并不了解。不过我们还是稍微说说看好的情况吧。因为我确实认为在受限的情况下它非常有用,这些情况是我们理解并能施加约束和规范的。但即使在这些情况下,你也可以说,一方面我们不会很快失业,因为你只是做了更多工作。关于上瘾这件事,我注意到我经历过 14 小时的 Claude Code 马拉松式会话,我真的感觉上瘾了。就像老虎机一样。真的。
I know. And Rachel, just to be clear, she also said that one of the hallmarks of gambling is that you kind of delude yourself that you have some awareness of what's going on, but actually you don't. But let's do the bull case a little bit, though. Because I do think in restricted cases it is very useful, and these are cases where we understand and we can place constraints and specification. But even in those cases, you could argue that on the one hand we're not going to be unemployed anytime soon because you just do more work. On the addiction thing, I've noticed that I've had 14-hour Claude Code marathon sessions, and I actually feel addicted to it. It's like a slot machine. It really is.
我也经历过。绝对如此。
Been there, too. Absolutely.
是的,我知道。我写代码从来没有这么疲惫过。
Yeah, I know. And I've never felt more drained writing code.
我事后真的需要休息,比如休息几天,因为整个过程糟透了,你懂的。
I actually need to take a rest afterwards, like a few days rest because it completely was crap, you know.
是啊,确实。
Yeah, definitely.
我确实有过一些成功经验,对吧?实际上,过去几年我们一直在构建一个完整的产品,基于我们确信会成功的领域——那就是处理足够小的、能完全理解的模块,你可以设计它们,并建立自己的抽象层,从而创造出比这些模块本身更大的东西。
I've had some successes, right? And so in fact, we've spent the last couple of years building a whole product based around where we know the successes are going to be, which is when you're working on reasonably small pieces that you can fully understand and that you can design and you can build up your own layers of abstraction to create things that are bigger than the parts that you're building out of.
最近我遇到一个非常有趣的情况,基本上算是个实验。我们非常依赖一个叫 IPython kernel 的东西,它是 Jupyter notebook 的底层引擎。IPython kernel 从版本 6 升级到 7 后,它停止工作了。我们试图用它配合的两个产品——一个是 NB Classic,也就是原始的 Jupyter notebook,另一个是我们自己的产品 Solara——都会随机崩溃。IPython kernel 有超过 5000 行代码,非常复杂,涉及多线程、事件、阻塞、与 IPython 和 ZMQ 的接口,还有各种不同的组件和 debugpy。我完全搞不懂,也看不出崩溃的原因。所有测试都通过了。我想知道 AI 能不能解决这个问题。我一直对一个问题感兴趣:目前 AI 能独立处理多大的代码块?结果答案是肯定的。我觉得它可以。我花了几周时间,过程中并没有对 IPython kernel 的工作原理有太多了解,但我花了不少时间拆解各个组件。答案是:两小时内,Claude 5.2——当时可能是 5.2,或者 3 刚出来——搞不定。然后如果我用了每月 200 美元的 GPT-5.3 Pro 来修复问题,它就能搞定。所以通过在这两个模型之间来回切换,我花了几周时间让东西跑起来了。就像你说的,这完全不好玩。非常累人,而且压力很大,因为我没有真正的掌控感。但有趣的是,现在我拥有一个据我所知能正确工作的 Python Jupyter kernel 实现,它适配了新的版本 7 协议改进。然后我就想,这太迷人了,因为我们没有一套软件工程理论来指导现在该怎么做。这里有一段没人理解的代码。
I had a very interesting situation recently where it was kind of an experiment basically. We rely very heavily on something called IPython kernel, which is the thing that powers Jupyter notebooks. There had been a major version release of IPython kernel from 6 to 7, and it stopped working. Both of the products that we were trying to use it with—one was called NB Classic, which is the original Jupyter notebook, and then our own product called Solara—would just randomly crash. IPython kernel is over 5,000 lines of code. It's very complex code, multiple threads, events, blocks, interfaces with IPython, with ZMQ, all kinds of different pieces, debugpy. I couldn't get my head around it and I couldn't see why it was crashing. The tests are all passing. I wondered if AI can solve this. I'm always interested in the question of how big a chunk AI can handle on its own right now. The answer turned out to be yes. I think it can. I spent a couple of weeks; I didn't develop a lot of understanding about how IPython kernel really worked in the process, but I did spend quite a bit of time pulling out separate components. The answer was: in two hours, Claude 5.2—I think it was 5.2 at that time, or maybe 3 had just come out—couldn't do it. Then if I got the $200 a month GPT-5.3 Pro to fix the problems, it could. So by rolling back between those two pieces of software, those two models, I could get things working over a couple of weeks period. And like you say, it wasn't at all fun. It was very tiring and it felt stressful because I wasn't really in control. But the interesting thing is I now am in a situation where I have the only implementation of a Python Jupyter kernel that actually works correctly as far as I can tell with these new version 7 protocol improvements. And now I'm like, well, this is fascinating because we don't have a kind of a software engineering theory of what to do now. Here's a piece of code that no one understands.
是啊。
Yeah.
我敢把公司的产品押在这上面吗?答案是不知道,因为我不知道现在该怎么办。没人遇到过这种情况。它会有内存泄漏吗?如果协议有微小改动,一年后它还能工作吗?有没有什么奇怪的边缘情况会毁掉一切?没人知道,因为没人理解这段代码。这真是个奇怪的局面。
Am I going to bet my company's product on it? And the answer is I don't know, because I don't know what to do now. No one's been in this situation. Will it have memory leaks? Will it still work in a year's time if there's some minor change to the protocol? Is there some weird edge case that's going to destroy everything? No one knows because no one understands this code. It's a really curious situation.
首先,我们应该承认这种控制权的侵蚀是隐蔽而有害的。一开始,你只有 10% 的 AI 生成代码,然后你看着它一点点增加,六个月后一个 PR 进来,现在 60% 的代码都是 AI 生成的。你看到发生了什么吗?你慢慢变得脱节。但乐观的看法是:在 AI 中有一个概念叫功能主义,我们不在乎智能体是由什么构成的。只要它做了所有正确的事情,我们就说它是 AI。软件也是一样。所以乐观的看法是:我理解这个领域。我不需要写——我不需要知道如何写快速排序算法。我只需要理解它,对吧?然后我只需要有所有这些测试,它需要部署,这些事情需要发生。到那时,我其实不在乎了。
First of all, we should acknowledge the pernicious erosion of control. So at the very beginning, you have 10% AI generated code and then you can just see how it creeps up and up, and then at some point 6 months down the line a PR comes in and now you know 60% of the code is AI generated. Do you see what happens? You slowly become disconnected. But the bull case for this is, in AI there's this idea called functionalism that we don't care what the intelligent thing is made out of. As long as it does all of the right things, then we would say it's AI. And it's the same thing with software. So the bull case is: I understand the domain. I don't need to write—I don't need to know how to write the quicksort algorithm. I just need to understand it, right? And then I just need to have all of these tests and it needs to go into deployment and these things need to happen. At that point, I don't actually care.
我挺喜欢这个框架,但它实际上表明:哇,软件工程确实很重要,因为软件工程就是关于找出这些模块是什么、它们应该如何表现,然后如何将它们组合成更大的模块,再如何组合成更大的模块。如果我们做得好,十年后我们就能拥有比今天能想象的任何东西都强大得多的软件。
I quite like that framing, but you know what that actually does is it says: wow, software engineering sure is important then, because software engineering is all about finding what those pieces are and how they should behave and then how you can put them together to create a bigger piece, and then how you can put them together to create a bigger piece. And if we do that well, then in 10 years time, we could have software that is far more capable than anything we could even imagine today.
但通过真正优秀的软件工程,你已经能实现这一点了。
But you're already going to get that with really great software engineering.
是的,你得小心。我认为最终,比如 IPython kernel,就是太大了,对吧?因为最终,制作原始 IPython kernel 的团队没能创建一套能正确测试它的测试集,因此现实世界的下游项目,包括原始的 NB Classic——IPython kernel 就是从它里面提取出来的——都不再工作了。所以这就是我们现在在 Answer.ai 开发方面的重点:找到合适大小的模块,并确保它们是合适的模块。知道如何识别这些模块、如何设计它们以及如何组合它们,实际上通常需要几十年的经验才能真正擅长。对我来说当然如此。我觉得我大概在 20 年经验后才变得比较擅长。是的,一个大问题是:如何培养这些软件工程技能,它们现在比以往任何时候都更重要?这是擅长编写计算机软件的人和不擅长的人之间的区别。这感觉像是一个具有挑战性的问题。
Yeah, you want to be careful. I think in the end, IPython kernel, for example, is just too big a piece, right? Because in the end, the team that made the original IPython kernel were not able to create a set of tests that correctly exercised it, and therefore real world downstream projects including the original NB Classic—which is what IPython kernel was extracted from—didn't work anymore. So this is kind of where our focus is now on the development side at Answer.ai: finding the right sized pieces and making sure they're the right pieces. Knowing how to recognize what those pieces are and how to design them and how to put them together is actually something that normally requires some decades of experience before you're really good at it. Certainly it's true for me. I reckon I got pretty good at it after maybe 20 years of experience. Yeah, it's a big question: how do you build these software engineering chops which are now even more important than they've ever been before? They're the difference between somebody who's good at writing computer software and somebody who's not. That feels like a challenging question.
我知道。而且还有一种观点认为,抽象和表示事物有无数种方式。世界非常复杂。也许我们一直以来的软件抽象和表示方式,很大程度上反映了我们自身的认知局限,对吧?即使在物理学等科学中,你也倾向于使用许多相当还原论的方法来建模世界。然后还有复杂性科学,它恰恰拥抱了事物那种建设性的、耗散的、棘手的本质。
I know. And there's also this notion that there are so many different ways to abstract and represent something. The world is a very complex place. And maybe the way we've been abstracting and representing software is mostly a reflection of our own cognitive limitations, right? And even in the sciences in physics, you tend to have a lot of quite reductive methods of modeling the world. And then you've got complexity science, which is just embracing the constructive dissipative, gnarly nature of things.
而且我认为当今很多软件我们并不理解。例如,有许多全球分布的软件应用使用了 Actor 模式,这基本上就是一个复杂系统。我们理解它的唯一方式就是通过模拟和测试,因为没有人真正知道所有这些部分是如何组合在一起的。所以你可以说,从乐观的角度看,也许我们在软件工程的顶层已经在做这件事了,而这正是我们最终想要做的。
And I think a lot of software today we don't understand. For example, there are many globally distributed software applications that use the actor pattern, and it's basically a complex system. The only way we can understand it is by doing simulations and tests because no one actually knows how all these things fit together. So you could argue as a bull case that maybe we are already doing this at the top of software engineering, and that is what we want to do eventually anyway.
是的,我可能不同意。你看像 Instagram 和 WhatsApp 这样的公司,只有 10 名员工却主导了各自领域,击败了谷歌和微软这样的公司。我认为这种在大型公司里构建软件的方式实际上正在失败。我们看到很多大公司变得越来越绝望。例如,微软 Windows 和 Mac OS 的质量在过去 5 到 10 年里明显大幅下降。当年 Dave Cutler 逐行检查 NT 内核并确保它优美时,那是一个优雅而神奇的软件。我不认为世界上有任何人会说 Windows 11 是一个优雅而神奇的软件。所以我确实认为我们需要找到这些我们完全理解的更小组件,并把它们构建起来。但问题在于:AI 并不擅长这个。我凭经验说,它们在软件工程上真的很差。而且我认为这可能永远都是这样,因为我们要求它们经常跳出训练数据。如果我们试图构建一个以前从未有过的东西,并以比以往更好的方式去做,那就是在说不要只是复制训练数据中的内容。这对很多人来说是一个困惑点,因为他们看到 AI 非常擅长编码,然后就想,哦,那就是软件工程,所以它一定擅长软件工程。但它们是不同的任务。它们之间没有太多重叠,目前也没有经验数据表明 LLM 在软件工程上获得了任何能力。每次你看它们做的一个软件工程作品,比如 Cursor 创建的浏览器或 Anthropic 创建的 C 编译器,我读过这些代码很多。Chris Lattner 比我更熟悉那个编译器的例子,但它们都是对已有东西的明显复制。所以这就是挑战:如果你想构建不只是复制的东西,你就不能把它外包给 LLM。没有理论理由相信你永远能做到,也没有经验数据表明你永远能做到。
Yeah, I'd say probably not. You see companies like Instagram and WhatsApp dominate their sectors while having 10 staff and beating companies like Google and Microsoft. I would argue this way of building software in very large companies is actually failing. And I think we're seeing a lot of these very large companies becoming increasingly desperate. For example, the quality of Microsoft Windows and Mac OS has very obviously deteriorated greatly in the last 5 to 10 years. Back when Dave Cutler was looking at every line of the NT kernel and making sure it was beautiful, it was an elegant and marvelous piece of software. I don't think there's anybody in the world who's going to say that Windows 11 is an elegant and marvelous piece of software. So I actually think we do need to find these smaller components that we do fully understand and build them up. And here's the problem: AI is no good at that. I say that empirically. They're really bad at software engineering. And I think that's possibly always going to be true because we're asking them to often move outside of their training data. If we're trying to build something that literally hasn't been built before and do it in a better way than has been done before, we're saying don't just copy what was in the training data. This is a confusing point for a lot of people because they see AI being very good at coding and then think, oh, that's software engineering, so it must be good at software engineering. But they're different tasks. There's not a huge amount of overlap between them, and there's no current empirical data to suggest that LLMs are gaining any competency at software engineering. Every time you look at a piece of software engineering they've done, like the browser which Cursor created or the C compiler which Anthropic created, I've read the source code of those things quite a bit. Chris Lattner is much more familiar with the compiler example than me, but they're very obvious copies of things that already exist. So that's the challenge: if you want to build something that's not just a copy, then you can't outsource that to an LLM. There's no theoretical reason to believe that you'll ever be able to, and there's no empirical data to suggest that you'll ever be able to.
是的。我认为这次对话的要点是,我相信你会同意,我们需要 AI 和人类的结合。因为人类提供理解以及我们之前谈到的所有关于知识的东西,但我们仍然可以把 AI 当作工具。我们需要设计运营模式或工作方式,确保我们不削弱自己的能力和理解。所以这是一条非常微妙的界线。
Yes. I think the punchline of this conversation is, and I'm sure you would agree, that we need to have the combination of AI and humans working together. Because the humans provide the understanding and all the stuff we were saying about knowledge, but we can still use AI as a tool. We need to design operating models or ways of working that make sure we don't diminish our competence and understanding. So it's a very fine line.
这一直是我们的重点,无论是教学还是我们自己的内部开发。我研究了 20 年的东西结果证明是让这一切运作的关键。功劳应该归于创建笔记本界面的人,尽管很多想法可以追溯到 Smalltalk、Lisp 和 APL。基本思想是,当人类能够实时操作计算机内部的对象、研究它们、移动它们并将它们组合在一起时,人类可以用计算机做更多事情。这就是 Smalltalk 关于对象的思想,APL 关于数组的思想也是如此。Mathematica 基本上是一个超强化的 Lisp,然后加上了一个非常优雅的笔记本界面,让你可以构建一种活文档。所以几年前我构建了一个叫 NBDEV 的东西,它是一种在这些笔记本界面内、在这些丰富的动态环境中创建生产软件的方法。我发现这让我作为程序员的生产力大大提高。今天,尽管我从未以全职程序员为职业,但当你查看我的 GitHub 仓库输出时,我认为 GitHub 产生了一些统计数据,我几乎是澳大利亚最高产的程序员。这很有效,我构建的很多东西都有很多人使用,因为这是一种非常丰富强大的构建方式。所以我们发现,如果你把 AI 和人类放在同一个环境中,同样是一个丰富的交互环境,AI 也会表现得更好。这也许并不令人震惊,但通常的方式,比如你使用 Claude Code(我知道你在用,它是一个非常好的软件),我们给 Claude Code 的环境与 40 年前人们的环境非常相似。它是一个基于行的终端界面。它可以使用 MCP 或其他东西,但大多数时候它只使用 bash 工具,这也很强大。我喜欢 bash 工具,我一直在用 CLI 工具,但它仍然只是用文本文件作为与世界的接口。这真的很贫乏。所以我们把人类和 AI 放在一个 Python 解释器里。现在突然之间,你拥有了一个非常优雅的表达性编程语言的全部力量,人类可以用它来与 AI 对话。AI 可以与计算机对话。人类可以与计算机对话。计算机可以与 AI 对话。你有了这个非常丰富的东西。然后我们让人类和 AI 实时构建彼此可以使用的工具。对我来说,这就是关键:创造一个环境,让人类能够成长、参与和分享。对我来说,当我使用 Solvit 时,它与你描述的 Claude Code 的体验完全相反。几个小时后,我感到精力充沛、快乐和满足。
That's been our focus, for teaching and for our own internal development. The stuff I've been working on for 20 years has turned out to be the thing that makes this all work. Credit should go to the guy who created the notebook interface, although lots of ideas go back to Smalltalk, Lisp, and APL. Basically, the idea that a human can do a lot more with a computer when the human can manipulate the objects inside that computer in real time, study them, move them around, and combine them together. That's what Smalltalk was all about with objects, and APL was the same with arrays. Mathematica is basically a superpowered Lisp which then added a very elegant notebook interface that allowed you to construct a kind of living document out of all this. So I built this thing called NBDEV a few years ago, which is a way of creating production software inside these notebook interfaces, inside these rich dynamic environments. I found that made me dramatically more productive as a programmer. Today, even though I've never been a full-time programmer as my job, when you look at my GitHub repo output, I think GitHub produced some statistics and I was just about the most productive programmer in Australia. It's working, and a lot of the stuff I build has lots of people using it because it's such a rich powerful way to build things. So it turns out we've now discovered that if you put AI in the same environment with the human, again in a rich interactive environment, AI is much better as well. Perhaps that isn't shocking to hear, but the normal way, like if you use Claude Code which I know you do and it's a very good piece of software, the environment we give Claude Code is very similar to the environment that people had 40 years ago. It's a line-based terminal interface. It can use MCP or whatever, but most of the time it just uses bash tools, which again very powerful. I love bash tools, I use CLI tools all the time, but it's still just using text files as its interface to the world. It's really meager. So we put the human and the AI inside a Python interpreter. Now suddenly you've got the full power of a very elegant expressive programming language that the human can use to talk to the AI. The AI can talk to the computer. The human can talk to the computer. The computer can talk to the AI. You have this really rich thing. And then we let the human and the AI in real time build tools that each other can use. That's what it's about to me: creating an environment where humans can grow and engage and share. For me, when I use Solvit, it's the opposite of that experience you described with Claude Code. After a couple of hours, I feel energized and happy and fulfilled.
我来谈谈我的看法。我认为你指出的关键是,拥有一个交互式的、有状态的环境能给你反馈,这其中有某种魔力。这是因为我们的大脑可以完成一定量的工作。
I'll give you my take. I think that the thing that you're pointing to here is there's something magic about having an interactive stateful environment that gives you feedback. And that is because our brains can do a certain unit of work.
所以,我们实际上是通过与现实进行反复打磨和测试来思考的。这就是为什么我在读博期间使用了 Mathematica 和 MATLAB。我们有这种 REPL 环境,你知道,这里有矩阵,做个图像绘图,改一下,它就会显示现在的样子。这实际上是一种很好的方式,可以不断优化我对某件事的心智模型。
So, we actually think through refining and testing with reality. That's why, during my PhD, I used Mathematica and MATLAB. We've got this REPL environment, and you know, here's the matrix, do an image plot, change this, and it shows what it looks like now. It's actually a wonderful way to just refine my mental model about something.
而 Claude Code 也做了很多这类事情。我认为这主要是技能问题。我觉得那些有效使用 Claude Code 的人就是这样做的。我写过一个内容管理系统。
And Claude Code does a lot of this stuff. I think it's mostly a skill issue. I think the people that use Claude Code effectively do this. I've written a content management system.
有可能。确实有可能。
It's possible. It is possible.
有可能。是的。所以,我写了一个叫 Rescript 的内容管理系统。当我在制作纪录片视频时,它可以拉取转录文本,然后我就能验证其中的说法。AI 素养的一部分就是理解语言模型的不对称性。当你给它们一个判别性任务时,它们其实相当擅长。所以如果我让一个子智能体去验证每一个说法,它比我在生成模式下生成一堆说法要准确得多。还有那个带状态的反馈机制,我可以有一些模式化的 XML 转储,旁边有一个应用在可视化,这就形成了一个反馈循环。对我来说,这是 AI 素养的问题。AI 领域的高手已经在这么做了。
It's possible. Yeah. So, I've written a content management system called Rescript. When I'm putting together a documentary video, it can pull transcripts and then I can verify the claims. Part of AI literacy is just understanding the asymmetry of language models. When you give them a discriminative task, they're actually quite good. So if I tell a sub-agent to go and verify every individual claim, it's much more accurate than if I was in generation mode generating a bunch of claims. And the stateful feedback thing again, I can have some schematized XML dump and an application on the side which is visualizing, and it's a feedback loop. For me, this is an AI literacy thing. The good people at AI are already doing this.
是的。所以我并不完全同意你的看法。我同意你可以在 Claude Code 里做到这一点,我也同意这确实是 AI 素养的问题,但 Claude Code 并不是为此而设计的。它并不擅长这个,也没有让它成为与它交互的自然方式。我不想说这是 AI 素养的问题,因为那就像在说‘哦,这是你的问题’。对我来说,如果一个工具没有让人类自然地变得更博学、更快乐、与工作有更深的理解和联系,那就是工具的问题。工具就应该这样设计。现在很多模型和工具被评估的标准是‘我能不能给它一个完整的工作,让它自己去做完?’这在我看来是个巨大的错误。而应该是‘你是否评估了人类在另一端是否对某个主题有了深刻理解,以便将来能轻松地构建东西?’
Yeah. So I don't fully agree with you. I agree you can do it in Claude Code and I agree it is an AI literacy thing as to whether you can, but also Claude Code was not designed to do this. It's not very good at it and it doesn't make it the natural way of working with it. I don't want to say it's an AI literacy problem because that's like saying, 'Oh, it's a you problem.' To me, if a tool is not making it the natural way for a human to become more knowledgeable, more happy, more connected with a deeper understanding and a deeper connection to what they're working on, that's a tool problem. That should be how tools are designed to work. So many models and tools expressly are being evaluated on 'Can I give it a complete piece of work and have it go away and do the whole thing?' which feels like a huge mistake to me, versus 'Have you evaluated whether a human comes out the other end with a deep understanding of a topic so that they can really easily build things in the future?'
我同意所有这些,但还有另一个有趣的视角:Joel Grus 有一个著名的演讲,我们会谈到这个,他说笔记本很糟糕。从软件工程的角度来看,它们真的很差。在当时,也许现在某种程度上,我同意他,因为我做过 ML DevOps。我在大型组织里工作过,试图找出如何弥合数据科学和软件工程之间的鸿沟。Claude Code 已经更偏向软件工程这边,这意味着它创建的是幂等、无状态、可重复的产物。所以正如你所说,从教学的角度来看,这种带状态的反馈很好,因为我能理解发生了什么,但我需要把它转换成可部署的东西。你能说说你回应 Joel Grus 的故事吗?那有点像个闹剧,对吧?但跟我们讲讲那个故事吧。
I agree with all of that, but then there's the other interesting angle: there was a famous talk by Joel Grus, and we'll talk about this, and he said that notebooks are terrible. They're really bad from a software engineering point of view. At the time, and maybe still now to a certain extent, I agree with him because I've done ML DevOps. I've worked in large organizations trying to figure out how to bridge data science and software engineering. Claude Code is already more towards the software engineering side, and what that means is it creates idempotent, stateless, repeatable artifacts. So as you say, from a pedagogical point of view, it's really good having this stateful feedback because I can understand what's going on, but then I need to translate that into something which is deployable. Can you tell us the story of you responding to Joel Grus? It was a bit of a fiasco, wasn't it? But just tell us about that story.
他做了一个非常好的视频,叫《我不喜欢笔记本》。非常搞笑,做得非常好。是的,我完全错了。他说笔记本做不到的所有事情,它们都能做到。他说你不能用笔记本做的事情,我一直在用笔记本做。所以那是一个非常棒、非常有趣的错误演讲。然后我做了一个模仿它的视频,叫《我喜欢笔记本》,基本上我复制了他的大部分幻灯片(并注明了出处),然后展示了每一张都是完全错误的。但我实际上认为你的评论触及了核心,那就是软件工程通常的做法与科学研究及类似工作的做法之间的差异。我同意存在这种二分法,而且我认为这种二分法真的很可惜,因为我觉得软件开发的方式是错误的。它现在的方式完全围绕可重复性和这些死的东西——全是死代码、死文件。我永远无法像 Brett Victor 在他的作品中那样清晰地表达这一点,所以我鼓励没看过 Brett Victor 的人去看看。但他一次又一次地展示,与你正在做的事情建立直接、发自内心的联系才是最重要的。那是他的使命,确保人们拥有那种联系,这基本上也是我的使命。所以对我来说,传统软件工程离那种状态远得不能再远了。我觉得它很恶心。我觉得人们被迫那样工作很可悲。那是不人道的,而且我认为它效果并不好。从经验上看,它效果并不好。而且它对 AI 和人类都不太好。但情况并非一直如此。艾伦·凯和 Smalltalk、艾弗森和 APL、Lisp、沃尔夫勒姆和 Mathematica——对我来说,那是黄金时代,人们专注于如何让人类进入计算机,尽可能紧密地与之协作。例如,鼠标就是由此而来的,可以点击、拖拽,把计算机中的实体可视化为可以移动的东西。所以我觉得我们已经失去了那种东西。我觉得非常遗憾。是的。对于 Claude Code 这类东西,默认的工作方式就是深入其中。就像,好吧,有一个装满文件的文件夹。你甚至都不去看它们。你与它的整个交互都是通过提示词进行的。
He did a really good video called 'I Don't Like Notebooks.' It was hilarious. It was really well done. And yeah, I was totally wrong. And all the things he said notebooks can't do, they can. And all the things he said you can't do with notebooks, I do with notebooks all the time. So it was a very good, very amusing, incorrect talk. So then I did a kind of parody of it called 'I Like Notebooks,' in which I basically copied, with credit, most of his slides and showed how every one of them was totally incorrect. But I actually think your comment about it does come down to the heart of it, which is this difference between how software engineering is normally done versus how scientific research and similar things is normally done. I agree there is a dichotomy there, and I think that dichotomy is a real shame because I think software development is being done wrong. It's being done in this way which is all about reproducibility and these dead pieces—it's all dead code, dead files. I will never be able to express this one millionth as clearly as Brett Victor has in his work, so I'd encourage people who haven't watched Brett Victor to watch him. But he shows again and again how a direct, visceral connection with the thing you're doing is all that matters. That's his mission, to make sure people have that connection, and that's basically my mission as well. So for me, traditional software engineering is as far from that as it is possible to get. I think it's gross. I find it disgusting and I find it sad that people are being forced to work like that. It's inhumane and I just don't think it works very well. Empirically, it doesn't work very well. And it's much less good for AI as well as it's much less good for humans. It hasn't always been that way. With Alan Kay and Smalltalk, Iverson and APL, Lisp, Wolfram with Mathematica—to me, these were the golden days when people were focused on the question of how to get the human into the computer to work as closely with it as possible. That's where the mouse came from, for example, to click and drag and visualize entities in your computer as things you can move around. So I feel like we've lost that. I think it's really sad. Yeah. With Claude Code and stuff, the default way of working with them is to go super deep into it. It's like, okay, there's a whole folder full of files. You never even look at them. Your entire interaction with it is through a prompt.
是的。
Yeah.
这真的让我感到恶心。我真的认为这很不人道,而我的使命和二十年来一样,就是阻止人们这样工作。
It literally disgusts me. I literally think it's inhumane, and my mission remains the same as it has been for like 20 years, which is to stop people working like this.
我知道。但回想起来,我以前和数据科学家一起工作过。他们用 Jupyter 笔记本。我发现通常,那时候你没法把它们提交到 git,因为效果不好。这些数据科学家大多不会用 git。他们会不按顺序运行单元格,这就导致不可复现。诸如此类的问题很多。但问题是,我同意你的观点,你可以在这个工作流中使用它们。
I know. But casting my mind back, I used to work with data scientists. They were using Jupyter notebooks. What I found was typically, back then you couldn't check them into git because it wouldn't look very good. Most of these data scientists didn't know how to use git. They would run the cells out of order, which means it wouldn't be reproducible. There were all sorts of things like that. But the thing is, I agree with you that you can use them in this workflow.
但这又回到我之前说的,呼叫中心是低智力工作。数据科学家做的是智力工作,因为他们创造不存在的东西,勾勒问题的轮廓,在理解不足的领域工作。但乐观的情况是,当数据科学家能简洁描述问题轮廓时,也许我们可以用 Claude Code 来正确实现。但如何在这两个世界之间架起桥梁?
But it comes back to what I was saying before about the call center being a low intelligence job. Data scientists do intelligent work because they create something that doesn't exist. They figure out the contours of a problem, working in a poorly understood domain. But you could argue the bull case is when data scientists can succinctly describe the problem contours, maybe we could go to Claude Code and implement it properly. But how do we bridge between those two worlds?
我认为那是个糟糕的主意。你不应该把人从探索环境中移开。研究和科学是通过建立洞察力来发展的。伟大的科学家,比如费曼,通过与学习对象互动来建立更深的直觉。费曼无法拿起旋转的夸克,但他确实研究过旋转的盘子。你必须找到与工作深度互动的方式。我见过数据科学团队被摧毁,因为软件工程师成为他们的经理,强迫他们停止使用 Jupyter notebook 来使用可复现环境。解决方案不是增加纪律和官僚主义,而是解决实际问题。例如,我们构建了一个 NB 合并驱动。Notebook 对 Git 很友好,但 Git 没有自带合并驱动。Git 只有基于行的文本文件合并驱动,但它是可插拔的。我们为 JSON 文件写了一个。现在差异显示单元格级别的变化,合并冲突也是单元格级别的。Notebook 始终可以在 Jupyter 中打开。NBDime 也做了同样的事。解决方案不是抛弃 Bret Victor 的想法,让人们远离探索工具,而是修复这些工具。所有软件开发者都应该使用基于探索的编程来加深理解,建立强大的心智模型,并提出更好、更渐进测试的解决方案。我很少使用调试器,因为我很少遇到 bug。这不是因为我是一个出色的程序员,而是因为我以小步骤构建东西,每一步都有效,我可以看到并与之互动。没有 bug 存在的空间。
I think that would be a terrible idea. You don't want to remove people from their exploratory environment. Research and science are developed by people building insight. Great scientists, like Feynman, build deeper intuition by interacting with what they're learning. Feynman couldn't pick up a spinning quark, but he literally studied spinning plates. You have to find ways to deeply interact with your work. I've seen data science teams destroyed when a software engineer becomes their manager and forces them to stop using Jupyter notebooks for reproducible environments. The solution isn't more discipline and bureaucracy; it's solving the actual problem. For example, we built an NB merge driver. Notebooks are git-friendly, but Git doesn't ship with a merge driver for them. Git only has a merge driver for line-based text files, but it's pluggable. We wrote one for JSON files. Now diffs show cell-level changes, and merge conflicts are cell-level. The notebook is always openable in Jupyter. NBDime did the same thing. The solution wasn't to throw away Bret Victor's ideas and push people away from exploratory tools, but to fix those tools. All software developers should use exploratory programming to deepen their understanding, build strong mental models, and come up with better, incrementally tested solutions. I rarely use a debugger because I rarely have bugs. It's not because I'm a great programmer; it's because I build things in small steps, each step works, and I can see and interact with it. There's no room for bugs.
我对此很纠结。我同意你的观点,但我对那些说组织会收敛到某种做事方式、不再需要进化的人持怀疑态度。创新就是适应性。我们应该尽可能增加适应性的表面积。我们需要人们不断测试新想法,发现约束。但我们也需要使用云、CI/CD,并将东西投入生产。
I'm torn on this. I agree with you, but I'm skeptical of people who say organizations converge on ways of doing things and no longer need to evolve. Innovation is adaptivity. We should increase the surface area of adaptivity. We need people constantly testing new ideas, finding constraints. But we also need to use the cloud, CI/CD, and get things into production.
是的,绝对如此。NBD 自带开箱即用的 CI 集成。测试就在那里,因为源代码是 notebook。整个探索——API 如何工作、调用时是什么样子、函数的实现、示例、文档和测试——都在一个地方。所以在这个环境中更容易成为好的软件工程师。两者都做。
Yes, absolutely. NBD ships with out-of-the-box CI integration. The tests are literally there because the source is a notebook. The entire exploration of how an API works, what it looks like when you call it, the implementation of functions, examples, documentation, and tests are all in one place. So it's much easier to be a good software engineer in this environment. Do both.
你还记得那份由 Hinton 和 Demis 签署的声明“存在风险应成为紧急优先事项”吗?你和 Aravind(那个蛇油推销员)一起写了反驳。跟我说说。你认为我们应该担心 AI 的存在风险吗?
Do you remember the statement 'Existential risk should be an urgent priority' signed by Hinton and Demis, and you responded with a rebuttal with Aravind, the snake oil guy? Tell me about that. Do you think we should be worried about AI existential risk?
那是一个特定时期,情况已经变了,谢天谢地。我觉得不仅是我和 Ara,而是我们所在的社区大体上赢了那场辩论。现在我们有其他问题要担心。当时的主流叙事是 AI 随时可能变得自主并毁灭世界。这来自 Eliezer Yudkowsky 的工作,我认为在很多层面上已被证明是错误的。他们当然会反驳,就像任何末日邪教一样,除非你给出一个日期并且日期过去了。我稍微更新了观点:我现在认为这些模型在受限领域可以说是有智能的。ARC 挑战证明了这一点。如果你施加约束,你可以更快地朝着已知目标前进。甚至智能体:你可以放一个规划器,如果你知道要去哪里,你可以更快到达。但如果你没有知识和约束,那也没用——你会更快地走向错误方向。他们似乎没有意识到这些模型实际上并不了解世界。这些与 Aravind 和我的观点无关,我们的观点过去和现在都是:它误解了真正的危险所在。真正的危险是,当一项强大得多的技术进入世界时,它可以使某些人变得强大得多。热爱权力的人会寻求垄断这项技术,技术越强大,那些渴望权力的人的冲动就越强烈。
That was a certain time, and things have changed, thank God. I feel like we—not just me and Ara, but broadly the community—kind of won that. Now we have other problems. At that point, the prevailing narrative was AI is about to become autonomous at any moment and could destroy the world. That comes from Eliezer Yudkowsky's work, which I think has clearly been shown to be wrong at many levels. They would refute that, of course, just like any doomsday cult unless you give a date and the date passes. I've updated a bit: I now think these models can be said to be intelligent in restricted domains. The ARC challenge showed that. If you place constraints, you can go faster toward a known goal. Even agency: you can put a planner on there and go faster if you know where you're going. But that doesn't help if you don't have knowledge and constraints—you go in the wrong direction faster. They don't seem to appreciate that these models don't actually know the world. None of that was relevant to Aravind and my point, which was and is that it's misunderstanding where the actual danger is. The real danger is that when a dramatically more powerful technology enters the world, it can make some people dramatically more powerful. People who love power will seek to monopolize that technology, and the more powerful it is, the stronger that urge from power-hungry people.
所以,问题来了。如果你说,我不管那些,我只关心自主 AI 的崛起,比如奇点、纸夹、纳米粘液之类的。显而易见的解决方案是,哦,让我们集中权力。这正是我们当时不断看到的。让非常富有的科技公司或政府,或者两者,拥有所有这些权力,并确保其他人没有。在我的威胁模型中,这是最糟糕的做法,因为你把控制权集中在一个地方,因此那些渴望权力的人只需要接管那个东西。
So to ignore people, so here's the problem. If you're like, I don't care about any of that. All I care about is autonomous AI taking off, you know, singularity, paperclip, nano goo, whatever. The obvious solution to that is, oh, let's centralize power. And this is which is what we kept seeing particularly at that time. Let's give either very rich technology companies or the government or both all of this power and make sure nobody else has it. In my threat model, that's the worst possible thing you can do because you've centralized the ability to control in one place and therefore these people who are desperate for power just have to take over that thing.
但我们能不能区分一下你说的权力是什么意思?因为我们刚才花了一些时间讨论它实际上并不像人们想象的那么强大。
Could we distinguish though what you mean by power? Because we've just spent some of this conversation talking about how it's not actually as powerful as people think it is.
但我的观点是“即使”的情况,对吧?我只是说,即使它变得极其强大,我也不想争论它是否会强大,因为那是推测。即使它变得极其强大,你也不应该把所有权力集中在一家公司或政府手中。
But I'm not even that's what but mine is an even if thing, right? So, like I'm just saying even if it turns out to be incredibly powerful, right? Like I don't even want to argue about whether it's going to be powerful because that's speculative. Even if it's going to be incredibly powerful, you still shouldn't centralize all of that power in the hands of one company or the government.
是的。
Yeah.
因为如果你这样做,所有权力都会被渴望权力的人垄断,并用来摧毁文明。基本上,你会看到所有财富和权力集中在那些希望集中权力的人手中。社会几百年来一次又一次地面对这个问题。比如,写作曾经只有最精英的人才能接触。同样的论点也被提出:如果让每个人都写作,他们会用来写我们不想让他们写的东西,那会很糟糕。印刷术也是如此,投票权也是如此。一次又一次,社会必须对抗那些拥有现状权力的人的自然倾向,他们说,不,这是一种威胁。所以当我们说,好吧,如果 AI 变得极其强大,对社会来说,是把它掌握在少数人手中更好,还是分散到整个社会更好?
Because if you do, all of that power is going to be monopolized by power-hungry people and used to destroy civilization. Basically, you'll end up with a case where all of that wealth and power will be centralized with the kinds of people who want it centralized. So like society for hundreds of years have faced this again and again and again you know so when it's like you know writing used to be something that only the most exclusive people had access to knowing about writing. And the same arguments were made. If you let everybody write, they're going to use it to write things that we don't want them to write and it's going to be really bad. You know, ditto with printing, ditto with the vote. Like, and again and again, society has to fight against this natural predilection of the people that have the status quo power to be like, no, this is a threat. So when we're saying like, okay, what if AI turned out to be incredibly powerful, would it be better for society to be that to be kept in the hands of a few or spread out across society?
我的观点是后者。还有一种观点是,别担心,它不会那么强大。我只是不想讨论那个,因为那不是一个容易赢的论点,因为你无法预测会发生什么。我们都在猜测。但我可以很明确地说,如果它发生了,只让埃隆·马斯克拥有它是个好主意吗?或者只让唐纳德·特朗普拥有它是个好主意吗?
My argument was the latter. Now, there's also an argument which is like, don't worry about it. It's not going to be that powerful anyway. I just didn't want to go there because it's not an argument that's easy to win because you can't really say what's going to happen. We're all just guessing. But I can very clearly say like, well, if it happens, would it be a really good idea to only let Elon Musk have it or would it be a good idea to only let Donald Trump have it?
Dan Hendrycks 谈到了这种攻防不对称。所以对我们来说,拥有制衡的防御非常重要。但让我们先假设这一点,因为显然当我们看 Meta 和 Facebook 时,权力不平衡很明显。他们控制了我们所有的数据。他们知道我们在用 OpenAI 和 Claude 做什么。所以情况并不像我们想象的那么好,因为实际上人类仍然需要参与。但例如,他们拥有我们所有的数据,对吧?你可能正在开发一些创新技术,你使用 Claude,你把所有信息发送上去,他们现在可以复制你。我的意思是,你具体在说什么风险?
Dan Hendrycks spoke about this offense-defense asymmetry. So it's actually very important for us to have countervailing defenses. But let's just take that as a given for a minute because obviously when we look at something like Meta and Facebook, it's quite clear what the power imbalance is. You know, they control all of our data. They know what we're doing with something like OpenAI and Claude. So it's not as good as we thought it was because actually humans still need to be involved. But for example, they have all of our data, right? and you might be working on some new innovative technology and you're using Claude and you're sending all of your information up there and they can now copy you. I mean what kind of risks are you talking about to be more concrete?
是的。不,我的意思是我不是在谈论那些事情,对吧?当时我在讨论这个推测性问题:如果 AI 变得极其强大怎么办?比如现在他们说这是新的生产资料,这对我来说完全是夸张的说法,但根据你最好的估计,如果存在风险,它们是什么?
Yeah. No, I mean so I was not talking about any of those things, right? So at the time I was talking about this speculative question of what if AI gets incredibly powerful? I mean like now for example they say that this is the new means of production and that seems completely hyperbolic to me but like in your best estimation now if there are risks what are they
如果当前技术存在风险,我认为其中一些是我们讨论过的,即人们通过基本上失去随时间变得更强的能力而削弱自己。那是我最担心的重大风险。隐私风险是存在的,但我不确定它比之前谷歌和微软的情况严重多少。你知道,你曾在微软工作,你知道他们拥有多少普通 Outlook、Office 等用户的数据,谷歌也一样,普通 Google Workspace 或 Gmail 用户的数据。这些隐私问题是真实的,尽管我认为围绕这些公司存在更大的隐私问题,政府可以将数据收集外包给它们。过去是 ChoicePoint 和 Axiom 这样的公司,现在可能更多是 Palantir 这样的公司。美国政府实际上被禁止建立关于美国公民的大型数据库,但公司不被禁止这样做,政府也不被禁止与这些公司签订合同。所以,那是一个巨大的担忧,但我不认为这是 AI 独有的。当然,你在英国,你知道,英国监控已经普遍存在相当一段时间了。它确实使利用监控变得更容易,但一个资源充足的组织可以投入一千个人力来解决这个问题。所以,是的,我不确定这些是隐私问题,也许它们比以前更常见了。
if there are risks with the current state of technology I mean I think some of them are the ones we've discussed which is people enfeebling themselves by basically losing their ability to become more competent over time. That's the big risk I worry about the most. The privacy risk, it's there, but I'm not sure it's much more there than it was for Google and Microsoft before. Like you know, you used to work at Microsoft, you know, how much data they have about the average Outlook, Office, etc. user, ditto for Google, you know, the average Google Workspace or Gmail user. Those privacy issues are real although I think there are bigger privacy issues around these companies which the government can outsource data collection to. So back in the day it used to be companies like ChoicePoint and Axiom, nowadays it's probably more companies like Palantir. The US government is actually prohibited from building large databases about US citizens for example, but it's not prohibited. Companies are not prohibited from doing so, and the government's not prohibited from contracting things to those companies. So, I mean, that's a huge worry, but I don't think it's one that AI is uniquely creating. It certainly, you're in the UK, as you know, in the UK surveillance has been universal for quite a while now. It certainly makes it easier to use that surveillance, but a sufficiently well-resourced organization could just throw a thousand bodies at the problem. So, yeah, I'm not sure these are due privacy problems as maybe more common ones than they used to be.
是的。
Yeah.
Jeremy,我刚注意到时间。我得去机场了。
Jeremy, I've just noticed the time. I need to get to the airport.
好的,这太棒了。
All right, this has been amazing.
谢谢你,先生。感谢你的到来。
Thank you, sir. Thank you for coming.
是的。希望你旅途愉快。非常感谢。
Yeah. Hope you had a nice trip. Thank you so much.