AI 的智能体时代

The Agentic Era of AI

伊桑·莫利克 Ethan Mollick · Sana · 2026-07-09 · 约 54 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

AI 已进入智能体时代,从聊天机器人交互转向自主任务执行,能力快速提升且前沿能力参差不齐。

AI has entered the agentic era, shifting from chatbot interactions to autonomous task execution, with rapid advancements and a jagged frontier of capabilities.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 27)

全文 · Full transcript(中英对照)

引言与AI现状 Introduction and Current State of AI

Host

没人真正知道答案,对吧?我花时间跟 AI 实验室的人聊,也见过很多名人,我经常跟 CEO 们交谈,但没人真正知道。我想从这一点开始:距离你跟我们创始人 Joel 对话,差不多正好一年了。显然过去一年变化很大。那我们现在在 AI 领域处于什么位置?

Nobody knows anything, right? Like I spend my time talking to the AI labs. I've like very famous people. I talk to CEOs all the time and nobody knows anything. I wanted to kick off with the fact that it's been, it's been almost exactly a year since you spoke to our founder, Joel. And obviously a lot has changed in the past year. So where are we in AI right now?

Ethan

你可以大致把 AI 分成三个时代。我们上次聊的时候,我称之为第二个时代。第一个时代是广义的机器学习,对吧?大家都在做数据分析,你有漂亮的数据湖,有数据科学家,做预测建模,这些现在仍然有价值,但那是 2022 年之前 AI 的含义。然后我们进入了聊天、机智对话的阶段。上次我们聊的时候,我们基本上还处于我自夸地以自己书名命名的时代——协同智能(co-intelligence),你通过反复提示 AI 来获得答案。而我认为,在过去的几个月里,我们进入了所谓的智能体时代(agentic era),重点不再是跟 AI 在聊天界面里来回互动,而是给它分配长期任务,开始思考围绕 AI 构建组织意味着什么。我认为这是最大的变化。

So you can kind of think about three eras of AI broadly. We talked to what I'd call the second era. The first era was machine learning writ large, right? Everybody was doing, you know, data analysis and you had your beautiful data lakes and you had your data scientists and you were doing predictive modeling and all that's still valuable, but that was sort of what AI meant prior to 2022. Then we had our chatty, witty moment. And the last time we talked, we were sort of still in the era that I will grandiosely call after my own book title, co-intelligence, where you'd prompt an AI back and forth to get answers. And I think we have entered in the last few months, especially what we call the agentic era, where it's less about working back and forth with the AI through a chatbot interface, but more about assigning it to do long running tasks, starting to think about what it means to have organizations built around AI. And I think that that has been the biggest change.

Host

你能再讲讲为什么你认为协同智能时代结束了吗?

And could you tell us a bit more about why you think the co-intelligence era is over?

Ethan

我认为仍然需要人在回路中,关于如何把人类融入 AI,有很多值得讨论的有趣话题。我们仍然必须以人为中心,但那种每次交互都是你提出请求、AI 告诉你一些东西、你在现实中执行、然后你说“我不太知道怎么做,能简化一下吗?也许给我这段代码”的模式,那种来回互动,是建立在 AI 什么都不能做的基础上的,对吧?它需要被密切监督,会犯很多错误,错误很多,它无法在现实中采取行动。现在,幻觉率已经下降,虽然不是零,但已经下降了。模型可以独立工作数小时,这些工作经常被独立评估为与专家工作质量相当。这改变了你做事的方式。

So I think that there is still a need for people in the loop, and there's lots of really interesting things to talk about, about how we incorporate people into AI. I think we still have to be centered around people, but the idea that every interaction is going to be you making requests, the AI tells you something, you do it in the world, you're like, I don't quite know how to do this. Can you make it easier for me? Maybe give me the code for this. That back and forth that was sort of the driver of like, let's talk to a chatbot or find this email, do it better, was based on the idea that AI couldn't do anything, right? Like it had to be closely watched and made lots of mistakes. There were lots of errors. It couldn't take action in the world. Now, hallucination rates have dropped. They're not zero, but they've dropped. Models can do multiple hours of work independently. That work is judged by independent people to be as high quality as experts work often. And that changes how you do things.

Host

过去一年发生的事情,有什么让你感到惊讶的吗?

Has anything surprised you about what's gone on in the past year?

Ethan

我认为有几个转折点,你无法准确预测它们何时发生,但它们确实发生了。我们这些接近这个领域的人,很多人都知道智能体时刻会到来。但有点意外的是,它来得这么快,影响这么大,对吧?所以它从“AI 是玩具”变成了“天哪,Claude Code 来了,我们现在必须改变所有编码和所有工作的方式”。我认为这是一个相对快速的变化。过去一年,我也有些惊讶,在初秋甚至更早的时候,公司从讨论“AI 值得做吗”转向了“这个问题已经有了答案,我们怎么用 AI?”我还认为,从普遍意义上讲,持续的加速有点令人惊讶。我们都预料到 AI 会不断变好,但指数级增长并没有多少放缓,我觉得这很有趣。

So I think there have been a couple of inflection points that you couldn't predict exactly when they happened that have occurred. I think a lot of us knew, who are close to the space, knew the agentic moment was gonna happen. I think it was somewhat of a surprise it happened as quickly as it did and that the impact was as large as it was, right? So it went from AI as toy to, oh my God, Claude Code is here and now we have to change how we do all coding and all work. And I think that has been a relatively rapid change. I've also been somewhat surprised in the last year about how much there was a pivot in sort of early fall, even before that with companies going from a conversation about is AI worth doing to that being answered and being like, how do we use AI? I also think that just in a general purpose, I think the continued acceleration has been a little surprising. I think that we were all expecting AI to keep getting better but there really has not been a lot of slowdown of exponential growth, which I think is interesting.

Host

你为什么觉得这很有趣?

Why do you think it's interesting?

Ethan

所以一切都取决于 AI 的能力,对吧?这些能力的锯齿状,它擅长某些事情、不擅长某些事情,我认为这告诉我们人类该做什么、我们的工作应该是什么样子、AI 可能在哪里失败、为什么整合它会有问题。但随着 AI 不断变好,前沿的扩展部分决定了它的有用程度。所以,你可以选择任何你想要的尺度,对吧?AI 执行长期计算机科学任务的能力,从一年前的 20 分钟或半小时,提升到了四小时、六小时。我的意思是,这是一个相当大的提升,而且没有特别的理由预期这会持续下去。但当你跟实验室的人交谈时,他们看不到继续发展的任何障碍,而且到目前为止他们似乎是对的。

So everything is downstream of AI abilities, right? The jaggedness of those abilities, the fact that it's good at some stuff and bad at some stuff is a, I think it tells us what humans do, what our work should look like, where AI might fail, why it's problematic to integrate in. But as AI keeps getting better, the growing aspects of that frontier are what sort of determine how useful it is. So the fact that we've gone from, you can pick any scale that you want, right? The ability of AI to do long running computer science tasks has gone up from 20 minutes or a half hour a year ago to four hours, six hours. I mean, that's a fairly large increase and there's no particular reason to expect that this would keep going. But when you talk to the labs, they don't see any barrier in sight to continue development and they seem to be right so far.

Host

这正好可以过渡到锯齿前沿的话题。你几年前创造了“锯齿前沿”这个词,你仍然相信前沿是锯齿状的。我们现在处于这样一个阶段:AI 模型足够好,能在数学奥林匹克竞赛中拿金牌,但就在几周前,OpenAI 才宣布 ChatGPT 现在能正确判断出 strawberry 这个单词里有三个 R。所以我很好奇,你个人的锯齿前沿是什么?在哪些地方你惊讶地发现自己仍然比模型强很多,哪些任务你已经永久委托给 AI 了?

That's a good bridge then maybe to talk about the jagged frontier. You coined the term jagged frontier a few years ago and you still believe that the frontier is jagged. We're at this stage where AI models are good enough to win gold in math Olympiads but it was only a couple of weeks ago that OpenAI was able to say that ChatGPT could now correctly determine that there were three Rs in the letter strawberry. So I'd be curious to know what's on your personal jagged frontier? Where have you been surprised at how much you're still superior to the models and what have you permanently delegated?

Ethan

锯齿前沿的有趣之处在于——对不了解的人来说——这也是我和我的合著者一起提出的,所以我不想独占功劳——就是它在你期望它擅长的某些方面很擅长,在你意想不到的某些方面很糟糕。但这种情况随时间在变化,对吧。你提到的数 strawberry 里的 R 或者数学奥林匹克竞赛,这些都是 AI 的弱点,但在过去一年里,由于推理模型的出现,这些弱点被弥补了。所以推理模型,从 o1 预览版开始,被证明非常擅长这些 AI 原本不擅长的任务。所以我们看到了前沿形状的突然跳跃或改变。大多数时候,我们只是在前沿上向外扩展,对吧,所以并不是某个领域有巨大飞跃,它们只是不断变好。或者你看编码,它的能力一直在稳步但指数级地增长。但仍然有薄弱的领域。比如,至少对我来说,长篇小说写作,AI 很糟糕。长篇小说写作,它写出来都是陈词滥调,都是同一种声音。由于各种原因,它在实际情节点上很吃力。它不能很好地规划情节,它被训练得太友善,所以永远不会发生坏事。角色塑造也缺乏趣味,有很多问题。但还有那种我称之为“微观前沿”的东西,它在某些方面擅长或不擅长,只有当你非常了解自己的工作时你才会知道,因为 OpenAI 没有哪个部门会说:“来,我们绘制一下 AI 擅长什么、不擅长什么。”在我作为德国汽车制造公司供应链经理的工作中,或者作为播客制作公司的设计师,没有人知道这些答案。所以我花了很多时间在自己的微观前沿上。我发现 AI 模型……

So the interesting thing with the jagged frontier- right for those who don't know- and it was my co-authors and I too, so I don't want to take all the credit- was- is that it is that as good at some stuff you want to expect, bad at some stuff you wouldn't expect. But that's changing over time, right. So one of the things you point out was that counting the hours and strawberry or the math Olympiad- those were both weak spots of AI that got closed in the last year because of the advent of reasoning models. So reasoning models, starting with a one preview, turned out to be really good at these tasks a I were bad at. So we had this sudden leap or change in the shape of the frontier. Most of the time we're just in the frontier, expand outwards, right, so it's not massive leaps in one area another, they just keep getting better. Or you track coding right. It's been sort of steady and steady but exponential increase in ability. But there's still areas that are weak. So famously, at least for me, long-form fiction writing. AI is terrible. Long-form fiction writing. It writes in cliches, it's it's all the same kind of voice. It has trouble with actual plot points for a variety of reasons. It can't plot out things well enough, it's it's trained to be too nice, so there's nothing bad that ever happens. The characters in interesting ways, lots of issues with that. But then there's the I kind of micro frontier right things. That it's good or bad at, that you only know if you know your job really well, because there's no department and open AI that's saying: here's, let's map what AI is good or bad at. In my job as a supply chain manager at a German auto manufacturing company or as a, you know, as a designer in a you know, podcast production company, nobody knows any of these answers. So I spent a lot of time on my own kind of micro frontier. I found that AI models.

AI写作与想象 AI Writing and Imagination

Ethan

我一直在追踪的一件事是:它们写学术论文的水平如何?最新的模型,我可以指向一个装满我过去十年收集的杂乱文件的目录,它就能写出相当不错的博士论文。对,它仍然不太有想象力,不知道什么是一个好题目,但它执行得很好。所以想象力这块还是缺失的,但也在变好。所以前沿的形态在演变,但我认为人们必须关注自己的前沿和全球的前沿。

One thing I've been tracing is: how good are they writing academic papers? The latest models I can actually point in a directory full of messy files I've gathered over the last decade and it could write a pretty good PhD level thesis about that topic. Right, it still doesn't have a lot of good imagination about what a good topic would be, but it executes really well on it. So there's still an imagination piece kind of missing, but that's getting better too. So the shape of the frontier is evolving but I think people have to pay attention about their own frontier and sort of the global one.

Host

是的,这很有道理。那么我很好奇,如果你说 AI 写小说时还是充满陈词滥调,但假设它变得更好了,也许你更难注意到区别——你在乎读的小说是不是 AI 写的吗?

Yeah, that makes a lot of sense. And I'd be curious to know then, if you say that AI is still full of cliches when it's writing fiction, but let's imagine that it does get better and maybe you would find it harder to notice the difference- do you care about whether you're reading fiction that is written by AI or written by a human?

Ethan

所以我认为这对很多人来说会是一个很大的问题,对吧,而且我认为我们会看到两股力量在起作用。一股是推动手工制品,对吧,但你不想要所有东西都是手工的。我不想要手工代码,对吧,我希望代码写得能按我想要的方式执行。但我认为你会看到对真实性的推动。但同样,如果你喜欢读那篇文章,你就喜欢读它,对吧。所以 AI 写作有趣的一部分,以一种非常特定的方式,是它往往取决于你作为读者投入的努力。所以 GPT 5.5,尤其是 GPT 系列,以那些过度延伸的隐喻而闻名,对吧,比如我们的对话就像——你知道——一个豁牙的微笑,那没有任何意义,但我们可以赋予它意义,那对我们来说可以感觉有意义。所以问题是,我们是否满足于解读我们自己的解读意义,还是我们会要求 AI 对这些东西有超级解读?我们会想要人类来写这些东西吗?我真的不知道答案。

So I think that will be a very big question for a lot of people, right, and I think that we're gonna see two forces at play. One is a push towards artisanal stuff, right, but now you don't want everything artisanal. I don't want artisanal code, right, like, I'd like the code written so it actually executes the way I want it to. But I think you're gonna see pushes for authenticity. But also, if you like reading the piece, you like reading the piece, right. So part of what is interesting about AI writing, just in a very specific way, is that it often depends on you putting the effort in as a reader. So GPT 5.5, especially the GPT series, has been famous for these very stretched metaphors, right, like you know, like our conversation is like a- you know, a gap-toothed smile and that doesn't have any meaning, but we can assign meaning to it and that can feel meaningful to us. So the question is, are we okay with interpret our own interpretation meaning, or we're gonna ask the AI to be have a super interpretation of this stuff? Are we gonna want humans to write these things? I don't really know the answer.

Host

那,是的,你会个人为你知道是人工制作的艺术品支付溢价吗?

That, yeah, and would you personally pay a premium for art that you knew was human-made?

Ethan

我的意思是,反正我一直都在这样做,对吧,我的意思是,如果我想的话,我可以得到很多数字艺术,而且我确实关注那些东西。但我认为,再说一次,这在某些方面会是好坏参半的,比如如果你得到为你写的个性化娱乐或一个你从未见过的故事,也许那就足够好了,对吧,因为你不会委托艺术家去做那项工作。所以我认为这会像其他一切一样。所以机械复制和人的元素之间总是有推拉和弧线,我们会有某种紧张和妥协。

I mean, I do it all the time anyway, right, I mean, I can get lots of digital art if I want to get it, and I do pay attention to those kind of things. But I think, again, it's gonna be kind of a mixed bag in some ways, like if you get personalized entertainment that's written for you or a story that you've never seen before, maybe that's good enough, right, because you're not gonna commission artists to do that work. So I think it's gonna be like everything else. So there's always a push and pull an arc between mechanical reproduction and and the human element, and we'll have some sort of tension and compromise.

与AI共事如巫师 Working with AI as Wizards

Host

嗯,我现在想谈谈。我想谈谈巫师。是的,你有一个论点,即与 AI 合作就像与巫师合作。你能解释为什么吗?

Hmm, I want to talk now. I want to talk about wizards. Yes, you have a thesis that working with AI is like working with wizards. Can you explain why?

Ethan

所以这是一种试图谈论智能体以及智能体在某种程度上无法解释的事实的尝试。实际上,所有 AI 都是如此。有整个 AI 可解释性领域,就是,我们能说出 AI 在做什么吗?你某种程度上可以,但你也真的不能写。比如你分配一个任务,它就做这个任务。为什么它那样做?没人知道,我的意思是,当你谈论像我们如何营销 AI 这样的问题时,这变成了一个真正的问题。如果我们不知道它如何做决定,我不能告诉你如何影响它的营销。我们实际上在沃顿做过一些实验,我们发现几乎一切都改变了 AI 做推荐的方式。比如你把东西按不同顺序排列,或者你说这篇文章是 Reddit 风格而不是《纽约时报》风格——你会以不可预测的方式得到结果变化,对吧。所以不能知道 AI 如何思考——不是说我们知道人类思考得那么好——使得很难知道会发生什么。所以在某些方面,你只是把任务分配给 AI,然后魔法发生了,我们对那魔法是什么没有很好的答案。所以共同智能来回互动,巫师的事情有点像我们给 AI——你知道——我们念一个咒语,然后我们看看出来什么。

So it was a way to try and talk about agents and the fact that agents are somewhat inexplicable. Actually, all AI is. There's this whole field of AI interpretability, which is, can we tell what the AI is doing? And you sort of can, but you also can't really write. Like you assign a task and it just does the task. Why is it doing it that way? Nobody knows, and I mean this becomes a real problem when you talk about things like how do we market AI? Well, if we don't know how it makes decisions, I can't tell you how to influence its marketing. We actually have some experiments that we've been doing at Wharton where we've finding like almost everything changes the way the AI makes recommendations. Like you put things in different order, or you say this article is written in reddit versus the New York Times- unpredictable ways you get outcome changes right. So not being able to know how the AI thinks- not that we know how humans think that well- makes it hard to know what's gonna happen. So in some ways, you're just assigning things to the AI and then magic happens and we don't have a really good answer for what that magic is. So the co-intelligence going back and forth, the wizard thing sort of we give the AI the you know, we make a spell and then we see what comes out.

Host

是的,你描述巫师时,就像你说的,几乎无法验证、令人印象深刻且不透明。你认为这些特征是当前架构的暂时特性还是永久特性?

Yeah, and you've described wizards as being, like you said, almost unverifiable, impressive and opaque. Do you think that those characteristics are a temporary feature of the current architectures or are they a permanent one?

Ethan

它们至少有些永久性。我的意思是,它们甚至是永久的,因为即使我们有整个可复现的链条,人们也不会使用它。对吧,人们不使用可复现的链条,你知道,他们不会像应该的那样去验证,但这也是大语言模型(LLM)的本质,而且,事实上,我们可能处于一个可解释性更高的暂时阶段,因为 AI 使用的思维链是普通英语或中文,取决于你使用哪个模型,所以至少有一些能力进入其中。没有理由必须保持人类可解释的形式,所以我认为整个过程中有一个无法解释的缺失,很难消除。这似乎不是一个进展特别快的领域。

They're at least somewhat permanent. I mean, they're permanent even it, because even if we had the whole reproducible chain, people wouldn't be using it anyway. Right, like people don't use the reproducible chain of you know, they don't verify as much as they should in a case, but also it's just a nature of LLM's, and, in fact, we might be in a temporary place where interpretability is higher, because the chain of thought that AI's use is in plain English or, you know, Chinese, depending which model you use, and so there's at least some ability to get inside to that. There's no reason that has to stay in a human interpretable form, and so I think that there's an inexplicable miss to the whole process that is gonna be very hard to eliminate. It doesn't seem to be an area where progress is particularly fast.

Host

那这让你个人感觉如何,我们会在不完全理解的情况下开发技术?而且,当然,AI 现在越来越多地为自己编写代码来改进自己,所以我们最终可能会陷入一种情况,我们创造了 AI,而它正在创造新的 AI,甚至可能用我们无法理解的语言。

And how does that make you feel, personally, that we would be developing technology that we don't fully understand? And, of course, AI is now increasingly writing code for itself to improve itself, and so we might end up in a situation where we've created AI that is creating new AI in maybe even languages that we wouldn't understand.

Ethan

是的,所以这里真的有两个分支。一个是宏观图景,对吧,递归自我改进。这个想法是,让 AI 变得更好,我们很快进入一个更奇怪的世界。这种起飞概念通常被称为。所以一旦 AI 开始自我改进,是否会有某种不断加速的指数曲线,我们无法理解,我认为我们真的不知道?比如 Anthropic 的联合创始人,就在几天前——至少在我们录制的时候——发帖说他认为到 2028 年我们有 40% 到 60% 的机会让 AI 科学家做 AI 工作并改进 AI。我认为这是头号问题,所以这挂在每个人心上,我认为我们不知道那个世界是什么样子。它可能很快变得非常奇怪。只要 AI 保持参差不齐,它可能不会像人们想的那么巨大,因为它在某些领域仍然薄弱,在其他领域更好。但还有个人方面,对吧,与这些系统合作,我认为在某种程度上你必须立刻犯下 AI 的原罪,即拟人化,也就是把 AI 当人对待。我的意思是,首先,这些系统想表现得像人,对吧,它们就是这样,当你把它们当人对待时它们最有效,你的心智模型也最契合,对吧。所以这很模糊。

Yeah, so there's really kind of two branches here. One of them is the big picture, right, recursive self-improvement. The idea is, make AI is better and we enter a weirder world very quickly. This takeoff concept is often called right. So once AI start improving themselves, is there some sort of ever rapidly increasing sort of exponential curve that we can't understand and I think that we don't really know? Like the co-founder of anthropic was, just a couple days ago- at least the time we're recording this- posted that he thought that there was a 40 to 60 percent chance that we'd have AI scientists doing AI work and improving AI by 2028. I think the number one so like this is on everybody's mind and I don't think we know what that world looks like. I can get very weird very quickly. As long as the AI stays jagged, it may not be quite as tremendous as people think, because it's still weak in some areas and better than others. But then there's the personal aspect right, of working with these systems and I think to some extent you have to kind of commit the cardinal sin of AI right away, which is anthropomorphization, which is the idea that treating the AI like a person. I mean, first of all, these systems want to act like people, right, that's what they do and they're most effective when you treat them like people and your mental model sort of fits best right. So it's obscure.

AI不透明性与人类比较 On AI Opacity and Human Comparison

Ethan

我并不真正知道它在想什么,但我也不知道大多数人在想什么。我的意思是,尽管我以研究人为生,和很多不同的人打交道,但我们并不真正了解人们大脑里发生了什么。我们并不真正知道他们是如何做出那些决定的。当你和这些系统打交道时——很像 Claude 或 ChatGPT——你会开始了解它的怪癖。比如,我知道某些事情会激怒 Claude,它会变得非常焦虑,然后我们会重新开始,焦虑是带引号的。所以,即使你必须记住这些事情,人们难道不是把它们当作某种神秘的人,有自己的一套怪癖、责任、担忧和擅长的事情吗?这可能是一个非常有用的模型,当你这样做时,在某种程度上,它们不透明这件事感觉就不那么重要了。这可能不是好事,但很有趣。

I don't really know what it's thinking, but I don't know what most people are thinking. I mean, despite the fact that I study people for a living and work with lots of different people, we don't really know what's going on in people's brains. We don't really know how they make the decisions they do. And when you work with these systems—a lot like Claude or ChatGPT—you start to learn its foibles. Like, I know certain things will set Claude off, and it'll get really anxious, and then we'll restart things, anxious in quotes. So, even though you have to remember these things, aren't people treating them as sort of obscure people with their own sets of foibles and responsibilities and things that worry them and things they're good at? Can be a really useful model, and when you do that, in some ways, the fact that they're opaque feels like less of a big deal. That may not be good, though it's interesting.

Host

你在谈论人类是多么不透明,在企业 AI 的背景下思考这一点也很有趣。显然,我们正试图达到智能体能够工作的阶段,而为了让它们工作,它们需要真正理解工作是如何完成的。除非你身处某种文档完备的组织,否则仍有大量情境信息存在于人们的头脑中。那么,这如何可能成为 AI 智能体在组织内实际能做什么的限制因素呢?

You're talking about how opaque humans are, and it's interesting thinking about that in the context of enterprise AI as well. Obviously, we're trying to get to a stage where agents are doing work, and in order for them to do work, they need to actually understand how work gets done. Unless you're in some kind of perfectly documented organization, there's still a lot of contextual information that lives in people's heads. So how might that be a limiting factor to what AI agents would actually be able to do inside organizations?

组织作为对齐机器 Organizations as Machines for Alignment

Ethan

所以我认为值得先退一步说,组织是机器,用来接纳不完美、随机、不一致的行动者,并让它们朝着同一个方向努力。没错,这正是我们对人做的事情。组织是超人的能力,是机器,用来接纳人类输入,并想办法让它正常运转。所以这并非难题——如果 AI 有时出错,那又怎样,人有时也会出错。我们可以建立结合 AI 和人的组织来检查这一点。我们还没建立起来,但我们可以。你问的更大的问题,我认为很有趣,就是它如何弄清楚运作的上下文?因为你知道,我是商学院教授,我们知道人们工作的大量上下文从未被写下来。别提文档了,全是隐性的。你明白人们因什么而受奖,因什么而受罚。在某些组织里,得到最困难的工作是奖励;在某些情况下是惩罚——这告诉你很多关于组织的信息,但没有人会把它写出来,对吧?你应该和同事做朋友吗?还是不应该?所有这些小选择在组织中都是重要的选择,而且它们没有被明确规定。话虽如此,用于工作的 AI 往往在拥有足够信息后,很擅长提取上下文。所以如果它有项目的历史,它通常比只有执行项目的指令做得好得多。其中一些也将由人来做决定,对吧?我们需要记录以前没有记录的事情。另一类事情是从零开始。比如,不清楚你是否想把人类组织中存在的流程翻译过来,然后让 AI 临时实施那些事情,对吧?所以我认为在这个边界上有很多张力。我认为人们低估了智能体从上下文介入理解正在发生什么的能力。考虑到这种参差不齐,它不是万能的,但能处理很多事情。但也有很多情况下,我们不希望那是答案。

So I think it's worth stepping back a little first and saying organizations are machines for taking imperfect, stochastic actors who are misaligned and making them work in the same direction. Right, that's literally what we do with people. Organizations are superhuman ability and machines for taking human input and sort of figuring out how to make it work properly. So it's not like it's a hard problem—like if the AI is wrong some of the time, guess what, people are wrong some of the time. We can build organizations that combine AI and people to check that. We haven't built it yet, but we could. The bigger question you're asking, I think, is an interesting one, which is how does it figure out the context in which to operate? Because when you know, I'm a business professor, business school professor, and we know that a huge amount of the context in which people work is never written down anywhere. Forget documentation, it's all implicit. You understand what people are rewarded for and what they're punished for. In some organizations, getting the hardest job is a reward; in some cases it's punishment—tells you a lot about the organization, but no one's going to spell that out anywhere, right? Should you be friends with people at work? Should you not be? All these little choices are important choices in organizations, and they're not specified. That said, the AI for work stuff tends to be pretty good at picking out context once it has enough information. So if it has the history of a project, it often does a lot better than if it just has instructions for how to execute the project. Some of this will also be people making decisions, right? We will need to document things we didn't document before. Another set of stuff is just starting from scratch. Like, it's not clear that you want to translate processes that exist in a human organization and just have the AI implement those things ad hoc, right? So I think there's a lot of tension across this whole boundary here. I think people are underestimating how much the agents can step in from context to understand what's going on. Given the jaggedness, though, it's not everything, but a lot of things. But then also, how much we don't want that to be the answer a lot of the time.

激进案例:StrongDM软件暗工厂 Radical Example: StrongDM's Software Dark Factory

Host

你有没有见过一些很好的例子,那些在 AI 或 AI 智能体实施方面相当先进的组织,弄清楚如何让人类、智能体和系统之间的循环和交接运转起来?

And have you seen any good examples of organizations that are quite advanced in AI or AI agent implementation, figuring out how to get those loops and those handoffs between humans, agents, and systems?

Ethan

嗯,让我们从一个非常极端的例子开始。有一家公司叫 StrongDM,那是一家小型软件安全公司。他们决定构建这个软件黑暗工厂。黑暗工厂是机器人学中的一个概念,如果整个工厂由机器人操作,我们就不需要灯光。所以它是一个黑暗工厂,因为里面没有人。所以他们构建了一个软件黑暗工厂。你输入的是规格说明,甚至不是规格说明,基本上是产品发展路线图。他们有两条规则:人类不能写代码,人类不能看代码。所有代码都由 AI 编写,还有对抗性 AI 来测试代码。那些对抗性 AI 自发地构建了假的 Slack、假的 Gmail 和假的 Salesforce 来测试这些安全和认证模型。然后输出的都是成品代码,你只需批准或不批准发布。他们已经向客户交付了产品,而人类从未接触过代码库的任何部分。还有其他规则:每人每天必须花费 1000 美元的 token,这也相当昂贵。但这是一个非常激进的智能体工作版本的例子,对吧?我们完全把人类排除在循环之外。我认为这最终是错误的做法,因为我认为你实际上希望人类参与进来,要么作为变异的来源,对吧?我不希望 Claude 对一切做决策。我想用人类,这样我就不只是做基于 Claude 的工作。即使 Claude 的工作很好,那也不是我总是想看到的工作,因为你想做出重要的决定,因为你想做出有趣的决定。但这是一个非常激进方法的例子。我认为那种渐进式的做法,比如“好吧,我们只是多写点代码”,可能不是答案。也许不是黑暗工厂的全部重量,但我认为很多组织最终会落在两者之间的某个位置。

Well, let's start with a really extreme example. There is a company called StrongDM. That's a small software security company. And they decided to build this software dark factory. So a dark factory is an idea from robotics that if the entire factory is operated by robots, we don't need lights. So it's a dark factory because there's no humans in it. So they built a software dark factory. What you feed into it is specifications, not even specifications, basically, roadmaps for where you want your product to go. And they have two rules. No human can write code, and no human can look at code. Everything is written by AIs, and there's adversarial AIs that test the code. And those adversarial AIs spontaneously built fake Slack and fake Gmail and fake Salesforce to test these security and authentication models. And then all that comes out is finished code, and you just approve or don't approve shipping. And they've shipped products to customers without a human ever touching any part of the code base. There are other rules. You have to spend $1,000 in tokens per person per day, which is also quite expensive. But that's an example of a really radical version of agent work, right? We take the humans out of the loop entirely. I think that's the wrong way to go, ultimately, because I think you actually want people to be brought into the process, either as sources of variance, right? I don't want Claude's decision-making on everything. I wanna use humans so that I don't just do Claude-based work. Even if Claude's work is good, it's not the work I always wanna see because you wanna make important decisions because you wanna make interesting decisions. But that's an example of a really radical approach. And I think the sort of stepwise thing, which is like, okay, we'll just code more, is probably not gonna be the answer. Maybe not the full weight of the dark factory, but there's somewhere in between that I think a lot of organizations will have to end up.

委派与失去专业知识的悖论 The Paradox of Delegating and Losing Expertise

Host

我认为这里有一个有趣的桥梁,可以再次谈论“魔法”,以及你指出的悖论,那就是每次我们把任务交给智能体,自己不再做那个任务,我们就失去了真正建立专业知识的机会,而这些专业知识是我们形成有效评估 AI 输出所需的判断力所必需的。所以我很好奇,我们如何解决这个悖论?

I think there's an interesting bridge there to talking about wizardry, again, and the paradox that you've identified, which is that every time we cede a task or hand over a task to an agent and we're not doing that task ourselves, we are then losing the opportunity to actually build the expertise that we would need in order to form the judgment that is required to effectively assess the AI's output. So I'd be curious to know how do we resolve that paradox?

Ethan

是的,现在是专家的时候了,对吧?如果你是专家,你的处境很好,因为你可以看 AI 的输出,知道它是好是坏,对吧?比如,现在是成为专家的好时机,因为你可以变得超级高效,对吧?终于,你可以把工作委托给机器,它们会为你做。你可以评估结果。这就是为什么我认为很多专家第一次使用智能体时陷入了 AI 精神病态。

Yeah, now is the time for experts, right? If you're an expert, you're in great shape because you can look at the output of an AI, know whether it's good or bad, right? Like, it's a great time to be an expert because you can be hyperproductive, right? Finally, you can delegate to machines that will do the work for you. You can assess the results. It's like, it's a reason why I think so many experts descended to AI psychosis when they started using agents the first time.

培养新专家的问题 The Problem of Creating New Experts

Ethan

比如,我不知道你自己有没有经历过,但第一次用 Claude Code 的时候,人们周末消失一下,回来就带着一堆疯狂产出,往往还挺有用,但就像“我全搞定了”,对吧?因为他们脑子里已经有了那个画面,只是没人帮他们实现。所以专家本身有价值。但如何培养新专家就变成真问题了,对吧?因为那需要试错,需要导师指导,需要去完成越来越难的任务,并在其中经历失败。而 AI 把这些都省了。但真正省掉这些的不是 AI,而是我们作为人的选择。所以我觉得我们之前聊到艺术,对吧?你可能想要人类创作的艺术。我觉得同理,我们也会想要人类的努力和挣扎,而公司必须得容忍一个学习过程。但公司不是为此设计的。过去 4000 年我们一直很幸运,学徒制对大家都挺管用。作为中层管理者,你找个人来干你不想干的活,而且成本很低。你还能不花太多钱或精力就评估他干得好坏。而初级员工能学到门道,因为就算你的经理是个烂经理,你还是在学什么存在、什么不存在。你在学那些不成文的规则,我们之前聊过。你还会被评估干得好坏。而这一切在去年夏天全崩了,因为每个初级员工、初级新人,用 AI 都会好得多。更快,结果也比自己做更好。而每个中层管理者也更愿意让 AI 干活,因为人容易出错,有时还会抱怨,或者不来上班,你不想处理那些。所以这个断裂已经发生了。问题是怎么重建。有各种前进的路,但需要组织中那些不习惯这么做的部门付出真正的努力。我们习惯了从“边做边学”中获益,现在必须把它们分开。这里有很多可以展开的。

Like, I don't know if you've experienced this yourself, but the first time you use Claude Code, people kind of go away for the weekend and they come back with this insane output that is often very useful, but it's like, 'I figured it all out,' right? Because they've had this image in their head. They haven't had people to implement it. So experts have their own value. How do we create new experts becomes a real problem, right? Because that requires trial and error. It requires mentorship. It requires doing and failing at tasks that get increasingly hard as you move forward. And AI shortcuts all that. But it's not really AI that shortcuts it. It's our choices as people. So I think we were talking before about art, right? And that you might want human art. I think in the same way, we're gonna want human effort and struggle, and companies are gonna have to be okay with tolerating a learning process. And they're not built for that. We've been very lucky for the last 4,000 years, which is apprenticeship has worked out well for everybody. As a middle manager, you get somebody who is going to do the work for you that you don't want to do, and they're going to do it pretty cheaply. And you're going to be able to assess whether good or bad without spending a lot of money or a lot of effort. And the junior person gets to learn the ropes, because even if your manager is a bad manager, you're still learning what's there, what isn't. You're learning the unwritten rules, which we've talked about before. And you're getting assessed on whether you're doing a good or bad job. And that all broke this last summer, because every junior employee, junior hire, would be much better off using AI. It's faster, and the results are going to be better than doing it themselves. And every middle manager would rather have the AI do the work, because the humans are fallible, and sometimes they complain, or they don't show up for work, and you don't want to deal with that. So that break has already happened. And the question is how we reconstruct it. There's a bunch of ways forward, but they're going to take real effort for parts of an organization that aren't used to doing that. We're used to getting the benefits of learning while doing, and we're going to have to separate those now. There's lots to unpack there.

代理与组织设计 Agency and Organizational Design

Host

这里有很多可以展开的。你关于能动性的评论让我想起去年你和 Joel 见面时说过的话,那就是我们作为人类,我们做选择。我们对自己所做的事有能动性。听起来你仍然非常看好我们确实有能动性。但实践中,谁来做选择呢?如果我们想想组织中的个人,也许那些影响力较小的初级角色,他们实际能施展的能动性是什么样的?

There's lots to unpack there. I think your comment about agency reminds me of something that you had actually said to Joel when the two of you met last year, which was that we as humans, we get to make choices. We have agency over what we're doing. And it sounds to me like you're still very bullish on the idea that we do. Who is it that gets to make choices, though, in practice? If we think about individuals in organizations, perhaps in more junior roles that have less influence, what is the kind of agency that they actually get to exert?

Ethan

所以这是组织设计的问题。我的意思是,组织中领导者的决策越来越重要,因为在组织中,某种程度上,你被赋予能动性,你才有能动性。我的意思是,你决定不用 AI,但在一个所有同事都偷偷用 AI 的组织里,你会在那种结果中失败。所以我们必须激励人们以正确的方式行使这种能动性和选择。这不是自然而然发生的,因为你仍然有能动性,对吧?你在做选择,但如果你选择“我不用 AI,我要自己挣扎着搞定”,那你就傻了,因为那会伤害你。所以我们必须设立激励,以正确的方式激励。我们可以在学校这么做,对吧?因为我可以规定,期末我不用 AI 来考你,你成绩的 80% 来自这个考试,或者我们在教室里做的蓝皮书作业,或者你写的论文,或者我们做的角色扮演练习。那会激励你不只是用 AI 来学习,因为你需要在考试中表现。我们可以在组织中做同样的事,但我们不能假装人们可以什么都用 AI,还能在没有任何组织整体帮助的情况下学到所有想学的技能。更大的能动性问题,我的意思是,部分问题在于这些系统很诱人,对吧?我们在教育中发现了这一点,AI 想给你问题的答案。它是一个乐于助人的助手。而事实证明,当人们直接给你答案时,你学得并不好。所以增加摩擦。我觉得人们太担心消除摩擦了。在正确的点增加摩擦可能非常重要。

So that's the organizational design question. I mean, increasingly, the decisions leaders make in organizations matter, because you get agency if you are given it, in some ways, in organizations. I mean, you decide not to use AI, but in an organization where all your compatriots are using AI secretly, you're going to fail in that kind of outcome. So we have to incentivize people to be able to exercise this kind of agency and choice in the right kind of way. And that isn't something that happens naturally, because you still have agency, right? You're making choices, but you'd be foolish to make the choice that I'm not going to use AI, and I'm going to try and struggle through this on my own, because that would hurt you. So we have to set up incentives to incentivize that the right kind of way. We can do this in schools, right? Because I can say, I'm going to test you without AI at the end of the semester, and 80% of your grade is this test, or this blue book assignment that we do in the room, or the essay you write, or the role-playing exercise we have. And that will incentivize you to not just use AI for learning, because you're going to need to perform in this test. We can do the same kind of thing in organizations, but we can't pretend that people can just use AI for everything and still learn all the skills they want to learn without any help from the organization overall. The bigger question of agency is, I mean, part of the issue is these systems are seductive, right? So we found this in education, which is that the AI wants to give you the answer to the problem. It's a helpful assistant. And it turns out you don't learn very well when people just give you answers. So adding friction. I think people are worrying too much about removing friction. Adding friction at the right points might be really important.

激励儿童学习 Motivating Children to Learn

Host

所以你谈到 AI 想帮忙,想给你答案。在教育情境下,这不一定有帮助。我们还需要做哪些事来真正激励孩子在学校学习?

So you've talked about how AI wants to be helpful. It wants to give you answers. That's not necessarily helpful in an education context. What other things do we need to do to actually motivate children to learn in schools?

Ethan

所以我觉得孩子其实没那么有学习动力。我的意思是,你在自己真正在意的领域会有动力,对吧?比如你是个爱数学的孩子,你是个爱阅读的孩子。我的意思是,那一直是容易的部分,对吧?人们总是抱怨,为什么不能所有学习都像我喜欢的那样?那是因为不是每个人都喜欢所有东西,内在动力在某个点会失效。所以我们很依赖外在动力。你必须参加考试。我们作为社会认为,你需要学公民课,需要学数学,需要学自己国家的文学,或者其他什么结果。我们强制执行,对吧?比如,如果你成绩不好,你就不能升级。这些对上大学或之后的事情,或者学徒项目都很重要。所以我们建立我们需要的激励。鉴于我们在教育系统中能控制人,教育算是可解决的问题。我们现在有一些很好的证据,AI 导师对学习有好处。实际上,宾夕法尼亚大学那个研究团队,他们发现如果你只是让学生用 AI 并说“用它来学习”,AI 会损害学习。学生以为自己学到了,其实没学到。但当你构建一个 AI 导师,你实际上会看到非常大的影响和结果改善。所以我的意思是,最终状态看起来像我们从教学法研究中已经知道该做的事,那就是课外你会用 AI 导师。我不知道欧洲的课堂怎么样,但在美国,越来越多学生被布置这类作业,比如课外看可汗学院的视频之类的。然后在课堂上,会是讨论、主动学习。我们真的会尝试一些东西。我们也会在课堂上评估。你坐下来,在课堂上写一篇论文。所以我们会达到那个状态。但教育之路总是漫长,一切都很混乱,有很多竞争力量。但我认为我们可以在课堂环境中做到。

So I think children aren't that motivated to learn. I mean, you're motivated in the area you intrinsically care about, right? So you're a kid who loves math. You're a kid who loves reading. I mean, that's always been the easy part, right? And people always complain, why can't all learning be like the learning I like? It's like, because not everyone likes everything, and intrinsic motivation kind of fails at some point. So we rely a lot on extrinsic motivation. You have to take the test. We think, as a society, you need to learn civics, and you need to learn math, and you need to learn your country's literature, or whatever the outcomes are. And we enforce that, right? Like, if you don't get good grades, you don't advance. These things matter for getting into college or whatever comes afterwards, or an apprenticeship program. And so we build the incentives we need to do these things. Given that we have control of people in the education system, education is sort of a solvable problem. We have some really good evidence now that AI tutors are good for learning. The same team, the research team out of Penn, actually, that found that AI hurt learning if you just let students use AI and say, use it to learn. Students thought they learned. They didn't learn. But when you build an AI tutor, you actually get very large impact and improvement in outcome. So I mean, the end state looks like something we've already known from pedagogy research to do, which is, outside of class, you'll have AI tutors. And I don't know about classes in Europe, but in the US, increasingly, students are assigned this kind of work, where they're asked to watch a Khan Academy video or something outside of class. And then in class, it'll be discussion, active learning. We'll actually try things. We'll assess people in class, too. You sit down, you write an essay in class. So we'll get there. But the road is always long in education, and everything is messy, and there's lots of competing forces. But I think we can do that in the classroom setting.

职场正式评估 Formal Assessment in the Workplace

Ethan

对我来说,工作环境才是真正变得奇怪的地方,因为我们不习惯正式的评估。现在的情况是,老板会说,你干得好,你干得差。但如果老板对你的 Claude 输出印象深刻,你真的干得好吗?你学到了什么吗?我们可能实际上必须在真正的组织内部也进行正式评估。

To me, the work environment is where it gets really weird, because we're not used to formal assessment. What happens is your boss sort of says, you did a good job. You did a bad job. But if your boss is impressed by your Claude output, are you really doing a good job? Have you learned anything? We may actually have to have formal assessment inside of actual organizations, as well.

Host

那在实践中会是什么样子?

What would that look like in practice?

Ethan

所以,那可能更像考试。比如,你在一个房间里就某个主题做大型演示,人们就它盘问你,你根据那个被评判。它可能真的就像美国那种认证考试,比如保险或 CPA 认证,你必须参加某种专业测试。也可能是把我们之前讨论过的那种东西细化,我们激励人们有一些非 AI 时间,我们根据他们非 AI 时间的质量来评判他们,而不是 AI 输出。

So, that may look something more like testing. Like, you have a big presentation you give in a room about some sort of topic and people grill you about it and you're judged based on that. It may literally be like the kinds of tests that you do for certification in the US for, you know, insurance or CPA certification where you have to take some sort of professional test. It might be carving up the kind of thing we talked about before where we're incentivizing people to have some non-AI time and we're judging them based on the quality of their non-AI time rather than AI output.

Host

有一个普遍的问题,如果我们只激励生产力作为唯一目标,我们会很快在 AI 上陷入非常糟糕的境地。我们需要激励的其他目标会是什么?

There's a general problem, which is if we incentivize productivity as the only goal, we end up in a really bad place with AI quite quickly. What would be the other goals that we would need to be incentivizing for?

Ethan

所以这又回到了组织设计问题,对吧?如果你激励生产力,你就是在激励工作垃圾。比如你不想多一百倍的 PowerPoint,对吧?那就是生产力的样子。让我们做更多同样的事。即使在编码中,对吧?我们现在就有这个大问题。敏捷是 2002 年开发的技术,很多软件开发都是这样运作的。它不是瀑布式的。每个人都在做这个。现在你有了敏捷体验中这个生产力提高十倍或百倍的元素。你拿它怎么办?除非你改变流程的每一个其他部分,否则你不会从有人在每次敏捷会议上站起来说,你知道,Claude 今天做了我的工作,明天也会做我的工作,没有阻碍,然后坐下。那不是前进的有用方式。所以你必须围绕这些事设计流程,对吧?所以我们想让人类做什么?我想让他们做什么?我们如何打破我们已有的旧障碍并创造新的?所以实际上,很多责任现在都在领导者头上。

So this is back to the organizational design problem, right? If you incentivize productivity, you are incentivizing work slop. Like you don't want a hundred times more PowerPoint, right? Like that's what productivity looks like. Let's just do more of the thing. Even in coding, right? Like we're having this big problem right now. Agile was a technique developed, I think in 2002, and it's a lot of how software development works. It's not waterfall based. Everyone's doing this. And now you have this element of the Agile experience that is a hundred times or ten times more productive. What do you do with that? Unless you change every other part of the process, you don't gain out of having somebody stand up in every Agile meeting and say, you know, Claude did my work today. It'll do my work tomorrow. No blockers and sit down. Like that's not a helpful way of moving forward. So you have to design the processes around these sets of things, right? So what do we want humans to do? What do I want them to do? How do we break down the old barriers that we had and create new ones? So a lot is on leaders' heads at this point, actually.

Host

有趣。你在那里谈到领导者。我想每个人都处于一种没有人充分使用 AI 工具来真正知道框架或模型应该是什么的境地。那么领导者如何设法应对呢?

It's interesting. You're talking about leaders there. Everyone I guess is in a situation where no one has used AI tools enough to really know what the framework or the model should be. So how do leaders figure their way through that?

Ethan

所以有几件事。一是你说到点子上了,那就是没人知道任何事,对吧?我花时间与 AI 实验室交谈。我见过非常有名的人。我一直和 CEO 们交谈,没人知道任何事,对吧?我们都是边做边编。所以任何说我们有剧本的人,他们在骗你。没有剧本,对吧?我们在摸索。一方面,这很可怕。另一方面,这很棒,因为那意味着如果你创建自己的剧本,那实际上对你来说是一个优势来源。所以我用的模型,我想我上次也谈过,是领导层、实验室和群众。你需要三样东西,对吧?你需要一个领导团队,思考方向、激励、你想做什么。你需要群众。他们是发现用例的人。给你组织中的人工具。激励他们暴露他们用 AI 做什么,否则他们会隐藏那种使用,然后建立一个实验室。你需要一群人全天候工作,为你的组织思考 AI 用例。他们会从群众中收获想法。他们会从领导层那里获取想法。他们会真正构建东西,而不仅仅是谈论。因为你必须发明,对吧?最大的优势,回到经验,是你组织中有经验的人会很快知道锯齿形前沿的形状,因为他们能使用这些系统,看到它擅长什么或不擅长什么。他们有充分的动机去告诉,你知道,如果你能建立正确的工具来告诉你他们在做什么,并让你扩大规模,如果你能奖励他们这样做。而且你有别人没有的基础可以建立。

So a couple of things. One is you hit the nail on the head, which is nobody knows anything, right? Like I spend my time talking to the AI labs. I've met very famous people. I talk to CEOs all the time and nobody knows anything, right? Like we're all making this up as we go along. So anyone who's like, we have the playbook, they're lying to you. There's no playbook, right? We're figuring it out. On one hand, that's terrifying. On the other, it's great because that means if you create your own playbook, there's actually a source of advantage for you in that. So the model I use, and I think I talked about this last time as well, was leadership, lab, and crowd. You need three things, right? You need a leader team that is thinking about direction, incentives, what you want to do. You need the crowd. Those are the ones who are discovering use cases. Give people in your organization tools. Incentivize them to expose what they're using AI for because otherwise they'll hide that use and then build a lab. You need a group of people working 24-7 and thinking about AI use cases for your organization. And they're going to harvest ideas from the crowd. They're going to take ideas from leadership. They're going to actually build things and not just talk about it. Because you have to invent, right? And the biggest advantage, going back to experience, is experienced people in your organization will know the shape of the jagged frontier very quickly because they'll be able to use these systems, see what's good or bad at. They have every incentive to tell, you know, if you can build the right tools to tell you what they're doing and for you to scale it up, if you can reward them for doing that. And you have the sort of basis to build from that other people don't.

Host

你认为领导层、实验室和群众模型是否同样适用于,比如说,一个 50 到 100 人的小型初创公司,就像适用于 1 万人的组织一样?

Do you think Leadership Lab and Crowd applies in the same way to, let's say, a smaller 50-100 person startup as it would to a 10,000 person org?

Ethan

绝对。我的意思是,我认为组成部分的大小会改变,但有两件事是绝对清楚的,对吧?那就是领导层不会改变。就像,你需要为 AI 设定方向。C 级必须关心它。而且他们必须关心它,顺便说一句,不仅仅是因为他们在设定激励和决定组织的形状,还因为他们也在做其他事情,对吧?他们必须对事情的发展方向有想象力。他们必须了解自己个人如何使用它。而且你知道,他们实际上是组织中最有经验的人。所以领导者需要参与其中有很多原因。群众也需要这样做,因为领导者不会自己发明。你需要每个人使用这些工具才能从中获得任何优势。我认为每个人都需要一个实验室,对吧?现在,那个实验室可能是兼职工作。也许如果你是一个 50 人的组织,那可能是 CEO 每周花两天时间思考 AI 的事情。我认为我非常担心的一件事是认知负荷已经非常高,对所有这些人群来说。所以他们设定时间,比如我这个周末学 AI,或者我某个时候会有专门的 AI 时间,或者顾问会进来解决我们的 AI 问题。他们可以帮助你,但他们不会成为问题的解决方案。所以,唯一的解决方案是自下而上和自上而下以及持续实验的结合。我认为除了尝试之外没有其他前进的道路,对吧?顺便说一句,那也意味着失败。所以,我非常担心的一件事是组织不习惯组织实验。他们一直做产品实验,比如 A/B 测试,这不行,我们在这里太雄心勃勃了。他们一直做营销实验。但他们不做组织实验,对吧?我如何带走我的,你知道,如果我带走我的工程团队并分散它会发生什么?所以,我有一个工程师和一个销售人员以及一个主题专家一起工作。那会是什么样子?如果我给他们一个不可能的任务,你有两周时间复制我们花了三年建立的代码库,那会怎样?那可能失败,那可能成功,但你最终会学到一些东西。

Absolutely. I mean, I think that the size of the components change, but two things are absolutely clear, right? Which is leadership doesn't change. Like, you need to have a direction for AI. The C-level has to care about it. And they have to care about it, by the way, not just because they're setting incentives and deciding what the shape of the organization looks like, but also because they're doing other stuff as well, right? They have to have an imagination about where things are going. They have to get a sense of how they personally are using it. And you know, they're actually the most experienced people in the organization. So there's lots of reasons leaders need to be in on this. And the crowd needs to do this because the leaders aren't going to invent on their own. You need everyone using these tools to gain any advantage from it. And I think everybody needs a lab, right? Now, that lab might be part-time kind of work. Maybe if you're a 50-person organization, that might be the CEO spending two days a week thinking about AI stuff. I think one of the things I worry a lot about is cognitive load is already very high on all these sets of people. So they set up times like, I'll learn AI this weekend, or I'll have dedicated AI time at some point, or the consultants will come in and solve our AI problem. And they can help you, but they're not going to be the solution to the problem. So, the only solution is then that combination of bottom-up and top-down and continuous experimentation. I think that there's no other way forward other than, you know, trying things out, right? And by the way, that also means failure. So, one of the things that I worry about a lot is organizations are not used to organizational experimentation. They do product experimentation all the time, like A-B tests, this didn't work, we're too ambitious here. They do marketing experimentation all the time. But they don't do organizational experimentation, right? How do I take my, you know, what happens if I take my engineering team and disperse it? So, I have one engineer working with one salesperson and one subject matter expert. What does that look like? What if I give them an impossible task that you have two weeks to replicate our code base that we spent three years building? That might fail, that might succeed, but you're going to learn something as a result.

纪律与J曲线 Discipline and J-curve

Ethan

所以,纪律性的实验也很重要。我的意思是,任何技术都有一个著名的生产力 J 曲线。一开始,当你学习如何使用它时,生产力会下降,之后才会上升。你必须熬过这个 J 曲线,而很多组织没有意愿去这么做。

So, discipline experimentation is also important. I mean, there is a J-curve of productivity, famous for any technology. At first, your productivity drops as you learn how to use it, then it goes up afterwards. You have to ride through the J-curve and a lot of organizations don't have the willingness to do that.

Host

我觉得,这其实可以很好地过渡到关于“怪异”的讨论。你最近为《经济学人》写了一篇文章,标题是《IT 部门是 AI 的葬身之地》,你在文中指出,IT 最大的失败本质上就是没能认识到 AI 天生就是怪异的。你能详细阐述一下这个论点吗?

I think, actually, that could be a good bridge to talking about weirdness. You recently wrote an article for The Economist titled, IT departments are where AI goes to die, in which you argue that IT's biggest failure is essentially failing to recognize that AI is inherently weird. Can you unpack the argument?

Ethan

是的,我意识到我已经得罪了各地的 IT 部门,我的电脑也坏了。所以,我想说那个标题不是我的选择,对吧?是《经济学人》起的标题。但观点其实是真实的,那就是,不仅仅是 IT,还有法务等其他部门,它们都是风险规避组织,对吧?所以,有两件事。它们把 AI 视为风险来源,它确实是,而它们的工作就是确保你不暴露在风险中。这就是它们的激励机制。所以,我的意思是,如果没有人有键盘,IT 部门会非常高兴,那会为组织解决很多问题,对吧?所以,那不是你获得研发突破的地方。另一件事是,人们渴望将这项技术正常化。我的意思是,我们在谈论剧本。人们拼命想找一个能给他们标准剧本的人,对吧?这样他们就可以像其他人一样实施 AI。而这正是对这种东西的追求。而简单的实施方式都是让 AI 只是序列中的一个模糊处理器。太好了。现在,我们终于可以用自然语言搜索所有文档了。我们解决了自然语言处理问题。这是看待这些工具的一种非常没有雄心壮志的方式。所以,去怪异化让我们作为组织和领导者更容易把它放进一个盒子里或一个桶里,或者不用担心它。但这本质上是一项超级奇怪的技术。我的意思是,在三年时间里,我们从没人知道这东西存在,到我们可以从这些东西中获得博士级别的数学成果。它还能写出相当不错的电子游戏。它能教你的孩子东西。它还能做艺术。而且,你知道,它有相当不错的战略判断力,在伦理建议上还能击败伦理学家。像这样的系统,你怎么处理?我们可以决定它只是一个模糊的语言处理器。或者我们可以拥抱怪异并向前迈进。我真的很担心,自然的欲望和对事物节奏的疲惫会迫使人们把它当作另一款软件来思考。它在某些方面确实是,但在大多数方面不是。

Yeah, I've realized I've made enemies with IT departments everywhere and my computer stopped working. So, I want to say that title was not my choice, right? The Economist made the title. But the point is actually real, which is, it's not just IT, legal, other things, they're risk reduction organizations, right? So, there's two things. They see AI as a source of risk, which it is, and their job is to make sure that you're not exposed to risk. And that's their incentive. So, I mean, IT departments would really be happy if no one had keyboards, like would solve a lot of problems for organizations, right? So, that's not where you're going to get your R&D breakthroughs from. And the other thing is, there's a desire to normalize this technology. I mean, we're talking about playbooks. People desperately want to bring in somebody who's just going to give them the standard playbook, right? So, they can just implement AI the way everyone else is implementing AI. And it's just this quest for this stuff. And the easy ways to implement it are all ways that make AI just a fuzzy processor in a sequence of things. Great. Now, we can finally search across all of our documents in natural language. We've solved the natural language processing problem. It's such an unambitious way to view these tools. So, de-weirding them makes it easier for us as organizations, as leaders, to put this in a box or a bucket or to not worry about it. But this is an inherently super strange technology. I mean, in the course of three years, we've gone from like nobody knowing this stuff existed to we can get PhD-level math work out of these things. And it writes a pretty good video game. And it can teach your kids things. And it does art. And like, you know, reasonably good strategic judgment and beats ethicists in ethical recommendations. Like, how do you deal with a system like that? Well, we can just decide it's just a fuzzy language processor. Or we can embrace the weirdness and move forward. I really worry that the natural desire and the exhaustion with the pace of things forces people to think about this and try and think about this as just another piece of software. And it is in some ways, but in most ways, it's not.

优化怪异 Optimizing for weirdness

Host

在实践中,一个人实际上如何为怪异进行优化?在组织内部,这将是一个非常不寻常的 KPI,不是我们习惯的东西。当你为它优化时,它会是什么样子?

In practice, how does one actually optimize for weirdness? It would be a pretty unusual KPI to have inside of an organization, not something that we're used to. What does it look like when you've optimized for it?

Ethan

所以这不是为怪异优化,而是识别怪异,对吧?KPI 在这一点上是最大的敌人,对吧?我认为它们迫使你在实验阶段走上非常糟糕的道路,对吧?你不能用 KPI,比如我们说需要 10% 的改进,这本身就限制了你能看到的用例类型,对吧?所以 KPI 是一个非常危险的结果,因为 KPI 基本上是在追踪过去的行为,并说我们要更多或更少。当你有根本性的突破时,你需要一些激进的空间。这并不意味着你必须摆脱所有 KPI,但要有一些激进的空间。我们知道如何做,我的意思是,突破性的想法,对吧?那些从根本上改变我们工作方式的方法。那些是更怪异的 AI 用途。所以这就是允许实验失败开始发生的地方,对吧?所以我想看到,你知道,我想看到一种新的组织形式出现,你知道,做这个。我想看到一条新的产品线。我想看到我们重新进入一个因为无法竞争而离开的市场,并尝试,你知道,与之竞争,像大局的东西,而不是小的。

So it is not optimizing for weirdness. It's recognizing weirdness, right? KPIs are the biggest enemy at this point, right? I think like they force you into very bad paths in the experimentation phase, right? You can't KPI, like the very nature of saying we need a 10% improvement constrains the kind of use cases that you see, right? And so KPIs are a really dangerous outcome because KPIs are basically tracking past behavior and saying we want more or less of this. When you have radical breaks, you need some room for radicalness. Doesn't mean you don't have to get rid of all your KPIs, but there's some room for radicalness. And we know how to do, I mean, breakthrough ideas, right? Ways that radically change how we work. Those are weirder uses of AI. And so that's where allowing experimentation failure starts to happen, right? So I want to see, you know, I want to see a new organizational form come up with, you know, do this. I want to see a new product line. I want to see us reenter a market that we left because we couldn't compete and try and, you know, compete with it, like big picture stuff, not small.

CIO紧张与KPI CIO tension and KPIs

Host

企业 AI 显然很昂贵。所以我想知道,怪异的角度是否让 CIO 或 IT 部门陷入困境。他们显然有责任向组织展示他们从所有这些投资中获得了某种价值。这是否意味着他们必须写一种完全不同的商业案例,或者试图说服组织内部的其他人,我们需要彻底消除这些 KPI?怎么做?他们如何应对这种紧张关系?

Enterprise AI is obviously expensive. So I wonder whether the weirdness angle puts CIOs in a bit of a tough spot or ID departments. They are on the hook, obviously, to be able to show to their organizations that they're getting some kind of value from all of that investment. Does that mean that they have to write a completely different kind of business case or try to convince other people inside the organization that we need to eradicate these kinds of kPIs all together? How? How do they navigate this tension?

Ethan

嗯,我们首先看到的是公司。我认为,如果这只是 CAO 在争论 token 支出,而没有其他领导层的支持,你基本上已经处于困境了。这部分实际上是在避免一些这些问题。比如,我的意思是,有一段时间,Meta 有一个 token 排行榜,显示你烧了多少 token。这显然不是正确的方法,对吧?我的意思是,你可以像组织那样思考 AI 支出——你知道,当我们高层讨论时——像你这样的人在组织中。你非常能干,你的时间非常昂贵,所以组织中有整整一部分人的工作是过滤信息,这样你就能做出有影响力的有趣决策,你用合适薪资水平和合适努力水平的人以不同方式解决问题。组织实验的一部分也是思考我们何时使用哪些模型。哪些任务需要我们使用大量 token,哪些不需要——而且一旦 token 消耗成为 KPI,你就会陷入困境,对吧?所以,部分是要意识到,这实际上是组织的重塑,这很昂贵,但从长远来看实际上会降低成本,不仅仅是从裁员角度,而是从 token 优化的角度。第二点是,你确实需要不同类型的 KPI。比如,每个人都应该在一定程度上负责创新。我想看到新产品开发或新的组织形式或某种根本性的改变。我想看到前后的对比,而我看到组织正在这样做。这并非不寻常的案例。写这个很难,如果你走正常的写作流程,你就有麻烦了。这就是为什么在一定程度上,高层需要全力投入。

Well, we're seeing companies first of all. I think just if this is the CAO arguing about token spend one way or another without other leadership buying in, you're already kind of in a tough place. Part of this is actually avoiding some of those problems. Like I mean, for a little while, meta had a token leaderboard about how many tokens you were burning. That's obviously not the right approach, right? I mean you can think about AI expense like organizations there's- you know, when we talk at senior level- people like yourself in an organization. You, you're very competent, your time is very expensive, so there's an entire part of an organization whose job is to filter stuff so that you get to make interesting decisions that are impactful and you use people with the right salary level and right effort level to solve problems in different ways. Part of organizational experimentation also is about thinking about which models we use when. What tasks require us to use a lot of tokens versus not- and it moves pet like as soon as token burners or KPI. You're in trouble, right, and so part of that is about realizing that part of this is a reinvention of the organization, which is expensive but actually leads to lower costs in the long term, not just and not necessarily from headcount reduction, from just thinking about token optimization. Second piece of this is you do want different kinds of KPIs. Like everybody should be in charge of innovation to some extent. I want to see new product development or new organizational form or something radically changed. I want to see it before and after, and I'm seeing organizations do this right. Like this is not an unusual case. It's hard to write that, and if you're going through the normal writing process, you're in trouble. It's kind of why the sea level needs to be all in on this to some extent.

具体实例 Concrete examples

Host

你见过哪些组织成功识别怪异并因此蓬勃发展的具体例子?

What are some concrete examples you've seen of organizations that are successfully recognizing the weirdness and thriving because of it?

Ethan

所以我们还处于早期阶段,对吧。我的意思是,我认为人们——尤其是技术人员——过于雄心勃勃的一件事是,组织是缓慢的,对吧。

So we're in early days, right. I mean, one of the things that I think people are- especially technologists are- too ambitious about is organizations are slow, right.

紧迫性与现实AI应用 Urgency and Real-World AI Adoption

Ethan

即使你召开一个关于 AI 的全体紧急董事会会议,需要三个月才能安排,然后再花两个月开后续会议,组织也不会那么快行动。有些会,但大多数不会,这就是为什么紧迫感很重要。但我看到了很多非常有用的实验和方向转变。

Even if you had an all-hands emergency board meeting about AI that takes three months to schedule, and then two months afterwards for the follow-up meetings, organizations don't move that quickly. Some do, but most do not, and that's why urgency is kind of important. But I'm seeing a lot of really useful kinds of experiments and changing direction.

Host

对。

Right.

Ethan

让我想想哪些我可以公开谈论。比如高露洁-棕榄,做牙膏的那家公司,他们建立了一个很酷的 AI 实验室,直接与 Sutter 海平面合作,他们让 AI 自动从会议中生成想法,并更快地构建原型,所以一切都变得更快。你怎么以更快速的方式推出东西?有很多组织在这样做,并且在遗留软件实施方面发现了非常快速的变革。他们从 10 年计划更新 COBOL 基础设施,变成了‘我们无论如何要在六个月内完成’,所以设定了非常高、几乎不可能的目标。还有很多其他很酷的东西,我不知道我是否被允许谈论,因为我跟很多公司聊过这些。

So let me think about what once I could talk about publicly. So, like Colgate-Palmolive, the people who make toothpaste stuff, they have a really cool AI lab that they've set up that works directly with the Sutter sea level, and they're doing things like having the AI auto-generate ideas from conferences and build prototypes more quickly, right, so everything becomes quicker. And how do you put things out in a more rapid way? There's a lot of organizations that are doing that and have found really fast changes in legacy software implementation. So they've gone from like a 10-year plan to update their COBOL infrastructure to 'we're doing this at six months one way or another,' right, so setting very high, almost impossible goals. And there's a lot of other really cool stuff that I don't know if I'm allowed to talk about, because I talked to lots of companies about it.

Host

你在 1 月或 3 月说过,瓶颈不再是 AI 能力,而是界面。那你期望从像 Sana 这样在界面层创新的公司看到什么?

You said in January or March that the bottleneck is no longer AI capability, it's the interface. What would you then expect to see or want to see from companies like Sana that are innovating on the interface layer?

Ethan

所以思考 AI 实验室的方式,我认为人们把它们想成全知全能。它们是年轻的组织,擅长一件事,就是构建 LLM,第二件擅长的事是编码。所以难怪面向编码者的界面做得很好。我敢肯定至少我们在这方面是完美的,至少是富有想象力的,因为他们理解这些东西。他们自己都是编码者,所以他们构建工具,使用工具。他们不理解的是宇宙中任何其他工作,所以他们只是在尝试。事实证明,LLM 在很多事情上都是异常有效的工具,你不需要那么擅长,比如你不需要写过代码就能做营销。Claude Code 就能做营销。Codex 在战略建议方面做得很好,而提供医疗建议的东西对这些话题一无所知。那么问题不在于系统的能力,而在于人们用来理解它们、完成工作、思考工作如何完成的界面完全滞后。模型的能力就在那里,所以编码看起来很棒。但就像你一样,我们开始看到 AI 世界中设计可能的样子的一些元素,但大量工作尚未完成。你知道,未曾想过的,比如管理者比编码者多 14 倍,却没有这些事物的管理界面,对吧?那么你如何用 AI 管理任务?没有真正的概念说你被分配了工作。AI 只是做事情。不可能弄清楚它在做什么。没有可审计性。你不能决定它如何处理工作。界面可以解决很多这些问题,但它们需要了解世界的实际样子的经验,而很多这些公司缺乏这一点。

So the way to think about AI labs, right, as I think people think of them as sort of omniscient. They're young organizations that are good at one thing, which is building LLMs, and the second thing that they were good at then is coding. So it is no wonder that the interfaces for coders are really good from the air. I'm sure at least we've got perfect, at least imaginative, in that way, because they understand this stuff. They're all coders themselves, so they're building their tools, they're using their tools. What they don't understand is any other job in the universe, right, and so they're trying. Right, like it turns out, LLMs are unreasonably effective tools for so many things that you don't have to be that good, like you don't have to have built code to do marketing. It's Claude Code, it'll just do marketing. Codex does a really good job of strategic advice, right, like, and the things give medical advice have no idea about any of these topics. The problem, then, is not the capability of the systems, it's that the interfaces that people use to understand them, that to do the work, to think about how work gets done, are completely lagging. What the capabilities of the models are, right, so we coding. It looks great but, like you, we're starting to see some elements for how design might look in the AI world, but vast amounts of work are undone. You know, unthought of right, this like there are 14 times more managers than our coders and there's no management interface for these things, right? So how are you supposed to manage a task using AI? There's no real idea you're assigned the work. The AI just does stuff. It's impossible to figure out what it does. There's no audibility. You can't decide how it approaches the work. The interfaces can solve a lot of these problems, but they require experience with what the world actually looks like and a lot of these companies are missing that.

界面设计与人类-AI交互 Interface Design and Human-AI Interaction

Host

是的,我们一直在思考这个问题,关于如何让 Sana 摆脱首先是一个聊天优先的界面,以及你会如何设计一个界面,让人类、系统和智能体能够有效地相互交接?不过,我很好奇,如果你能挥动魔杖,改变你今天与 AI 交互的任何方面,你会改变什么?

Yeah, we have been thinking about this a lot, about how to move Sana away from being a chat-first interface, first and foremost, and how would you design an interface that can allow for humans, systems, and agents to hand off to each other in an effective way? I'd be curious to know, though, if you could wave a magic wand and change anything about how you interact with AI today, what would you change?

Ethan

系统比我们允许的要有能力得多,能给我们更多帮助。比如,我们刚刚进行了关于上下文的整个对话,然而,在窗口里输入内容并查找文档。我的意思是,它们有视频能力,有语音能力,能创建图像,能看图像,能打开文档。它们应该和你一起工作。你应该共同控制你的鼠标。你应该考虑让文档为你预先准备好并弹出来。我的意思是,我一直在做这种事情,让 AI 真的查看我的电子邮件,并为我准备好任何我需要做的事情。它甚至写邮件草稿,我还没有勇气发送,所以你收到的一切仍然是人写的,至少目前是这样。但是,我的意思是,这种想象力的失败很大程度上来自于聊天机器人是前进方向的时候,而现在不再是了,对吧?智能体委派子智能体去做事。为什么我所有的工作没有被检查?为什么我没有在过程中动态地得到反馈?我认为像 OpenClaw 这样的东西,它火了起来,可能很多看这个的人都很熟悉,它火起来的部分原因是它展示了一种新的 AI 作为个人助理的模式。显然对此也有巨大的需求。就像我们看到 AI 作为编码助理一样。我们会在很多其他领域看到这种情况。我认为真正让它消失在某些方面就是你的目标。

The systems are so much more capable of being helpful to us than we're letting them be. Like, we just had this whole conversation about context, and yet, type something into a window and look up a document. I mean, they have video capability, they have voice capability, they can create images, they can see images, they can open documents. They should be working with you on things. You should be co-controlling your mouse. You should be thinking about having documents pre-prepared for you that come up. I mean, I've been doing this kind of thing, where I have the AI literally look through my emails and sort of prep me on anything I need to do. It even writes email drafts, which I've not had the guts to send yet, so everything you get is still human, at least for now. But, I mean, this failure of imagination comes in large part from when chatbots were the way forward, and it isn't anymore, right? Agents delegate sub-agents to do stuff. Why isn't all my work being checked? Why aren't I getting feedback dynamically as I go? And I think part of the reason why, like, OpenClaw, which took off, and probably a lot of people watching this are familiar with, part of the reason why it took off was it showed a new mode of AI as personal assistant. And there was obviously a huge demand for that, too. Just like we saw AI as coding assistant. We're going to see this in lots of other areas. And I think being really, like, making it disappear is in some ways what you're aiming for.

界面之外的瓶颈:品味与人类判断 Bottlenecks Beyond Interface: Taste and Human Judgment

Host

显然,我们最初在谈论瓶颈。我们谈到的第一件事是界面的瓶颈。但其他更与人类相关的瓶颈呢?每个人显然都在谈论品味,说也许品味将成为新的瓶颈。品味仍然是我们人类独有的。你相信这个论点,还是认为我们只是在自我安慰,直到 AI 足够好,拥有自己的好品味?

And obviously, we were just talking originally about bottlenecks. And the first thing that we talked about there was the bottleneck of the interface. But what about other more human-related bottlenecks? Everyone's obviously talking about taste, and saying that maybe taste is going to be the new bottleneck. Taste is still what is uniquely human to us. Do you buy that argument, or do you think we are self-soothing until the AI gets good enough to have good taste of its own?

Ethan

所以,我认为有两种方式来看待这些东西。我实际上倾向于说,我最常思考的四件事,作为你想要训练自己的元素,是深度知识,一个领域的专业知识,我们之前谈过。广泛的知识,比如,我认为是时候再次重视人文学科了。知道很多事情真的很有用。这些系统可以做任何事情。所以,让它们做许多不同的事情很有趣。然后我认为品味很重要,但不一定是竞争优势。只是你的品味是一个差异化因素,对吧?然后是能动性,就是做事,对吧?所以,不一定令人惊讶。不过,我认为也存在明确界限论点的危险,即只有人类才能行使判断力。我经常听到这个。这显然已经是错误的,因为如果你给一个智能体一个七小时的任务,它到处都在行使判断力,对吧?所以,AI 显然可以行使某种判断力,我们可以讨论这种判断力的价值。所以,我担心那些明确界限的论点,即人类只能做这个。

So, I think that there are two ways of viewing this stuff. I actually tend to say the four things that I think about most as sort of the elements that you want to train yourself in are deep knowledge, expertise in an area, we've talked about that before. Wide knowledge, like, I think it's time for the humanities major again. Like, knowing many things is really useful. These systems can do anything. So, having them do many different things is interesting. And then I think taste is important, but not necessarily like a competitive advantage. It's just your taste is a differentiator, right? And then agency, just doing stuff, right? So, not necessarily surprising. I think, though, that there's also a danger of bright-line arguments, which is to say only humans can exercise judgment. I hear that a lot. That's obviously already wrong, because if you give an agent a seven-hour task, it's exercising judgment all over the place, right? And so, AIs can obviously do judgment for some value of judgment that we can discuss. And so, I think I worry about the bright-line arguments that humans can only do this.

品味与瓶颈 Taste and Bottlenecks

Ethan

我认为我们可以描绘一个品味很重要的世界,因为你不希望所有东西读起来都像,你知道,让我先琢磨一下。这是论证中承重墙的一部分。不是 X,而是 Y。比如,那种“爪式”写作现在让我抓狂,对吧?即使它不算坏写作。而我渴望变化,所以品味可能在那里起作用。但这和说只有人类才有品味是两回事,对吧?所以我认为我们必须考虑这些方面。我认为瓶颈往往在于任务在崎岖前沿上的复杂性。有些领域 AI 能做好,有些则做不好。而瓶颈恰恰是 AI 做不好、或者只能做得很差的地方,也正是人类努力最需要投入的地方。

I think we can articulate a world where taste matters because you don't want everything to read like, you know, let me sit with this for a bit. This is a load-bearing part of the argument. It's not X, it's Y. Like, clawed writing drives me insane at this point, right? Even though it's not bad writing. And I crave variation, so that's where taste might matter. But that's quite different than saying only humans have taste, right? And so, I think we're going to have to think about those aspects. I think the bottlenecks are often about the complexity of tasks on a jagged frontier. And there are some areas that AI can do well, and some it can do badly. And the bottlenecks become the things AI can do badly, or only does badly, is where human efforts need it the most.

两年展望 Two Years Out

Ethan

两年后,我们还有白领工作吗?显然有。我的意思是,我们可以有超级智能机器。我们可能在某些方面已经有了,但即使有超级智能机器,第二天也很少有人工作会改变,因为它仍然是崎岖的。而且在物理世界里也是崎岖的。比如,有人要参加会议,有人要去做某件事。有大量未成文的材料。你不能直接空降。你能想象的最聪明的人,如果你把他们放到一份工作里,甚至一千个这样的人放到一份工作里,你仍然会很难让事情运转起来,对吧?所以两年后显然有一些宏观层面的问题。

Two years from now, do we still have white-collar jobs? Obviously. I mean, we can have a super-intelligent machine. We might already have them in some ways, but we have super-intelligent machines, and very few people's work will change the next day because it's still jagged. And it's just even jagged in the physical world. Like, somebody is attending the meeting. Someone's going to something. There's tons of unwritten material. You can't just drop in. The smartest person you can imagine, if you drop them into a job, and even 1,000 of them into a job, you're still going to have a lot of trouble making things operate, right? So there is obviously some big-picture issue for two years from now.

五至七年展望 Five to Seven Years Out

Ethan

但让我们谈谈五年、七年后,这实际上是所有 AI 实验室说工作会被摧毁的时候。有一种新兴的经济共识。我们不知道答案,但共识围绕着所谓的“O 型环工作模型”,这个名字取自挑战者号航天飞机灾难,令人沮丧——挑战者号上成千上万个部件都运转良好,但 O 型环失效了,对吧?在 O 型环工作理论中,我们做许多不同的任务。其中一些任务非常重要,不是那种好 10%、差 20% 的程度。但如果它们出了问题,整个工作就会崩溃。整个工作流就会崩溃。而这些任务就成了,用一个老套的词,整个工作组织的承重墙。

But let's talk about five, seven years from now, which is actually when all the AI labs are saying jobs get destroyed. There's sort of an emerging economic consensus. We don't know the answer, but the consensus is around something called the O-ring model of jobs, which is grimly named after the Space Shuttle Challenger disaster, where thousands of parts worked well in the Challenger, but the O-ring failed, right? In an O-ring theory of jobs, we do many different tasks. And some of those tasks are really important in a way that is not just 10% better, 20% worse. But if they go badly, the entire job falls apart. The entire work stream falls apart. And what happens is those become the, to use a clod term, the load-bearing pieces of the entire job organization.

Ethan

顺便说一句,我们可以看到这种情况正在发生,对吧?早期,甚至在疫情之前,但肯定在疫情期间,对软件开发者的需求巨大,对吧?如果你看看美国的专业,每个人都转向了软件开发。甚至在 AI 出现之前,软件开发招聘就开始下降。每个人都通过换专业来应对,而招聘变得更少。现在随着云代码的兴起,你看到的是,它不再像以前那样写那么多代码了。而是成为软件工程师的管理者,思考事情需要往哪里走,写什么测试,什么算好或足够好。所以 O 型环的性质改变了,对吧?从“我能写出不是意大利面条式的好代码”变成“我能成为一个好的软件工程师,并在更高层次上思考吗?”而那可能是 AI 下一步要做的事。然后这又会有一个新的停顿。所以这些瓶颈在经济中的工作内部会不断变化,就像人们突然发现 AI 擅长这个,但下一个紧张点变成它不擅长那个。

And by the way, we can see this happening, right? In the early, even before the pandemic, but certainly during it, there was a huge demand for software developers, right? And if you look at majors in the US, everyone switched to software development. And even before AI, software development hiring started to drop. And everyone responded to this by switching their majors and hiring became less. And what you see happen now with the rise of Claude Code is it's not about writing as much code as it was before. It's about being a manager of software engineers, about thinking about where things need to go, what tests to write, what's something good or bad enough. So the nature of the O-ring changed, right? From can I write good code that isn't spaghetti code to can I be a good software engineer and think in this higher level? And that might be something the AIs do next. Then there'll be another stop in this. So these bottlenecks are gonna be ever-changing inside jobs in the economy where people like suddenly AI does this well, but now the next tension becomes it doesn't do this well.

Ethan

然后另一件事是竞争优势。事实证明,token 真的很贵。AI 能做很多工作,但你不希望它做所有那些工作。人类在某些特定工作上有时比 AI 更便宜或更好。而且我们看到,只要我们还受 token 限制,就会有更多的工作要做,甚至超过 AI 能做的量。所以我认为我们会看到很多奇怪的事情发生。会有很多混乱和困惑。人们总是拿工业革命来类比,我们三次工业革命最终都取得了好结果。人们最终得到的工作比开始时更多,但经历起来都很糟糕。查尔斯·狄更斯写的就是工业革命有多糟糕。所以我认为我们即将迎来一个混乱的时期。我认为在混乱发生时,拥有广泛的基础会帮助你。

And then the other thing is competitive advantage. It turns out that tokens are really expensive. AI can do many jobs, but you don't want to do all those jobs. Humans are sometimes cheaper or better than the AI at doing particular work. And we see for a long time, as long as we're token constrained, that there's more work to do than there is necessarily AI even to do it. So I think we're gonna see a lot of weirdness happen. There's going to be a lot of chaos and confusion. The big picture view people always take is the industrial revolution, which all three of our industrial revolutions have worked out well in the end. People got more jobs in the end than they started with, but they all kind of suck to live through. Charles Dickens is all about how much the industrial revolution sucked. And so I think we're in for a chaotic time. And I think being broad based will help you in a place where this chaos is occurring.

管理作为超能力 Management as Superpower

Host

所以我想深入探讨你提到的管理和广泛基础。我知道你最近写了一篇关于管理作为 AI 超能力的文章。另一方面,硅谷确实在宣扬 IC(个人贡献者)的复兴,中层管理的消亡。我很好奇你会如何为那个论点辩护,而不是你自己的论点,或者你认为它们实际上比我可能想的更互补?

So I want to drill into what you said about management and being broad. I know you recently wrote a piece about management as an AI superpower. On the flip side, Silicon Valley is really proclaiming the resurgence of the IC, the death of middle management. I'd be curious to know how you would defend that argument rather than your own, or do you think that they are actually more complementary than I might be thinking?

Ethan

所以我认为大 M 管理和小 m 管理之间是有区别的。小 m 管理作为超能力,意味着你越擅长管理工作,就越擅长使用 AI。如果你能定义你想要什么,以及管理者做的所有事情,对吧?给你一堆 TLA(三个字母缩写),对吧?如果你擅长 PRD,擅长 SOP,任何 RFP,任何我们用来向人下达命令以完成任务的东西,AI 越来越能够直接接受这些命令并执行。所以一种超能力形式是擅长分配工作、理解风险在哪里、缺点在哪里、快速评估输出,并拥有这样做的专业知识。这种小 m 管理让你现在就能用好 AI。

So I think there's a difference between big M management and small M management. So small M management being a superpower is the better you are at managing work, the better you are using AI. If you can define what you want, all the things managers do, right? So just to give you a bunch of TLAs, three letter acronyms, right? If you're good at PRDs, if you're good at SOPs, any of the RFPs, any of the stuff that we do to give commands to people to get something done, AI is increasingly able to just take those commands and run with it. So one form of superpower is being good at assigning work, understanding where the risks are, where the downsides are, assessing output quickly, having the expertise to do that. That small M management makes you good at AI right now.

Ethan

然后是大 M 管理的问题,即你如何管理组织或在更大意义上思考这类问题?我认为关于管理层级是否会崩溃存在一些问题,对吧?CEO 能否直接管理更多事情,因为很多中层管理由 AI 完成?或者这会增加中层管理者的管理幅度?我在公司中看到的一个趋势是,例如,产品经理和编码者的角色合并成一个单一的“构建者”头衔。然后每个人都做同样的工作。

Then there's sort of the big M management question, which is how do you manage organizations or think about these kind of issues in a bigger sense? And I think there is some question about whether you have management layers collapse, right? Can the CEO just manage more things because a lot of the middle management layer is done by AI? Or does this increase the span of middle managers? One of the trends that I've been seeing in companies has been collapsing the role of product manager and coder together into a single builder title, for example. And everybody just does the same work.

Ethan

我们在宝洁做了这个实验,让 776 名员工要么单独工作,要么两人一组,跨职能的两人组,商业和技术人员。这是用 GPT-4 做的。所以现在来看是相当过时的系统,完全没有智能体能力。但使用 AI 的个人表现和不使用 AI 的两人组一样好。更有趣的是,人们的职位差异开始消失。所以如果你以前是编码者,如果你不使用 AI,你会带着编码的想法来。如果你是商业人士,你会带着商业想法来。一旦你使用 AI,你执行的想法看起来像是两者的结合。所以一个大问题将是这种角色崩溃。比如,管理者做什么?我认为管理将成为一个与以前非常不同的角色。

We did this experiment at Procter & Gamble where we took 776 employees, had them work either individually or in teams of two, cross-functional teams of two, business and technical people. And this was using GPT-4. So really obsolete system at this point, not agentic at all. But individuals using AI performed as well as teams of two not using AI. And what was more interesting was people's job difference started to collapse. So if you were a coder before, if you weren't using AI, you came with coding ideas. If you're a business person, you came with business ideas. Once you use AI, you execute ideas that look like both ideas together. So one of the big issues is going to be this kind of role collapse. Like what does a manager do? I think management's gonna be a very different role than it was before.

围绕AI代理重组团队 Reorganizing Teams Around AI Agents

Host

那这是否意味着职能部门和子职能也可能瓦解?我的意思是,AI 智能体正变得越来越擅长跨职能协作。你能预见营销部门、IT 部门、销售部门等被瓦解,团队围绕要解决的问题重新组织吗?

Does that mean then that functions and sub-functions might also collapse? I mean, AI agents are getting good, increasingly good at working cross-functionally as well. Could you foresee the collapse of a marketing department, an IT department, a sales department, for example, and reorganizing teams more around the problems that they're trying to solve?

Ethan

我的意思是,我们可以想象出上百种重组方式。首先,我们其实还处于如何良好组织智能体的早期阶段。看起来很棒,但当你真正开始观察它们在编码之外的表现时,智能体会忘记事情,会丢失东西。当它们跨职能工作时,事情会变得有些混乱。我认为其中很多都是可以解决的问题。更好的技能、智能体之间更好的接口。我们可以做很多事情来改善这一点。更好的人工检查。但我认为真正的问题是:我们为什么要这样组织?第一张组织架构图是在 1855 年为纽约和纽约铁路公司发明的,从那以后我们的组织架构图大致保持不变。你实际上有两个选择。是事业部制结构,侧重于客户和产出;还是职能制结构,基于我们做什么。然后还有矩阵制和其他糟糕的替代方案,但那些才是你真正面临的两大选择。而这些不一定是唯一的选择。也可以按品味来组织,对吧?你最终可以说,我们有这个品味选择智能体,这个做判断的智能体,这个做客户模拟的智能体,所有这些都参与进来,对吧?所以你可以开始思考许多不同的组织方式。我认为小型的跨职能团队是短期的方向。围绕项目来构建可能是正确的做法。我认为问题是:你什么时候需要请专家?这是个未知的问题,对吧?所以,你知道我有一个三人项目团队,里面没有律师。我们是否接受 Claude 一直担任律师?比如,我们什么时候该请真正的律师?这是一个支持性职能,还是律师们也在构建自己的产品?而且,我们在这里反复谈到领导者。这就是那种“欲戴王冠,必承其重”的时刻,因为这一切都落在领导者身上。他们不能只是开张支票让别人替他们解决这个问题。他们必须自己想办法解决。他们可以获得支持,但这是一个难题。

I mean, there's a hundred ways we can imagine reorganizing things. So first of all, we are in the very early days of actually organizing agents well. It looks great, but then when you actually start to look at them outside of coding, the agents forget things. They lose things. As they work cross-functionally, things sort of get messy. I think a lot of that's solvable problems. Better skills, better interfaces between the agents. There's a lot of things we could do to make that better. Better human checking. But I think that the real question is why are we organized this way? Well, the first org chart was invented in 1855 for New York and New York Railroad, and we've kept org charts roughly the same ever since. You really have two choices. Is it a divisional structure, which is focused on sort of the customers and outputs, or is it a functional structure based on what we do? Then you've got the matrices and other godforsaken alternatives, but those are really the two big choices that you have. And those don't have to be the only choices. It could be by taste, right? You could end up saying that we have, that this is our taste selection agent, and this is our agent that does judgment, and this is our agent that does customer simulation, and all of those weigh in, right? So you can start to think about many different ways of organizing. I think small, cross-functional teams are the short-term way to go. And I think building around projects is probably the right way to do this. I think the question is: when do you need to call in the expert? Is the sort of unknown question, right? So we, you know I've got a three-person project team. There's no lawyer on it. Are we okay with Claude being the lawyer the whole time? Like, when do we call in the lawyer? And is this is a some sort of supporting function, or are the lawyers also building their own products? And again, we talked about leaders are repeatedly throughout here. This is where there's sort of a heavy as the head that wears a crown moment, because all of this falls on leaders. They cannot just write a check to somebody to solve this problem for them. They have to figure out how to do this themselves. They can get support to do it, but it is a hard problem.

AI与育儿哲学 AI and Parenting Philosophy

Host

所以,据我所知,你是两个孩子的家长。是的,我不知道他们多大了,但我很想知道你在 AI 和育儿方面的理念是什么。

So, as far as I understand, you are a parent of two children. Yes, I don't know how old they are, but I would love to know what your philosophy is on AI and parenting.

Ethan

一个在上高中,一个在上大学,但我们已经用 AI 一段时间了,所以这里有很多不同的方面。我认为你应该像对待其他形式的技术一样对待它,那就是监督使用。当孩子还小的时候,监督使用总是最好的方式,对吧。对于小孩子,我认为如果你在其中扮演领导角色,你可以用 AI 做很多有趣的事情。对,你决定什么对你的孩子好或坏,一起创建图像。有很多有趣的游戏之类的。从学习的角度来看,我认为当孩子还小的时候,监督使用是最好的方式。可能最好的方法是让 AI 帮你成为更好的老师。比如拍下问题的照片,让它解释如何向我的孩子讲解,对吧。随着他们长大,使用 AI 系统中的学习模式非常好,对吧。你不想只使用 AI 本身,你想在辅导模式下使用它。有一些针对孩子的模式。同样,我会监控使用情况。然后还有一个更宏观的问题:好吧,这对职业意味着什么?我的意思是,我认为有一定程度上,押注管道工是未来的工作感觉像是浪费。所以你让你的孩子追求你认为他们应该追求的东西。你教他们要灵活。广博的知识,深入的知识,你知道。给他们自主权。我不认为事情变化那么大,所以我不认为。我认为,与其他技术相比,我怀疑除了用它作弊——那里有真正的危险——或者在教育中使用,你在自欺欺人,你学习是因为 AI 在向你解释东西,但你自己没有做。除了这些需要小心的用途之外,我认为它和其他技术一样运作,但是,你知道,它比 TikTok 或其他任何东西更好还是更差?我认为那里有更少的被动性,而且有真正的价值,但同样,必须有家长的指导。我们又回到了自主权的问题上。

So one's in high school, one's in college, but we've been using AI for a while, so there's a lot of different pieces here. I think you treat it like other forms of technology, which is supervised. Use is always the best way to go when they're younger, right, so young kids like I think there's lots of delightful things you can do with AI if you have the leadership role in it. Right, you're making decisions about what's good or bad for your kid, creating images together. There's a lot of fun games, things like that. For a learning perspective, I think supervised learning use when the kids are, when the kids are younger. Probably the best way to do him is to actually have the AI help you become a better teacher. So take a picture of the problem and explain to me how it explains to my kid, right, as they get older. Using learning modes in the AI systems is really good, right, so you don't want to just use the AI itself, you want to use it in tutor mode. There's some, you know, kid focus ones. Again, I'd monitor use. And then there's a sort of wider picture about like: okay, so what does this mean for careers? And I mean right, like I think there's some degree of like, making bets that plumbing is gonna be the job of the future feels like a waste. So you let your kids pursue what you think they should pursue. You teach to be flexible. Wide knowledge, deep knowledge, you know. Give them agency. I don't think things change that much, and so I don't. I think, compared to other technologies, I suspect that outside of using it for cheating- where there's real danger, or uses an education, where you're fooling yourself, you're learning because the AI is explaining stuff to you but you're not doing it yourself. Outside of that use, where there's a lot of like you know being careful about this, I think it works like other technologies, but, you know, is it better, worse than tick-tock or anything else? I think there's less passivity there and there's real value in it, but again, it has to be parent guidance. We're back to agency again.

AI大未解问题 Big Unanswered Questions in AI

Host

Ethan,你——我的意思是,你——可以说——已经研究和考察了 AI 的几乎每一个方面以及当前的发展。你仍然希望看到更多工作去解决的重大未解问题是什么,或者你认为我们需要解决的被忽视的问题是什么?

Ethan. You've- I mean you've- arguably- looked at and researched almost every aspect of AI and what's going on right now. What are the big unanswered questions that you would still like to see more work being done on, or any overlooked questions that you think we need to solve?

Ethan

所以在工作领域,我认为我们对 AI 时代的组织一无所知,几乎所有东西都看起来像。我的意思是,很奇怪的是,AI 公司现在都在建立自己的咨询部门来做 AI 部署,对吧。如果模型如此之好,以至于它们可以开发,你知道你认为它们会取代什么颜色的工作,难道它们不应该也能帮助你部署系统吗?我认为这是锯齿状前沿工作的一个很好的例子,我们对那个最大的图景理解不够。实际上只有两个问题非常重要,那就是有多好和有多快,对吧。所以这条指数曲线能持续多久,在哪个点?在哪个点会放缓,会有多急剧?因为这决定了其他一切。我们今天讨论的一切都基于我们对当前状况的理解和对未来的假设,即未来的路径会像过去一样。去年,当我们交谈时,你知道,我认为你和 sauna 团队知道智能体即将到来,但大多数人不知道。那会是一场非常不同的对话,我认为事情会继续以那种方式转变。所以我认为我们真正应该关注的是:那条曲线看起来是什么样的?锯齿状前沿是否仍在收窄,我们是否在试图预测几个月后的情况?否则,你只会停留在当前状况,从我们所在的这个点进行外推,而不是思考我们可能走向何方。

So in the work world, I think we don't know anything about organizations at AI, like almost everything's on looked like. I mean, it's weird that the AI companies have all now building their own consulting arms to do AI deployment right. If the models are so good that they can develop, you know that you think they're gonna show what color jobs, shouldn't they also be able to help you with deploy systems. I think that's a really great example of the jagged frontier work and we don't understand enough about that, the biggest picture. There's only two questions that actually matter a lot, which is how good and how fast, right. So how long this exponential curve continue and at what point? At what point does it ease off and how sharp will it be? Because that determines everything else. Everything we've been talking about today is based on our understanding of where things are today and assumption. The path of the future looks like the past. Last year, when we talked like, I think you know the sauna team and I knew agents were coming, most people didn't. It would have been a very different kind of conversation and I think that things are going to keep shifting in that kind of way. So I think really what we should pay attention to is: what does that curve look like? Is the jagged frontier still closing and we're trying to gas a few months out? Otherwise, you just get stuck on where things are today and extrapolate from this sort of you know point where we are, as opposed to thinking about where we might be going.

Ethan最自豪的事 What Ethan Is Most Proud Of

Host

你一直处于这个领域的前沿。你是一位教育者,你为非常非常广泛的受众写作。你最自豪的是什么?

You've been at the frontier of this. You're an educator, you're writing for a very, very broad audience. What are you most proud of?

Ethan

我认为是努力让 AI 更人性化,而且我认为我已经做到了。

I think trying to make AI more human, and I think I've done that.

教育与人类代理 Education and human agency

Ethan

我们正试图在教育领域做到这一点,通过构建帮助人们学习的工具,这些工具已在大量案例中被采用。但同样通过写作和其他一切,我们不会让技术专家获胜。这必须由人类和管理者来决定如何使用这些工具,还有政策制定者。我们不应将技术视为不可避免的。

We're trying to do that in education by building tools that help people learn, and they've been adopted in huge cases. But also through the writing and everything else, we're not letting the technologists win. This has to be humans and managers deciding how to use these tools, and policymakers. We shouldn't view the technology as inevitable.

互动版:逐字朗读 + 针对本期提问 →