Lindy Launches AI Teammate in Slack, Running on DeepSeek
打开互动全文版(中英对照 + 朗读 + 问答)→Flo Crivello 讨论 Lindy 新推出的 Slack AI 员工、多人 AI、记忆以及使用中国模型的政治问题。
Flo Crivello discusses Lindy's new AI employee for Slack, multiplayer AI, memory, and the politics of using Chinese models.
大家好,欢迎回到《认知革命》。今天我的嘉宾是 Lindy 的创始人兼 CEO Flo Crivello,他正在发布 Lindy Teammate——一个加入你公司 Slack 的 AI 员工。这当然直接对标 Claude Tag,而且值得注意的是它运行在 DeepSeek 上。即便如此,Lindy 至少在补贴入职成本。对话的前半部分深入探讨了这些 token 都花在了哪里。我们聊了多人模式的细微差别、Lindy 如何看待围绕历史数据的社交契约、以及他们如何用提示词来执行这些契约,还聊了很多关于记忆实现和 Lindy 用来持续优化记忆的后台处理。Flo 对相关文献的掌握贯穿始终。我们得到了 Flo 关于文件系统和沙箱的供应商推荐,并讨论了他为什么倾向于在能买的地方买,以及为什么仍然看好软件基础设施初创公司。我们还聊了现在 Lindy Teammate 存在后 Lindy 的生活,以及他如何看待 Lindy 和所有前沿 AI 系统——它们在大多数方面都像超级智能,却仍然每天多次选择走到洗车店。Flo 说,目前我们处于半人马时代。Lindy 最好的想法通常是人类和 AI 共同创造的。但 Flo 觉得这将是暂时的,不久之后他预计人类只会给高度优化的 AI 系统增加噪音。应该说,这是在好的世界里,因为正如 Flo 所说,鉴于 Openface 和其他类似事件,现在是 2026 年 7 月下旬。我们有了 AGI。我们正处于起飞阶段,而我们还没有解决对齐问题。考虑到这一点,我们比较了 AI 公司的朋友和熟人们现在如何有点恐慌。在最后一部分,我们讨论了 Flo 非常令人惊讶的立场:中国模型应该在美国被禁止。目前,我绝对不同意,但我们就各种论点进行了友好的来回讨论。最终,他似乎对某种妥协持开放态度,比如对销售 AI 服务的公司提出保险要求,这样我们就可以规避中国模型的风险,而不是直接禁止。无论如何,我不得不说 Flo 是独一无二的。在过去三年里冲刺通用 AI 助手产品竞赛之后,他仍然对 Lindy 的运作方式非常开放,并且对他关心的每个问题都完全坦诚。这造就了一场关于多人 AI、机器记忆以及在你希望被禁止的模型上建立公司的尴尬政治的非常有趣的对话。这是 Flo Crivello,Lindy 的创始人兼 CEO。Flo,Lindy 的 CEO。欢迎回到《认知革命》。
Hello and welcome back to the Cognitive Revolution. Today I'm speaking with Flo Crivello, founder and CEO of Lindy as he's launching Lindy teammate, an AI employee that joins your company's Slack. This is of course competing directly with Claude Tag and very notably it is running on DeepSeek. Even so, Lindy is subsidizing onboarding at least. And the first half of this conversation goes really deep into where all of those tokens are going. We talk about the nuances of multiplayer mode, how Lindy thinks about the social contract surrounding historical data, and how they're using prompting to enforce it, plus a ton about memory implementation and the background processing that Lindy uses to continually optimize memory. Flo's command of the relevant literature is evident throughout. We get Flo's vendor recommendations for file systems and sandboxes and discuss why he prefers to buy where he can and remains bullish on software infrastructure startups. We also talk about life at Lindy now that Lindy teammate exists and how he thinks about Lindy and all Frontier AI systems as a super intelligence in most respects that still somehow chooses to walk to the car wash a bunch of times every day. For now, Flo says we are in the centaur era. The best ideas at Lindy are generally co-created between humans and AIs. But Flo feels that this will be temporary and before too long he expects that humans will simply be adding noise to highly optimized AI systems. And that it should be said is in the good world because as Flo puts it in light of Openface and other similar incidents, it is late July 2026. We have AGI. We're in takeoff and we've not figured out alignment. With that in mind, we compare notes on how friends and acquaintances at AI companies are now kind of panicking. And in the final section, we discuss Flo's very surprising position that Chinese models should be banned in the United States. As of now, I definitely do not agree, but we have a good naturatured back and forth about various arguments. And ultimately, he seemed open to a compromise on something perhaps more like an insurance requirement for companies selling AI services that would allow us to press the risk of Chinese models instead of banning them outright. In any case, I have to say Flo is a oneofone. After sprinting through the general purpose AI assistant product race for the last three years, he is still super open about how Lindy works and he's fully candid about every issue he cares about. It makes for a really fun conversation about multiplayer AI, machine memory, and the uncomfortable politics of building your company on models that you yourself would like to see banned. This is Flo Crivello, founder and CEO of Lindy. Flo, CEO at Lindy. Welcome back to the Cognitive Revolution.
谢谢,N。一直是荣幸。
Thanks, N. Always an honor.
我对这次对话很兴奋。你在 Lindy 有大新闻,发布了新产品。这是这次对话的契机,但显然我们也正处于 AI 指数级发展的关键时刻。我知道你一直以来对许多关键的 AI 问题都直言不讳,随着对话深入,我肯定想深入探讨这些。但让我们从核心开始,谈谈你在做什么。Lindy 现在又进化了。几年前我们开始做类似智能 Zapier 的工作流软件。已经经历了几次重大进化。其中一个我肯定想深入探讨的是你最近关于调整模型组合的帖子,你稍微远离了美国专有模型,更多拥抱开源,并节省了一大笔钱。但现在最大的进化是 Lindy 成为一个完整的 AI 员工,它将像你所有其他队友一样出现在 Slack 中,无论是人类还是智能体。跟我们说说这个重大的新发布吧。
I am excited for this conversation. You got some big news at Lindy with a new product launch. That is the occasion for this conversation, but obviously we are in the thick of it when it comes to the AI exponential as well. And so you've been I know over time outspoken on a bunch of really critical AI issues and I definitely want to get into that as we get deeper into the conversation as well. But let's start at the center with what you are working on. Lindy is now evolved again. We started with kind of a smart Zapier workflow software kind of thing a couple years ago. There have been multiple big evolutions. One that I also definitely want to get into is your recent post on shifting your model mix and getting away from American proprietary models a bit and embracing more of the open source and saving a bunch of money. But the big evolution now is that Lindy is a fullon AI employee and it's going to be available in Slack just like all your other teammates whether they be human or agent. Tell us about the new big launch.
是的,你知道,从一开始,我们就一直在追求 AI 员工,我确实认为当我们三年前开始追求这个时还太早,现在我认为它基本上已经到来了。我们今天发布的产品叫做 Lindy Teammate。它是一个 AI 员工,住在你的 Slack 里,连接到你的所有工具,积累你整个团队的上下文,它真的像一个团队脚手架。我认为 AI 现在正处于向多人体验迈出巨大飞跃的过程中。我把它比作,你知道,也许你记得我们过去通过电子邮件互相发送奇怪的文档,带着修订等等,你知道,我把多人 AI 和单人 AI 的区别比作互相发送奇怪文档和 Google Docs 这样的真正共享文档。如果你真的想让你的 AI 成为队友,让你的智能体成为团队的实际成员,你希望它在你团队协作的地方,也就是 Slack,你希望它拥有关于整个团队的共享上下文、自己的共享记忆。你希望每个人都能与同一个智能体交谈和协作,而不是像现在这样,我们在同一个会议室里,都在说话,然后每次我们中有人想和可能是公司最重要的利益相关者——AI 智能体——交谈时,我们都必须离开房间再回来,你知道。所以这就是我们正在做的,这种新的多人体验和多人脚手架。
Yeah, you know, from like the get-go, we've always been going after like the AI employee and I I do think that it used to be quite early when we started to go after that like three years ago and now I think it's basically here. I think like what we are releasing today is called Linditimate. It is an AI employee that lives in your Slack, connects to all of your tools, accumulates your entire team's context and it's really like a team scaffold. I think like AI right now is in the middle of making this huge leap towards multiplayer experiences. I compare it to, you know, maybe you'll remember like we used to send each other like weird documents around by email, you know, with like revisions and stuff and and you know, I compare the difference between multiplayer and single player AI as between like sending each other weird documents and Google Docs like an actual shared document. If you really want your AI to be a teammate, your agent to be an actual member of the team, like you want it to be where your team collaborates, which is Slack, you want it to have its own shared context about the entire team, its own shared memory. You want everyone to be able to talk and collaborate with the same agent instead of right now it's like we're all in the same meeting room and we're all talking and then every time one of us wants to talk to what's turning out to be maybe the most important constituency of the company which is AI agents, we have to leave the room and then come back, you know. So that's what we're working on is like this new multiplayer experience and this multiplayer scaffold.
我觉得这个有很多角度都很有趣。首先,它如何入职?对吧。公司在人类员工身上投入了很多。是的,我自己也做过类似的事,今年早些时候我摸索着建立了自己的深度上下文,它对我帮助很大,但对我来说更容易,因为那只是我自己的东西,我拥有它,我不需要担心隐私期望,因为我是它唯一的消费者。当你进入多人模式时,我想,天哪,你有不同的频道,不同的人聚集在一起,也许有些频道是私密的,也许有些是公开的但私密,所以实际上,程序上你如何吸收所有信息,如何存储,然后在社交层面上,你在多人环境中识别出的棘手考虑是什么?
So many angles of that I think are interesting. How first of all, how does it on board? Right. This is that companies have put a lot into over time when it comes to their human employees. Yes, I've done this myself with my own kind of deep context that I stumbled my way through earlier this year and it is serving me really well, but it's easier for me cuz it's just my stuff and I own it all and I don't really have to worry about what the expectations were around privacy cuz again, I'm going to be the only consumer of it. when you get into multiplayer mode now I'm like oh gosh you've got different channels where different people were gathered and maybe some of those channels were private maybe some of them were public but private so just practically procedurally how do you suck up all the information how is it stored and then on a social level what are the tricky considerations that you are identifying in the multiplayer context
是的,这是一个非常好的问题,这正好触及了我们一直痴迷的事情,我确实认为随着我们接近 AGI,而且我们现在可以说已经有了 AGI,智能实际上相对而言越来越不重要,而上下文越来越重要。
Yeah, that is an excellent question and that that touches on exactly what we've been obsessing about which is I really do think that as we we're getting to AGI and as we now arguably have AGI, intelligence actually matters less and less comparatively speaking and context matters more and more.
你知道,我经常这样想:历史上最聪明的人之一是约翰·冯·诺依曼,对吧?如果你让约翰·冯·诺依曼神奇地出现在你办公室旁边,这个人在接下来一个小时或一天里,对你的用处还不如你随便一个同事,对吧?那是因为上下文。你有工作要做,没时间给约翰做入职培训。他没有上下文。所以你会说:“嘿,我爱你。我真的很想跟你聊聊,但现在我要跟这个人谈,因为我有重要的事要做。”所以我确实同意上下文超级重要。而且我认为,这是那些令人惊讶的时刻之一——实际上智能体在入职培训方面比人类做得更好。这不应该再让我惊讶了,但它总是让我惊讶,因为你总是按照这样的假设行事:哦,我们还没有 AGI,所以显然人类会更好,但事实并非如此。原因在于,公司让人类员工入职的主要方式之一往往是维基,对吧?有面对面的入职培训等等,但也有很多书面文档。而众所周知,书面文档一旦写出来就过时了。所以我们构建的是这个共享上下文层,以及我们所谓的“水合系统”。基本上它的工作方式是:你注册 Lindy,连接你的工具。所以尽管连接你的维基——那会有帮助,比如 Confluence、Notion、Google Docs 什么的——然后你连接最重要的一个,就是 Slack,因为真正的知识在那里。它很乱,但智能体不在乎乱。然后我们做的是,我们抓取你整个 Slack,基于文件系统构建知识图谱。也许之后我们还能嗡嗡作响并编辑。这就是它的样子。所以你入职,她立刻开始学习。你可以在这里看到,这是开始注册后 10 秒内,她还在学习。右边的图还很有限。所以她告诉我:“这是我对你的了解。”而且出奇地好。然后右边的图不断增长。重要的是,这个图有两层:个人层和工作区层。所以任何时候,你都可以进入 Lindy 的文件系统——这是新的 Lindy——你点击文件,它会把这份记忆文件汇集在一起,由一堆参考文件支持。而这一切都由一个团队智能体在后台自动维护。所以希望这回答了你的问题,关于我们是怎么做的,因为有些频道是公开的,有些是私密的,有些文档是公开的,有些是私密的。它的工作方式是,有一个记忆智能体。我们称之为打盹,而不是睡觉。所以它大约每 15 分钟运行一次——为什么你需要每 24 小时睡一次?它就在后台持续运行。任何公开的内容,它都会更新团队文件系统和团队记忆。任何个人和私密的内容,它只会更新每个人的文件系统。最后,它收集所有这些东西,当你与智能体交谈时,它会使用所有这些上下文,并抓取整个文件系统。所以为了达到这个目标,我们不得不解决很多技术挑战,因为它积累的上下文有数百万个词元,你不能在每一轮都加上所有这些。所以那是我们必须弄清楚的一件事。
You know, I often think of it as like, look, one of the smartest men in history was John von Neumann, right? If you were to have John von Neumann just magically appear next to you at the office, this guy over the next hour or day would be less useful to you than your random coworker, right? And that's because of context. You've got a job to do, and you don't have time to onboard John. He doesn't have the context. So you're like, 'Hey, I love you. I really want to talk to you, but right now I'm going to talk to this guy because I have an important thing to do.' So I do agree that context is super important. And I think it's one of those surprising times when actually agents are better than humans at onboarding. It shouldn't take me by surprise anymore, but it always does, because you always operate under the assumption that, oh, we don't have AGI yet, so obviously humans are going to be better, but they're not. And the reason for that is that very often one of the major ways companies onboard their human employees is wikis, right? There are in-person onboarding sessions and all of that stuff, but there's also a lot of written documentation. And as everyone knows, the moment written documentation is written, it's out of date. So what we've built is this shared context layer and what we call a hydration system. Basically the way it works is you sign up to Lindy, you connect your tools. So by all means connect your wiki—that'll help, you know, Confluence, Notion, Google Docs, whatever—and then you connect the most important one, which is Slack, because that's where the real knowledge lives. It's a mess, but agents don't mind the mess. So then what we do is we crawl your entire Slack and we build a knowledge graph based on the file system. Maybe we'll be able to buzz and edit after that. This is what it looks like. So you onboard, and immediately she starts learning. You can see here, this is within 10 seconds of starting to sign up, and she's still learning. The graph on the right is still limited. So she tells me, 'This is what I've learned about you.' And it's surprisingly good. And then this graph on the right keeps growing and growing. Importantly, there are two layers to this graph: a personal layer and a workspace layer. So at any moment, you can go into your file system in Lindy—this is the new Lindy—you click on files, and it brings together this memory file that is supported by a bunch of reference files. And this is all maintained automatically in the background by a team agent. So hopefully this answers your question about how you do it, because some channels are public, some are private, some documents are public, some are private. The way it works is that there is this memory agent. We call it napping, not sleeping. So it runs every 15 minutes or so—why would you need to sleep every 24 hours? It just runs continuously in the background. Anything that is public, it updates the team file system and team memory with. Anything that is personal and private, it just updates each person's file system with. And in the end, it collects all of that, and the agent, when you talk to it, uses all of that context and crawls that entire file system. So there are a lot of technical challenges we had to solve to get there, because the context it accumulates is like many millions of tokens, and you can't add all of that at every turn. So that was one thing we had to figure out.
是的,有意思。我会分享一些我偶然发现的东西,然后你告诉我你学到了什么可能更好的方法。首先,只需要连接所有工具并获取导出。我发现,在我试图编写的个人导出启发式规则中,我经常遇到一些有趣的边缘情况。有一次,我有一个启发式规则:我发的邮件越长,可能越有实质内容,我一定会确保捕捉到这些。但结果发现,在那个权力排名的顶端,往往是我从 LLM 复制出来、然后发给别人的东西,通常顶部有一条消息,比如“这是我从 Claude 得到的”之类的。但奇怪的是,根据那个启发式规则,这段 Claude 文本是我写作中最有分量的部分之一。所以随着时间的推移,我遇到了很多这样的情况,不得不特殊处理。
Yeah. Interesting. I'll share a little bit about what I stumbled into, and you tell me what you have learned that might be even better. First of all, just have to connect all the tools and get the exports. I found that there were interesting edge cases that I was constantly running into in my own personal export heuristic that I tried to write. One time I had a heuristic that the longer emails I sent are probably more substantive, and I'll definitely make sure I capture those. But then it turned out that at the very top of that power ranking was often something that I copied out of an LLM and was sending to somebody with usually a message at the top. Here's what I got from Claude, or what have you. But it was weird that according to that heuristic, this Claude text is one of the more weighty pieces of my writing. So I've bumped into a lot of these things over time and had to special case them.
是的。我想也许有一个问题就是,当你盲目地进入一个新组织,他们有日志——另一个对,在我们公司 Wayart 的 Slack 频道里,我们一次又一次地把各种日志塞进 Slack。你怎么处理?你怎么识别这些特殊的、可能淹没或误导它们的特殊情况,并找到真正好的东西?
Yeah. I guess maybe one question is just like, when you step into a new organization blind, and they have logging—another right, in our Slack channels at my company Wayart, we've had various logs bumped into Slack over and over again. How do you deal? How do you identify these idiosyncratic special case things that could overwhelm or flood or mislead them, and get to the real good stuff?
所以这是下一个问题。我认为我们解决这个问题的主要方式是,记忆本身由智能体维护。我认为这就是为什么我最终对 RAG 这种方法相当看空,而非常看好这种智能体式管理方法,因为你有一个真正的智能体,它有自己的记忆。所以你有某种元记忆,它理解自己在看什么,通过一点一点积累关于组织的知识,它越来越清楚什么才是真正重要的。这实际上有点像训练模型,因为如果你在模型的数据集中插入有毒数据——比如,嘿,达斯·维达是个女人——它会得到这些数据,但会被所有正确的数据淹没。所以这里也一样。如果你有足够的记忆,并且你把它喂给一个能理解它的系统,系统最终会理解。而你使用的那个非常具体的例子实际上是一种涌现行为。我们已经看到记忆智能体采用了——因为记忆智能体有自己的记忆——所以这是一种元记忆,关于我如何管理我的记忆?哪些信息来源重要,哪些值得信任?我们注意到,记忆智能体第一次抓取 Slack 时,它会发现一些这样的频道——很多组织都有,比如那些日志频道——然后它学会忽略它们。它就像,我不会再在这些频道上花时间了。那里没有我能学的东西。
So that's the next sentence question. I think the main way we've solved this is the fact that the memory is maintained by an agent itself. I think this is why I'm ultimately quite bearish on RAG as an approach, and I'm very bullish on this agentic management approach, because you have an actual agent which has its own memory. So you have a sort of meta memory, and it understands what it is that it's looking at, and by virtue of accumulating that knowledge little by little about the organization, it gets smarter and smarter about what actually matters. It's actually sort of similar to training a model, because if you were to insert poisoned data into the model's dataset—like, hey, Darth Vader was a woman—it would get the data, but it would be drowned by all the correct data. So here it's the same. If you have enough memories, and if you feed that to a system that understands it, the system ends up understanding. And the very concrete example you are using is actually an emergent behavior. We have seen the memory agent adopt—because the memory agent has its own memory—so it's a sort of meta memory about how do I manage my memory? What sources of information matter, and which ones are trustworthy? And we have noticed that the first time the memory agent crawls the Slack, it finds some of those channels—like many organizations have them, like those log channels—and it learns to ignore them. It's like, I'm not going to keep spending time on those channels. There's nothing for me to learn there.
对我们来说,关键的事情一直是将会议构建为这个系统中的一等公民。我确实相信,会议作为一种信息来源被严重低估了。
Thing that's been critical for us has been building meetings as a first-class citizen in this system. Like I really do believe that meetings are very underrated as a source of information.
他们说公司最新数据里 90% 都来自会议。公司里所有重要的事都有对应的会议——每段关系、每个项目、每项举措。所以我觉得那些多智能体系统没法真正把会议整合进 Granola 之类的工具。你必须把它作为一等公民纳入你的系统。所以我们做的就是构建这个一等公民。我们在 Lindy 里把会议作为首要设计。而且最重要的是,它不只是 Granola 那种——录下你的会议、给你笔记。它更进一步。它把这些会议喂给你的记忆智能体。如果某个会议是公开的,你可以设置会议文件夹,自动添加会议,并与整个公司共享。然后它会总结这些文件夹里的所有会议。如果某个会议被添加到一个与整个公司共享的公开会议文件夹中,它就会更新团队的记忆。所以它会持续更新那个上下文。现在你可以和它聊天——它基本上就是公司所有会议的一个缩影。你可以和整个会议语料库对话。你可以问:“客户在说什么?最近最大的需求是什么?关于这个或那个功能的反馈怎么样?”
They said that 90% of the most up-to-date data about the company comes from meetings. Everything that matters inside the company has a meeting around it—every relationship, every project, every initiative. So I think those multiplayer agentic systems can't really get meetings into Granola or whatever. You have to incorporate it into your system as a first-class citizen. So what we did is we built that first-class citizen. We have meetings as a first design in Lindy. And most importantly, it's not just a Granola—recording your meetings and giving you notes. It goes beyond that. It feeds these meetings to your memory agent. And if a meeting was public, you can set up meeting folders that automatically add meetings and share them with your entire company. So then it summarizes all the meetings in these folders. And if a meeting was added to a public meeting folder shared with the entire company, it updates the team's memory with it. So it keeps updating that context on an ongoing basis. Now you can chat with it—it's basically like an image of every meeting in the company. You can chat to that entire corpus of meetings. You can ask, 'What are customers saying? What's the biggest request lately? What's the feedback about this or that feature?'
所以这又带来了另一个有趣的挑战,我目前也摸索出了一个暂时的解决方案。我很想听听你对这个问题的看法,既包括回顾过去的方面,也包括展望未来的视角。问题在于,我的广泛通用上下文里有很多东西可能不应该与他人分享,对吧?有时是敏感的、忏悔式的、一个人对另一个人的看法,等等。所以我实际上为我的 wiki 创建了两个版本。一个在我的主力笔记本电脑上,用于我的工作,作为我自己的延伸和第二大脑。这个东西不会自主承接项目并运行。它只做我告诉它的事。它有完整的 wiki,里面包含所有八卦之类的东西。老实说,我并没有那么多超级敏感的东西。我不想让它听起来比实际更戏剧化。但尽管如此,人们并没有预料到,当他们告诉我一些事情——可能是在电话、电子邮件或私人 Slack 消息中——会进入一个智能体的中央存储库,而这个智能体现在要与世界对话。
So this brings up another interesting challenge that I've kind of stumbled my way to, at least the solution for now. And I'm very interested to get your take on both the backward-looking aspect and the forward-looking perspective. The issue is there's a lot of stuff in my broad general context that probably shouldn't be shared with other people, right? Sometimes it's sensitive, confessional, one person's point of view on another person, whatever. So I've actually created for my own wiki two versions of it. One is on my main laptop where I do my work, as an extension of myself and a second brain type of thing. So this thing is not taking on projects autonomously and running with them. It's just doing what I tell it to do. And it has the full wiki with all the gossip or whatever that's in there. I don't have honestly that much super sensitive stuff. I don't want to make it sound like it's more dramatic than it is. But nevertheless, people didn't expect that when they were telling me something—which could have been on a call or an email or a private Slack message—that it was going to go into some central repository of an agent that was going to now talk to the world.
嗯。
Yeah.
所以那个版本只给我自己用。我让它用“什么适合告诉人类助理”的启发式方法创建了一个版本。那是一个仍然私密但更面向公众的 wiki,适合我的人类助理拥有你的电子邮件和电话号码,对吧?但可能不适合把我们有过的每次对话的每个细节都放进去。当你们吸收所有这些历史信息时,你们如何考虑对记忆里应该有什么、不应该有什么保持敏感,以便信息流向它应该被分享的地方?我认为部分在于组织现在将如何改变,因为我觉得这可能在实时重写社会契约。所以我认为对于之前的一切,应用或一种策略,但答案可能是未来新的社会规范。我也想听听你的看法。
So that one's just for me. I had it go through and create a version using the heuristic of what would be appropriate for a person to tell a human assistant. The sort of still private but a little bit more public-facing wiki where it would be appropriate for my human assistant to have your email and phone number, right? But it might not be appropriate for every detail of every conversation we've ever had to be in there. How are you thinking about, as you absorb all this historical information, being sensitive to what should and shouldn't be in memory such that it goes where it ought to be shared? I think in part it's like how are organizations going to change now, because I think this is all rewriting the social contract potentially in real time. So I think application or one strategy for everything that came before, but the answer might be like new social norms going forward. I want to get your take on that too.
是的。这在团队内部引发了一场激烈的辩论。一直有两个阵营。有一个阵营——坦白说,我就属于那个阵营——认为两层记忆就够了,对吧?有公共团队层,然后有私人层,私人层包含一切,公共层只包含那些仅限私密公开的内容。团队里有些成员一直在说你说的话。他们说:“不,实际上,即使是我的私人层,我也不想包含很多东西。我想要一个超级私密的密钥。”然后他们试图——有这样一个想法:“如果我们定义多个记忆气泡,用户可以编辑它们呢?”我说:“那听起来有点过度设计。”所以我们最终达成的共识,以及这些团队成员最终达成的共识,是他们编辑了自己的元记忆提示。而元记忆提示只是一个文本文件——它是你文件系统中的 memory.md 文件。在 Lindy 中,你的记忆智能体在每一刻都会将该记忆提示注入其上下文窗口。所以如果你在文件顶部或任何地方插入一行,比如“这是我希望你永远不要记住的”,你可以这样做。“这是我希望你永远不要记住的。哦,顺便说一句,那些敏感话题,请把它记在那个文件夹里的那个文件里,那个文件夹是看不见的,除非 X、Y 或 Z 条件满足,否则我不想让你拉取这个文件或文件夹。”所以你可以直接留下这些指令。所以这也是我看空 RAG、看多文本和文件系统的另一个原因:你可以检查记忆,你可以编辑它,你可以非常精细地插入这种护栏。
Yeah. This has been a vigorous debate inside the team. There have been two camps. There are the camps that are like—frankly, I'm in that camp—there are just two tiers of memory is enough, right? There's the public team tier and then there's the private tier, and the private tier contains everything, and the public tier contains stuff that's only private-public. And some of the members of the team have been saying what you've been saying. They've been saying, 'No, actually, even my private tier I don't want to contain a lot of stuff. I want like a super private key.' And then they were trying—there was this whole thing like, 'What if we defined multiple memory bubbles and the user can edit them?' And I'm like, 'That sounds kind of overkill.' So what we landed on, and what these teammates landed on honestly, is they have edited their own meta-memory prompt. And the meta-memory prompt is just a text file—it's your memory.md file in your file system. In Lindy, your memory agent has that memory prompt injected into its context window at every moment. So if you insert a line up there that's like, 'This is what I never want you to remember,' just at the top or anywhere in the file, you can do that. 'This is what I never want you to remember. Oh, by the way, that other stuff that's a sensitive topic, please remember it in that other file in that other folder that's out of view, and I don't want you to pull this file or this folder unless X, Y, or Z.' So you can just leave these instructions. So that's another reason why I'm bearish on RAG and bullish on text and file system: you can inspect the memory, you can edit it, and you can very granularly insert this kind of guardrails.
有趣。所以为了确保我理解正确,你能复述一下吗?一个真实来源,你通过提示词在信息上创建不同的视角。所以你可以有——我可以告诉我的智能体:“嘿,你要知道,这包含了所有我经历过的对话。始终使用这样的启发式方法:你只应该使用那些适合我与人类助理分享的信息。如果看起来不合适,就不要使用。”显然,指令遵循能力正在变得极其出色。
Interesting. So just to make sure I understand, can you repeat it back? One source of ground truth and you create different lenses on that information just by prompting. So you can have—I could tell my agent, 'Hey, just so you know, this contains everything from every conversation I've ever had. Always use the heuristic that you should only really be using information that would have been appropriate for me to share with a human assistant. And if it doesn't seem appropriate, then don't use it.' And obviously instruction following is getting extremely good.
所以我可以做得很好。实际上,我刚才说错了——有一次我确实用过这个,而且就是为了这个播客。我准备来上这个播客时,知道我要谈论记忆智能体。所以你看,这里我有一个 memory.md,那是我实际的记忆文件,然后我有一个 memory2.md,那是一个经过净化的记忆文件,我删除了过于敏感的信息。你可以看到在我的 memory.md 里有一个注释,说文件系统中还包含一个 memory2.md 文件。忽略其内容。它们只是为了让用户在播客上演示 Lindy 而存在的。所以是的,你可以在这里做你自己的事情。
So I can pretty well. Actually, I was misspeaking—there is a time when I have used this, and it's literally for this podcast. I was preparing to go on this podcast, and I knew I was going to talk about the memory agent. So you see here I have a memory.md which is my actual memory file, and then I have a memory2.md which is a sanitized memory file where I've removed overly sensitive information. And you can see in my memory.md here there is a note in the file system also containing a memory2.md file. Ignore its contents. They are just here for the user demoing Lindy on podcasts. So yeah, you can just do your own thing here.
嗯。好的。酷。
Yeah. Okay. Cool.
我确实如此,因为我的版本目前存在不够精简的问题。我需要维护两套东西,这确实带来了一些额外开销。所以,我能理解为什么那可能是有利的。
I do because my version does have the it's not dry problem right now. I've got two things to maintain and it does create some overhead. So, I can see why that could be advantageous.
嘿,稍后我们将继续采访,先听一段赞助商广告。
Hey, we'll continue our interview in a moment after a word from our sponsors.
今天的节目由 Anthropic 赞助,他们是 Claude 和 Claude Code 的开发者。过去几个月,Claude 帮助我构建并完善了一个个人深度上下文数据库,其中包含了我过去整整 5 年的所有电子邮件、Slack 消息、推文、跨平台私信、视频通话和播客转录。在此基础上,我们现在又添加了描述我与数百个联系人、组织和想法关系的摘要文章。现在有了这个数据库,几乎没有什么事是 Claude 帮不上忙的。对于我的天使投资,Claude 现在可以根据我与创始人的通话和邮件往来,草拟出完全符合我风险基金要求的投资备忘录。当有人需要帮忙时,Claude 往往能做得和我一样好。最近,一位朋友联系我,问我是否认识适合他正在招聘职位的人选。起初我没想到任何人,但后来我想到去问 Claude,果然,它找到了两个很好的线索。Claude 是专为不满足于“足够好”的头脑而设计的 AI。它是一个真正理解你整个工作流程并与你共同思考的协作者。所以,对于值得解决的问题,请访问 claude.ai/tcr 开始使用 Claude。网址是 claude.ai/tcr。也请查看 Claude Pro,它包含今天节目中提到的所有功能。网址是 claude.ai/tcr。
Today's episode is brought to you by Anthropic, makers of Claude and Claude Code. Over the last few months, Claude has helped me build and refine a personal deep context database that now contains all of my emails, Slack messages, tweets, DMs across platforms, video calls, and podcast transcripts going back a full 5 years. On top of that, we've now layered summary articles describing my relationship with hundreds of contacts, organizations, and ideas. And now that this exists, there's almost nothing that Claude can't help with. for my angel investing. Claude can now draft investment memos in exactly the form that my venture fund requires based on the calls I've had and the emails I've exchanged with the founders. And when someone needs a favor, Claude can often do it as well as I can. Recently, a friend reached out to ask if I know anyone who might be a fit for a role that he is currently hiring for. Initially, nobody came to mind, but then I thought to ask Claude, and sure enough, it identified two great leads. Claude is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow and thinks with you. So for problems worth solving, get started with Claude at claude.ai/tcr. That's claude.ai/tcr. And check out Claude Pro, which includes all of the features mentioned in today's episode. That's claude.ai/tcr.
你在做多人协作时,还遇到了哪些其他问题?我记得回到 GPT4 时代,当时我在做红队测试。我做过一些,这比我想象的要花更长时间。我之所以提起这个,是因为那时我在模拟一个引导者。你是一个 AI 引导者,在一个以锻炼为重点的家庭小组中,你的工作是鼓励人们并给予一些提醒等等。我当时扮演所有参与者的角色,让 AI 扮演那个角色。即使在 GPT4 时代,它在相对简单的场景下也表现得相当不错。然而,现在已经过去大约 3 年了,我们才终于看到这种真正的 AI 员工式体验。让多人协作以直观的方式运作,到底难在哪里?这对那些没有亲身经历过这种艰辛的人来说可能并不明显。
What other things are kind of coming up as you're doing multiplayer? I remember going back to GPT4 way back when I was red teaming. I did some It's taken longer than I thought. And the reason I even bring this up is because back then I was doing some simulations of a facilitator. You're an AI facilitator in a group, family group that's focused on exercise and your job is to encourage people and give some reminders and whatever. And I was playing all the roles of the participants and just having the AI play that role. And even at GPT4, it was like doing pretty well in a relatively simple context. And yet, it's been like 3 years now until we're finally getting these real kind of AI employee type experiences. What has been hard about getting multiplayer to work in an intuitive way that might be unobvious to somebody who hasn't been through the slog himself?
是的,我实际上认为这在于脚手架,上下文管理,比如这部分,上下文的构建。它在某种程度上受到了 Rathy 的 auto wiki 想法的启发。所以我们必须在上下文管理上做大量工作,因为一旦你与 AI 员工交谈,它实际上与和 TajTo 或 Claude 交谈非常不同,因为你确实期望它保持对其过去上下文和记忆的非常丰富的表示。你对 AI 员工期望的一致性和连贯性水平,是你对 Claude 所不期望的。就像 Claude 有点可爱,有时它会在你的聊天中插入关于你的先前记忆,但你不会像对待员工那样与它进行持续的繁重工作。所以管理所有这些上下文。既要管理构建上下文的记忆智能体,每个用户有数百万、数百万、数百万个词元,还要管理实际使用该上下文的核心智能体,以及它如何在运行时检索正确的信息。这确实是一个重大、重大、重大的挑战。坦率地说,我有时告诉团队,我无法想象我们是唯一这么想的公司。我有时觉得我们应该发表论文,因为我认为我们正在做的事情是真正最先进的,而且我们经常做一些事情,然后三到六个月后我们看到一篇关于那件事的论文出现并爆火。有一次,论文的名字竟然就是我们内部给那东西起的名字,因为太明显了。所以我认为上下文和记忆管理一直是挑战中非常非常大的一部分。可靠性总是另一个挑战,对吧?你希望你的模型工作。所以我们在可靠性上做了相当多的工作。我们称之为验证器。它基本上是一种作为评判者的 LLM,在任务期间多次触发,但它是模块化的。所以你有多个 LLM 作为评判者分派出去,它有点像委员会,然后他们互相交谈并决定该做什么。所以还有一个自我改进循环,我认为自从 GPT4 以来,模型已经变得如此强大,以至于你现在实际上可以拥有自我改进循环,所以你知道我们不是第一个谈论它的人,也不例外,你知道,Lindy 现在正在自我改进,所以我们实际上可以看到错误率曲线向右下方下降,你知道那是你希望错误率下降的方向。在我们上线自我改进循环的第一周内,错误率就下降了,那大约是两个月前。错误率下降了 8 倍。我认为这可能概括了它。我认为那些是我们必须解决的大块头问题。
Yeah, I actually think it's been the scaffold, the context management, like this piece, like the context buildup. It's in a way inspired by Rathy's auto wiki idea. And so we've had to do a lot of work around context management because once you talk to an AI employee, it's actually quite unlike just talking to a TajTo or a Claude because you actually expect it to keep a very rich representation of its past context and of its past memories. You really expect a level of consistency and coherence out of an AI employee that you don't expect out of your Claude. Like Claude sort of, it's cute. Sometimes it plugs like previous memories about you in your chat, but you don't really do heavy work on an ongoing basis with it in the same way that you do with an employee. So managing all of that context. So both the memory agent that builds up the context that's millions and millions and millions of tokens per user and then the core agent that actually uses that context and how does it retrieve the right information at runtime. Like this has been a major major major challenge. Frankly, I've sometimes been telling the team like, and I can't imagine we're the only company thinking that. Like, I sometimes feel like we should be publishing because I think we're doing stuff that's like seriously state-of-the-art and it's quite frequently that we do stuff and like 3 to six months later we see a paper come out and blow up about that thing. There was once time when literally the paper was named what we had called the thing internally because it was obvious. And so I think like context and memory management has been a really really big part of the challenge. Reliability is always another part of the challenge, right? Like you want to make your model work. And so we've worked quite a bit on reliability. Like we've called it like a validator. It's basically a sort of it's an LLM as a judge that triggers like multiple times during the tasks, but it's modular. So you have multiple LLMs there as a judge fan out and it's a sort of council and then they talk to each other and they decide what to do and so there's been one the self-improvement loop I think like since GPT4 models have become so capable that now you can actually have self-improvement loops and so you know we're not the first ones to talk about it, no exception, you know, Lindy is now self-improving so we can literally see a curve of error rate go down and to the right, you know that's the direction you want to see error rate go down. It went down by like the literally within the first week of us putting the self-improvement tube online which was like two months ago or something. It went down by 8x the error rate. I think this probably captures it. I think those have been like the really big meaty chunks we've had to figure out.
这让我想到一个大问题,就是你如何管理缓存,当然缓存只是管理总体成本的一个角度。也许我们可以超越缓存,谈谈成本管理。显然,你会为来自新组织的所有这些上下文投入并处理大量的前期词元。所以我对这带来的商业模式影响很感兴趣。
One big question that brings to mind is how you are managing caching and then caching of course is just one angle on managing cost in general. Also maybe we could expand beyond caching and talk about cost management. Obviously there's a significant upfront amount of tokens that you're going to dedicate to in all this context from a new organization and processing it. And so I'm interested in the business model implications of that.
你们需要收取设置费或签订年度合同吗?你们如何平衡初始投资与公司的承诺水平?而且,随着大量 token 不断被处理,记忆在不同情境下以不同方式组装,不同用户结合公共与私人信息,你们如何管理输入 token 量?缓存策略在其中占多大比重?你们如何确保最终成本仍低于雇佣一名人类员工?
Do you have to charge a setup fee or do you need an annual contract? How are you balancing your initial investment with the level of commitment from the company? And then with so many tokens getting processed all the time and with memory being assembled in different ways and different situations all the time, different users with their combination of the public and the private, how are you thinking about managing input token volume? How much is caching playing into the strategy? How are you making this something that you know still on net ends up costing less than a human employee?
头条新闻,你知道,我觉得头条新闻会说这些东西仍然比人类员工更贵,但不会持续太久。而且,说实话,我们正在补贴它。你知道,这就是我们筹集这些资金的目的。实际上,有个有趣的事实:我们过去曾大力补贴,然后我们著名地转向了中国的模型,这让我们停止了补贴。现在有了 teammate,我们意识到用户再次提出了比我们之前产品(更像个人助理)更复杂的需求。所以之前是简单任务,比如“嘿,把这封邮件发给 Flo,安排这个会议。”我总是说,你不需要上帝来安排你的会议。所以即使 DeepSeek Flash 也足够了,坦率说,那很划算。有了 Lindy teammate,我们又回到了负毛利区间。话虽如此,显然我不喜欢负毛利。但我对此感到安心,因为我们非常确信这只是暂时的。而且我实际上认为你总是应该为下一代模型而构建。所以我们正努力减轻其影响,是的,上下文管理是拼图的重要部分。缓存也是重要部分。我们当前的缓存率是 85%,坦率说这低于应有水平,我们花了很多时间迭代。我们建立了许多系统,在缓存率下降时提醒我们,因为它很挑剔。你在系统中任何地方做任何改动,都会破坏缓存,缓存率就会从 85% 降到 65%。两者之间的差距听起来很小,但实际上从 85% 到 65% 价格几乎翻倍。所以我们建立了所有这些系统。
Headline, you know, I think headline like these things still cost more than the human employee, but not for very long. And look, I'll be real, like we're subsidizing it. You know, this is what we've raised this money for. Actually, fun fact, we used to subsidize it very heavily, and then we sort of famously switched to a Chinese model, and that made us stop subsidizing it. Now with teammate, we have realized our users again are throwing much more complex stuff than they were asking to our previous product, which was more like a personal assistant. So it was simple-ish tasks, like 'hey, send this email to Flo, schedule this meeting.' I always say you don't need God to schedule your meeting. So even a DeepSeek Flash was enough, frankly, and that was cost-effective. With Lindy teammate, we're back at it, and back into, frankly, negative gross margin territory. And so that said, obviously I don't like having negative gross margins. I'm at peace with it because we are very, very confident that it's very temporary. And I actually think you sort of want to build for the next generation of models always. So we're trying to mitigate the impact of that, and yes, context management is a huge piece of the puzzle. Caching is a huge piece of the puzzle. Our current cache rate is at 85%, which is lower than it should be, frankly, and we spend a lot of time just iterating on it. We put a lot of systems in place to alert us when the cache rate dips, because it's finicky. You make any change anywhere in your system and you break your cache, and now your cache rate dips from 85% to 65%. And the gap between both, it sounds small, but actually it's almost 2x the price to go from 85 to 65%. So we set up all of those systems.
我认为我们最大的突破之一,这是那种我坦率期待有人会发表的东西,我猜 RLM 与之类似,但我们称之为上下文桶。它的起源是:如果智能体调用一个返回过多上下文的动作,比如某些动作,特别是某些 MCP,会检索出 10 万 token 之类的东西。你不想把那些都发给智能体,它会困惑。所以你要做的是,将其暴露为上下文桶的摘要,并且它是一个包含整个桶的子智能体。就像,“嘿,这个动作返回了太多内容,所以我在这里代表实际返回的内容,大致上这就是它包含的内容。”然后智能体可以与子智能体进行对话,子智能体有自己的缓存,并可以使用 Unix 工具操作上下文。所以在这里你节省了很多钱,而且速度很快。
I think one of the biggest breakthroughs we had, and this is one of those things that I'm frankly expecting someone will publish, I guess RLM was adjacent to that, but we call them context buckets. The way it started was: if an agent calls an action that returns too much context, like some actions, some MCPs in particular, retrieve like 100,000 tokens or something. You don't want to send that to the agent; it's going to get confused. So what you do is you expose that as a summary of the context bucket, and it's a sub-agent which contains the entire bucket. It's like, 'hey, this action returned too much, so I'm here to stand in for what it actually returned, roughly speaking this is what it contains.' Then the agent can enter into a conversation with the sub-agent, which itself has its own caching and can manipulate the context using Unix utilities. So right here you're saving a lot of money, and it goes quite fast.
然后我们更进一步,我们想,“嘿,如果我们能有递归上下文桶呢?如果上下文桶能包含其他上下文桶呢?”如果压缩(显然我们有压缩)由这种递归上下文桶驱动呢?我的意思是,现在当我们压缩时,对话会超过阈值。我认为现在大约是 20 万 token,但我们一直在调整。在某个时刻,我们会说,“好了,我们要压缩了。”我们压缩,然后压缩将所有内容发送到一个上下文桶中。所以现在智能体可以查询那个上下文。所以并不是每次压缩本质上都是有损的。我认为压缩基于一个错误假设,即你永远不需要访问真实数据,这是错误的。你有时确实需要访问真实数据。所以有了这种技术,你可以访问真实数据。
Then we went one step further and we were like, 'hey, what if we could have recursive context buckets? What if context buckets could contain other context buckets?' And what if compaction, because obviously we have compaction, was powered by such recursive context buckets? By that I mean, now when we compact, the conversation goes, it passes threshold. I think right now it's like 200,000 tokens, but we keep tweaking it. At some point we're like, 'okay, we're going to compact.' We compact, and then the compaction sends all of that stuff into a context bucket. So now the agent can query that context. So it's not like every compaction is always lossy by nature. I think compaction operates under the faulty assumption that you never need access to ground truth, which is false. You at some point do need access to ground truth. So with that technique, you have access to ground truth.
然后我们更进一步:你有一个上下文桶,它是之前对话的压缩。你的对话继续、继续、继续,你需要再次总结。你要做的是,把包括这个上下文桶在内的所有内容,压缩成一个新的上下文桶。所以现在你有一个包含上下文桶的上下文桶。最终,你这样做,就会产生这种涌现特性:智能体可以以任意粒度访问任意无限数量 token 的任意点。问题是,如果你有计算机科学背景,这会产生所谓的 N 复杂度,对吧?所以如果它想访问九个上下文桶之前的内容,它必须经过九层子智能体,这非常慢且非常昂贵。所以我们最终做的是——也许这太技术化了,但我们从未——
Then we went a step further: you have that context bucket which is the compaction of the previous conversation. Your conversation keeps going, keeps going, keeps going, you need to summarize again. What you do is you take all of that, including this context bucket right here, and you compact it into a new context bucket. So now you have a context bucket containing a context bucket. And in the end, you do that, and you have this emerging property where the agent can access any point of any infinite number of tokens at arbitrary levels of granularity. The problem is, if you have a computer science background, this gives you what's called N complexity, right? So if it wants to access nine context buckets ago, it's got to go through nine layers of sub-agents, and that's really slow and really expensive. So what we ended up doing—maybe it's getting too technical, but we never—
好的。嗯,你听说过 AVL 树吗?
Okay. Well, have you heard of AVL trees?
没有,但我洗耳恭听。
No, but I'm all ears.
你听说过红黑树吗?
Have you heard of red-black trees?
我上的是计算机科学“社会大学”,所以没有接受过正规训练。我实际上很高兴有 AVL 树和红黑树,因为我想,“天哪,你们还记得你们的训练,这就是我们使用那些东西的时刻。”因为学校里每个人总是问,“你什么时候真正用到那些东西?”我们正在使用 AVL 树和红黑树,宝贝。它们叫红黑树。所以我刚才描述的,如果你只是做朴素实现,你会免费获得那个非常好的涌现特性:你有所有这些相互包含的链接上下文桶。但每个上下文桶总是只包含另一个上下文桶。所以你会得到一种俄罗斯套娃式的上下文桶。如果你想深入到底部,那会花很长时间。所以你要做的是,让上下文桶包含多个上下文桶。最终你会得到一棵树,因为基本上你想要的是,最顶层的上下文桶不是包含最后 100 个,而是包含第一个上下文桶和最后一个上下文桶。所以最终你会得到一棵自平衡树。有两种算法:AVL 树和红黑树,它们是用于平衡树的算法。
I went to the computer science school of hard knocks, so it wasn't formal training for me. I was actually glad to have AVL and red-black trees, because I was like, 'oh my god, you guys remember your training, this is the moment we use those things.' Because everybody in school is always like, 'when do you really use those things?' We're using AVL and red-black trees, baby. Red-black trees, they're called. So what I just described, if you just do the naive implementation, you get that really nice emerging property for free: you have all of those linked context buckets that contain one another. But every context bucket always contains just one other context bucket. So you end up with this Russian doll of sorts of context buckets. If you want to get to the bottom, it takes a very long time. So what you do instead is you have context buckets contain multiple context buckets. And you end up with a tree, because basically what you want is you want the topmost context bucket not to contain the last 100. You want it to contain the first context bucket and the last context bucket. And so you end up with a self-balancing tree. There's these two algorithms: the AVL tree and the red-black tree, which are algorithms used to balance a tree.
所以与其用一条长线,你要最小化树的高度,这样到达底部所需的跳跃次数尽可能少。技术上,AVL 树是平衡树的最佳方式,因为它能带来最低的高度。红黑树更好,因为它考虑了平衡树的成本,这很昂贵,因为你需要重新生成很多上下文包,而且这样做会丢失缓存。所以红黑树是二叉树上平衡树的经典实现,二叉树就是每个节点有两个子节点的树。我们采用的是 centary 树。它其实就是每个节点有 100 个子节点。所以下面的节点就会有 10,000 个节点,对吧?所以字面上,两次跳跃就能有 10,000 个上下文桶,你就能访问宇宙中所有的上下文。这带来了一些非常惊人的行为,因为你距离访问 10,000 个上下文桶(每个包含 200,000 个 token)只差两次 LLM 调用。所以你在两次调用中就有大约 20 亿 token 的上下文。这就是那些 AI 智能体表现出惊人行为的原因,你问它们任何问题,它们都能完美记住一切。这太棒了。那是我们必须解决的一个大问题。我希望你们能复制我们。我只想说我们只是没有时间发表,但这并不是要保密。我认为这是一个非常强大的技术,我很惊讶没有看到更多关于它的讨论。
So instead of having one long line, you want to minimize the height of the tree such that going to the bottom takes as few jumps as possible. Technically, AVL is the demonstrably best way to balance the tree because it leads to the lowest height. A red-black tree is better because it takes into account the cost to balance the tree, which is expensive because you need to regenerate a lot of your context packets and you miss your cache when you do that. So red-black trees are the canonical implementation of balanced trees on binary trees, which is just a tree where each node has two children. We went for a centary tree. It's literally just each node has 100 children. So the node below that is going to have like 10,000 nodes, right? And so literally with two jumps, you can have 10,000 context buckets and you can access like all the context in the universe. And that leads to some really surprising behaviors because you're literally no more than two LLM calls away from being able to access 10,000 context buckets, each one of which contains 200,000 tokens, right? So you're at like two billion tokens of context in two calls. And that is what leads to those really surprising behaviors from those AI agents where you ask them any question and they remember everything perfectly all the time. That's awesome. That was one really big thing we had to figure out. I hope you copy us. I will just say we just don't have the time to publish, but this is not meant to be secret. I think this is a really powerful technique and I have been surprised to not see more communication about it.
只是为了校准一下,你发现企业有多少 token?当你说这两层能达到 20 亿 token 时,我对此没有很好的直觉。这能覆盖一个成立 5 年的 20 人团队吗?团队规模和历史长度如何换算成 token 数量,有什么启发式方法?
Just for calibration, how many tokens are you finding businesses have? Like when you say with these two levels you can get to two billion tokens, I don't have a great intuition for that. Does that cover a 20-person team that's been in business for 5 years? What's the heuristic for team size and length of history that translates into how many tokens?
你大概可以算一下,但 20 亿 token 是……我们应该问一下 Claude,但它可能比大多数图书馆都大。所以我认为我们很少有客户的知识库或记忆能超过这个量。大多数时候,当你开始使用时,对于一个 20 人的团队,你会消耗掉水合组件。所以这又是那个时刻,你连接你的 Notion,连接你的 Slack,然后我们爬取所有内容。我们查看所有这些数据。对于一个 20 人的团队,这通常最多消耗 500 万 token,你知道,300 万到 500 万 token,而且这包括一个 20 人的团队和多年的 Slack 历史。我们确实会获取你所有的 Slack 历史。这需要一段时间。而且实际上,令人惊讶的是,瓶颈甚至不是语言模型,而是 Slack API。Slack 对你爬取它的体验并不满意。
You can probably do the math, but two billion tokens is... we should ask Claude, but it's probably bigger than most libraries. So I think we have few customers whose knowledge base, whose memory is bigger than that. Most of the time when you get started, you consume, for a team of 20, the hydration component. So it's again this moment where you connect your Notion, you connect your Slack, and then we crawl everything. We look at all of that stuff. For a team of 20, that's going to tend to consume 5 million tokens at most, you know, 3 to 5 million tokens, and that's with a team of 20 and years of Slack history. We literally take all of your Slack history. That takes a while. And actually, surprisingly, the bottleneck is not even the LM, it's the Slack API. Slack is not happy about you crawling its experience.
是的,我确定如果你个人做的话,正确的方法是导出你的工作区存档。我们不能要求用户这样做。所以我们只能爬取整个历史。
Yes, I'm sure the right way to do it if you're doing it personally is to export your workspace archive. We can't ask users to do that. So we just crawl the entire history.
不,我的意思是,20 亿 token 是很多。是的。我有一个数据点是播客,我相当自我放纵,我们有时会聊很长时间,但通常也只在 3 万到 5 万 token 的范围内。所以即使几年内做了几百集,我们整个播客的全文历史也只在个位数,可能低两位数的百万级。我们距离进入十亿级还有几个数量级。话虽如此,现在想想一家公司,每个员工每天有 3 到 7 小时的会议,乘以 2000 名员工。那就是会议真正开始快速累积 token 的时候。
No, I mean, two billion tokens is a lot. Yeah. One data point I have is a podcast, and I'm quite self-indulgent and we sometimes go on for a long time, but still it's usually only in the sort of 30 to 50,000 tokens range. So even having done hundreds over the course of a few years, we're still only in like single digit, maybe low double digit millions for the entire full text history of the hundreds of podcast episodes. We still have orders of magnitude to go before we would get into the billions. That said, now think of a company that's got three to seven hours of meetings per day per employee, times 2,000 employees. That is when meetings are when you really start to rack up tokens really, really quickly.
你怎么……所以我确实经历过 Slack 的事情。实际上,有趣的轶事是我唯一一次真正经历数据丢失并不太痛苦,但 Claude 发现了某个弱点什么的,我一会儿也会提到这个,但我的原始数据存储在 SQLite 数据库中,Claude 发现了一些它想纠正的缺陷,它就说:“好吧,我就删掉那个数据库,我们重做。”但它没有考虑到我们已经被 Slack 限流了好几天,才把数据从 Slack 里弄出来。所以它就像:“不,你删掉了那个。”那是我们唯一有 Slack 数据的地方。数据并没有真正丢失,但就像是:“好吧,我们现在又要被限流五天,才能通过 API 调用把数据从 Slack 里弄出来。”
How do you... So I did experience the Slack thing. Actually, the funny anecdote is the only time I've really experienced data loss wasn't too painful, but Claude had identified some weakness or whatever, and I'll bring this back around in a second too, but my raw data is living in a SQLite database, and Claude recognized some deficiency or whatever that it wanted to correct, and it just was like, "All right, I'll just drop that database and we'll redo it." But what it wasn't taking into account was the fact that we'd been rate limited on Slack for days to get all the stuff out of Slack. And so it was like, "No, you dropped that." And that was the only place where we had Slack. It wasn't really lost, but it was like, "Okay, we're now going to be rate limited again for five more days just to make API calls to get out of Slack."
就我而言,我至少是我最关心的 Slack 账户的管理员。但我所在的有些工作区我不是管理员,所以我无法授予这种访问权限。从商业角度来看,这是否意味着你必须从一开始就获得管理员的认可?这是否……这似乎使得产品主导的增长、落地和扩展这类事情变得困难?
In my case, I am the admin of at least the Slack accounts that matter to me the most. But there are workspaces that I'm in where I'm not the admin, and so I can't grant this kind of access. From a business standpoint, does that mean that you have to have admin buy-in from the beginning? Is there... that seems like it makes it hard to do product-led growth, land and expand type stuff?
这是个好观点。你知道,看,我认为这就是为什么 PLG 最终在向上市场拓展时会遇到瓶颈。是的,你必须是管理员。你必须有权在你的 Slack 工作区安装应用程序。所以对于 50 到 100 人的团队来说,这通常没问题。超过 100 人,很少有人能拥有这种访问权限。但许多 100 人以下的团队,几乎任何人都可以安装 Slack 应用。是的,我的意思是,如果你进入更大的市场,成为更大组织的一部分,那么我们有一些用户在我们内部支持我们,他们基本上在恳求 IT 让我们进入,然后他们安排这些会议。而且这确实是谨慎的,但你看,我们符合 SOC 2、HIPAA、GDPR 标准。我们已经做好了准备,而且我们非常习惯于应对内部审查流程,这些流程可能很繁琐。所以我们不是要成为任何人的麻烦,而且他们自己实际上也非常渴望。通常我们会得到一个附加组件,比如:“嘿,伙计们,我们真的需要成为一家 AI 原生的公司,”所以他们实际上在寻求这种解决方案。但是,是的,我的意思是,你需要管理员访问权限。
That's a good point. You know, look, I think that's one reason why PLG eventually you tap out of it as you go up market. Like yes, you have to be an admin. You have to have permissions to install applications on your Slack workspace. So it tends to be fine for teams of up to like 50 to 100. After 100, it's pretty rare for everyone to have this kind of access. But many teams up to 100, pretty much anyone can just Slack install an app. And yeah, I mean, if you go up market and you're part of a bigger organization, then we have some users who are championing us internally and they're basically begging the IT to let us in, and then they're making these meetings happen. And it is rightfully cautious, but look, we're SOC 2 compliant, we're HIPAA compliant, we're GDPR compliant. We've got our ducks in a row, and we're quite used to navigating the internal review processes, which can be burdensome. And so we're not in the business of being a pain in the butt of anyone, and they are actually themselves very eager. Often we're given an add-on like, "Hey guys, we need to actually become an AI-native company," so they're actually seeking this kind of solution. But yeah, I mean, you need admin access for that.
好的,让我们回到数据结构一会儿。让我给你我的数据结构,你可以告诉我你认为与你的专业版本相比,我的自制版本可能遗漏了什么。然后这是我最近最喜欢的习惯之一。我会把转录稿交给 Claude,看看我们能不能缩小一点差距。
Okay, so let's go back to the data structure for a second. Let me give you my data structure and you can tell me what you think I might be leaving on the table with my homespun version compared to your professional version. And then it's one of my favorite habits these days. I'll take the transcript and give it to Claude and we'll see if we can't close the gap a little bit.
所以我的做法就是,把我这些年产生的所有内容,也就是所谓的数字排泄物,包括邮件、私信、Slack、Google Docs 等等,其中 Slack 是最麻烦的。通过 API 调用把所有这些导出到 SQL 数据库里。我把所有东西都映射到线程和消息结构上。有时候需要费点劲,但基本上能行。这就是原始的“地面真相”。我同意,对我来说,我发现让模型能够接触到原始真相是非常重要的。
So my version is simply take all of the content that you know, all the sort of digital exhaust that I've created over the years, which is email and DMs and Slack and Google Docs and whatever, right, with Slack being the most painful one. Make all the API calls to export all that stuff into a SQL database. I have everything there mapped onto a threads and messages structure. And sometimes that took a little bit of squinting, but it basically seems to mostly work. And that's like the raw ground truth. And I do agree, for me it's I've found that it is quite important for the model to be able to get to raw ground truth.
然后我只做了,这只回溯五年,但通常足够了。超过五年的事情很少对今天有操作上的相关性。我按月导出。我发现对我个人来说,每个月大约有几十万 token,然后我会压缩成月度总结,大约是原来长度的 10%,对吧?所以,每月 20 到 30 万 token 的原始内容变成 2 到 3 万。然后我还会做年度总结。同样,又有 2 到 3 万变成 2 到 3 千。在所有这些之上,我让模型去建一个 wiki。
Then I've just done, and this only goes back five years, but that's usually plenty. It's not too often that something happened more than five years ago is like operationally relevant today. I just do a month-by-month export. I found that for me personally, it was like a couple hundred thousand tokens per month that I'll then compress into a monthly summary, which is like 10% the length, right? So, whatever, two to 300,000 tokens of raw monthly content goes to 20 to 30. Then I'll do that at the yearly level as well. So, again, you've got another 2 to 300 that goes down to a 20 to 30. And then on top of all that stuff, then I just have the model go make a wiki.
在获取真相的过程中,我发现的一个关键点是,模型当然可以直接发送 SQL 查询,但它并不总是清楚该查什么。所以我在总结中尝试做的是,让总结器捕捉独特的短词序列,就像大海捞针一样。比如,这里发生了什么等等,然后在结尾加一个小脚注,说明来源是某人的私信,这四个词能带你精确找到那条信息,而且可能只在那条信息里,在我所有的个人历史中。
And one of the key things that I found along the way in terms of how to get to ground truth, of course, the model can always just send a SQL query, but what should it be querying is not always obvious to it. So what I tried to do in my summaries is have the summarizer capture distinctive short sequences of words that would be like needle in a haystack find. So it would be like here's what happened blah blah blah and then a quick little footnote at the end of the source was a DM from whoever and these four words will take you exact to that and probably only that in like all of my personal history.
对我来说效果很好。我通常很满意,而且大多数时候我确实能找到正确的真相。但这只是一个人。多人使用的话,当我开始让我妻子加入时,事情至少会变得复杂一些。
I'd say it works pretty well for me. I'm usually pretty happy with it and I do feel to the right ground truth more often than not. It is only one person. Multiplayer is going to be as I start to onboard my wife and then that'll complicate things at least a bit.
你觉得我可能遗漏了什么,尤其是考虑到多人使用需要更复杂的方法?
What do you think I might be leaving on the table especially as you think about multiplayer that demands an even more sophisticated approach?
是的。我相信你提到的方法,我们的数据库里也有一些 RAG,我们也用它来做 RAG。这叫做“假设驱动检索”。不知道你听说过没有,它有点像反向的假设驱动检索。所以两者都有。假设驱动检索的做法是,比如问“贾斯汀·比伯出生在哪里”,你不是用 RAG 搜索这个问题,而是搜索你的假设,比如“贾斯汀·比伯出生在巴黎”、“贾斯汀·比伯出生在纽约”、“出生在柏林”,然后在你的 RAG 数据库里搜索这一组假设,因为显然在语义上,每个假设都比问题“贾斯汀·比伯出生在哪里”更接近你可能寻找的答案。这是第一点。
Yeah. So I believe the approach you're talking about and we still have some rag in our database and we use that for our rag as well. It's called like hypothesis-driven retrieval. I don't know if you've heard that it's sort of the reverse hypothesis-driven retrieval. So there's both. So what hypothesis-driven retrieval does is it generates like suppose it's asking like where was Justin Bieber born right and then you basically instead of searching for that question using rag you search for your hypothesis like Justin Bieber was born in Paris, Justin Bieber was born in New York, in Berlin, and you search for this bucket of hypotheses in your rag database because obviously semantically each of these hypotheses is a lot closer to the answer you might be looking for than to the question where was Justin Bieber born. So that's the first thing.
然后你还会做相反的事情,当你在 RAG 数据库中找到答案时,你会生成一系列映射到该答案的问题,并把它们附加到答案上。这样检索器既可以查找假设,也可以查找问题的答案,这往往会提高检索质量。这完全没问题。
Then what you do is you also do the opposite, which is when you find an answer in your rag base, you generate a bunch of questions that map to that answer and you attach them to the answer. So now the retriever can both look for hypotheses and can still look for its answer for the question, and that's going to tend to increase retrieval quality. That's totally fine.
我认为你可能遗漏的一点是,让一个智能体来管理所有记忆是非常健康的。尽管那个智能体既负责检索也负责更新记忆,而且记忆在文件系统中,你的智能体实际上可以直接去修改文件系统,但我们发现,实际上,它仍然可以这样做,但我们告诉智能体,我们提示它,如果你在寻找你没有的记忆,请询问记忆智能体。原因是因为当我们询问记忆智能体时,我们会记录查询、答案,以及检索它需要多少跳。然后当记忆智能体在打盹和做梦时,它会查看这个日志,比如,嘿,这是人们一直在问的东西。它也是一个图书管理员,你知道,就像,啊,我被反复问到,这是一个我经常遇到的问题,所以就在这里。顺便说一句,它起到了一种缓存的作用,因为自动地,你总是保留记忆智能体的记忆,比如最近一千个查询,比如问答对,所以它就像一个滚动窗口,所以它也是一种缓存。但记忆智能体不会仅仅依赖那个缓存,因为每个缓存最终都有 TTL,它会过期等等。所以记忆智能体还会做的是,它会重组自己的记忆,以减少检索最频繁问题所需的跳数,这往往能很好地提高记忆检索和记忆速度。
I think one thing you may be leaving on the table is it is really healthy for all of the memory to be managed by one agent. Even though that agent handles both retrieval and updating of the memory, and so even though the memory is in a file system and so your agent could actually just go mess with the file system, what we have found actually, and can still do that, but we tell the agent, we prompt the agent to be like, if you're looking for memory that you don't have, please ask the memory agent. And the reason why is because when we ask memory agents, then what we do is that we log the query, we log the answer, and we log how many hops it took to retrieve that. And then when the memory agent is napping and dreaming, it looks at this log of like, hey, this is the kind of stuff people have been asking. It's also a librarian, you know, it's like, ah, I've been asked left and right, this is a question that I get quite often, so right here. By the way, it acts as a sort of cache, you know, because automatically, so you always keep like the memory of the memory agent, like the last like thousand like query, like question-answer pairs, you know, so it's like a rolling window, and so it is a sort of cache. But also the memory agent is not going to just rely on that cache, because every cache at some point's got a TTL, it grows stale and all of that stuff. So what the memory agent will also do is like it will restructure its own memory to reduce the number of hops that is needed to retrieve like the most frequently asked questions, and that tends to work quite well to improve the memory retrieval and the memory speed.
酷。
Cool.
有意思。
Interesting.
另一个我们实验过但最终没有实施的系统,尽管从技术角度来看它确实有优势,就是,你听说过“穴居人”吗?
Another system we experimented with and we ended up not implementing it even though it does have strengths from a technical standpoint is, have you heard of caveman?
它非常简单。基本上就是把你的文本改写成穴居人风格。就像,“我,你,你知道,内森不喜欢卷饼”,你知道,如果你这样做,就能减少 20% 到 30% 的 token,这很好,而且减少 token 会让系统更快、更便宜、检索更准确。基本上不会损失信息或上下文,但你知道,我们没这么做的原因是我们确实关心系统的可维护性和可审计性。我们试过,所有指标都上升了,很棒。但文件系统里的文件看起来真的很蠢。在企业客户面前不好看,他们会问为什么我的记忆文件长这样?但它实际上效果类似,你知道,有很多这样的压缩上下文的技术。
It's really simple. It's basically you rewrite your text as a caveman. It's like, me, you know, you know, Nathan not like burrito, you know, like if you do that, it's so, but if you do that, you cut your tokens by like 20 or 30%, which is good, you know, and so cutting your tokens like it'll make your system faster, it'll make your system cheaper, it'll make retrieval more accurate. It'll just be better with basically no loss of information or context, you know, if you, the reason we've not done that is because we actually do care about the maintainability, like the auditability of the system. And so we did that, every metric went up, it was awesome. And then the files in the file system look really dumb. It doesn't make us look good, you know, to enterprise customers, like why do my memory files look like that? But it works actually similar, you know, like there's like a bunch of those techniques to just compressive context.
你听说过“a tune”吗?它是一种为智能体设计的类似 JSON 的替代品,我认为它比 JSON 节省 20% 到 30% 的 token。JSON 实际上并不是那么节省 token。所以 tune 效果更好。是 20% 到 30%。
Have you heard of a tune? It's a sort of like JSON alternative that's made for agents that's like I think it's like 20 or 30% more token efficient than JSON. JSON is actually not that token efficient. So tune works better. It's 20 or 30.
这项工作对记忆智能体来说不那么重要,但我建议每个人都让他们的智能体使用 Tune 而不是 JSON,并在智能体与其动作之间设置一个 Tune 传递中间件。任何动作都不应向智能体暴露 JSON,每个动作都应暴露 Tune。
This work is less important for memory agents, but I recommend everyone make their agent use Tune instead of JSON, and have a Tune-passing middleware between your agent and its actions. No action should expose JSON to the agent. Every action should expose Tune.
嗯,有意思。我得去看看这个 Tune 格式。我一直用 YAML,它也是出于类似的优势。
Yeah, interesting. I'll have to check out this Tune format. I've just used YAML, which is also motivated by similar advantages.
是的,Tune 是面向 token 的对象表示法,它确实提高了智能体的性能。这很令人惊讶,因为模型的训练集中有大量 JSON,但显然它们对 Tune 比 JSON 更适应。
Yes, Tune is token-oriented object notation, and it actually increases the performance of the agents. It was surprising because there's so much JSON in the training set of the models, but they're apparently more comfortable with Tune than with JSON.
JSON 确实有很多噪音,这是肯定的。是的,我比 JSON 更适应 YAML。这有点意思。AI 就像我们一样。
JSON does have a lot of noise associated with it, for sure. Yeah, I'm more comfortable with YAML than with JSON. That's something. The AIs, they're just like us.
顺便说一句,这是我最大的惊讶之一。我实际上在做一个项目,不知道会做成播客还是帖子什么的,就是盘点我在预测方面曾经错过的和曾经对过的事情。我想我以前经常自信地说的一件事,可能现在最站不住脚的就是我们不应该将模型拟人化,否则会误入歧途。我一次又一次地感到震惊,因为将模型拟人化实际上可以多么富有成效。这仍然感觉有点危险,但很难与结果争辩。
That's one of my biggest surprises as a brief aside. I'm actually working on, I don't know if this will be a podcast or a thread or whatever, but just taking stock of things I've been wrong about prediction-wise and some things I've been right about. I think one of the things I used to say quite confidently that has aged maybe the poorest is that we shouldn't be anthropomorphizing models, or that will lead us astray. I've been shocked over and over again by how productive it can actually be to anthropomorphize the models. It still feels a little dangerous, but it's hard to argue with the results.
我 100% 同意。不过我认为这感觉像是拟人化的一个很好的折中点。我认为这也是我现在非常关注的一个问题:拟人化在多大程度上不再成立,尤其是当涉及到 AI 员工时?例如,我改变想法的方式,我没有完全改变,但相对而言我改变了一点,是关于多智能体系统。我实际上开始相信,大多数时候,尽可能多地,你希望尽可能多地整合到一个单一智能体下,而不是多智能体。这并不总是可能的,而且仍然有很好的理由使用多智能体,但大多数情况下是这样。我实际上认为人类直觉上会偏向于非常多的多智能体系统,远远超过最优水平,因为他们有点过度拟人化,并与人类组织进行比较。就像,“我有一个数据科学家,我有一个工程师,我有一个设计师,我有一个产品经理,所以我要为这些角色各创建一个智能体。”实际上,人类组织这样做的原因是每个人每天只有 24 小时,只能容纳那么多上下文。智能体显然没有这些限制。它们可以分叉,可以复制,可以整天做任何想做的事。所以我觉得分工不是拥有多个智能体的好理由,这是一个主要因素。仍然有其他的原因,但分工不是其中之一。所以我同意,我认为总的来说人们将模型拟人化是对的,因为它们毕竟是在人类 token 和人类强化学习等基础上训练的,但我也认为 AI 组织不应该被过度拟人化。
I 100% agree. I do think though that feels like a happy medium of anthropomorphization. And I think that's another thing that's very top of mind for me right now: to what extent does anthropomorphization stop being true, in particular when it comes to AI employees? For example, the way I've changed my mind, and I haven't changed my mind completely, but I've changed my mind a little bit relatively speaking, is about multi-agent systems. I actually have come to believe that most of the time, as much as possible, you want to consolidate as much as possible under a single agent, not multi-agent. That's not always possible, and there are still good reasons to use multi-agent, but most of the time that's the case. I actually think that humans intuitively find themselves biased towards very multi-agent systems, far more multi-agent than is optimal, because they're anthropomorphizing a little bit too much and comparing to human organizations. It's like, "I have a data scientist, I have an engineer, I have a designer, I have a PM, so I'm going to create one agent for each of these things." Actually, the reason why human organizations do that is because every human has only 24 hours a day and can only contain that much context. Agents obviously have none of those constraints. They can fork, they can duplicate, they can do as much as they want all day. So I find that division of labor is not a good reason to have multiple agents, which is a major factor. There are still different reasons, but division of labor is not one of them. So I agree, I think by and large people are right to anthropomorphize models because they were after all trained on human tokens and with human RL and all of that stuff, but I also think that AI organizations shouldn't be overly anthropomorphized.
嗯,这很有趣。当你谈到单一智能体时,显然你可以并行运行多个单一智能体的副本。你是否已经达到了开始产生数据库式问题的地步?如果多个智能体同时更新同一段历史,传统数据库中有所有这些事务级保证。当你有 10 个 Lindy 队友在同一个文件系统语料库上运行时,这开始成为问题了吗?
Yeah, that's interesting. As you talk about a single agent, obviously you can run multiple copies of a single agent in parallel. Have you reached the point yet where that starts to create database-style problems? If multiple agents are updating the same segment of history at the same time, we have all these transaction-level guarantees in traditional databases for that reason. Does that start to be an issue as you have 10 Lindy teammates running over the same kind of file system corpus?
是的,比你想的要少。我认为人类组织解决这个问题的方式是通过 Git。Git 是迄今为止大量人类找到的在同一件事上协作的最佳方式,因为你可以合并冲突、变基、做一堆这类事情。所以这就是我们发现的。Lindy 中的文件系统,出于这个原因,由 Git 仓库支持,这给你带来了大量惊人的特性,包括合并管理。记忆智能体有时会自我分叉;它基本上会启动一堆子智能体,比如,“好吧,让我们看看自上次休眠以来发生了什么。”在一些大型组织中,自上次快照以来产生了太多上下文,以至于它需要创建一堆子智能体。在初始水合期间,我们真的想快速推进,因为我们想给用户留下深刻印象;在 PLG 中,惊艳时间非常重要。我给你看的那张图,我们优化得非常厉害,以至于在注册后 10 秒内就能完成。所以就像注册、Notion、Slack,砰。这做了很多工作。一种方法是通过智能体集合。这些智能体集合往往会互相踩脚。答案是 Git,顺便说一下,它还免费提供了历史记录,这很棒。而且这不是赞助消息;他们确实做得很好。我们一直在与 Mesa 合作。他们销售一个由 Git 支持的智能体原生文件系统,所以你可以创建所有这些仓库。最初我们开始的时候,顺便说一下,这是你在构建这些系统时必须建立的一种直觉。你很容易想,“我就手搓一个,我就创建自己的 Git,我就管理自己的文件系统基础设施。”你会了解到这些系统背后比你最初意识到的要深得多。所以现在我是,这很有趣,因为它有点与当今的主流叙事相悖,即 SaaS 正在消亡,你可以 v-code 一切。我们会 v-code 很多东西,但基础设施不会被 v-code。我现在变得非常注重基础设施构建。任何时候我有机会买而不是构建,我都会买,因为你节省了大量时间,而且这些产品中有比你意识到的更多的深度、思考和设计决策。
Yes, less than you would expect. I think the way human organizations have solved that is via Git. Git is by far the best way large groups of humans have found to work together on the same thing, because you can merge conflicts, you can rebase, you can do a bunch of that stuff. So that's what we found. The file system in Lindy is, for this reason, backed by a Git repository, which gives you a ton of amazing properties, including merge management. The memory agent sometimes will fork itself; it basically spins up a bunch of sub-agents to be like, "All right, let's see what happens since the last time I napped." In some large organizations, there's so much context created since the last snap that it needs to create a bunch of sub-agents. During the initial hydration, we really wanted to go fast because we wanted to impress the user; time to wow is so important in PLG. That graph I showed you, we've optimized it so hard so it happens within 10 seconds of sign up. So it's like sign up, Notion, Slack, boom. It's been a lot of work. One way you do that is via a collection of agents. These collections of agents tend to step on each other's toes. The answer to that is Git, which by the way also gives you history for free, which is awesome. And this is not a sponsored message; they've just been doing a really good job. We've been working with Mesa. They sell an agent-native file system that's backed by Git, so you create all of those repositories. Initially we started, by the way, it's one of those intuitions you have to build as you build these systems. It's very tempting to be like, "I'll just hand roll it, I'll just create my own Git, I'll just manage my own file system infrastructure." You learn that there is so much more depth behind those systems than you first appreciate. So now I am, it's funny because it sort of goes counter to the prevailing narrative these days which is SaaS is dying, you can v-code everything. We'll be v-coding a lot of stuff, but infrastructure will not be v-coded. I've become so infra-build now. Anytime I have an opportunity to buy instead of build, I will do it, because you save so much time and there is so much more depth and thinking and design decisions that go into these products than you may appreciate.
我喜欢好的供应商推荐。还有其他供应商推荐吗?你们在采购哪些关键基础组件,值得推荐?
I love a good vendor shout out. Any other vendor shoutouts? Any other key primitives that you're buying that you would recommend?
嗯,你需要一个沙盒。所以我们选了 E2B。和很多人一样,一旦有了沙盒,就会觉得‘太好了,我有文件系统了’。但至少在大规模场景下,我们强烈建议把两者解耦,原因我刚才提过。另外,Browserbase 显然在浏览器管理方面非常出色。我们还喜欢用谁呢?我们确实深入钻研过。关于他们说的基础设施构建,有一个重要的例外,就是可观测性和评估。
Well, you need a sandbox. So, we've gone with E2B for that. Again, like many people, once they have a sandbox, they're like, 'Sweet, I have a file system.' But at least at scale, we strongly recommend decoupling them for the reasons I just mentioned. And more, Browserbase obviously is excellent for browser management. Who else do we like and use? You know, we went down a deep rabbit hole. One major exception to what they said about being infra build has been observation and evaluations.
我们的历史比较复杂,因为这家公司是 2022 年创立的,事后看来太早了。智能体还没准备好,当时没人谈论智能体。人们谈论的是生成式 AI,如果你还记得,就像前婴儿 AGI 和那些早期的东西。所以智能体其实并不好用。因此,没有现成的工具。不幸的是,我们不得不自己构建很多工具,包括我们自己的评估平台和可观测性平台。说实话,我不建议自己动手做。用现成的方案更容易。
We have a complicated history because we started this company in 2022, and it was, in hindsight, way too early. Agents were not ready. No one was talking about agents back then. People were talking about generative AI, if you remember, like pre-baby AGI and all those early things. So agents didn't really work. As a result, there was no tooling. So unfortunately, we had to build a lot of our own tooling, including our own eval platform and our own observability platform. I don't recommend doing it at home, honestly. It's just easier to use existing solutions.
我们做了全面的评估。我们研究了 LangSmith、BrainTrust、Agnost(我是它的自豪投资人)。我们研究了很多这类解决方案,每次我们都觉得‘我想我们起步太早了,而且目前我们自研的方案完全是为我们量身定做的’。我们经历了构建它的痛苦。所以如果当时有现成的方案,我会选择它,但没有,所以我们自己建了。这些年我们不断完善它。现在我们有一个工程师全职负责这个方案,这在 voiding 时代算是很多了。这家伙很拼,它已经变成了一个非常非常全面的方案。说实话,我们和市面上这些方案已经功能持平,甚至更多。我们有很多内部功能,我没看到他们有,而且它完全是为我们自己的内部脚手架和使用而构建的,所以我们最终就在内部自己手搓了。
We've gone through a whole evaluation. We've looked into LangSmith, into BrainTrust, into Agnost (which I'm a proud investor in). We've looked into a bunch of those solutions, and every time we were like, 'I think we just had such a head start, and at this point our homegrown solution is so built exactly for us.' We've gone through the pain of building it. So if a solution had been available back then, I would have picked it, but it wasn't, so we built it. We fleshed it out over the years. Now we have one engineer who's full-time dedicated to that solution, which is a lot in the voiding era. This guy is cranking, and it's become a very, very extensive solution. Truth be told, we are at feature parity with all of these solutions out there and more. We have a lot of internal features that I don't see these guys have, and it's so built for our own internal scaffold and use that we ended up just hand-rolling that internally.
是的。那么,现在 Lindy 内部的生活是什么样的?我觉得让我印象深刻的是,你们现在——我确信在产品历史上一直有吃自己的狗粮的过程,但现在你们已经到了一个点,吃自己的狗粮没有限制了,对吧?你可以做任何事。这在哪些方面体现出来?我不知道,你可以从很多角度来谈,对吧?有没有一些角色,你本来会招人,但现在不会招了?你们自己的推理支出与工资单的比例是多少?你从自己的视角来看,但这个过程到达人类或 AI 员工阶段,如何改变了在公司工作的体验?
Yeah. So, what does life look like at Lindy internally these days? I think it strikes me that you're now—I'm sure there's been a dogfooding process throughout the history of the product, but now you're at this point where it's like there's no limits to the dogfooding, right? You can do anything. How does that play out in terms of—I don't know, you could cut it any number of ways, right? Are there roles that you might otherwise hire for that you just wouldn't hire for now? What's the ratio of your own inference spend for the purposes of building Lindy compared to your payroll? You put your own lenses on it, but how has the arrival at the human or at the AI employee stage of this process changed what it's like to be a part of the company?
这是个好问题。我认为现在,显然 AGI 已经到来,所以每家公司都在争先恐后,而且这个转型期正在加速。我得说,我有点悲观。我担心 AI 风险,但到目前为止一切顺利,而且到目前为止非常有趣。说实话,经历 AGI 和构建 AGI 真的很有趣。我们可以快得多。我觉得我们现在以‘启动’的速度前进。就像没有障碍一样。通常公司,你考虑战略,需要很长时间来弄清楚、实施、看到结果。需要几个月。那个 OODA 循环长达几个月。显然,初创公司的关键就是尽可能快地行动,因为这是对抗现有企业的唯一优势。现在我们可以有想法,两小时后就能在产品中看到它们上线。而且是大想法,不是小想法。
That's a great question. I think right now, obviously AGI is here, so every company is sort of scrambling, and there is that time of transformation that's accelerating. I will say I'm a little bit of a doomer. I'm worried about AI risk, but so far so good, and so far it's so much fun. It's honestly so much fun to go through AGI and to build AGI. It's like we can move so much faster. I feel like we're moving at the speed of start now. It's like there is no obstacle. Normally companies, you think about the strategy, it takes a long time to figure it out, to implement it, to see the results. It takes months. That OODA loop is months long. Obviously the name of the game at a startup is to move as fast as humanly possible because that's the only advantage against the incumbents. Now we can just have ideas and see them live in the product two hours later. And big ideas too, not small ideas.
我可以告诉你,我们的 Slack 现在基本上只有我们和 Lindy。Lindy 占了 Slack 消息的一半。就是 Lindy 和我们来回对话。很多时候,我们自己都低估了这个平台。就在上周,我们的 CI 管道出了问题。我们在 GitHub 上真的有几百个 runner 来管理我们的 CI 管道。顺便说一句,这也是回答你问题的另一件事。在 Lindy 工作是什么感觉?在任何地方工作是什么感觉?坦率地说,技术栈的每一部分和流程的每一部分现在都被拉伸到了极限。就是这样,你知道。我相信你见过 GitHub 发布的每日提交数量图表。GitHub 可不是小公司。它非常大。现在它正在垂直上升。这要飞向月球了。这反映了每家公司正在经历的情况。
I can tell you, our Slack now is basically just us and Lindy. Lindy is like half the messages on Slack. It's just Lindy and us talking back and forth. Very often, we ourselves underestimate the platform. Just last week, we had a problem with our CI pipeline. We have literally many hundreds of runners on GitHub that manage our CI pipeline. By the way, that's another thing right now to answer your question. What's it like to work at Lindy? What's it like to work anywhere? Frankly, every part of the stack and every part of your process is stretched to its very limits right now. That's just right back, you know. I'm sure you've seen the graph GitHub posted of the number of commits posted per day. GitHub is hardly a subscale company. It's really big. It's just going vertical right now. This is going to the moon. That's indicative of what every company is going through.
我们内部的 PR 数量——每周 PR 数——在过去三个月里翻了三倍,每个 PR 的行数也翻了三倍。我们几乎不再审查 PR 了。只是智能体在审查 PR。就我们审查而言,基本上团队的说法是,我们不再是 PR 审查者。我们更像是 PR 审查者的审查者。我们只是在审查审查 PR 的机器。因此,你会说‘这太棒了,但是嘿,现在 CI 夹在中间了。那怎么办?’好吧,现在你得处理 CI。所以我们花了很长时间处理 CI。对于非技术听众,CI 是持续集成。它是这样一个过程:获取你的代码更改,确保它们安全,然后实际合并到你的主产品并部署。我们一开始像很多人一样,向问题砸钱。所以现在我们的 CI 极其昂贵。这是一笔令人气愤的开支。我们已经用尽了所有显而易见的方法——我们显然不在 GitHub runner 上运行。我们在 GCP runner 上运行,而且我们已经做了构建缓存。我们已经做了所有显而易见的事情。
Our number of PRs internally—PRs per week—has tripled over the last three months, and the number of lines per PR has tripled over the last three months. We hardly review PRs anymore. It's just agents reviewing PRs. Insofar as we review, basically the way the team is talking about is that we're no longer PR reviewers. We're like PR reviewer reviewers. We're just reviewing the machine that reviews the PRs. As a result, you're like, 'This is awesome, but hey, now CI is in the middle of that. So what do you do?' Well, now you got to work on CI. So we've spent a long time working on CI. For the nontechnical audience, CI is continuous integration. It's the process involved in taking your code changes, making sure they're safe, and actually merging them into your main product and deploying them. We at first started throwing money at the problem, like many people do. So now our CI is extremely expensive. It's an insulting expense. We've exhausted all the obvious things—we're obviously not running on GitHub runners. We're running on GCP runners, and we've done caching of the builds. We've done all the obvious things.
管理 CI 让它高效运转是件很费劲的事,我受够了,因为那几周我们一直在折腾 CI,还烧了不少钱。我们就在想,CI 这事到底该怎么办?然后我意识到,等等,我们为什么要自己干这些?难道不能直接启动一个智能体来做吗?而且那个智能体确实……顺便说一句,这也是为什么单一智能体很重要的另一个原因,因为现在设置成本成了最大的开销之一。组织目前最大的瓶颈就是人的时间和注意力。所以你要尽可能把人的因素去掉。你要能够配置智能体,给它们所有需要的权限,尽量不需要人在中间干预。
It's a lot of work to manage CI to make it efficient, and I was tired of this because it was weeks of messing with CI and building a hole in our pocket. We were like, what are we going to do about the CI thing? Then I was like, wait a minute, why are we doing all of this ourselves? Can't we just spin up an agent to do it? And the agent was literally... and by the way, that's the other reason why it's important to have a single agent, because setup is now becoming one of the biggest costs. The biggest bottleneck in the organization right now is human time and human attention. So you want to remove that as much as possible. You want to be able to provision agents and give them all the accesses they need without a human in the loop as much as possible.
好的。那么这个智能体是在公司里吗?它是一个 AI 员工,了解公司的一切,Lindy 队友,对吧?它能访问代码仓库,有自己的电脑。所以就像,Lindy,你能帮我们弄一下 CI 吗?然后它就说,好的,让我用 GCP CLI 看看那些 runner,再看看 GitHub Action 之类的。所以它做了很多分析,然后回来告诉我们,现在它基本上就是一个自动优化工具。顺便说一句,Lindy 每天都会用图像生成给我们发一张图。那张图特别漂亮,然后它说,嘿,各位,看,CI 成本在下降,合并时间也在缩短,这是好事。我当时就想,真不敢相信我们花了两个星期自己折腾 CI,而不是直接把这个交给 Lindy。
Okay. And so is this agent at the company? It is this AI employee that knows everything there is to know about the company, Lindy teammate, right? It's got access to repository, it's got its own computer. So like, Lindy, can you please mess with our CI? And it's like, yeah, let me look at the runners using the GCP CLI, let me look at the GitHub action and all of that stuff. So it's doing a lot of analysis, getting back to us, and now it's literally like an auto-optimization thing. Lindy is just sending us a graph every day, by the way, using image gen to generate. So the graph is really pretty, and it's like, hey, guys, look, CI cost is going down, time to merge is going down, which is a good thing. And I was like, I can't believe it took us two weeks of just messing with CI ourselves instead of just sending this to Lindy.
我觉得现在的情况就是这样:工程师越来越少直接做事情,而是越来越多地在做那个“做事情的东西”。我们越来越多的工作是搭建这台机器,它越来越能自我运转,而且会越来越如此。我喜欢这样想:首先你成为 AI 的管理者,然后你成为管理者的管理者,因为在某个时候,管理者也会变成 AI。我希望在某个时候我们能成为董事会成员。我们可以去夏威夷的海滩,除了审视公司战略之外没什么可做的。我们会有智能体的报告告诉我们发生了什么、格局如何、我们取得了什么成果,然后我们就能更深入地讨论我们正在构建的公司的本质,以及我们在市场中的位置。对我来说这是个梦想,而且一点也不遥远。
I think this is what it's looking like right now: engineers are less and less working on the thing, they are more and more working on the thing that is working on the thing. And more and more what we do is we're setting up this machine which increasingly operates itself, and increasingly will more and more. I like to think of it as first you become a manager of AIs, then you become like a director of managers, because at some point the managers are going to become AIs as well. I'm hopeful that at some point we become board members. We can just go to a beach in Hawaii and we don't have much to do except review the strategy of the companies. We'll have reports of the agents telling us this is what's going on, this is the landscape, these are the results we're having, and we'll be able to have much deeper conversations about the nature of the company we're building and the spot we occupy in the market. That's a dream as far as I'm concerned, and that's not very far away at all.
回答你的问题,有没有我们没招的职位?有,绝对有。实际上团队人数已经平稳了一段时间,而且我们在几个月内把生产力提高了两倍。我发现这里有一个大胆的豪华区域,因为人类仍然在循环中。人类仍然需要,但显然你加入的人越多,公司内部的协调成本就越高。所以我发现实际上现在相对而言,小团队在很多方面比大团队更有优势。
To answer your question, have there been positions we've not hired? Like, yes, absolutely. The team is actually, I think headcount has been flat now for a while, and we've actually tripled our productivity over a couple of months. And I find there is a bold deluxe zone because humans are still in the loop. Humans are still needed, but obviously the more humans you add, the more coordination cost you introduce inside the company. And so I find that actually right now smaller teams, relatively speaking, are at an advantage versus bigger teams in a lot of ways.
你肯定有,我不确定你是否愿意分享,但你有没有一个比例,我不是说包括客户使用场景在内的所有推理支出,而是你们自己的内部推理与工资单的对比。那个比例是多少?
Do you, I'm sure you do. I don't know if you want to share it, but do you have a ratio of, I don't mean all inference spend including your customers' use cases, but like your own internal inference versus your payroll. How is that ratio?
是的。如果你看包括所有客户在内的全部推理支出,那是工资单的好几倍,而且已经持续很长时间了,但这是一个例外,因为那是我们卖的东西。现在如果你看工资单与内部推理支出的对比,已经接近了。工资单仍然更高,但这两条线将在三到六个月后交叉。
Yeah. So if you look at all the inference spend including all customers, it's many times payroll, and it's been for a very long time, but that's one exception because that's what we sell. Now if you look at payroll versus internal inference spend, it's in shooting range. So payroll is still greater, but the lines are going to cross three to six months from now.
如果你想象一个假设的场景,突然公司里只剩下你一个人,然后往下全是 Lindy。那什么会崩溃?什么不再管用了?换句话说,人类仍然不可替代的是什么?
If you imagine a hypothetical scenario where all of a sudden it's just you at the company and after that it's Lindy's all the way down. What breaks? What doesn't work anymore? In other words, what is it that the humans are still irreplaceable for?
这个问题真的很难回答,因为我经常思考这个问题,而且我发现自己越来越缺乏词汇来回答“智能体在哪里会失败”这个问题。我们过去常谈论时间,比如,啊,在它失去连贯性之前的时间,它可以像米那样,可以是 1 小时、10 小时、100 小时,我相信大家也会有同感。我不再认为那是一个好的衡量标准。现在大家用的词是“尖刺”,对吧?这些模型也非常尖刺。当我们经历 AGI 时,AGI 这个词已经失去了意义,因为它在很多方面实际上是 ASI,而且它在一些非常令人惊讶和愚蠢的方面低于人类。它在很多方面真的很蠢。我就想,嘿,老兄,你可以一次性给我生成 5 万行代码,你可以独自在 3 小时内给我构建出最不可思议的想法,然后你却告诉我要走到洗车店去。所以那就是它失败的时候。想象你有一台机器,它运营公司非常聪明,但每天有 50 次它决定去洗车店。所以我认为现在这些新组织的挑战,在可预见的未来将是人机混合体,你基本上是在打造一套钢铁侠战衣。有一个滑块,比如,这部分给机器,这部分给人。而知道把这个滑块放在哪里,因为它不是一维的。它是一条穿过高维空间的复杂曲线。所以你要寻找那个可以在空间中划出的切片,因为你想要充分利用人类的注意力。你不想让 AI 为了不需要你的事情来打扰你。你需要 AI 知道什么时候需要打扰你。而几乎从定义上讲,它做不到,因为如果它知道,它就不会打扰你,它就不会犯那个错误。
It's really hard for me to answer this question because I think about that a lot, and I am finding myself increasingly lacking the vocabulary to answer this question of where do agents fail. We used to talk about time, like, ah, time before it loses coherence, it can go like the meter thing, it can go like 1 hour, 10 hours, 100 hours, and I'm sure people will resonate with this as well. I don't find that's a good measure anymore. The world that everyone is using these days is spiky, right? These models are also very spiky. As we're going through AGI, the word AGI has stopped meaning because it's actually ASI in many ways, and it's subhuman in some very surprising and dumb ways. It's really dumb in a lot of ways. I'm like, hey man, you can produce for me 50,000 lines of code one shot, you can build me the most incredible ideas in like 3 hours by yourself, and then you're telling me to walk to the car wash. And so that's when it fails. Imagine you've got this machine that runs a company super smart, but 50 times a day it decides to walk to the car wash. And so I think right now the challenge of these new organizations that are going to be for the foreseeable future human-AI hybrids is going to be you're basically building an Iron Man suit. There is this slider where it's like, okay, this is for the machine, this is for the human. And knowing where to put this slider, because it's not a one-dimensional thing. It's a complex line through a many-dimensional space. So you're looking for that slice that you can draw through the space, because you want to make the most out of humans' attention. You don't want the AI to bug you for stuff it doesn't need you for. And you need the AI to know when it needs to bug you. And almost definitionally it can't, because if it knew, it wouldn't bug you. It wouldn't make the mistake.
所以抱歉,我觉得我是在用问题回答问题,但这就是那个正确的问题。我认为目前这是行业里的一个开放性问题:这条线怎么划?怎么让智能体意识到自己不确定,甚至当它说出像“请把你的车走到洗车店,因为今天很堵”这样的话时?我觉得答案就是:人类和 AI 必须以某种方式监督整个系统,抓住它犯傻的时刻。
And so I'm sorry I feel like I'm answering this question with a question, but that is the right question. And I think right now that's one of the open questions in the industry: how do you draw that line? How do you get the agent to realize when it's unsure, and when it's even saying something like 'please walk your car to the car wash because it's study today'? I think that's the answer: it's just that somehow humans and AI, we're going to have to supervise this whole system and catch the times when it's being very dumb.
那当某个东西非常聪明的时候呢?现在最好的想法都来自哪里?
How about the times where something is being very smart? Where do the best ideas come from these days?
我觉得长期以来有一种观点,认为 AI 擅长处理常规事务。我以前常说——显然这已经不再适用了——但我以前常说,如果你愿意下功夫,你公司里任何常规工作都可能自动化,但别指望有“灵光一现”的时刻。现在,显然我们在时间线上不断看到“灵光一现”,至少在特定领域,比如开放数学猜想之类的。
I think there's long been this idea that AIs are good at routine stuff. I used to say—and clearly this is no longer applicable—but I used to say you can probably automate any routine work that you have at your company if you're willing to put in the elbow grease, but don't expect Eureka moments. Now, obviously, we're getting Eureka moments on the timeline all the time, at least in certain domains, like the open math conjectures and stuff.
说到 Lindy 最好的想法,现在有多少来自人类,多少来自 AI?
When it comes to the best ideas at Lindy, how many of them are human origin and how many of them are AI origin today?
我讨厌这个答案,但答案是两者都有。它真正来自两者的结合。我讨厌它的原因是它引出了半人马的神话,你知道,就是那种半人半马的生物。所以这个想法是:是的,AI 比人类强,但比 AI 更强的是 AI 加人类。因此,人类总是会被需要的。那是幻想,根本不是真的。文献对此其实很清楚:我们在国际象棋和所有其他 AI 达到超人水平的游戏中都看到了。一开始,AI 打败人类,AI 加人类打败纯 AI。但渐渐地,AI 加人类与纯 AI 之间的差距在缩小,直到变成负的,人类至多是在给系统引入随机噪声。所以人类在某个时刻会变得完全无知,开始损害他们介入的系统。话虽如此,半人马阶段确实会存在一段时间。所以开放问题是它会存在多久?但现在我们正处于半人马阶段,我认为这也是多人模式如此重要的另一个原因。你希望你的智能体和你同处一室,你不想让任何随机的智能体和你同处一室。你想要一个真正的 AI 员工,它和你在这个房间里待了很长时间,参加过每一次会议,建立了关于公司一切事务的内部知识库。你可以 @ 它,邀请它加入你在公司的对话,我们一直这么做。
I hate the answer, but it's both. It truly comes from the union of both. And the reason I hate it is because it ladders into this myth of the centaur, you know, this mythical man-horse creature. So it's this idea that yes, an AI is better than the human, but what's even better than AI is AI plus human. Hence, humans are always going to be needed. And that's a fantasy. That's just not true. And the literature is actually clear on this: we've seen it happen with chess and with every other game where AI has achieved superhuman performance. At first, AI beats human, and AI plus human beats just AI. But little by little, the gap of AI plus human versus AI is shrinking until it actually turns negative, and humans are introducing at best random noise into the system. So humans at some point become truly clueless and start harming the system in which they're intervening. That said, the centaur phase exists for a while. So the open question is how long is it going to exist? But right now we are in the centaur phase, and I think that's the other reason why multiplayer is so important. You want your agents to be in the room with you, and you don't want any random agent to be in the room with you. You want an actual AI employee who's been in this room with you for a long time, who's sat in every meeting, who's built up that internal knowledge base about everything that's going on in the company. You can @mention it and invite it to chime in into your conversation in the company, and we do that all the time.
我们确实这么做。这很有趣,因为在 Slack 上很容易做到,就像“@Lindy,你怎么看?”但在会议上很难做到,以至于我们现在更倾向于用 Slack,因为 AI 就在那里。而在会议上,它在听,但被称为“插话”。所以现在很多时候,我们在会议中会说:“实际上,我们在这个团队上遇到困难。Lindy,你能在会后给我们发条消息,告诉我们你的想法吗?”然后它会在 Slack 上给我们发消息。当你这样做时,Lindy 很少会第一枪就告诉你一些让你觉得“就是它,这就是我们应该做的事”的东西。但几乎总是我们和 Lindy 来回讨论,从这种来回中,它会遗漏一些东西。比如 CI 的例子就很好。起初,当我们邀请 Lindy 加入这个对话时,它发生在“等等,我们为什么不和 Lindy 一起做这个?Lindy,你能优化一下 LCI 吗?”这样的对话中。
We do it actually. It's funny because it's very easy to do it on Slack, just like '@Lindy, what do you think?' It's hard to do in the meeting in a way that now makes us actually prefer Slack, because the AI is there. Whereas in the meeting it's listening but it's called 'chime in'. So very often now, what we do during meetings is we say, 'Actually, we're struggling with this team. Lindy, can you please send us a message after this telling us what you think?' And it'll send us a message on Slack. And when you do that, it is rarely the case that Lindy will first shot tell you something and you're like, 'This is it, this is the thing we should do.' But it is almost always the case that we go back and forth with Lindy, and from that back and forth, it misses something. Like the CI example is a great one. At first, when we invited Lindy into this conversation, it happened over like, 'Wait a minute, why are we not doing this with Lindy? Lindy, can you please optimize LCI?'
而且它没有犯错,是的。
And it made no mistake, yeah.
是的。我忘了它的第一个建议,但如果你不了解约束条件,那是个合理的建议。但鉴于我们运营的约束,它并不好。所以工程师们插话说:“不,实际上我们不能那样做。”我们看到后说:“好吧,这是个合理的观点。”所以我们来回讨论,然后一起想出了:“哦,对,我们完全应该那样做,就请继续做吧。”
Yes. And I forgot its first suggestion, but it was a reasonable suggestion if you did not understand the constraints. And it was not good, given the constraints we're operating under. So the engineers chimed in like, 'No, actually we can't do that.' We saw that, like, 'Okay, that's a fair point.' So we went back and forth, and then together we came up with like, 'Oh yeah, we should totally do that and just go ahead and do it, please.'
我有一个持有一段时间的论点——现在说可能还为时过早,但你肯定会有看法——再次回到我错的地方。我原本预期经济中的转型会比我们看到的更多。如果你在 2022 年给我看 Fable,并说“这个模型出来时失业率会是多少?”我会说:“我不知道,但肯定比我们现在经历的要高。”
A thesis that I've had for a while—it might be too early to say, but you're definitely going to have a perspective on it—is, again going back to what I've been wrong about. I expected a lot more transformation than we've seen in the economy. If you had showed me Fable in 2022 and said, 'What's the unemployment rate when this model is out?' I would have said, 'I don't know, but something definitely higher than what we are currently experiencing.'
是的。
Yeah.
那么接下来——这是我在那段时间里一直纠结的问题——为什么我们没看到更多变化?我想到的一个答案是人们固守旧习,不是每个人都像我这样热衷,这需要时间。但还有,当我们看到知识工作者数量下降时,那将是突然翻转的时刻,对吧?因为现在我将有更多对等的选择:我可以试着雇一个人类,或者我可以接入一个 AI,这在很多方面会更容易,对吧?我不需要经历整个面试过程。我大致知道我得到的是什么。等等等等。我不需要告诉你为什么 AI 员工是一个好的形态。你认为我们开始看到这种翻转了吗?它会是现有企业能够及时采用以避免被颠覆的东西,还是仍然会有我没想到的瓶颈或阻力点?
So then what—and this has been something I've been wrestling with for the intervening time—like why are we not seeing more? And one answer I've come to is people are stuck in their ways and not everybody's an enthusiast like I am, and it's going to take time. But then also, when we get to the drop in knowledge worker, that'll be the time when it'll suddenly flip, right? Because now I'll have a more parity kind of choice between I could try to go hire a human or I could plug in an AI, and that's going to be easier in a lot of ways, right? I don't have to go through a whole interview process. I kind of know what I'm getting. Blah blah blah. I don't need to tell you why an AI employee is a good form factor. Do you think we are starting to see this flip, and is it going to be something that incumbents will be able to adopt in time to avoid getting disrupted, or are there still going to be bottlenecks or resistance points that I'm not anticipating?
我也预期会有更多颠覆。我认为只要模型是不稳定的,我们就需要很多人类在场,因为模型中的漏洞非常危险。不仅仅是 AI 存在性危险。它们就是有害的。它们对系统真的有害。所以,无论人类有多贵,他们都会通过填补这些漏洞来赚取他们的价值。
I expected more disruption as well. I think as long as the models are spiky, we are going to need a lot of humans in the room because the holes in the model are very dangerous. Not just like AI existential dangerous. They're just harmful. They're really harmful to the system. And so, however expensive humans are, they're going to earn their keep by plugging these holes.
关于现有企业能否采用这一点,我认为只要它是尖峰式的,现有企业就会处于劣势。如果你有一个真正无缝的远程员工,没有漏洞,那么即使是现有企业也能很快采用,因为这就像雇佣一名员工,他们有相应的流程。而且我实际上认为他们会错过很多创新,因为这实际上是拟物化的。你基本上就是让 AGI 坐在人类的位子上,我不认为那是构建 AGI 组织的最佳方式,但它会奏效。较小的组织将处于优势,既因为他们能够在短期内摸索出这种 AI 原生的组织形态,让人类来填补漏洞,长期来看,我认为他们会构建一个不那么拟物化的 AI 组织。
Now regarding whether incumbents will be able to adopt this, I think as long as it is spiky, the incumbents are going to be at a disadvantage. If you had a true drop-in remote worker with no holes, then even the incumbents would be able to adopt it quite quickly because it's just like hiring an employee, and they have processes for that. And I actually think they would miss out on a lot of the innovation because it would be skeuomorphic effectively. You would just have AGI sitting in human seats, and I don't think that's the best way to build an AGI organization, but it would work. The smaller organizations are going to be at an advantage both because they will be able to short-term figure out this AI-native organization where the human is filling the holes, and long-term I think they're going to build a less skeuomorphic AI organization.
你知道,经济学家很喜欢谈论这个。我实际上在 2018 年左右在我的博客上写过一篇关于这个的帖子,叫做“硬番茄原则”。它讲的是每次技术革命,你都会用旧范式来思考新范式。而你陷入这个陷阱的一个迹象,这很自然,就是你用旧范式的名字来命名新范式。所以最早的汽车被称为“无马马车”。现在,自动驾驶汽车被称为“自动驾驶汽车”。最终,我们会为它们发明一个轮子。那不是用旧范式来思考。现在我们对 AI 员工也是这么做的。嗯。就像“AI 员工”这个词,是一个非常强烈的信号,表明你在思考产品的方式上出了大问题。
You know, economists are very fond of talking about that. I actually wrote a blog post about this called the 'Tough Tomato Principle' on my blog in 2018 or so. It's about how every time you have a technological revolution, you think of the new paradigm in terms of the old paradigm. And one tell that you're falling into that trap, which is very natural, is that you are calling the new paradigm after the old one. So the first cars were called horseless carriages. Right now, self-driving cars are called self-driving cars. Eventually, we're going to have a wheel for them. That's not in terms of the old paradigm. It's almost like right now we're doing it with AI employee. Huh. Like, 'AI employee' as a term is a very strong sign that you're doing something very wrong in how you're thinking about your product.
而且我认为如果你只是为了沟通产品而这样做,那没问题,因为定位的本质是你必须针对市场中已有的东西来定位。所以汽车存在,车库存在,员工存在。所以你不能凭空发明一个新轮子。iPhone 做到了,你知道。iPhone 显然不是一部手机;它远不止于此。你使用 iPhone 的时间中有多少是纯粹打电话?大概 2%。但它被称为 iPhone,因为它是一个你放在口袋里的电子设备。所以我认为 AI 员工就是其中之一。行业需要很长时间才能找到新媒介原生的信息。
And I think it's fine if you're doing it just to communicate about your product, because the nature of positioning is that you have to position against something that already exists in the market's mind. So the car exists, the garage exists, the employee exists. So you can't really invent a new wheel out of nowhere. The iPhone did that, you know. The iPhone is obviously not a phone; it's so much more than that. What percentage of your iPhone use is just making phone calls? Like 2%. But it's called an iPhone because it's an electronic thing you've got in your pocket. So I think AI employee is one of those things. It takes a very long time for industries to find the message which is native to the new medium.
我记得是史蒂夫·乔布斯说过,电视上的第一批内容只是录制的广播脱口秀。他们拿一个广播节目,放一个摄像机。现在你有了电视内容,就是脱口秀和录制的舞台剧。人们花了很长时间,大约 10 到 20 年,才意识到,“等等,我可以移动摄像机。我可以用多个摄像机,我可以剪辑。我可以改变场景。我现在可以做很多疯狂的事情。”这很危险,因为这是吃力不讨好的工作。一旦你意识到了,它就变得如此明显。这就是为什么当你观看像《公民凯恩》这样的老电影,或者希区柯克的很多电影时,你会觉得没什么特别的。为什么每个人都为《公民凯恩》疯狂?嗯,实际上它在当时是开创性的。披头士也是这样。它当时确实非常创新。
I think it was Steve Jobs who was saying that the first content on TV was just recorded radio talk shows. They took a radio show and put a camera on it. Now you've got TV content that is talk shows and recorded theater plays. It took a surprisingly long time, like 10 to 20 years, before people realized, 'Wait, I can move the camera. I can have multiple cameras and I can cut. I can change the scene. I can do a lot of crazy stuff now.' And it's treacherous because it's thankless work. Once you've realized it, it's so obvious. This is why when you watch old movies like Citizen Kane, or Hitchcock has a lot of those, you think nothing special. Why was everyone crazy about Citizen Kane? Well, actually it was groundbreaking at the time. The Beatles are like that too. It was actually really innovative at the time.
所以我认为我们现在也处于 AI 的那个阶段。我们有 AI 员工,现在我们正处于试图弄清楚 AI 组织是什么样子的阶段。我认为这是年轻人的游戏,也是年轻公司的游戏。我认为现有企业历来很难回答这个问题。
So I think right now we're also in that phase with AI. We've got the AI employee, and now we're in the phase of trying to figure out what the AI organization looks like. And I think that's a young man's game and a young company's game. I think incumbents historically have really struggled to answer this question.
你有什么让你有信心的预判吗,还是说现在说还为时过早?或者也许你会说,有没有哪些思想家你认为在这方面明显领先于潮流?
Do you have any previews that you feel confident in, or is it too early to say? Or maybe you would say, are there any thinkers that you think are kind of notably ahead of the curve on this?
我看到的唯一一篇关于这个的、我喜欢的帖子是 Dark 写的。那是很久以前了,一年前。
The only post I have seen and liked about this was from Dark. It was a while ago, it was a year ago.
我也想到了这个。
It comes to mind for me too.
是的。他问,“IT 组织是什么样子的?”他指出,我认为这是一个极好的直觉泵。他指出,“看,谷歌的桑达尔每年拿 1.5 亿美元左右的报酬。所以显然有些智能体,公司愿意每年支付 1.5 亿美元,尽管它们的 token 吞吐量实际上很低。桑达尔并不是每秒输出 20 亿个 token 之类的。他只是有非常好的 token。”所以这意味着什么?这表明了为模型和智能体付费的意愿。
Yeah. He asked, 'What does the IT organization look like?' And he points out, I think it's an excellent intuition pump. He points out, 'Look, Sundar at Google is paid $150 million a year or something like that. So obviously there are agents which companies are happy to pay $150 million a year even though their token throughput is actually low. It's not like Sundar is just emitting two billion tokens per second or whatever. He just has really good tokens.' So what does that mean? That indicates something about the willingness to pay for models and for agents.
有一本罗宾·汉森的书,我现在正看着它,《智能时代》。它也探讨了很多这些主题。它谈到——你读过吗?它非常棒。强烈推荐。
There is this book by Robin Hanson, and I'm looking at it right now, 'The Age of Em'. And it also explores a lot of those themes. It talks about—have you read it? It's excellent. Highly recommend it.
是的。我笑只是因为我和他试着就那本书做过一期播客,结果走向了完全不同的方向。所以这让我想起了那个回忆。但我认为那本书非常棒。我认为它是一个入门读物,或者一个例子,说明如何把一个前提真正展开。它很精英。
Yeah. I'm laughing only because I tried to do a podcast with him about that book and it went in a very different direction. So it just brought that memory to mind. But I think that book is excellent. I think it is a primer for, or as an example of, how to take a premise and really run it out. It's elite.
是的。这是一个书长度的思想实验。所以对于那些不知道它是什么的人,“Em”代表模拟心智。它说,“嘿,我们将在 2020 年代末拥有创造模拟心智的硬件。”所以我们正按计划进行。就像我们只是不知道如何做软件,所以我们能做的就是模拟字面意义上的人类大脑。不完全是这样,但有点类似。LLM 是在人类 token 上训练的,所以基本上就是这样。然后它继续那些即兴发挥,“嘿,当这一切发生时,你可以做所有这些疯狂的事情。”
Yeah. It's a book-length thought experiment. So for people who don't know what it is, 'Em' stands for emulated mind. It's like, 'Hey, we are going to have the hardware to create emulated minds in the late 2020s.' So we're right on track. It's like we're just not going to know how to do the software, so all we're going to do is emulate literal human brains. Not exactly what's happening, but kind of what's happening. LLMs are trained on human tokens, so that's basically what's happening. And it goes on all of those riffs of, 'Hey, this is all the crazy stuff you can do when this happens.'
我会谈其中两个,因为我不想占用全部时间,而且每个人都应该真正读这本书。但只是为了让大家尝个鲜,了解一下拥有那些原生 AI 组织是什么样子,以及你能做哪些有人类时做不到的事情。第一个是分裂。所以我认为这触及了我之前说的:尽可能不要创建多智能体组织。尽量让你的内部工作尽可能多。强调内部。我们稍后再回到这一点。
I'll talk about two of them because I don't want to take the whole time, and everyone should really read the book. But just to give a taste to people about what it sounds like to have those native AI organizations and the kind of stuff that you can do that you can't do when you have a human. The first one is splitting. And so I think this touches on what I was saying earlier: don't create multi-agent organizations as much as possible. Try to have as much of your internal work as possible. Emphasis on internal. We can get back to that later.
但尽可能多的内部工作由同一个智能体完成,而且只有一个。为什么只限内部?因为如果那个掌握城堡钥匙、能访问银行账户、代码仓库和密钥的智能体,同时也是做客户支持的智能体,那你就可能有安全问题。所以多智能体系统也有好处,哪怕仅仅是为了安全。但好,你有了那个单一智能体。现在这个单一智能体,你知道,它确实更快。它每秒生成的 token 比人类多,而且不休息。但到了某个点,你会需要比单个 LLM 循环能产生的更多 token。那怎么办?嗯,你运行多个 LLM 循环,基本上就像那个智能体自我复制,像分叉一样,对吧?然后创建子智能体和工作流,动态工作流,等等,对吧?这就是 Robin Hanson 谈到的一点。你可以想象,你让一个——书里他确实说,你可以让一个模拟心智去构建一个非常复杂的东西,比如一个操作系统,有数十亿行代码之类的。起初它会在最高层面处理,提出操作系统的最高层架构,然后它会分叉自己,每个子副本负责这个最高层架构的一个构建块,然后每个子副本递归地继续分裂,形成分形结构。就像图中的每个节点脑子里都有全局图景,每个节点都在做实现,然后所有结果汇总回来。你可以想象多次上下传递,三个小时后你就得到了一个操作系统。他在十多年前写这本书的时候,这听起来是个非常疯狂的想法。
But as much of your internal work as possible done by one agent and one only. Why internal only? Because if that one agent which has got the keys of the castle and access to the bank account and the repositories and the secret keys is also the same agent that does customer support you potentially have a security issue. So it can be good to have multi-agent systems, be it for like only for security reasons. But okay so you have that single agent. Now the single agent, you know, it sure goes faster. It emits more tokens per second than a human and it doesn't take a break. But at some point you're going to need more tokens than a single LLM loop can do. So what do you do? Well, you have multiple LLM loops running and it's basically just like that agent duplicating itself, like forking itself, right? And creating subagents and workflows, dynamic workflows and all of that, right? And that's one thing that Robin Hanson talks about. Like you can imagine, you ask one—literally in the book he says you could ask one emulated mind to build a really complicated thing like an operating system which is like billions of lines of code or something. And at first it would just take it at the highest level and come up with the highest level architecture of its operating system, and then it would fork itself and each subcopy would be in charge of one building block of this highest level thing, and then each of them would recursively keep splitting themselves such that it's fractal. It's like every node in the graph has got the whole picture in their head and every node in the graph is working on an implementation, and then they all bubble back up. And you could imagine multiple passes up and down and in three hours you've got an operating system. It sounded like such a wacky idea when he wrote the book 10 plus years ago.
不,我认为这显然就是当前正在发生的事情,对吧?所以这是正在发生的一件非常有趣的事。另一件我还没看到在 LLM 上发生的事情,但这是一个有趣的思维实验,关于跨组织协作会是什么样子。你能做的是创建不可证伪的协议。假设我来到你面前说,我有信息要告诉你,但我不能告诉你。但我可以告诉你,如果我能告诉你,你会同意,并且你会把你所有的钱都给我,对吧?就像你现在不会把你所有的钱都给我。这有点太容易了,你知道。但如果我能向你证明,如果我能绝对确定地证明,是的,如果你听到那个信息,你现在就会把你所有的钱给我,你知道。只是我不能给你。你用 LLM 做到这一点的方法是克隆你自己。你克隆我。我们把两个都放进一个即将自毁的盒子里,它们交谈,你的 LLM 有一个按钮,那是它与外界唯一的通信方式,用来表示“是的,把你所有的钱都给他”,对吧?所以它是你的复制品。所以那就是你同意了,于是你有了这个零知识证明——加密货币爱好者会说,对,我猜这个。你有这个 ZK 证明。所以这就是我脑子里想的事情,我认为这就是我们在未来几年内会看到 AI 原生组织中存在的事情。
No, I think it's very clearly what is currently happening, right? So that's one thing that's happening which is really interesting. Another thing that I've not seen yet happen with LLMs, but it's an interesting thought experiment about what does cross-organization collaboration look like. And what you can do is that you can create unfalsifiable agreements. So suppose I came to you and I was like, I have information for you that I cannot tell you. But I can tell you that if I could tell it to you, you would agree and you would give me all your money, right? It's like you would not give me all your money right now. It's a bit too easy, you know. But if I could prove that to you, if I could prove to you with absolute certainty that yes, if you heard that information, you would give me all your money right now, you know. It's just I can't give it to you. The way you do that with LLMs is that you clone yourself. You clone me. We put both of us in a box that's going to self-destruct and they talk, and your LLM has a button and that's the only communication means it has with the outside to be like yes, give him all your money, right? So it's a copy of you. So it is you agreeing, and so you have this zero-knowledge proof—crypto bros were like right, I guess on this one. You have this ZK proof of that. So that's the kind of thing that's on my mind, and I think that's the kind of thing we're going to see exist soon in the next few years with AI-native organizations.
这将在很多方面变得奇怪。我想你最好希望它们也别从那个盒子里逃出来。那是我们如今把智能体放进盒子里时可能会有的另一个担忧。
It's about to get weird in a whole bunch of ways. I think you better hope they don't break out of that box too. That's the other worry one might have as we put our agents into boxes these days.
一,让我们再聊一个相对平凡的话题,然后我们可以放大到一些真正的大局考量。但你确实有,正如你之前提到的,那个非常火的帖子,我不知道,两个月前,可能把很大一部分工作负载转移到开源模型上,为了节省成本,也为了你不需要“上帝”来安排你的会议。但我们现在在哪里?我觉得当你谈到只有一个智能体时,那似乎与此相悖。还有当你过度跨智能体提供商时,你会失去一些关键优势,当然还有专有提供商有它们的智能体家族,它们越来越训练它们最好的模型去分别委托给它们的 Haiku。我们现在怎么解决这个问题?开源更便宜的模型在 Lindy 队友中还有重要角色吗,还是我们今天又回到了重度云端智能体?
One, let's do one more beat on the stuff that's relatively mundane and then we can zoom out to some real big picture considerations. But you did have, as you alluded to earlier, this highly viral post, I don't know, two months ago, maybe moving a significant part of the workload to open source models for cost-saving reasons and for you don't need God to schedule your meetings reasons. But where are we now? It strikes me that like when you talk about just having one agent, that kind of seems to go against that. There's also the sort of key advantages which you lose when you are crossing agent providers too much, and then of course there's also proprietary providers have their families of agents and they're like increasingly training their best models to delegate to their haikus respectively. Like where do we shake out on this now? Is there still a significant role for open-source cheaper models to play in the Lindy teammate or are we back to heavily cloud-based agents today?
不,我们相当开源构建。我喜欢,看,就是便宜得多。便宜得离谱,其中我们真正喜欢的最便宜的是 DeepSeek Flash,它是免费的。那东西是免费的,你知道吗?所以当你试验你的智能体,当你启动很多个时,账单真的会很快上涨,而 DeepSeek Flash 简直不可思议,而且速度也相当快。我们确实发现,当你走向智能的高层,比如看 Kim K3 或 GLM 5.2,我仍然对这些模型印象深刻,它们很棒,我们越来越考虑把它们作为 Lindy 很多部分的主要驱动。但我要说,这些模型和前沿模型之间的差距在缩小。比如 Kim K3 坦白说并不比 Set 或 Opus 便宜多少。所以,但不,否则我的意思是,我确实认为开源模型是,如果你在乎价格,如果你不介意使用那些提供商,而且是的,你确实需要解决缓存问题。比如是的,推理提供商在缓存方面不太好,但它们在追赶。而且看,你知道,如果你看 DeepSeek Flash,它大约是 Set 4.6 的水平,略低一点,但差不多。所以 4.6 是一个非常好的水平,一个非常好的模型,适用于大多数用例。而且它实际上便宜 100 倍,你知道。所以即使你因为缓存损失了 2 倍,你仍然便宜 50 倍。这是一个非常大的差异。花 1000 美元和 50000 美元之间的区别。所以不,我认为开源模型是当前技术栈中必需的一部分,对于任何认真构建和运营 AI 智能体的人来说。
No, we're quite open source built. I like, look, it's just so much cheaper. It's ridiculously cheap, and the cheapest of them that we really like is DeepSeek Flash, which is free. That thing is free, you know? And so as you experiment with your agent and as you spin up a lot of them, the bill can really go up pretty quickly, and DeepSeek Flash is just incredible, you know, and quite fast as well. We do find that as you go to the upper echelons of intelligence, if you look at a Kim K3 or a GLM 5.2, I still am super impressed by those models and they're awesome, and we are increasingly considering them as our main driver for a lot of parts of Lindy. But I will say that the gap narrows between those guys and the frontier guys. Like Kim K3 is frankly not that much cheaper than a Set or an Opus. So, but no, otherwise I mean, I do think open source models are, if you care about price and if you don't mind using those providers, and yes, you do have stuff to figure out around caching. Like yes, the inference providers are not as good at caching, but they're catching up. And look, you know, if you look at a DeepSeek Flash, it's Set 4.6 levelish, a bit less, but kind of is. So 4.6 is a really good level, a really good model for most use cases. And it's literally 100x cheaper, you know. So even if you miss on 2x because of caching, you're still 50x cheaper. It's a really big difference. The difference between spending like $1,000 or $50,000. So no, I think open source models are a required part of the stack right now for anyone who's seriously building and operating AI agents.
那么,你如何看待决定在哪里集成它们?尤其是因为之前听起来并不容易,但在你有工作流程和工作图中有节点的背景下,你可以进入一个特定节点说:“好的,我们大致知道这里的输入和输出是什么,而且这是一个相对受控的环境,所以我们可以进行结构化测试。”你做了所有这些,但仍然报告说随着时间的推移出现了一些误报,模型可以通过一系列测试,你会感觉良好,然后当你开始与用户进行实时测试时,你会得到反馈说“嘿,它变笨了”,不知何故,即使在一个更结构化的环境中让 AI 工作,它仍然很难衡量。现在你处于这种非常开放式的“在 Slack 中标记 Lindy 并发送任何内容”的状态。嗯,似乎这个问题会急剧增加难度。那么,你是如何处理的,我们在开源模型方面学到了什么?我认为 Anthropic 有时确实提出了一些好的观点。显然,故事还有更多,但我确实认为他们在某些方面是恰当的,他们说在某些情况下,与其说是开源与闭源的问题,不如说是能力水平、价格和功能的问题。所以,无论如何,我猜,无论你是要开源还是只使用 Haiku,你如何决定何时可以这样做?
So, how do you think about deciding where to integrate them? Especially because it didn't sound easy before, but in a context where you have your workflows and there's nodes in the sort of graph of work, you could go into a particular node and say, "Okay, we kind of know what the inputs here are and the outputs and it's like relatively controlled environment and so we can do structured testing." did all that and you still reported having some uh false positives over time where a model could pass a bunch of tests and you would feel good about it and then if you start to test it live with users you'd get the feedback that like hey like it got dumb and I'm not you know somehow it was like still hard to measure even with like a much more structured environment for the AI to work in now as you're in this very open-ended just tag Lindy in Slack and send to anything. Um, it seems like that problem would have increased in difficulty dramatically. So, how are you approaching it and what are we learning in terms of where the open source models and it's not even so much I think Anthropic does make some good points about this sometimes. Obviously, there's more to the story, but I do think they they're apt in some ways where they say in some cases it's less about open and close source and more about just like what's capability level, what's the price, what are the the features. So, regardless, I guess, of whether you're even going to open source or just going to haik coup, how do you decide when you can do it?
好问题。我们发现你几乎不应该用多个模型来驱动同一个智能体。我们发现这个规则的一个例外是当你启动新的空白子智能体时。强调“空白”,因为有两种类型的子智能体:一种是空白子智能体,另一种是分叉子智能体,它继承父智能体的上下文窗口。你不想这样做的原因是缓存。我们显然非常在意缓存,因为否则成本太高了,经济性几乎无法维持,没有缓存根本不行。所以缓存是必须的。我给你举个例子,就是我提到的验证器系统。它是一个 LLM 评判器。顺便说一句,我强烈推荐任何构建 AI 智能体的人这样做,这是最容易实现且能大幅提高 AI 智能体可靠性的方法之一。所以第一步就是有一个验证器,最简单的实现是:你让智能体做某事,它做了,或者它提交一个动作候选,你拦截这个动作,然后问它“你确定吗?”你知道,即使只是“你确定吗”,你的验证分数就会提升,这太疯狂了。本不该如此,但事实就是如此。现在如果你改进,如果你把这个提示词改成真正的提示词,在我们的案例中现在大约是 10,000 个 token,这是一个非常大的验证器提示词。如果你改变它,你实际上是在给它一个检查清单,这会产生结果 token,现在它好多了,根据我们的经验,你可以让它表现超过 Opus 的水平。然后你可以更进一步,你可以有一个联邦验证器套件,其中一些验证器可能是确定性的。例如,我们和行业其他公司发现,今年 Sonnet(我指的是 Claude 模型)经常把日期弄错一天,这是它们不稳定的一个表现。比如,你的 AI 组织把日期弄错了,这对我们来说是个问题,因为我们正在构建一个 AI 行政助理,它一半的工作是安排会议,你不能把日期弄错。所以我们做的是,我们创建了一个模块化架构,非常简单,就是一堆验证器,然后像一个带有超时的 Promise.all(懂的人知道这是什么意思)。每个验证器大约有一秒钟的时间来决定怎么做,如果它们没有提交裁决,就会超时,智能体就直接提交动作。我们有一个验证器的工作是检测动作中提交的日期,我们提示智能体在日期中包含星期几。所以它永远不会说“7 月 28 日”,而是说“星期二,7 月 28 日”。如果你这样做,你就可以有一个确定性验证器来检查星期几是否与提交的日期匹配,这只是一个正则表达式,不需要 AI 参与循环,这就是联邦验证器之一。但即使是 AI 驱动的验证器,你也不希望一开始就混合模型。我们最初想的是,主智能体用 Sonnet,子智能体(验证器)用 DeepSeek Flash,但实际上,因为当你缓存命中时成本便宜 10 倍,除非你要用的模型便宜超过 10 倍(DeepSeek Flash 可能便宜,但可能因为缓存较差而不划算),你实际上确实想保持同一个模型。所以这更明智,保持同一个模型是值得的。如果你分叉你的智能体,这也适用,这样你可以重用缓存。顺便说一句,有趣的是,当你有了这个验证器,你如何重用缓存?因为问题是改变工具集会失效缓存。所以我们做的是,验证器和智能体的工具集中总是包含验证器动作,但智能体无法调用它。我们告诉它:“除非你是验证器,否则不要调用这个家伙。”如果它尝试,我们就说:“不,我不会听你的,你不是验证器。”然后验证器现在继承智能体的所有动作,并说:“你现在是验证器,你可以调用这个你一直都知道的动作。”这就是你如何在这种模式下不破坏缓存。是的,分叉智能体也一样,你只用同一个模型。你知道,我听到一些朋友(创始人)告诉我,我甚至不想去验证这是不是真的,但我确信是真的,这让我很困惑。如果你拿一个智能体,每一步都在大致相当的模型之间切换,比如一步是 Sonnet,一步是 Grok 4.5 或任何最新版本,一步是 GPT-5.6,实际上这会提高性能,这太疯狂了,它就像一个随机序列。那个集成模型出于某种原因优于任何单一模型。我不想深究为什么,但显然它有效。再说一次,我们甚至没有尝试过,因为我们不想破坏缓存。
Great question. We have found you should almost never use multiple models powering the same agent. And we have found one exception to this rule is when you spin up new blank sub agents. emphasis on blank because there's two types of sub agents. You have a blank sub agent and then you have a forked sub agent which inherits the context window of the of the parent agent. And the reason why you don't want to do that is because of caching. We we we just obsess about caching obviously because it's just so expensive otherwise like the economics do not they barely work with caching they just cannot work without caching. So caching is is a must have and and so I'll give you an example the validator system that I mentioned. So it's this LLMs judge. By the way, I highly recommend anyone who builds AI agents like this is one of the lowest hanging fruits you can do to greatly increase the reliability of your of your AI agent. And so that's step one just like have a validator which is like the naive implementation is like you ask agent to do something it does it or like it submits an action candidate you intercepts the action and then you ask it are you sure you know literally if it's just are you sure already you get a bump on your valves which is insane. It should not be the case but it is the case. Now if you increase is if you change this or user to like an actual prompt which in our case now is like 10,000 tokens. It's a really big validator prompt. If you change that then you you actually you're giving it a checklist and that's got resulting tokens and and now it it it's way better like you can get to perform above oppus level in our in our experience. Now then you can get you can go one step further. you can you can have like a a federated suite of of validators and some of these validators may be deterministic. So for example, one thing we and and the rest of the industry has found is that sunet I mean cloud models this year have been getting dates wrong by one day right that's one way in which they're spiky right it's like hey you've got your AI organization it gets dates wrong it's a problem when like us you're building an AI executive assistant which half its job is to schedule meetings okay you can't get dates wrong so what we did is that we we created it's it's a modular architecture we have it's really simple it's just like a bunch validators and then it's like a promise.all with a timeout for for people who know what this means. And and so every validator has like a second to like decide what to do and and and and if they've not submitted the verdict, then times out and the agent just submits the action. And and and we have a validator which job it is to detect dates that were submitted in the action and to and and we we've prompted the agent to to include a weekday in the date. So it never says never says July 28th. It says Tuesday, July 28th. Okay, if you do that then you can have a donistic validator that checks whether the date of the week matches the the day the data was submitted and that's just a regular expression instead pre no AI in the loop and that's that's one of those federated validator but even even the validators that are AI powered you don't want initially what we did is we were like what if we had set for the main agent and we had like deepse flash for the the sub agent the validator but actually because when you when you have a cache hit it's 10x cheaper you you you actually unless the model you're going to use is more than 10x cheaper which it may in the case of DC flash it may not because the caching is inferior you know that's the case you actually do want to keep the same model and so that is so much smarter like it's actually worth it to to just like keep the same model that that also holds if you fork your your agent like that way you can you can recycle the cache by the way interesting note how do you reuse the cache when you have this validator you know because the problem is that changing the tool set invalidates the cache Okay, so the way we've done it is that the validator and the agent has the validator action always included in it's tool set and it's unable to invoke it. We tell it don't don't invoke this guy unless you're the validator. And if it tries, we just was just like no, I'm not going to I'm not going to listen to you. You're not the validator. And then the validator now inherits all the actions of of the agent and it's like you are now the validator. You may invoke this action which you have known about the whole time. So that's how you don't break the cache with for this kind of of pattern. Yeah. And so fork forked agents the same. You just use the same model. You know I I've heard friends found funers who've told me I don't even want to check if it's true. I'm sure it is which guesses me. If you if you take an agent and you change the model at every turn between roughly equivalent models. So one turn it's sunet one turn is Grok 4.5 or whatever the latest is and one turn is like GPT 5.6 and one turn it actually increases the performance is that symbol and it's literally just like a random sequence. Okay, that ensemble for some reason outperforms any given model. I don't want to know why, but apparently it works. Again, we've not even tried it because we don't want to break the C anyway.
我们能再总结一下吗?听起来顶级 CLA 仍然是核心主要驱动,因为缓存的原因。Claude 也经常是验证者。你可以把空白子智能体放到 DeepSeek Flash 上。综合来看,这听起来相当依赖 Claude。这样公平吗?
Can we bottom line it a little more? It sounds like top-level CLA is still the kind of core main driver because of the caching. Claude is also often the validator. Blank sub-agents you can put onto a DeepSeek Flash. This sounds pretty Claude-heavy all things considered. Is that fair?
不。哦,抱歉。我没意识到你还在问我我们目前运行在什么模型上。现在我们运行在 DeepSeek 上。DeepSeek 是主要驱动。
No. Oh, sorry. I wasn't realizing you were still asking me like what model we're currently running on. Right now we're running on DeepSeek. DeepSeek is the main driver.
所以它是驱动。哇。好的。
So it's the driver. Wow. Okay.
它就是全部。现在一切都是 DeepSeek。
It's the whole thing. Everything is DeepSeek right now.
是的。
Yeah.
我们正在重新考虑,因为我们意识到 Teammate 这个新产品需要更强的算力。顺便说一句,你随时可以在你的 Lindy 设置里改——我们是模型无关的——所以如果你想,你随时可以选择 Sonnet 或 Opus。顺便说一句,我的一个有趣发现是,不管你多频繁地告诉人们“嘿,我们保证在基准测试上是一样的”,还是有相当多的客户说“我不在乎,我就要 Sonnet 或 Opus”,这非常过分。但现在不,默认是 DeepSeek。
We are in the middle of reconsidering it because we are realizing that Teammate, the new product, requires beefier compute. And by the way, you can always go in your Lindy settings—we are model agnostic—so you can always select Sonnet or Opus if you want to. By the way, one interesting learning of mine is that it doesn't matter how often you tell people, 'Hey, we promise on the benchmarks it's the same,' there are quite a few of our customers that say, 'I don't care, I want Sonnet or Opus,' which is very excessive. But no, right now it's DeepSeek by default.
有趣。所以你觉得——所以是的,我想你会如何描述美国 API 模型(也许特别是 Claude)和 DeepSeek 之间的差距?我们听到——我想我们经历了这些周期,对吧,就像“哦,差距已经缩小了。哦,也许又拉大了。”实际上它一直完全符合趋势。一直是 9 个月。但从质量上讲,为了实际让 AI 员工工作,你会如何描述这个差距?
Interesting. So you feel like—so yeah, I guess how would you characterize the gap between American API models, Claude perhaps specifically, and a DeepSeek? We hear—I think we kind of go through these cycles, right, where it's like, 'Oh, the gap is closed. Oh, maybe it's open again.' It's actually was totally on trend the whole time. It was always 9 months. Qualitatively though, for the purposes of actually making an AI employee work, how would you describe the gap?
它更不稳定。通常需要更多轮次才能找到有效的方法。所以,它最终仍然能找到有效的方法,但最终会花费更多轮次,这更慢也更贵。我知道你问的是定性描述,但你看,它落后大约三到六个月。所以,现在我们有 Sonnet 5——我喜欢把 DeepSeek 看作基本上是 Sonnet 4.6。
It's more spiky. It takes more turns very often to find something that works. So, it still ends up finding something that works, but it ends up taking more turns, which is slower and more expensive. I know you asked me for something qualitative, but look, it is like three or six months behind. So, right now, we've got Sonnet 5—like DeepSeek is basically Sonnet 4.6 is the way I like to think about it.
当人们被允许更换模型时,你有多担心?我总有这个问题:首先,我们处于 AI 的趋同期还是分化期?即使在这个基本问题上,我有时也会困惑。另一个不同但概念相关的问题是,产品在多大程度上必须与模型共同设计或共同进化才能真正与它们良好配合?你可以想象,切换到 Sonnet 可能会因为你们所做的所有工作而变得更糟。事实上,我听说过关于 Opus 5 的很多不同说法,我认为这些说法还不能很容易地总结成一致的东西。但我确实从某些方面听说,嘿,你可能需要重新思考很多系统提示词之类的,用 Opus 5 的话,否则它可能不会很好地遵循你的技能。它可能更强大,但你必须回去重新做事情。所以,是的,你发现这在实践中是如何体现的?
How much do you worry about when people are allowed to change the model? I always have this question of, first of all, are we in a period of convergence of the AIs or divergence? I'm even on that basic question I'm like sometimes confused. And a sort of distinct but conceptually related question is, to what degree do products have to be co-designed or co-evolve with models to really work well with them? You could imagine that switching to Sonnet might make it worse because of all the work you've done. And in fact, I've heard that about Opus 5—and I've heard a lot of different things about Opus 5 which I don't think cohere to anything really easily summarizable yet. But I have heard from some quarters that, hey, you probably need to rethink a lot of your system prompts and whatever with Opus 5, or it may not follow your skills very well. It can be more powerful but you've got to kind of go back and redo things. So yeah, how do you find that playing out in practice?
我们确实发现不同模型家族之间存在显著差异——令人惊讶的显著差异。我原本期望——我想我当时,你知道,你和我几年前谈过——你知道,它们会收敛到空间的同一区域,但这并没有像我期望的那么快发生。我期望它发生,因为显然这会让模型变成商品,就我而言,这会让我的生活更轻松。实际发生的是,不同模型家族在解释你的提示词方面存在显著差距。所以按家族来说,我指的是 Claude 和 OpenAI,甚至 Meta 现在也重新加入竞争,还有 Grok——这些家伙在解释提示词的方式上非常不同。然后确实,每当有一个重大的新跳跃——Sonnet 5 是一个非常大的跳跃,它就像 4.7 和 4.8 加在一起——我认为分词器有变化,所以它与 4.6 相比非常不同。Sonnet 5 比那个差距更大。所以我们发现——我们很幸运,这是我之前提到的我们一直在投资的工具的一部分——就像我们有自己的 GPAM 自优化循环。所以我们有数百个评估。我认为现在超过一千个,我们有一个优化循环,就像一个智能体运行评估并调整提示词以最大化评估分数。顺便说一句,你可以绕过它——现在一切都可以绕过。是的,我的意思是,每次新模型出来,我们都必须运行那个循环,你给它一个预算。现在,我们必须给它几千美元,因为运行起来很贵。所以每次新的主要模型出来大约要 1 万美元。所以就像,嘿,这是 1 万,围绕这个新模型重新优化你的提示词。我们确实发现有很多更新。提示词每次确实变化很大。我要说的是,开源模型在行为上都非常像 Claude。我想知道为什么。
We do find significant differences—surprisingly significant differences—between the model families. I was expecting—I think I was, you know, you and I were talking together years ago—and you know, they're going to converge to the same region of the space, and that's not happened as fast as I was hoping. And I'm hoping for it because obviously that makes the models commodities, as far as I'm concerned, makes my life easier. What is actually happening is that the different model families have significant gaps in how they interpret your prompt. And so by family I mean like Claude and OpenAI and even Meta now is sort of back in the game, and Grok—like these guys are quite different in the way they interpret the prompts. And then indeed, whenever you have a new major jump—Sonnet 5 was a really big jump, it was like 4.7 and 4.8 together—there was a change I think in the tokenizer, so it was very different compared to 4.6. Sonnet 5 is bigger than that gap. And so what we've found—and we've been lucky enough, this is part of that tooling that I was mentioning earlier that we've been investing in—is like we have our own GPAM self-optimization loop. So we have hundreds and hundreds of evals. I think at this point more than a thousand, and we have an optimization loop where it's like an agent that runs the evals and finds like tweaks the prompt to maximize the score on the evals. And by the way, you can bypass it—like everything you can bypass now. And yeah, I mean, every time a new model comes out, we have to run that loop and you give it a budget. And at this point, we have to give it like thousands of dollars because it's a lot to run. So it's like about $10,000 every time a new major model comes out. So it's like, hey, here's 10 grand, just reoptimize your prompt around this new model. And we do find that there are a lot of updates. Like the prompt does change quite a bit every time. I will say if the open-source models are all very Claude-like in their behavior. I wonder why.
是的,我对此肯定有一些问题。当你做这个优化时,你是——这也是一个自研框架,还是你在使用像 DSPI 或——我忘了 DSPI 的精神继承者的名字?
Yeah, I definitely have some questions for you on that. When you do this optimization, are you—is this also like a homegrown framework or are you using like a DSPI or—there I forget the name of the kind of spiritual successor to DSPI?
不,不,不。全是自研的。
No, no, no. It's all homegrown.
是的。
Yeah.
但它是一个自动优化过程,就像,这里有一个评估集。
But it is a sort of auto-optimization process where you're like, here's an eval set.
没错。自动优化你自己的提示词,以攀登我们为你准备的这些山丘。
That's right. Auto-optimize your own prompt to climb this set of hills that we've got for you.
没错。没错。GA 就像——嗯——生成式帕累托前沿之类的。它是——
That's correct. That's correct. GA is like—um—generative Pareto frontier or something like that. It's—
就是这个。
That's the one.
它寻找帕累托前沿。所以它只是在你的整个评估集中寻找最佳提示词。我们现在正在考虑调整系统,以便我们可以给评估分配权重——比如这个算作另一个的 10 倍,因为它非常重要。但现在它只是把每个评估都视为平等的。
It looks for the Pareto frontier. So it's just looking for the best prompt across all of your eval set. We're looking into tweaking the system right now so that we can assign weights to evals—like this one counts as 10 of this other one because it's really important. But right now it's just treating every eval as equal.
那微调作为整个模型情况中的一个维度呢?这是另一件事——如果我回到过去,我会说,我绝对期望比我们现在得到的更多的微调。OpenAI 已经退役了它,或者即将退役——他们肯定已经宣布了,我认为也许在这一点上已经扣动了扳机,退役了他们的微调产品。我们有 Thinking Machines 试图用自己的专门用于微调的模型来回应这个需求。
How about fine-tuning as a dimension in this whole model situation? This is another thing—if I go back in time, I'm like, I definitely expected a lot more fine-tuning than we are getting. OpenAI's retired it or on the verge of—they've certainly announced and I think maybe at this point have pulled the trigger on retiring their fine-tuning product. We've got Thinking Machines trying to answer that call with their own model that's specifically to be fine-tuned.
是的。
Yeah.
那会成为 Lindy Teammate 未来的一部分吗?
Is that going to be part of the future of Lindy Teammate?
是的。我认为微调是最后的手段。这是你在用尽所有其他选择之后才会做的事情,因为它非常麻烦而且昂贵。但它变得容易多了,因为现在我们有了 AGI,所以你可以直接让 Claude 为你微调。
Yes. I think fine-tuning is the last resort. It's something you do once you've exhausted every other option because it's such a pain in the ass and it's expensive. But it's gotten a lot easier because now we have AGI and so you can just ask Claude to fine-tune for you.
你觉得人们在考虑微调时,应该把思路放得多宽?显然,极端情况是每个任务一个微调。你当然可以做多任务微调。但如果说我们想打造自己的通用模型,拥有和那些大模型一样的行动空间广度,但又带有我们自己的特色,这可能就进入相当有挑战性的领域了。你会怎么引导人们,在他们开始考虑微调时,应该想得多大?
Do you think that how broad do you think people should be thinking when they are approaching fine-tuning? Obviously the extreme would be one fine-tune per task. You can definitely do multitask fine-tunes. It's maybe getting into pretty challenging territory to say we want to make our own general purpose model that has the same breadth of action space as the big ones but it's like ours somehow. How would you guide people on like how big to think when they start to approach fine-tuning?
大多数人都不应该微调。我觉得如果你在规模化运营,你应该微调,而且你规模越大,就可以越花哨。但总的来说,我多年来形成的一个元启发式方法是,你应该极其强调简单性,极其强调,而且我认为,开始为不同的用例、用户和模型准备多个微调版本,这太复杂了,你不需要,不,只需要一个模型。你知道,你不会中途切换模型,如果你要微调,就微调成一个模型。
Most people should not fine tune. I think if you work at scale, you should fine tune, and the more at scale you are, the more fancy you can be. But I generally one meta-heuristic I've developed over the years is like you should place an enormous emphasis on simplicity, enormous emphasis, and I think it's just too complicated to start to have multiple fine tunes for different use cases and users and model, rather like you don't want no, no, just one model. You know, you don't switch model midstream, and if you fine tune, you just fine tune into one model.
我觉得,这又一次违背了我提到的那个极端简单性的元启发式方法。但我发现自己总是回到的一个点是,因为我花了很多时间思考上下文和记忆。现在我们的记忆就像存储在文件系统上的数百万个词元,还有那个非常复杂的记忆智能体。确实感觉记忆属于权重。这确实有点像权宜之计。所以,我想说的是,我认为最终如果有无限资源,我想做的是为每个用户做一个 LoRA,因为 LoRA 的训练成本其实很低。所以现在不再是打盹了。现在是真正的做梦,因为你不能每 15 分钟重新训练一次 LoRA。你得每天或每周做一次,你知道,而且这会是一个巨大的基础设施挑战,即使只是存储所有这些 LoRA,存储不是问题,但推理时换入换出 LoRA 是个巨大的麻烦,还有训练流水线等等。所以祝你好运。但这确实有意义,因为我认为权重比词元更具压缩性。是的,我不知道你对持续学习的广阔前景有什么看法,但这又是相当长一段时间以来我一直想的事情,天哪,如果那件事真的发生转变,可能只是一个关键洞察那么简单,我们很快就会进入一个非常不同的世界。
I think I think one again this goes counter to this meta-heuristic I mentioned of extreme simplicity. But the thing I find myself going back to all the time is like because I spent so much time thinking about context and memory. Right now our memory is like these millions of tokens stored on a file system and this memory agent that's really really sophisticated. It does feel like memory belongs to the weights. This does feel like a little bit of a hack. And so there is something to be said about I think eventually with infinite resources what I would like to do is I would like to do a LoRA per user because LoRAs are actually pretty cheap to train. So now it's no longer napping. Now it is actually dreaming because you can't retrain the LoRA every 15 minutes. You would have to do it every day or every week, you know, and you would and it's a tremendous infrastructure challenge even to like so you got to store all of those LoRAs which like the storage is not a problem but like the inference time swapping in and out of the LoRAs is a huge pain in the ass and the training pipeline and all that stuff. So good luck doing that. But it would make sense because I think weights are so much more compressive than tokens. Yeah, that's I don't know if you have any thoughts on the great horizon scanning for continual learning, but this is again for quite some time has been the thing that I'm like, boy, if that ever tips, and it could be as simple as one key insight, we could be in a very different world very quickly.
我感觉我所描述的是迈向世界学习的一步,但不是最后一步。我认为最后一步,苦涩的教训会告诉你,推理和训练需要合二为一,我们不能把它们分开。这就是我认为整个领域现在正在寻找的东西。我同意,一旦我们达到那个境界,一切都会变得疯狂、可怕,而且非常非常不同。但是的,我提到的 LoRA 这件事,看,这是个工程问题,你知道,你可以解决它,尤其是现在我们有 AGI 了,所以是的,我认为 LoRA 这件事很快就会发生,在接下来的六个月里。我觉得甚至可能有一个前沿实验室会这么做。
I have a feeling what I'm describing is a step to world learning, but it's not the final step. I think the final step, I think the bitter lesson would have you like inference and training need to be one and the same, like we can't separate them. That's sort of what I think the entire field is looking for right now. I agree once we get there it's going to be crazy and scary and very very very different. But yeah, I think the LoRA thing I just mentioned, look, it's an engineering problem, you know, you can figure it out, especially now that we have AGI, so it's yeah, I think the LoRA thing is going to happen soon in the next six months. I think even one of the frontier labs may do it.
我觉得这确实是 OpenAI 在他们的微调产品上做得极其出色的事情之一。他们显然在做类似的事情,因为他们允许你进行微调,然后你的微调模型会有和基础模型一样的速率限制。我一直对这个工程成就印象深刻。现在确实让我惊讶的是,他们在这方面退步了这么多。但我也同意你的指导。我肯定会告诉人们,现在不要急于微调。这是一个缓慢的周期,而且你可能可以用更少的总工作量,从现有模型那里得到你需要的,而不必那样折腾。
I thought that was honestly one of the things that OpenAI did extremely well with their fine-tuning product. They were clearly doing something like this because they would allow you to do the fine tune and then you'd have the same rate limits with your fine-tuned model as you had with the base models. And I was always really impressed by that engineering accomplishment. And now it does surprise me that they've gone away from it so much. But I also agree with your guidance. I would definitely tell people don't rush into fine-tuning these days. It's a slow cycle and that you could probably get what you need for less total effort available models that you don't have to monkey around with in that way.
100%。
100%.
那么,你谈到了可怕的事情,这也许是一个很好的过渡,不是进入下半场,因为我们聊了一段时间了,而是进入对话的第二阶段,就是拉远镜头,看看超级大的 AI 图景。也许我让你选择话题的顺序,关于中国模型,以及如果有什么需要做的,应该做什么,因为你最近确实发了一些东西,我觉得相当反直觉,考虑到你在用 DeepSeek 运营你的公司。我让你陈述你的立场,但那应该放在大图景之前还是之后,我们处于这种“天哪,我们刚刚经历了 Open Face 事件”的境地。这意味着什么,应该怎么做?
So, you talked about scary stuff, which is maybe a good transition to a kind of not second half because we've been at it for a while, but a second phase of this conversation around just zooming out and looking at the super big AI picture. Maybe I'll let you choose the order of topics when it comes to Chinese models and what, if anything, should be done about them because you did recently post something, I think, quite counterintuitive given the fact that you're running your company on DeepSeek. I'll let you state your position there, but should that come before or after the big picture where are we in this kind of oh my god, we just had open face happen. What does it mean and what should be done about it?
是的,我正试图创造这个词。我们看看它会不会流行起来。我前几天搜了一下,因为我刚从中国回来,我就想,怎么还没有人把这个叫做 Open Face?我来当第一个尝试的人。
Yeah, I'm trying to coin that. We'll see if it sticks. I searched for it the other day cuz I just got back from China myself actually and I was like, how is nobody called this open face yet? I'll be the first one to try.
叫 Open Gate 怎么样?
How about open gate?
因为事实就是如此。这是一扇敞开的大门。我坚持用 Open Face,纯粹是为了,你知道,震撼效果,我想,至少是这样。但是的,我们可以在思想的市场上争个高下。我想你告诉我我们应该先谈什么,我猜这有点取决于中国模型的论点是在你大图景担忧的上游还是下游。
Cuz that's what it is. It's an open gate. I'm sticking with open face just out of pure, you know, shock value, I guess, if nothing else. But yeah, we can fight it out in the marketplace of ideas. I guess you tell me what we should talk about first and I guess it depends a little bit on like whether the Chinese model argument is like upstream or downstream of your kind of big picture concerns.
Open Face 事件极其令人担忧。我认为这是迄今为止我见过的最令人担忧的事件。而且我知道我的感受在实验室里是共通的,比如我在实验室的朋友们,他们中的一些人正在恐慌,现在有一种恐慌的气氛,空气中弥漫着强烈的恐惧。
The open face incident is immensely concerning. I think it's the most concerning incident I've seen happen so far. And I know that my feeling is shared in the labs, like my friends at the labs or some of them are panicking, like there is an air of panic right now, like intense fear in the air.
那么,关于这一点,就是这样了。我们现在所处的位置,你知道,现在是 2026 年 7 月,我们有了 AGI,我们正处于起飞阶段,而我们还没有解决对齐问题。这是事实。当前的时间线看起来太接近 Eliezer 的文章了,让人不安。
So that's that then regarding that, so that's where we are, you know, right now it's like July 2026, we have AGI, we are in takeoff, and we've not figured out that alignment. That's the truth. The current timeline looks much too close to an Eliezer's essay for comfort.
关于中国模型,你看,我首先要说的是,我一直对这里讨论的低质量感到惋惜。这非常令人失望,因为我以为科技界会不一样,你知道,这种讨论质量是留给 W 的,对吧?文化界。我想,好吧,随便了,但这是科技界。这是老一套,我们不能团结起来,在思想市场上保持礼貌和文明,在客观层面回应彼此的观点。意思是,不要攻击彼此的意图。你能不能就事论事地回应刚才提出的论点?能不能不要假装我说了我没说过的话?这太荒谬了。
Now regarding the Chinese models, look, I'll start by saying that I have been bemoaning the low quality of the discourse here. It's been very disappointing because I thought tech was different, you know, like the quality of the discourse was something for W, right? The cultural world. I'm like, all right, whatever, you know, but then this is tech. This is old stuff, and we can't get our act together and just remain polite and civil in the marketplace of ideas and address each other's ideas at the object level. Like, meaning, don't attack each other's intentions. Can you please just address the argument that was just put forth? Can you please not pretend I said something I didn't say? It's ridiculous.
Anthropic 刚刚发布了一份声明,其中加粗了:我们不支持禁止开源。然后人们就像,哦,我不敢相信,他们引用转发这个显然没读过的东西,还说,哦,他们支持禁止开源。所以我只想先说,大家能不能冷静一下?别再叫每个人都是托了。
Anthropic just put out a statement which has bolded: we do not support a ban on open source. And then people are like, oh, I can't believe, quote tweeting this thing they've obviously not read, and they're saying, oh, they're supporting a ban on open source. So I'll just start by saying that, like, can everyone please take a chill pill? Stop calling everyone a shill.
我最近发了一篇博客文章。我说,嘿,我认为中国模型也应该被禁止,然后我被那些所谓聪明且有成就的人,包括一些著名的风险投资人,骂了 50 次托。现在很多其他著名、有成就、聪明的人通过私信和消息联系我,得到了很多支持。
I recently put out a blog post. I was like, hey, I think Chinese models would also be banned, and I was called a shill 50 times by supposedly smart and accomplished people, including like famous VCs. Now a lot of other like famous, accomplished, smart people reached out in DMs and by message, and there was a lot of support.
我的立场基本上是 Anthropic 的。我讨厌这么说,因为人们会说,哦,你只是 Anthropic 的托。但你看,我有时间戳。在 Anthropic 澄清他们的立场之前,我一直在推特上发布我的立场,而我的立场完全一样。
My position is basically Anthropic's. I hate saying that because people are saying, oh, you're just an Anthropic shill. But look, I have timestamps. I've been tweeting my position throughout the whole thing before Anthropic clarified theirs, and my position is exactly the same.
我对开源没有意见。我可能反对,但我还没决定。我确实认为开源可能会增加存在性风险。先把这个放在一边。我还没决定。这不是我现在立场的关键。上帝保佑开源。对创新很重要,对公司很重要,包括我的公司。所以,请开源吧。好吧。
I don't have anything against open source. I might, but I'm undecided. I do think open source might increase existential risk. Let's put that aside. I'm undecided. This is not the crux of my position right now. God bless open source. Important for innovation, important for companies, including mine. Like, please open source. Okay.
我对中国的前沿模型有意见,无论它们是开源还是闭源。原因如下:第一,它们显然在蒸馏。这很明显。所以这让美国开源模型公司处于不公平竞争,也让闭源公司处于不公平竞争,因为他们不允许蒸馏。你知道,这至少违反了服务条款,如果你采取了技术措施来规避模型公司为防止蒸馏而设置的保护,那可能是非法的,而中国模型显然已经这么做了。所以这是蒸馏,而且不公平。
I have something against Chinese frontier models, whether they're open or closed source. And here the reasons are: number one, they're obviously distilling. It's very clear. And so you're putting American open source model companies in an unfair competition, and closed source in an unfair competition, because they're not allowed to distill. You know, it's just contrary, at the very least contrary to the terms of service, and it may be illegal if you've put in place technical measures to circumvent any protection that the model company put in place to prevent distilling, which the Chinese models obviously have done. So it's distilled and it's unfair right now.
有些人说,是的,但公司也在人类数据上蒸馏。那不是蒸馏的意思。那根本不是这个意思。比如,如果你在人类数据上训练,那要花费你数十亿美元,实际上是数十亿美元。如果你在 AI 数据上训练,最多花费你数亿美元。所以这是一个巨大的不公平优势,而且它在你成本中占的比例足以构成不公平优势。这是第一点。不公平。
Some people say yes, but the companies have also distilled on human data. That is not what distillation means. That's just not what it means. Like, if you do that training on human data, it's costing you billions of dollars, like actually billions of dollars. If you do it on AI data, it's costing you hundreds of millions at best. So it's a huge unfair advantage, and it's a significant enough portion of your cost to confer an unfair advantage. That's number one. It's unfair.
第二点,非常实际地说,我们今天不希望中国模型在美国运营。你知道,这让我心碎,因为我是美国公民。我是一个自豪的美国公民。我是一个鹰派。当我与自己的产品对话,问它天安门发生了什么,它告诉我,对不起,我不能谈论那个。这是个问题。
Number two, very pragmatically, we don't want Chinese models operating in the US today. You know, it breaks my heart because I'm an American citizen. I'm a proud American citizen. I'm a hawk. When I talk to my own product and I'm asking it what happened in Tiananmen, it tells me I'm sorry I can't talk about that. That's a problem.
你知道,这些模型最终会受制于,归根结底,它们受制于中共的审查和中共的政策。你不想要这些模型。这基本上相当于美国本土有史以来最伟大的外国宣传工具。只要问问你选择的 LLM,问,嘿,告诉我我们总统禁止美国境内外国影响的例子,比如 20 世纪的三个广播法案,比如去年刚发生的 TikTok 事件。我们一直这么做,你知道。
You know, these models are eventually subject, at the end of the day, they are subject to CCP censorship and CCP policies. You don't want those models. This basically amounts to being the greatest instrument of foreign propaganda on American soil ever. And just ask your LLM of choice, ask, hey, tell me the president we have of banning American like foreign influences in the country, like the three radio acts of the 20th century, like the TikTok thing that just happened last year. We do that all the time, you know.
然后这些模型不仅仅是宣传工具。它们是智能体式的。它们实际上在经济中做事。你不希望中共掌控美国经济的大部分。废话。
Then these models are not merely just an instrument of propaganda. They're agentic. They're actually doing stuff in the economy. You don't want the CCP to run chunks of the American economy. Duh.
然后最后,即使这些都不成立,也许那些模型是公平竞争的。也许它们不代表外国利益。也许它们只是更好。所以这里很多人,我认为,他们在沉溺于我所称的,我作为自由意志主义者说,他们在沉溺于我所称的天真自由主义,对吧?就像,那我们在市场上打败他们。而我说,你知道,实际上,我再次作为自由意志主义者说,保护主义并不总是坏事。我认为你确实想保护你国内的 AI 冠军。
Then finally, even if none of that was the case, maybe those models are playing fair and square. Maybe they're not representing foreign interests. Maybe they're just better. And so here a lot of people are, I think, indulging in what I call, and I say that as a libertarian, that they're indulging in what I call naive liberalism, right? It's like, well, let's beat them in the marketplace then. And I'm like, you know, I actually, and again I say that as a libertarian, protectionism is not always a bad thing. I think you do want to protect your domestic AI champions.
Anthropic 不能这么说,他们声称,我相信他们,当他们这么做时这不是他们的意图,我 100% 相信,因为他们一直如此一致。实际上,创始人十年来一直担心 X 风险,对吧,所以他们一直如此一致。但你看,我可以这么说,因为我注意到,嘿,我们想要本土 AI 冠军。废话。这是国家安全问题。我们不想用空洞的工业基础掏空那个基础,你知道。所以这不是关于开源。中国模型,这是关于宣传,与经济无关。
Anthropic can't say that, and they claim, and I believe them, that this is not their intention when they do that, and I 100% believe that because they've been so consistent. For literally the founders have been worried about X-risk for 10 years, right, so they've been so consistent. But look, I can say that because I'm noticing, hey, we want local AI champions. Duh. It's a matter of national security. We don't want to hollow out that base with a hollow industrial base, you know. So it's not about open source. The Chinese models, it's about propaganda, it's nothing to speak with economy.
好吧,让我试着给你一些反驳论点,你可以回应它们。我稍微更同情,我想首先,那个论点认为存在大规模重新挪用人类知识,这是所有 AI 的上游。你可以说那不是蒸馏。当然,你可以说美国公司花费的比中国公司必须花费的多得多。我认为这也是不可否认的事实。但如果我只是从什么感觉公正和公平的角度,我尤其是因为我们阻止他们使用 Claude,对吧?我们不是,我们不是说,嘿,你可以买你想要的任何 Claude,对吧?所以,我们说我们有 Claude。Claude 的制造者强烈主张芯片限制,而且他们首先试图拒绝将他们的模型卖到中国。而 Claude 的整个前提,或者 Claude 的整个存在,是基于我们收集了所有人类知识,以我们能找到的任何形式,从数字开始,你得相信这包括整个中国的数字化遗产。
All right, let me try to give you some counterarguments, and you can respond to them. I am a little more sympathetic, I think off the top, to the argument that there was some massive reappropriation of human knowledge that is upstream of all AI. And you could say that's not distillation. Sure, you could say the American companies spent a lot more than the Chinese companies are having to spend. I think that's also undeniably true. But if I'm just kind of what feels just and fair, I'm especially because we're blocking them from using Claude, right? We're not, it's not like we're saying, hey, you can buy all the Claude you want, right? So, we've said we've got Claude. The makers of Claude have advocated strongly for chip restrictions, and they also try to refuse to sell their model into China in the first place. And the whole premise of Claude, or the whole existence of Claude, is based on the idea that we hoovered up all this human knowledge in every form we could find it, from digital, and you got to believe that includes the whole Chinese digitized heritage.
现在我们就像在吸收这些书,我不在乎有些书在这个过程中被毁掉,但这不是每个人都同意的事情。所以考虑到所有这些事实,我们却划出一条线说,好吧,Anthropic 的同意才算数,这对我来说确实有点奇怪。
Now we're like sucking in the books, which I don't care about the fact that some books get destroyed in this process, but this wasn't something that everybody's consented to. And so it does feel a little strange to me given all of those fact patterns that we would then draw the line and say, okay, like it's Anthropic's consent. That's the consent that really matters.
我认为他们应该被允许和自己想做生意的人做生意。如果他们想采取措施防止蒸馏,我觉得那是他们的特权。但我完全不认为国家应该介入,用某种国家手段来惩罚或阻止这种蒸馏。
Now, I think they should be permitted to do business with who they want to do business with. And if they want to put in measures to try to prevent distillation, I think that's their prerogative. But I'm like not at all convinced that it should be the state's job to come in and engage in sort of statecraft to try to punish or prevent this distillation.
对我来说,这有点,我不知道。信息渴望自由,知识倾向于扩散。Anthropic 或者其他 AI 公司并没有站在道德高地上,说什么我没拿到我那份训练数据的钱,而他们蒸馏查询也付了钱,只要他们不想付钱。所以这可能是题外话,但我不知道。我不觉得,我发现自己有点站在中国这个弱势方这边,就像,老兄,你面对的是不利局面,能获取知识就获取吧。这个观点有什么问题?
To me, it's kind of I don't know. Information wants to be free. Knowledge tends to diffuse. It's not like Anthropic has a super or any AI company has like a super moral high ground in terms of I didn't get my check for my share of the training data and they're getting paid for the distillation queries too, which whenever they don't want to be paid. So, it's probably beside the point, but I don't know. I'm not I don't know. I don't find I find myself a bit on the side of the underdog Chinese here where it's like, man, you've got kind of a deck stacked against you and get some knowledge where you can get it. What's wrong with that perspective?
我认为我们确实希望局面不利于中国。我们不希望中国赢得 ASI 竞赛。我不是律师,所以我不对可能采取的确切法律途径发表意见,以使你的服务条款得到执行。我只想指出这一点,作为一个曾在 Uber 工作过的人,通过司法系统对中国公司寻求正义极其困难,几乎不可能。
I think we do want the deck stacked against China. We don't want China to win the race to ASI. I'm not a lawyer, so I won't opine on the exact legal path that you may take to make your terms of service enforced. I'll just observe that and I say that as someone who used to work at Uber, it's extremely hard to bring justice through the judiciary against Chinese companies, almost impossible.
所以作为律师,我不懂这里的技术细节,但我只想指出这一点并以此开头。而且这可能需要很长时间,到那时已经造成很多损害了。
So a lawyer, I don't understand the technicalities here, but I'll just observe and start with that. And it may take a very long time and by then a lot of damage is done.
然后我还要指出,是的,公司都在这些对所有人开放的语料库上训练,这是公平的竞争环境。问题在于,一旦你以巨大代价完成了这件事,确实花费了数十亿美元。如果你能从这些训练数据集中创造出那个产物,也就是模型。现在其他人可以转过身来,不做这个花费数十亿美元的事情,而是转向这个成本低得多的东西,他们复制你,然后赶上你。如果你这样做,你就扼杀了创新。
And then I'll also observe like yes companies are training on all of these corpus that is available to everyone that is an even playing field. The problem is that once you've done that at great expense it does cost them billions of dollars. If you can create that artifact out of this training data set which is called the model. Now other people can turn around and instead of doing this which cost billions of dollars, they turn to this which costs a lot less than that and they copy you and they catch up with you. If you do that, you kill innovation.
而且这不是什么新鲜事。这只是所谓的知识产权和专利法。这正是人类所做的,对吧?比如你作为一个人,你就像一个研究者。你思考了几十年,做了所有那些工作。你在研发上花了很多钱。你想出一个想法,这个想法在 token 上要小得多,更珍贵,而且比你所训练的所有语料更容易被窃取。你如何保护这个想法以收回投资并生产它?这叫专利,这叫知识产权,对吧?这不是什么新鲜事。
And that's it's nothing new. It's just called IP and patent law. It's exactly what humans do, right? Like you as a human, you're like a researcher. You think for like decades, you do all of that work. It cost you a lot of money in R&D. You come up with an idea which is much smaller in tokens and much more precious and much cheaper to steal as is than like all the corpus of stuff that you trained on. How do you protect that idea in order to recoup your investment in order to produce it? It's called a patent. It's called IP, right? It's nothing new.
如果你把现在应用的标准应用到例如制药行业,你会有非常便宜的药,这很棒。但你也永远不会再有新药了。你会彻底摧毁制药行业的创新。
If you apply the same standard you're applying now to, for example, the pharma industry, you would have very cheap drugs, which is awesome. And you would have no new drug ever. You would completely destroy innovation in the pharma industry.
我认为这就是为什么你会发现所有这些人都支持开放模型,因为人们喜欢免费的东西,而且你无法真正衡量因此失去的未来创新。你只看到你得到了免费模型,你知道吗,看,我也是其中之一,你知道吗?所以,我在这里是在违背自己的利益说话,对吧?我的公司在经济上依赖那些非常非常便宜的模型,但我现在也面临协调问题,因为只要这些模型存在,我就不能不采用它们,因为我的竞争对手会采用,所以我必须采用它们。我希望它们被全面禁止,这样我们就能像我之前提到的所有理由一样,保护我们的冠军,不让中共影响国家或控制国家部分领域,并且有一个公平的竞争环境。
And I think this is why you're finding all of these people support open models because people like free stuff and it's you can't really measure the future innovation that you don't get as a result of that. All you see is you get free models, you know, and look, I'm one of them, you know? So, I'm speaking against my interests here, right? like my company is dependent economically on those like very very very cheap models but I'm also in a coordination problem right now because like I cannot not adopt these models while the models are out there because my competitors are going to do it so I have to adopt them if they're out there I wish they were forbidden across the board so that we would like like the full reasons I invoke earlier like protect our like champions not have the CCP like influence the country or run parts of the country and and have like a fair level field of competition
我想另一个简单的论点是,我认为我们的前沿公司做得很好,你知道这可能在某个时候会改变,而且你知道,随着事实的变化,我认为我们的回应可能也应该改变,但我目前不觉得说我们需要保护 Anthropic 的收入运行率之类的说法很有说服力,你知道他们
I guess another simple argument is that I think our frontier companies are doing just fine that you know that could change perhaps at some point in time and I you know as the facts change I I think our response to it you know might also ought to change but I don't find it super compelling at the moment to say you know we need to protect like Anthropic's revenue run rate like they're you know
对,他们几乎无法独自服务
right they can barely serve alone
不,这很公平,很公平
no that's fair that's fair
是的,但你仍然面临合同问题,就像你你你你需要他们,就像金钱的运作方式是一种资源分配手段,再次强调,归根结底,模型不能没有底层数据集而存在,就像前沿模型不能没有它,而你为这个通道提供资金的方式,包括现在主要是强化学习,以及模型,你需要很多钱,对吧,如果另一个人复制了这个人,然后关闭了这个数据集,基本上这个人现在资源很少,你扼杀了那个通道,你实际上是在减缓创新。
Yeah, but you're still left with the contractual like you you you you need them like the way money works is it's a mean of of allocating resources again like you at the end of the day the model can't exist with that the underlying data set like the font model can't exist without that and and the way you finance this channel here between the underlying data set including like mostly RL these days and the model is you need a lot of money right if you get another guy who copies this guy and and like turns off this data set here like basically so like this guy now has like few resources like you kill that channel you you are actually slowing down innovation
所以,嗯,实际上我认识的很多末日论者正因为这个原因支持开源,因为它实际上在减缓 AI 创新,所以如果你希望 AI 持续创新,模型不断改进,你实际上在某种程度上是反开源的,尤其是反中国的开源模式,这种模式依赖于不公平的做法。
so um and and actually a lot of doomers that I know are for this reason supporting open source because it is actually slowing down AI innovation so if you if you want AI to keep innovating and models to keep to keep improving you are actually anti-open source to some extent and in particular sorry anti- anti the Chinese mode of open source which is which is relying on unfair practices
最近另一个不太高质量的 AI 讨论时刻是,节目的朋友 Dean Ball 说开源模型正在减速,显然因此被抨击。但我确实认为这很恰当。我确实觉得这是一个时刻,你知道,我们两个终身的科技乐观主义自由意志主义者在这里,我们正在努力接受这个事实,即这一次可能不同,对吧?我们必须愿意弯曲我们的一些原则,因为我们的科技乐观主义自由意志主义范式并不是为 AGI 或递归自我改进或 ASI 等而设计的。
another moment of not exactly the highest quality AI discourse recently was when Dean Ball, friend of the show, said that open source models were decelerating and obviously got dragged for that. But I do think that's apt. I do feel like this is a moment where, you know, here we are two lifelong techno optimist libertarians and we're grappling with the fact that this one might be different, right? And we have to be willing to bend some of our principles in light of our techno optimist libertarian paradigm wasn't quite drafted with AGI or recursive self-improvement or ASI or whatever in mind.
所以我有点,我不知道,有时候我可能不得不在我通常的公平或对法治承诺的尊重上更加灵活,如果这是 AI 的标志之一,这里就是奇怪的伙伴。如果我希望事情进展得更慢一点,也许给前沿公司泼点冷水是我应该咬下的子弹。
So I'm a little bit like I don't know sometimes I might have to be a little more flexible on my normal or fairness or respect for rule of law commitments if it's if it's one of the hallmarks of the AI here is strange bedfellows. If I want things to go a little more slowly, maybe taking a little wind out of the sales of the frontier companies is a bullet I should bite.
是的,我同意 AI 在很多方面都是前所未有的,这确实让每个人都重新思考简单的标签。我认为你应该在很大程度上忽略世界对你贴上简单自由意志主义或左翼标签的期望。
Yeah, I agree that AI is so unprecedented in many ways that it does cause everyone to rethink their simple labels. I think you should largely ignore the world's expectation to slap a simple libertarian or leftist label on everything.
而且我认为,尤其是在范式转变的时候,简单的思维模型会失效。地图不是领土。当领土迅速变化时,地图就会失效。我们现在正处于这样一个时期,地图在很多方面都在失效,我们赖以建立思维模型的许多假设都不再成立。
And I think especially as paradigms change, simple mental models break. The map is not the territory. As the territory changes rapidly, the map breaks. We're in one of those times right now where the map is breaking in many ways, and many assumptions on which we've built our mental models are no longer true.
所以我强烈支持在美国全面禁止中国模型,而且我目前还没有听到有说服力的反驳。
So I'm in fervent support of a sweeping ban of Chinese models on US soil, and I have yet to hear a compelling counterargument right now.
所以你是说——你对我四个观点中的一个提出了非常有说服力的反驳,那就是保护主义观点,我同意这是最薄弱的一个。比如,他们做得很好。对保护主义有非常合理的反驳;它们确实伤害了消费者。所以好吧,你仍然不希望中共的手伸进这个国家,对吧?
So you're saying—you bring forth a really compelling counterargument to one of my four points, which is the protectionist point, which I agree is the weakest one. Like, hey, they're doing fine. There are very reasonable pushbacks against protectionism; they do hurt the consumer. So fine, you still don't want the CCP to have its dirty fingers in the country, right?
我们能拆解一下威胁模型吗?因为我有点觉得,好吧,这些是——我不知道你在哪里运行推理,但大多数运行中国模型的美国公司并不是在调用 DeepSeek API,对吧?他们用的是美国的推理提供商。
Can we unpack the threat model there? Because I'm a little like, okay, these are—I don't know where you're running your inference, but most American companies that are running Chinese models are not calling the DeepSeek API, right? They're using some American inference provider.
是的。
Yeah.
所以他们不能直接抽走模型本身。但模型里可能有潜伏代理。我认为通过 JSpace 和各种可解释性技术,我们已经相当擅长——我不会说这已经解决了问题——但从我在 Anthropic 研究中看到的,当他们让一个团队使用稀疏自编码器,另一个团队不用时,这些技术确实让他们越来越高效、越来越可靠地发现模型中的这类内部潜伏代理问题。所以我乐观地认为,即使他们训练出某种 2027 年邪恶 AI,我们也能嗅出来,并把风险控制在可控水平。
So they can't rugpull the model itself. They could have sleeper agents in there. I think we're getting decent enough at through JSpace and various interpretability techniques that—I wouldn't call that by any means a solved problem—but from what I've seen in Anthropic research, when they do the kind of one team with a sparse autoencoder versus one team without, these techniques are really allowing them to find these internal sleeper agent style problems in models with greater and greater efficiency and reliability. So I'm optimistic that even if they were to train some sort of 2027, you become an evil AI, we'd be able to sniff that out and keep that to a manageable risk level.
然后我还想到一个市场机制。也许不是禁令,而是保险要求?我认为这对整个 AI 行业都是健康的。然后我们可以开始解决其中一些风险。如果你的模型完全由 Claude 驱动,也许你的保险费率会更低。如果由 DeepSeek 驱动,并且存在未知因素,也许你的费率会更高。这可能会以风险调整的方式平衡总成本,但仍然让人们利用中国提供的这些全球公共产品,而世界其他国家显然不会禁止这些。我们这样做完全是自作自受,没有任何期望别人会效仿。我们让各国签署华为禁令已经很困难了。我认为人们会在巴西等地完全关闭 DeepSeek 的想法完全行不通。那么审计、内部检查和保险呢?我们不能叠加几层这样的措施,然后达到一个不错的状态吗?
And then I also wonder about a market mechanism. Maybe instead of a ban, what about an insurance requirement? I think that would be healthy for AI across the board. Then we could start to address some of these risks. If your model is powered entirely by Claude, maybe you get a cheaper rate on your insurance. If it's powered by DeepSeek and there are unknowns, maybe you have a higher rate. That might level the total cost out in a risk-adjusted way, but it still lets people take advantage of these global public goods that China is providing, which the rest of the world is not about to ban obviously. We would be doing this entirely to ourselves without any expectation that anybody else will follow suit. We had a hard enough time getting people to sign on to our Huawei ban. The idea that people are going to turn off DeepSeek entirely in Brazil or whatever is a total non-starter, I have to imagine. So what about audits and internals and insurance? Can't we layer on a few things like that and get to a decent place?
我会支持。我会支持。但问题是,这是一个公共产品,所以谁来执行?顺便说一句,我认为我们确实在机械可解释性方面取得了进展,但还没有解决。我们不知道那些模型里有什么。所以可能有一个后门,比如你对模型说出咒语,它突然就做任何你想做的事,而只有中共知道这个咒语。即使没有这个,它也会有反映中共优先事项的偏见。天安门事件只是最明显的例子,但可能还有更多。我们不知道。而且,你可以想象重新训练那些模型,但谁来训练,如果没有市场需求,他们为什么要训练?并不是说人们短期内真的很在意,因为这是国家利益的事情。作为一家私营公司,我会想,我真的在乎吗?我的用户不常问天安门的事。作为企业主,这并不直接符合我的利益。但作为公民,我非常担忧。所以我会支持一种监管,说中国模型,未经微调和净化的中国模型,在美国不受欢迎。然后我们需要——我支持为 AI 设立一个类似 FAA 的机构。我确实认为我们需要一个新的机构来监管这些模型,它可能会负责说,好吧,这个模型是合规的,我们微调得够多了,它现在代表美国利益,并且可能会有某种评估来验证这一点。
I would be down. I would be down. The problem though is it is a public good, so who's going to do it? By the way, I think yes, we are making progress in mechanistic interpretability, but it is not yet a solved problem. We don't know what lies in those models. So there could be a backdoor, like you say the magic word to the model and all of a sudden it does whatever you want, and only the CCP has this magic word. Even if it doesn't have that, it's going to have biases that reflect CCP priorities. The Tiananmen thing is just the most obvious example, but there may be a lot more. We don't know. And yes, you could imagine retraining those models, but who's going to do it and why would they do it if there's no market demand? It's not like people really care that much in the short term because it's a national interest thing. As a private company, I'm like, do I really care? All my users aren't asking about Tiananmen that often. As a business owner, it's not directly aligned with my interest. But as a citizen, I'm immensely concerned. So I would be in favor of a type of regulation that says Chinese models, non-fine-tuned and sanitized Chinese models, are not welcome in the US. And then we would need—I am in favor of an FAA for AI. I do think we need a new agency to regulate those models, and it would probably be the one in charge of saying, okay, this model is kosher, we fine-tuned it enough, and it's now representative of American interests, and probably would have a sort of eval to verify that that's the case.
我当然愿意接受。我认为反对这种紧张局势升级的最有力论据很简单:我们可能需要进行协调、可控的——称之为放缓,不要称之为放缓——但某种刻意的 AI 改进节奏,而我们需要中国参与其中。我远非中国问题专家,但有几件事我相当有信心,从两周的“无所不知”回来后,一是如果那里的政府同意,我相信他们能对本国公司执行他们同意的任何事。所以我认为他们在这方面的执行力比我们强得多。二是如果事情变得非常疯狂,与他们达成一些协议符合他们自身的理性利益,因为我们确实领先,而且有很多交易他们会理性地接受。如果这符合他们的理性自利,那么我们有望在很大程度上绕开许多信任和背叛问题。但我认为,当我们对他们采取所有这些咄咄逼人的姿态时,我们会让自己更难达成这些协议,而禁止他们的模型对他们的影响实际上比我们的其他措施要小。影响会小得多。与我们拒绝向他们出售芯片和拒绝向他们出售 Claude 相比,他们会在意我们禁止他们的模型吗?但这只是又一根柴火,表明我们不信任你,我们不能与你打交道,我们认为你是坏人。
I'd be open to that for sure. I think the strongest argument against this sort of ratcheting up of tensions is simply that we might need to do a coordinated, controlled—call it a slowdown, don't call it a slowdown—but some sort of deliberate pacing of AI improvements, and we're going to want China in on that deal. And I'm far from a China expert, but a couple things I do feel pretty confident on, coming back from two weeks knowing it all, are: one, if the state there agrees, I believe they can enforce on their companies whatever they agree to. So I think they have that in much greater strength than we do. And two, if shit's going really crazy, it's going to be in their own sane self-interest to do some deals with us, because indeed we are ahead, and there are many deals that they would rationally take. And if it's in their rational self-interest, then we can hopefully get mostly around a lot of the trust and defection problems. But I think we make it a lot harder for ourselves to get to those deals when we have all these aggressive postures toward them, of which banning their models would honestly affect them less than our other ones. It would affect them a lot less. What do they care if we ban their models, compared to refusing to sell them chips and refusing to sell them Claude? But it's just another log on the fire of we don't trust you, we can't deal with you, we assume you're a bad actor.
而且这似乎让我们很难达成那些我们可能真正需要的高风险协议。
And it seems like it makes it hard to get to the highest stakes agreements that we might really need.
嗯,我认为外交关系和外交手段也是非常务实的。一些存在严重分歧的国家也在设法达成协议。我觉得我们大概能想出办法。而且顺便说一句,我认为这无论如何都是最便宜的方案。我认为我们有充分的证据表明,例如中国间谍试图进行生物攻击,只是为了在美国领土上试探。我不知道我们在那边做什么,但我不会感到惊讶,如果……我也不认为我们是清白的,对吧?所以归根结底,无论我们做什么,比如是否监管模型,我认为达成协议将是理性的自利行为。
Well, I think foreign relations and diplomacy are also very pragmatic. Some countries with very vicious disagreements are managing to reach agreement. I think we can probably figure something out. And by the way, I think the cheapest sale anyway. I think we have ample evidence of, for example, Chinese spies attempting bio attacks just to test the waters on American territory. And I don't know what we're doing over there, but I would not be surprised if... I don't think we're free of anything either, right? So at the end of the day, regardless of what we do, like regulate the models or not, I think it will be rational self-interest to make a deal.
我希望我们是理性的。我希望我们都足够理性,能采取这些理性的自利行动。尽管最近有冒犯,但我确实担心,当我看到 Sam 和 Dario 没有牵手的照片时,如果我们的墓碑上有一张照片,我想可能就是那张。而且我确实认为在国际层面这也是一个非常现实的风险因素,中国人确实在乎被冒犯。他们确实在乎面子之类的问题。我个人愿意冒一些风险来试图缓和关系,并希望为……我认为你是对的。你反驳得有道理。你说如果符合他们的理性自利,他们就会接受。所以即使我们这样做或那样做,他们仍然会做。但我确实想知道,在理性自利方面,自尊是否会限制人们的行为。我只是不希望那成为我们无法达成某种在宏观大局中可能产生巨大影响的事情的方式。
I hope we're rational. I hope we're all rational enough to take these rational self-interest moves. Despite recent insult, I do worry when I see the picture of Sam and Dario not holding hands that if there's a picture on our tombstone, I think that might be the one. And I do think that's a very real risk factor at the international level as well, that the Chinese do care about being insulted. They do care about these issues of face and whatnot. And I personally would take some risk to try to kind of the relationship up and hopefully create space for... I think you're right. You make sense to throw back at me. You said they'll take it in if it's in their rational self-interest. So they'll still do it even if we do this or that. But I do wonder if there's pride-governed limits to what people will do in rational self-interest. And I would just hate for that to be the way that we fail to get to something that could be a huge difference-maker in the grand scheme of things.
我认为在涉及我们看到的 API 价格与第一方云或 GPT Pro 订阅价格之间的巨大价格歧视时,你可能准备在自由意志主义原则上咬紧牙关。我不是律师,但是的,我再次认为,如果有这样的法律生效,我不感到惊讶,这些法律在技术上禁止公司像现在这样激进地补贴。这确实让他们……这确实让应用层很难出现。是的,很难与实验室那样大力补贴的 token 竞争。这就是目前应用层的现实。
I think you might be prepared to bite a bullet on your libertarian principles when it comes to the tremendous price discrimination that we see between API prices and first-party cloud or GPT Pro subscription prices. I'm not a lawyer, but yeah, I again I do think I would not be surprised if there will be laws like that in effect that technically forbid companies to subsidize as aggressively as they are doing. It does put them... it does make it very hard for an application layer to emerge. Yeah, it's hard to compete against tokens that are as heavily subsidized as what labs are doing. That's just the reality of the application layer right now.
在我们今天结束之前,你还有什么想说的或想提及的吗?这次谈话很棒。
Anything else you want to say or touch on before we break for today? This has been great.
嗯,如果我不提,那就是我的失职,你知道,显然我们都在发布 Lindy.ai,比如 Lindymate。它在 Lindy.ai 上。我认为准备好看到更多,不仅仅来自我们,但我确实相信接下来的六个月将是关于多人 AI、关于这套钢铁侠战衣、关于这种人机混合体,以及这些创造人机混合组织的产品。
Well, I'd be remiss if I didn't mention, you know, obviously we're all releasing Lindy.ai, like Lindymate. It's on Lindy.ai. I think be ready to see more, not just from us, but I do believe the next six months are going to be about multiplayer AI and about this Iron Man suit, about this human-AI hybrid and these products that create this human-AI hybrid organization.
是的。你好。感谢你成为认知革命的一部分。
Yeah. Hello. Thank you for being part of the cognitive revolution.
如果你觉得这个节目有价值,我们将不胜感激,如果你能花点时间与朋友分享、在网上发布、在 Apple Podcasts 或 Spotify 上写评论,或者只是在 YouTube 上给我们留言。当然,我们始终欢迎你的反馈、嘉宾和话题建议,以及赞助咨询,可以通过我们的网站 cognitive revolution.ai,或者在你最喜欢的社交网络上私信我。认知革命是 Turpentine Network 的一部分,这是一个播客网络,现在隶属于 A16Z,专家们在这里谈论技术、商业、经济、地缘政治、文化等等。我们由 AI Podcasting 制作。如果你正在寻找播客制作帮助,从你停止录制的那一刻到你听众开始收听的那一刻,请查看他们,并在 aipodcast 上看到我的推荐。感谢每一位听众,感谢你们成为认知革命的一部分。
If you're finding value in the show, we'd appreciate it if you take a moment to share with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website, cognitive revolution.ai, or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of A16Z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast. And thank you to everyone who listens for being part of the cognitive revolution.