AI Coding Wars and the Future of Agents
打开互动全文版(中英对照 + 朗读 + 问答)→跨界对话,探讨 AI 编程现状、智能体框架及市场动态。
A crossover episode discussing the current state of AI coding, agent harnesses, and market dynamics.
这难道不疯狂吗?这个数字简直令人难以置信。今天 AI 编程战争的状况如何?
Isn't that crazy? That number is just mind-boggling. What is the state of the AI coding wars today?
我们正处于某种能力探索的阶段。我一直在追求的总论点就是,就像 2025 年是编程智能体的一年,2026 年则是编程智能体突破遏制去做其他一切事情的一年。
We're in a phase of sort of like capability exploration. The general thesis that I have been pursuing now is that the same way that 2025 was a year of coding agents, 2026 is coding agents breaking containment to do everything else.
你担心基础模型会蚕食掉很多这类创业公司的类别吗?
Do you worry about the foundation models just eating into a bunch of these startup categories?
中等规模的创业公司,是的。
Mid-size startups, yes.
你认为这个市场的最终状态是什么?
What do you think the end state of this market is?
要让市场结构发生显著变化,那将会是……
For the market structure to significantly change, there would be...
今天在 Unsupervised Learning 上,我们播出了一集有趣的节目,这已经真正成为一年一度的传统,与我们在 Latent Space 的朋友们的跨界合作。Swyx 和我坐下来,讨论了当今 AI 生态系统中发生的一切,我们对模型层各种变化的看法,基础设施领域正在发生的事情,编程战争,以及其他很多事情。和我真正尊敬的人以及另一位优秀的播客主持人一起做这件事非常有趣。那么,闲话少说,这就是我们的节目。
Today on Unsupervised Learning, we had a fun episode in what's really become an annual tradition, a crossover episode with our friends at Latent Space. Swyx and I sat down and we talked about everything happening in the AI ecosystem today, what we thought of the various changes at the model layer, what's happening in the infra world, the coding wars, and a bunch of other things. It's a ton of fun to do this with someone I really respect and another great podcaster in the game. So, without further ado, here's our episode.
嗯,Swyx,能再次参加 Unsupervised Learning 和 Latent Space 的跨界节目,真是太有趣了。
Well, Swyx, this is super fun to be back with another Unsupervised Learning Latent Space crossover episode.
是的。我觉得我们可以从很多地方开始,但你知道,我总觉得你花时间的方式有一点特别吸引人,那就是你显然处于这个工程运动和社区的中心,你举办这些活动和会议,呈现这些精彩的演讲,而且我觉得你对正在发生的事情的时代精神有很好的把握。
Yeah. I feel like a lot of places we could start, but you know, one thing I always find fascinating about the way you spend your time is you obviously are at the epicenter of this engineering movement and community and you run these events and conferences and put on these awesome talks and I think just have a great pulse on the zeitgeist of what's going on.
也许先开始谈谈,现在人们思考的最大话题是什么?
Maybe to start just what are the biggest topics people are thinking about right now?
是的,我刚从伦敦回来,我们在那里举办了 AI Europe,现在我们大约每季度举办一次,这加快了节奏。这是在努力跟上 AI 的速度,你知道的。
Yeah, so I just came back from London where we did AI Europe and we're doing roughly one per quarter now, which is up the pace. It's trying to match the AI speed, you know.
我想,演讲内容会完全不同,你知道的。
The talks will be completely different, I imagine, you know.
我确实策划了这些主题,当你看到主题列表和我邀请的演讲者时,你就能看出我的想法。显然,Open Claw 是过去四五个月的故事。然后,在其之下,我认为 harness 工程和上下文工程是智能体和 RAG 中两个相关的主题。然后,还有长尾的常青内容,比如评估、可观测性、GPU 和 LM 基础设施,以及一般性的内容。我们还有其他更新,比如多模态和生成式媒体,我们姑且这么称呼。但肯定的是,我提到的前三个是人们最关心的。
I definitely curate the tracks, like you can see what I think when you see the track list and the speakers that I invite. Obviously, open claw is like the story of the last four or five months. And then, just below that, I would consider harness engineering and context engineering to be two related topics in agents and RAG. And then, there's the long tail of evergreen stuff, like evals, observability, GPUs, and LM infra and just general in general. We also have other updates on like multi-modality and generative media, let's call it. But definitely the first three that I mentioned are top of mind for people.
我认为 harness 特别有趣。是的,最近 LangChain 的首席执行官 Harrison Chase 发了一条推文,引起了我的注意,他说,你知道,终于感觉我们在 AI 基础设施方面有了稳定性。我认为他基本上是在暗示,看看过去两三年,作为一家处于 AI 基础设施中心的公司,这有点像打地鼠游戏,对吧?你不断地随着构建模式的发展而移动。
I think harness is particularly so interesting. Yeah, there was this tweet from Harrison Chase, the LangChain CEO, that caught my eye recently where he said, you know, it finally feels like we have stability around the infrastructure for AI. And I think what he basically was implying is like look over the past 2-3 years as a company at the epicenter of AI infrastructure, it was a bit like playing Whac-A-Mole, right? You were constantly moving around with however the building patterns were evolving.
当然,对吧?自从他创办 LangChain 以来,他基本上每年都得重塑公司,对吧?先是 LangChain,然后是 LangGraph,现在是 Deep Agents。而且我认为他是这方面最灵活、最熟练、最敏锐的人之一。
For sure, right? He's basically had to reinvent the company every year since he started LangChain, right? It was LangChain, LangGraph, and now Deep Agents. And I think he's like one of the most nimble, adept, sharp people about this.
是的。就像现在终于到了稳定的时候。你相信这一点吗?或者你对这种看法有什么看法?
Yeah. It's like now now is finally the time for stability. Do you buy that or what do you kind of make of that take?
嗯。我认为说“这次不一样”有时代价很高,但当你只是在写代码时,实际上尝试做出判断是可以的。而且我认为这个判断是否正确甚至可能不重要。就像我根本不太在乎,因为你可以在一个论点上是对的,但如果你没有找到如何从这个论点中获利,那么谁在乎你是不是第一个说的呢?话虽如此,确实感觉例如我们经历了很多不同的方式来打包与智能体的集成,感觉我们已经落在了技能上,这就像是最小可行格式,只是一个带有一些脚本的 Markdown 文件,我看不出有什么比这更简单的了。因此,围绕 harness 的稳定性是有一些理由的。我觉得在实时元素、子智能体、记忆或任何这些我们称之为智能体工程中的智能体纪律方面,可能会有更多的适应。但如果论点是,好吧,你只是想要智能体是带有工具循环的 LLM,带有文件系统,它们可以用技能进行检索,以及所有这些现在似乎相对共识的标准工具,那么这可能有道理。我只是觉得没有必要把你的声誉押在这个我们已经到达的论点上,因为如果它再次改变,就随它改变。没关系。
Mhm. I think that it's very expensive to say this time is different sometimes, but when you're just writing code, like it's actually okay to just like try to make a call. And I think it may not even matter if this call is right or not. Like I just don't even care that much because you can be right on a thesis, but if you don't figure out how to monetize the thesis, then who cares if you said something first? That said, it does feel like for example we went through a lot of different ways of packaging integrations up with agents, and it feels like we've landed at skills, which is like the minimal viable format, which is just a markdown file with some scripts attached to it, and I don't see how it can be more simple than that. And so there is some justification for the stability around harnesses. I feel like there may be more adaptation with regards to maybe like the real-time elements or sub-agents or memory or any of those like agent disciplines let's call it in agent engineering. But if the thesis is that okay, you just want agents are LLMs with tools in the loop with a file system where they can do retrieval with skills and all these like standard tooling that now seems to be relatively consensus then probably that makes sense. I just think like there's no point trying to stake your reputation on this thesis that we're there because if it changes again, just change with it. It's fine.
是的。你知道,我一直对基础设施公司和应用公司面临的挑战更大这一点感到震惊。显然,我想你知道,在应用方面,你看到了 Sierra 的 Bret Taylor,Lagora 的 Max。他们就像是在说,“看,我们构建的是模型之前的东西,我们愿意每三个月就把一切推倒重来,你知道,随着模型越来越好。”但你至少在那里有一个最终客户,对吧?这有点像粘性很强。你知道,他们大多会留下来,至少会给你一个构建这些东西的机会。我一直觉得在基础设施层每三个月重塑自己更具挑战性,你知道,开发者肯定比会计师事务所或银行更挑剔。所以,必须不断重塑自己肯定是一个更具挑战性的位置。
Yeah. That's always you know, I've always been struck by how that is much more challenging for infrastructure companies and application companies. Like obviously, I think you know, on the application side you've seen, you know, Bret Taylor from Sierra, Max seems from Lagora. Like they're like, "Look, we build you know, what's ahead of the models and we're willing to throw everything out every 3 months, you know, as the models get better and better." But the thing you at least have there is you have an end customer, right? That's like decently sticky. You know, they will mostly stick you know, they'll give you a shot at least of building these things. What I've always found more challenging at the kind of like, you know, reinvent yourself every 3 months of the infrastructure layer it's like, you know, developers are definitely a pickier audience maybe than an accounting firm or a bank. And so it's definitely a more challenging position to be in to have to constantly reinvent yourself.
是的,是的,当他们转变时,那是非常彻底的。他们会离开去追求热门的新事物,因为我觉得没有防御性。就像即使你是一个数据库,人们也可以将工作负载从数据库迁移出去。这是众所周知的事情。所以我认为基本上我们谈论的是 AI 创业公司中垂直与水平的争论。我对此的看法也是,当你像 Lagora 一样,当你是一座桥梁时,你就是外包的 AI 团队,对吧?你的工作是应用任何最先进的 AI 方法。是的,就像在模型能力和最终客户之间有一个翻译层。而且,如果他们不找你,他们就得内部招聘,但他们不会内部招聘,所以他们找你。
Yeah, yeah, and like when they turn it's like very complete. Like they'll leave to the hot new thing because there's like no defensibility, I guess. Like even if you are a database, like people can migrate workloads off databases. Like it's a known thing. So I think like basically what we're talking about is the vertical versus horizontal debate in AI startups. And the way I think about it also is just that like when you're Lagora, when you are a bridge, like you are the outsourced AI team, right? Your job is to apply whatever state-of-the-art AI methods. Yeah, like there's this translation layer between model capabilities and your end customers. And like, well, if they didn't have you, they would have to hire in-house, and they're not going to hire in-house, so they have you.
而且我觉得,这是一个合理且稳健的做法,无论工程层面出现什么趋势和发现。我确实认为有一些有用的横向公司在被建立,但它们都是经典云在 AI 时代的重塑。最主要的就是沙盒,这其实是另一种形式的算力。大家别太兴奋了。但工作量是巨大的。
And like, I think that's a reasonable, robust approach regardless of trends and discoveries in the engineering layer. I do think there are useful horizontal companies being built, but they're all reinventions of classic cloud in the AI era. The primary one being sandboxes, which is another form of compute. Let's not get too excited about it. But the workloads are enormous.
是的,这很有趣。作为其中的一部分,人们围绕基础设施提出的问题包括:公司应该拥有多大的 AI 团队,以及应该内部做些什么。有人问是否应该训练自己的模型,或者基于自己的数据内部进行强化学习。我觉得每三个月就得更新一次看法,但你今天怎么看?
Yeah. It's interesting, and as part of this, the questions folks are asking around infrastructure include the extent to which companies should have their own AI teams and what they should do in-house. There are questions about whether people should train their own models, or do RL in-house based on their data. I feel like one has to evolve their takes on this every three months, but where are you at on this today?
实际上,自有模型的数量上升了。显然我参与了 Cognition,Cursor 也在大量训练自有模型。我认为这是我所说的“智能体实验室剧本”的一部分:你从大实验室的最先进模型开始,针对你的领域进行专业化,但一旦你有足够的工作负载和高质量的用户数据,就可以训练自己的模型,节省成本和延迟。你还能获得营销红利,给它起个花哨的名字并发布研究。
Actually, own models have gone up. Obviously I'm involved in Cognition, and Cursor is also doing a lot of own model training. I think that's part of what I've been calling the agent lab playbook: you start with state-of-the-art models from the big labs, specialize for your domain, but once you have enough workload and high-quality data from your users, you can train your own models and save on cost and latency. You also get a marketing bonus of calling it a fancy name and putting out research.
从我的角度看,我分不清其中有多少是提供给最终用户的实际价值,有多少是营销红利。似乎是两者的结合。
From my seat, I can't tell how much of it is actual value provided to the end user versus the marketing bonus. It seems some combination of both.
我认为两者都有。是的。不,确实有实际价值。原因有很多。第一,即使没有补贴,人们也会把它选为前四或前五的模型。这包括 Composer 2,我们有 1.6,是前五的模型之一。在公平市场、自由市场中。
I think it's both. Yeah. No, there is real value. For a number of reasons. One, even when it's not subsidized, people do choose it as one of the top four or five. This is both Composer 2 and we have 1.6, one of top five models. In a fair market, in a free market.
是的。
Yeah.
在自由市场中,人们确实会选择它,而且没有补贴。这是最好的情况。但除此之外,领域特定模型,例如两家公司的搜索模型,绝对非常有意义。每个人都说我们应该一直这样做。老实说,我认为这方面的基础设施正变得越来越容易,比如 Thinking Machines、Tinker Thing,以及 Prime Intellect 的实验室产品。
In a free market, people do choose it, and it's not subsidized. That's as good as it gets. But beyond that, domain-specific models, for example for search with both companies, absolutely make a ton of sense. Everyone says we should always do this. And honestly, I think the infrastructure for that is becoming easier with Thinking Machines, Tinker Thing, as well as Prime Intellect's lab stuff.
是的,这是“苦涩教训”的逆转之一:你首先在大规模通用模型上启动,以做大。一旦你有了定义明确、数量大但方差小的工作负载,你就提炼出一个更小的模型,自己运行。这完全合理。我不太清楚的是 DIY 强化学习的用例,我认为这主要围绕改进不同事物的质量。显然,可能有更高效的方法来获得一个更快更便宜的小模型。有趣的是,就像两三年前,那些预训练并声称在各自领域有更好结果的公司,随着每次模型迭代改进而被淘汰。我想知道类似的故事是否会在强化学习领域重演。
Yeah, this is one of those reversals of the bitter lesson where you first bootstrap on large general-purpose models to get big. As you get well-defined workloads that are high quantity but not high variance, you distill down to a smaller model and run that on your own. That makes total sense. What I'm less clear on is the DIY RL use case, which I think is mostly around improved quality for different things. Obviously, there are probably more efficient ways to get a smaller model that's faster and cheaper. It'll be interesting to see whether, like two or three years ago, companies that were pre-training and claiming better outcomes in their domains got cooked as each model iteration improved. I wonder if a similar story plays out in the RL space.
是的,如果专注于纯粹的结果和质量,而不是成本方面,那么显然自有模型在规模上降低成本非常有意义。我认为这是同一枚硬币的两面。你基本上总是希望保持质量不变,或者用一点质量换取大幅降低成本,这对每个人都是如此。
Yeah, for the focus on pure outcomes and quality, not the cost side, which clearly your own models for cost at scale makes a ton of sense. I think there are two sides of the same coin. You basically always want to hold quality constant or trade off a little quality for a drastic decrease in cost, and that's true for everyone.
我想提一个非常支持开放模型的元素:定制芯片。这包括 Cerebras,还有 Talas,以及介于两者之间的各种产品。过去一年这已经成为一个大新闻:所有非英伟达的硬件都在被抬高,包括 Matt X,这让我非常欣慰。但我认为,随着替代硬件数量的增加,你能获得的推理速度高得惊人——我们说的是每秒数千个 token,而不是不到 100 个。所以质量上的权衡不再那么重要,因为速度太快了。
One element I wanted to bring out, which is very much in favor of open models, is custom chips. This would be Cerebras but also Talas, and there's a huge range of stuff in between. This has been a huge story this past year: everything non-Nvidia is getting bid up, including Matt X, which is very rewarding for me. But I think one of those things where suddenly, because the number of alternative hardware is increasing, the inference you can get is insanely high—we're talking thousands of tokens per second instead of less than 100. So the trade-off for quality doesn't hold as much anymore because the speed is so high.
你看到很多公司全力投入替代芯片吗?
Have you seen a lot of companies go all in on the alternative chips?
Cognition 在 Cerebras 上投入了,OpenAI 也是。所以除此之外,我不认为有很多公司。这主要是因为,显然,是的。我以前是个怀疑者:好吧,如果我的推理速度从每秒 100 个 token 提升到 200 个,那又怎样?只是快了两倍,没什么大不了的。但我认为每 10 倍都会解锁不同的使用模式,我们在 Talas 和其他一些产品上证明了可以大幅提高推理速度。之后会发生什么我甚至不知道。当整个应用突然出现时,很难预测。而且,这不贵吗?所以这是那种我认为投资周期会持续多年的东西,我会提醒人们不要过早否定它。
Cognition has on Cerebras, and so has OpenAI. So no, I don't think so beyond that. That's mostly because, clearly, yeah. I used to be a skeptic: okay, so what if I get my inference at 100 tokens per second sped up to 200? It's only 2x faster, not a big deal. But I think every 10x does unlock a different usage pattern, and we have proof in Talas and some others that you can drastically improve inference speed. What happens from there I don't even really know. It's so hard to predict when entire applications just appear at once. And also, isn't that expensive? So this is one of those things where I think the investment cycle is going to be multi-year, and I would caution people not to dismiss it too quickly.
是的。我想听听你对一个基础设施问题的看法:似乎越来越多前沿基础设施公司正在为智能体构建产品,让智能体购买或使用,对吧?
Yeah. One of the infra questions I was curious to get your thoughts on: it seems increasingly a lot of the cutting-edge infra companies are building for agents to buy or use their product, right?
哦。另一个大主题。
Ooh. Another huge theme.
是的,是的。我想弄清楚向智能体销售需要做哪些不同的事情。它们只是终极理性的开发者,还是说……
Yeah, yeah. And I'm trying to figure out what you have to do differently about selling into agents. Are they just the ultimate rational developers, or is there...
不,绝对不是。我认为它们很容易被提示注入,而且非常倾向于强化现有的赢家。所以如果你在 2023 年之前幸运地进入了训练数据,那么你在可预见的未来就被安装在那里了。但 Vercel 的 CTO Malte Ubl 在我的会议上透露了一个数据:现在 Vercel 用于配置应用的管理架构中,60% 的流量是机器人,不是人类。所以你的主要客户现在是智能体,主要是编码智能体,主要是使用 CLI 的人,等等。
No, absolutely not. I think they are easily prompt injected and very attuned towards basically compounding existing winners. So if you won the lottery for getting into the training data before 2023, now you're installed there for the foreseeable future. But one stat that Vercel CTO Malte Ubl dropped at my conference: there are now 60% of traffic to Vercel's admin app architecture for configuring Vercel applications is bots. It's not human. So your primary customer is agents now, mostly coding agents, mostly people using a CLI on CP, whatever.
但没错,我的意思是,我认为第一步是,如果它不存在一个智能体可以使用的 API,那它就不存在,对吧?我认为这无论如何都是一种良好的卫生习惯,让所有东西都通过 API 可用,但这不是对产品人员的额外推动,让他们不只专注于 UI。你可能应该专注于 CLI 相关的东西。除此之外,我认为老实说,有……我来自这样的观念:你现在为智能体体验所做的一切,也就是 Netlify 的 Matt Billman 试图创造的那个术语,和你应该一直为开发者体验所做的事情是一样的。你应该有好的文档。你应该有一个一致的、大多无状态的 API。你应该有可发现性或渐进式披露或搜索之类的。所以,现在人们有精力去找到这些客户来做这件事,这很好。我是否相信要超越这个,进入像 AEO 这样的领域来操纵聊天机器人?不一定,但显然,那些找到短期胜利的人会有巨大的优势。而短期胜利可以复利。
But yeah, I mean, I think step one, if it doesn't exist as an API that agents can use, it doesn't exist, right? Which I think is a good hygiene thing anyway, to make everything API available, but not as an extra push on product people to not only work on the UI. You should probably work on the CLI stuff. Beyond that, I think honestly, there is... So, I come from the sensibility that everything you are trying to do for agent experience now, which is the term that Matt Billman at Netlify is trying to coin, is the same thing that you should have been doing for developer experience. Like you should have had good docs. You should have had a consistent API that is mostly stateless. You should have discoverable or progressive disclosure or search or whatever. And so, now that people have energy in finding these customers to do that, it's great. Do I believe in extending beyond that into something like AEO for gaming the chatbots? Not necessarily, but obviously there's going to be huge advantages from people who figure out the short-term wins. And short-term wins can compound.
你是否认为这些复利优势,就像预训练数据截止日期对公司的优势一样,你知道,显然在一段时间内,我想这不会持续。所以当你思考,我不知道,三四年后,最终的选择标准会是什么,你认为它是否仍然完全反映你之前所说的,就像你一直应该做的,向开发者销售好产品?
Do you view these compounding advantages like the pre-training data cut-off companies like, you know, obviously over some period of time I imagine that doesn't persist. And so as you think about, I don't know, 3 or 4 years from now, what the selection criteria end up being, do you think it still mirrors exactly what you were saying before, like it's exactly what you should have been doing all along to sell a good product to developers?
可能是这样,但我觉得三四年后,我们可能会有更好的记忆和个性化。所以那时一般的 AEO 或 GEO 就不那么重要了。所以我认为我们最终采用的任何记忆或个性化系统,可能比现在的情况更能决定你最终的选择,现在的情况只是提及频率,我们这么说吧。
It could be, except that I think in 3 or 4 years, we'll probably have much better memory and personalization. So then general AEO or GEO doesn't really matter as much. So I think whatever memory or personalization system we end up with will probably determine what you end up choosing much more than what is currently the case, which is just frequency mentions, let's call it.
是的,是的。所以你就是刷数量。
Yeah, yeah. So you just spam quantity.
而这是我期待的事情。我确实认为,你自己需要解决的基本练习是,如果你现在开始一家新的颠覆性公司,有一个大家都知道的大玩家。比如如果你想开始像新的 Supabase,你会如何与他们竞争?我不一定有答案,但我确实认为像 Resend 这样的公司,相对较新。我想他们是 2023 年成立的,最近有一项调查,人们检查 Claude 默认推荐什么,如果你不提示它任何东西,只说给我一个邮件提供商,它在 70% 的情况下会说 Resend。你能在这么短的存在时间内进入那里,我认为是令人鼓舞的。
And I think that's something I'm looking forward to. I do think that the fundamental exercise to work through for yourself is if you start a new sort of disruptor company now, there's a big incumbent that everyone knows. Like if you want to start like new Supabase, how would you compete with them? And I don't necessarily have the answer, but I do think people like Resend, relatively new. I think they were started like 2023, and there was a recent survey where people checked what Claude recommends by default if you just don't prompt it with anything, just say give me an email provider, and it says Resend in like 70% of cases. The fact that you can get in there with such a relatively short existence, I think is encouraging.
是的。我确实认为你想做任何事来进入那些非常简短的提及,因为不会只有 20 个。会是大约三个。这肯定感觉比以往任何时候都更整合,或者比过去市场推广的物理可能实现的更赢家通吃。
Yeah. I do think you do want to do whatever it is to get in those very short mentions, because it's not going to be 20 of them. It's going to be like three. That definitely feels like more consolidation than ever, or a winner-take-most market than maybe the physics of go-to-market in the past might have enabled.
另一件事是语义关联将非常重要,在某种意义上,你想做组合文章,比如“用我的东西搭配销售”之类的,所有这些都会被语料库捕捉到。所以这可能是你想做好的一件事。我不知道还有什么。这是那种我觉得自己落后的事情。我不知道你怎么看,但我认为 AI 就是每个人都不断觉得自己落后。
The other thing also is that semantic association is going to be very important, in a sense that you want to do the combo articles where you're like use my thing with for sale with blah blah blah, and that all gets picked up in the corpus. So that's probably one thing that you want to do well. I don't know what else. It's one of those things where I feel I'm behind. I don't know how you feel about this, but I think AI is just everyone constantly feeling like they're behind.
是的,我想见见那个不觉得落后的人。
Yeah, I want to meet the person that doesn't feel behind.
但就像 AXE 一样,对吧?所以我的立场就是我之前说的。你应该为智能体做的一切,都是你本来就应该为人类做的。所以如果你只是获得更多精力为智能体做事,那很好。但很难说清楚除了更多你应该做的垃圾内容之外,还有什么新东西。那是我现在的看法。我确实认为这会有更多的转折。我认为即将到来的个性化转折会很大,我不知道那会是什么样子,因为基本上我们在记忆方面感觉有点枯竭了。
But like with AXE, right? So my stance was exactly what I said before. Everything that you should do for agents is something that you should have done for humans anyway. And so to the extent that you're just getting more energy to do things for agents, great. But it's hard to articulate what new thing apart from just more spam you should be doing anyway. That would be my take right now. I do think there will be more turns at this. I think the personalization turn that is coming will be big, and I don't know what that looks like because basically we feel kind of tapped out on the memory side of things.
是的。我想自从我们上次聊天以来,你在 Cognition 担任了这个角色,你显然对今天的 AI 编码领域有前排座位。我觉得编码在很多方面,人们把它看作,我的意思是除了是所有市场之母和这个巨大的机会之外,我认为它有点像许多其他领域未来的预览。两者,你知道,我觉得智能体在编码方面最先进。我也觉得基础模型和应用公司之间的竞争反映了我们可能在其他领域看到的。所以,也许对我们的听众来说,你能概述一下今天 AI 编码战争的状况吗?
Yeah. I guess since we last chatted, you took this role over at Cognition, and you obviously have a front row seat to the AI coding space today. I feel like coding in many ways, people view it as this, I mean besides being the mother of all markets and this massive opportunity, I think it's kind of a preview of what's to come for many other spaces. Both, you know, I feel like agents are most advanced in coding. I also feel like the competition between foundation models and application companies mirrors what we may see in other spaces. And so, maybe for our listeners, can you just lay out what is the state of the AI coding wars today?
嗯,它非常庞大,对吧?我不认为上次我们谈论这个时,我们意识到了它的规模。
Um, it is massive, right? And I don't think necessarily last time we talked about this, we appreciated the size of it.
不,我希望我们意识到了。
No, I wish we did.
嗯,今天 AI 编码战争的状况,OpenAI 和 Anthropic 都把编码作为他们的 P0 来竞争。Anthropic 仅从 Claude Code 就达到了大约 25 亿美元的 ARR。他们确认 ARR 的方式还有争议。OpenAI,我不认为公开数字是已知的,但我们也说大约 20 亿。然后 Cursor 据传是 20 亿。这些是已知的公开数字。所以过去一年刚刚创造出的巨大市场。比如 Anthropic 最近刚刚庆祝了 Claude Code 的一周年。
Um, the state of AI coding wars today, both OpenAI and Anthropic have made it their P zeros to compete in coding. Anthropic is at like 2.5 billion in ARR just from Claude Code. The way they recognize ARR is up for debate. OpenAI, I don't think a public number is known, but let's call it 2 billion as well. And then Cursor is rumored to be 2 billion. Those are the public numbers that are known. So huge markets that have just been created in the past 1 year. Like Anthropic just celebrated Claude Code's 1-year anniversary recently.
是的,相当不错。
Yeah, pretty neat.
嗯,所以然后我认为我看到的另一件事是,有些其他人会说,“哦,这是 Claude 用例的相对渗透率,对吧?就像编码占 50%,然后法律、健康之类的占剩下的。”有一条非常受欢迎的推文说,“好吧,看看所有这些其他用例中的空白空间。”
Um, so then I think the other thing that I see is there are some other people who are like, "Oh, here's the relative penetration of Claude use cases, right? And it's like coding 50% and then legal whatever health is like the remaining ones." And there was a very popular tweet that was like, "Okay, well, look at the empty space in all these other use cases."
如果你今天是一位新创始人,你应该押注其他东西,因为按照追赶理论来说。
If you are a new founder today, you should be betting on the other stuff because on a catch-up theory.
而我的反驳和当年对苹果对谷歌的反驳一样,就是:“嗯,为什么这次不一样?比如,如果过去一年它从 10% 涨到 50%,为什么不能继续涨?” 而搞错这一点其实非常痛苦,因为你本可以押注动量,而不是均值回归。
And my pushback is the same pushback that I had on Apple versus Google, which is like, "Well, why is this time different? Like why if it went from, let's say, 10 to 50% in the past year, why can't it keep going?" And getting that wrong is actually a very painful one because you could have just did the momentum bet instead of the mean reversion bet.
所以,我认为现在的状况是,人们非常陷入这种狂热。他们因为花得更多而不是更少而获得回报,我认为我们不在效率阶段,而是在一种能力探索阶段。所以我认为那些更疯狂、更有创造力的人会相对获得更多回报。
And so, I think that is the state of things now that people are very much into the psychosis. They're getting rewarded for spending more rather than spending less, and I think we're not in that phase of efficiency. We're in a phase of sort of like capability exploration. So I think people who are more crazy, who are more creative, get rewarded comparatively.
嗯,这很有意思。我的意思是,感觉在这些 token 最大化排行榜背后,从劳动力角度来看,这是这个转型的第一阶段,你只需要向你的雇主展示:“嘿,我用了这些工具。这是我消耗的 token 数量。” 就这样。他们现在不在乎质量。这可能让那些在乎手艺的人反感。但从方向上看,每个人都只希望你往上走,不管怎样。所以这并不挑剔,可能也很粗糙,但我认为总体上是好的,因为我们可能仍然没有充分利用 AI。所以我觉得这非常有趣。
Well, that's interesting. I mean, it feels like behind these token-maxing leaderboards and whatnot, is this the first phase of this transition from a workforce perspective is you just got to show your employer like, "Hey, I use these tools. Here's my number of tokens I cost." And that's it. They don't care about the quality right now. It is maybe distasteful to someone who cares about the craft and all that. But directionally everyone just wants you to go up regardless. And so there is not very discerning and it's probably very sloppy, but I think it's net fine because we're still probably underusing AI just in general. And so I think that's like very interesting.
比如我们播客请过 OpenAI 的 Ryan LePopolo,他每天消耗十亿 token。是的。对在家算的听众来说,这大约价值 1 万美元。如果按市场价计算,每天 1 万美元的 API token。而我们大多数人都负担不起。
Like we had on the podcast Ryan LePopolo from OpenAI who spends a billion tokens a day. Yeah. And that's for those counting at home, it's like something like $10,000 worth. $10,000 worth a day of API tokens if they did market rates. And like most of us can't afford that.
但可能他做的很多事都是垃圾。对吧。但他会这样,他说如果有一种新能力,他会比你先发现。因为他在尝试,而你没有。对吧。而且你只做有效的事情。那很好,但那些会发现下一个热门事物的人正生活在边缘。
But like, and probably a lot of what he does is slop. Right. But like he's going to this, he's like if there were a new capability, he would discover it first before you. Because he was trying and you were not trying. Right. And like you only do things that work. Well, good for you, but like the people who are going to discover the next hot thing are living at the edge.
对。而且越来越生活在边缘,就是有算力预算来运行这些实验。我的意思是,有点像研究方面一直以来的边缘状态。你知道,它在很多方面受到你运行这些实验所需算力的限制。现在在构建者或实际使用这些工具方面,感觉也类似。
Right. And increasingly living at the edge is just having the compute budget to run these experiments. I mean, kind of similar to what living at the edge on the research side has always been. You know, it was constrained in many ways by the amount of compute you had to run these experiments. It feels similarly on the almost on the builder or like actualizing these tools now.
另一件非常明显的事情是,Anthropic 有点像高价位高端玩家,限制额度甚至限制模型发布就是游戏规则,而 Codex 则是:“来吧,伙计们,用我们的 SDK,用我们的登录,我们不在乎,我们会重置限制,随便。” 你确实想尽可能利用补贴,而 Codex 现在绝对是被高度补贴的。Gemini 也被高度补贴。相比之下,我认为你应该趁这个机会大赚一笔。作为能力探索者,仅仅使用 Claude Code 或 OpenAI 每月 200 美元的计划,也没那么糟糕。而且我的感觉是,人们甚至还没到那一步。
The other thing that's very obvious is Anthropic is kind of like the high-price premium player where restricting limits or restricting model releases even is like the name of the game, whereas Codex is like, "Come on in guys, use our SDK, use our login, we don't care, we're going to reset limits, whatever." You do want to try to exploit the subsidies where you can get it, and definitely Codex is super subsidized right now. Gemini also very subsidized. And comparatively, I think you should make hay, I guess, while that's going on. It's not that bad to be a capabilities explorer on just the $200 a month plan from Claude Code or from OpenAI. And my sense is that people aren't even there yet.
你认为这个市场最终会如何发展?我的意思是,这显然是一个巨大的市场,任何一块蛋糕对追求它的人来说都很有趣。但我认为,编码市场之所以特别有趣,是因为它感觉像是其他应用市场未来发展的预兆,基础模型最终会转向这些市场,并用所有模型对抗它们,围绕它们收集数据。那么,你认为最终会有很多不同类型的玩家的空间吗?或者,你认为这个市场的最终状态是什么?这适用于其他市场吗?
How do you think this market ultimately plays? I mean, it's obviously such a big market that any slice of that market is interesting for anyone going after it. But I think what makes people so interesting in the coding market particularly is it feels like it's kind of this foreshadowing of what will happen in other, you know, any other kind of application market that the foundation models eventually turn to and are all their models against and gather data around. And so, how do you think, you know, does there end up being room for lots of different kinds of players? Or like, what do you think the end state of this market is? And is that applicable to other markets?
我觉得会有,我的意思是,现状可能是最可能的结果,即有两个大玩家,还有一小部分长尾玩家,他们适合两个大玩家不覆盖的其他用例。这对我来说感觉是对的。我认为,要让市场结构发生重大变化,需要看到经济、品牌建设或相关公司价值主张等方面的变化,而在过去 6 个月里,我还没有看到任何实质性改变这些故事的东西。所以,我觉得他们会继续下去,直到发生其他事情。
I feel like there will be, I mean, status quo is probably the most likely outcome, which is there're two big players and there's a small range of longer tail people that fit other use cases that the two big players don't. That feels right to me. I think that for it to for the market structure to significantly change, there would need to be seen even change in like the economics or like the brand building or like the value propositions of the companies involved, and I haven't seen any in the last 6 months that have really changed the stories materially. So, I feel like they would just keep going until something else happens.
其他事情发生,意思是像微软醒过来,然后说:“伙计们,我们有 GitHub。我们有……”
Something else happens meaning like Microsoft wakes up and goes like, "Guys, we have GitHub. We have..."
你知道,我们会在这里做比 Copilot 大得多的事情。那将是一个巨大的变化。MSL 现在推出了一个模型,我和 Alex Wang 吃过早餐,他们说:“是的,我们真的很想进军编码用例。” 他们还没有做任何事,但不要低估他们,对吧?同样,对于中国实验室,我认为他们也在尝试。比如 ZAI 在做事情,GLM……ZAI 和 GLM 是同一家公司。所以每个人都在试图分一杯羹。我觉得现状在过去差不多一年里一直相当稳定,我会这么说。
You know, we'll do something much bigger here than just Copilot. And that will be a big change. MSL has put out a model now, and I was in a breakfast with Alex Wang where they were like, "Yeah, we really want to go after the coding use case." They haven't done anything yet, but don't underestimate them, right? And similarly for the Chinese labs, I think they're trying to go after it. Like ZAI is doing stuff, GLM... ZAI and GLM are the same thing. And so everyone's trying to get a piece of that pie. I feel like the status quo has been pretty stable for the past almost a year, I will say.
是的。那么,对于更偏向企业端的应用公司,还有空间吗?或者说,模型公司为应用公司留下了哪些服务领域?
Yeah. And is there room for the application companies more on the enterprise side or like where do the model companies leave for application companies?
是的,这是个好问题。这还在不断演变。我要说的是,因为 OpenAI 一年前并没有对编码给予如此高的关注,我们没有那么多历史,对吧?而且看起来,比如 OpenAI 现在的大推动是超级应用。这是消费级的事情吗?这是产品组合合理化的事情吗?在他们确实想投入更多编码的时候,这会在多大程度上分散对编码的注意力?我认为这很不清楚。所以,我确实认为在两大实验室都有这些……抱歉,在 OpenAI 和 Anthropic……而 minus 和 NXI 是独立的案例。他们正在尝试看到其他时间扩展领域。所以,Claude Code 用于金融。是的。Claude 协同工作,所有这些。而我认为 Cursor 和 Cognition 相对而言只专注于编码。所以,我确实认为他们留下了空间。
Yeah, that's a good one. It's very much evolving. I will say because OpenAI did not have this level of attention on coding a year ago, we just don't have that much history, right? And it seems like for example, the big push that OpenAI now is the super app. Is that a consumer thing? Is that like a product portfolio rationalization thing? How much is that going to take away attention from coding at the time when they actually do want to put more coding? I think it's very unclear. So, I do think there's all these like at both big labs, there's... sorry, at both OpenAI and Anthropic... and the minus and NXI are separate cases. They are trying to see the other time extension areas. So, Claude Code for finance. Yep. Claude co-work, all those things. Whereas, I think Cursor and Cognition are like comparatively just focused on coding. And so, I do think they leave space.
而且我确实认为对其他垂直领域来说也是如此,对吧?他们不会那么专注于那个领域。不过我觉得我会把金融和医疗保健列为他们接下来明确要攻的方向。我认为相比之下医疗保健似乎更棘手。虽然有一些相关公告,但我更看好金融业务,因为变现路径清晰得多。
And I do think for the other verticals that also means the same thing, right? That they're not going to be that intensely focused on that domain. Except for I think I would mark out finance and health care as the next ones that they're clearly going after. I would say comparatively health care seems more thorny. There've been some announcements about it, but I would respect the finance work a lot more just because the path to money is a lot clearer.
不,我的意思很明显,我觉得可能与其他领域留下的空间类似,在企业中实际部署这些工具显然需要很多工作,而不是仅仅开箱即用地给人们提供模型访问权限。
No, I mean obviously, I think maybe similar to the space that's being left in these other domains, there's obviously a lot that's required to actually implement these tools in enterprises versus maybe just giving model access to folks out of the box.
是的,是的,是的。所以智能体实验室的做法是“我们为你完成最后一公里”,而我认为模型实验室往往只是信任模型本身,采取极简方式。两者都有效。我不认为一种方式在所有用例中都优于另一种。我只知道大型企业似乎确实想要一个专门的合作伙伴,而不仅仅是模型实验室,这有点意思。
Yeah, yeah, yeah. So the agent lab thing is like we'll do the last mile for you, whereas I think the model labs tend to just trust the model and be minimalist about it. Both of them work. I don't necessarily think one beats the other for every use case. All I do know is that it does seem like the large enterprises do want a dedicated partner that isn't just the model labs, which is kind of interesting.
我们一直处于纯粹的能力探索阶段,所以我认为这对大型实验室来说再好不过了,对吧?我的意思是,他们总是处于能力探索的前沿,所以我认为他们与许多企业关系良好。但最终,随着时间推移,这些实验室的激励机制总是会朝着为最终客户最大化 token 消耗的方向发展,而且我认为真正达到大规模的公司少之又少。也许编码领域又是最有趣的,因为那是第一个真正完全消失的空间,你知道吗,你每天都得面对它。简直疯狂。
We've been in this phase of pure capability exploration and so I think nothing has been better for large labs, right? I mean they're always going to be at the frontier of capability exploration and so I think they have a very good relationship with a lot of these enterprises. But ultimately over time the incentive structure of these labs is always going to be maximal token consumption for the end customers they work with, and there's just I think so few companies that have actually gotten to massive scale. Maybe coding again is the most interesting since the first space that really is just completely gone, you know, yeah, you must live it every day. Like absolutely insane.
而且我认为当你达到甚至……好吧,我的意思是我们对 Cruise Automation 评价不错,但 Inflection 和 OpenAI 的爆发式增长,因为它们有独立的估值。我的意思是,把 xAI 也算进去,IP 估值达到 1.2 万亿。
And I think when you get even okay, I mean like I think we say good things about Cruise Automation, but the sheer lift-off of both Inflection and OpenAI because they have independent evaluations. I mean let's throw in xAI in there, IP wing at 1.2 trillion.
这个数字简直令人难以置信。我觉得在正常的投资或创业中,会有一个市值或估值的上限,你达到后就会想,好吧,从现在开始就稳定了。而这些家伙并没有放慢脚步。
That number is just mind-boggling. I feel like in normal investing or normal startups, there's kind of a ceiling market cap or valuation that you reach and you're going like, all right, it's going to be chill there from now on. And these guys are not slowing down.
不。但我也认为,这些后期公司中一个引人入胜的动态是,在过去,我觉得在风投界,如果你达到一定规模,围绕你的问题更多是估值问题。这就是为什么有不同类型的风投人士。比如,后期增长型的人非常擅长,你知道,既关注公司的最终市场机会,也关注如何正确估值。我们知道它处于某个结果区间,当然有一些波动,但那个区间相对明确,然后可能随着时间推移,会有上行惊喜。
No. But also I think the dynamic that's fascinating about some of these later stage companies is, in the past I feel like in venture world if you got to a certain level of scale, the question around you was really more of a valuation question. This is why there were different types of venture people. Like, the late stage growth people were just incredible at, you know, a little bit of what's the ultimate market opportunity of this company, but also what's the right way to value it. Like, we know it's in some bands of an outcome that is like, sure there's some variance to it, but it's relatively understood what that band is, and then maybe you get over time surprised to the upside.
然而,即使是实验室本身,任何后期公司,其当前价值区间,甚至一两年后的价值区间,都如此巨大,因为生态系统变化太快,以至于即使是后期公司,每 3 个月都可能发生一个关乎存亡的事件,无论是上行还是下行。
Whereas, any kind of even the labs themselves, any later stage company, the bands of which that company might be worth right now, even in a year or two years, are so massive because of how fast the ecosystem changes that it's like, even for later stage companies, every 3 months could be an existential level event to the upside, to the downside.
而且我认为你在编码领域显然看到了积极的一面,如果你想想像 Anthropic 这样的公司,有一段时间不清楚他们是否能获得足够的资本来真正留在竞争中,对吧?然后编码恰好适时爆发,他们拥有完美的模型,执行得非常出色,现在我们是世界顶级公司之一。
And I think that you're obviously seeing it in the positive with code, which, if you think about a company like Anthropic, for a while it was unclear if they were going to have access to enough capital to really stay in the race, right? And then coding hit at the exact right time, they had the perfect model for it, they executed brilliantly, and now we're one of the top companies in the world.
与此同时,我对 OpenAI 毫无同情,因为他们做得很好,而且他们都很富有。这就像是一个高级香槟问题,但在编码或其他方面排名第二。谁在乎呢?你们做得很好。
At the same time, I have zero sympathy for OpenAI because they're crushing it, and they're all rich. This is like a high class champagne problem to have, but to be number two at coding or whatever. Like, who cares? You're doing great.
是的。不过有趣的是,我甚至不能……我的意思是,你更接近这个领域,因为你身处 AI 编码领域,但很多和我交谈的人认为 Codex 即使不比 Claude Code 好,也至少一样好,对吧?我认为让我非常惊讶的一点,也许 Claude Code 在某些方面是更好的产品,我很好奇你的想法,就是在消费者 AI 领域,ChatGPT 展现了巨大的先发优势,对吧?诚然,今天,Claude、Gemini 都是好产品。我不确定 ChatGPT 是否明显更好,但人们坚持使用 ChatGPT。那是第一个向他们介绍的产品。
Yeah. It's funny though, I can't even, I mean, you would be closer to this given that you're in the AI coding space, but it's like a lot of people I talk to think Codex is just as good if not better than Claude Code, right? I think one thing that I've been really surprised by, and maybe Claude Code is a better product in some ways, I'm curious your thoughts, is just in consumer AI with ChatGPT, you saw this big first mover advantage, right? Where admittedly today, like I don't know, Claude, Gemini, great products. Not sure it's abundantly clear ChatGPT's any better, but people stick with ChatGPT. It's the first thing that introduced them.
他们表示不再增长了。我不知道你是否看到了,但对我来说,这更多是一个产品问题,而不是他们失去了市场份额。我的理解是,当今消费者 AI 的整体问题更多是如何将这个工具,对于我们这样的知识工作者来说,它是一个不可思议的神奇工具,但对当今世界上的许多人来说,它不一定是日常活跃使用的工具。
They state that they're not growing anymore. I don't know if you've seen that, but that to me is more of a product problem than it is, like, they've lost share to someone else. My understanding is the overall problem with consumer AI today is much more of a how do you take this tool and, for folks like us, knowledge workers, it's like this incredible magic tool, but it's not necessarily a daily active use tool for a lot of people around the world today.
而且产品方面,这有点像一个类别范围的问题。比如在编码领域,整个空间呈抛物线式增长。其他消费者 AI 玩家可能有一些相对增长,但并不是说消费者 AI 作为一个类别呈抛物线式增长,而他们没有抓住大部分。我认为更大的问题是,嘿,这个类别有点达到了平台期。人们还没有弄清楚如何吸引更多用户或提高用户使用频率。所以这似乎更多是一个类别范围的问题,而不是巨大的市场份额变化。
And what are the product it's kind of a category-wide problem. Like in coding, for example, the entire space has gone parabolic. There may be some relative growth in other consumer AI players, but it's not like consumer AI as a category is going parabolic and they're not capturing most of that. I think the larger problem is much more hey, the category has kind of hit a bit of a plateau. People haven't figured out how to bring tons more users on board or increase the frequency of those users. And so it seems more of a category-wide problem than it is a massive market share change.
我本来想与编码领域做比较,Claude Code 显然是第一个向人们介绍这种神奇体验的产品。你知道,据各方面说,Codex 即使不更好,也相当接近。但那个第一个产品,你本以为它不会是一个超级粘性的产品表面,但实际上它确实如此,结果我发现,第一个向你介绍某种体验的实验室确实能保持很多关注。
I was going to draw the comparison to the coding space where Claude Code was the first product obviously to introduce people to this magical experience. You know, by all accounts, Codex is pretty damn close to as good, if not better. But still that first product you would have thought that would not be a super sticky product surface area and it actually has, it turns out I feel like the first lab to introduce you to an experience really does keep a lot of the focus.
我想也许还为时过早,你知道,ChatGPT 已经 3 年多了,而 Claude Code 才 1 年。所以给它点时间吧。
I think maybe it's still early days, you know, ChatGPT is like 3 plus years old and Claude Code is only 1. So just give it time.
是的,我的意思是,确实有很多人从 Codex 切换过来了。也许这种情况会持续下去。真的很难说。我确实认为,因为我们正处于这种高波动、高温度的阶段,对先行者和品类开创者的忠诚度和粘性,可能不像我们在职业生涯中看到的其他一些领域那么高。
Yeah, I mean, definitely a lot of people have switched from Codex. Maybe that will keep going. It's really hard to tell. I do think that because we're in this high volatility, high temperature phase, the loyalty and stickiness to first movers and category creators isn't as high as it might be in some other areas we've looked at in our careers.
是的,不过我对 Claude Code 这件事感到惊讶。我本来以为……在很多方面我一直担心……你觉得它现在应该已经消失了吗?不是消失,但我一直担心这些公司的消费者业务会相当有粘性,而企业 API 业务实际上在某种程度上是你最不忠诚的买家——他们会转向……
Yeah, though I've been surprised by the Claude Code thing. I would have thought that... in many ways I always worried about the... You think it would have been gone by now? Not gone, but I always worried that the consumer business of these companies would be quite sticky, and then the enterprise API business was actually, in some ways, your least loyal buyers—they would move to...
对,对,对。但他们发现那不是企业 API,而是企业产品。
Right, right, right. But they worked out that wasn't the enterprise API; it was enterprise product.
完全同意。也许这就是秘密——在这个领域发生的锁定效应或默认行为的程度,超出了我的想象,尤其是两个产品据各方面来说都相当相似。
Totally. And maybe that was the secret—that the amount of lock-in or just default behavior that has happened in that space is more than I might have imagined, with two products that by all accounts are pretty damn similar.
是的,这没什么好争的。我要说的是,我确实认为 Codex 在个人体验方面仍处于追赶阶段。我从 Codex 中唯一喜欢的是 Spark 和技能集成——我觉得它稍微好一点。我觉得速度也快一点,也许是因为它是用 Rust 写的之类的。这些非常细微的东西,你几乎是在自我暗示,而不是客观地评估两者。我确实觉得在感觉上是这样。整个辩论中缺失的问题是,为什么这仅仅集中在两个名字上,对吧?Gemini 的存在在哪里?xAI 的存在在哪里?他们也在尝试,只是没有取得那么多进展。
Yeah, no fight there. I will say I do think that Codex is still in a catch-up phase in terms of personal experience. The only thing I like out of Codex is Spark and the skills integration—I feel like it's a little bit better. I feel like the speed is a bit better, maybe because it's written in Rust or whatever. Very minor things that you're almost telling yourself rather than objectively assessing between the two. I do think vibes-wise that's going on. The missing question in this whole debate is why is this so concentrated in only two names, right? Where is the Gemini presence? Where is the xAI presence? They are trying; it's just they haven't made that much progress.
但我认为 Claude Code 的时刻确实表明,而且在某些方面让你对其他人追赶的潜力更加乐观,因为确实感觉如果你是第一个引入某种神奇的全新产品体验的人,那实际上可能比人们想象的更有粘性。
But I think the Claude Code moment does show, and it actually in some ways makes you a little more bullish on the potential for someone else to catch up, because it does feel like if you're the first person to introduce some magical net-new product experience, that actually might be stickier than one might have imagined.
对,对,对。好的。是的,所以每个人都可以相信……你认为那个新产品体验可能是什么?比如,这是我想象力的失败。我总是想知道——人们总是说:“嗯,能拯救我们的是成为下一个新事物的先行者。”那是什么呢?
Right, right, right. Okay. Yeah, and so everyone can believe... What do you think that new product experience might be? Like, it's a failure of imagination on my part. I always wonder—people always say, 'Well, the thing that will save us is being first to the next new thing.' What is it?
我不知道。可能是围绕消费者智能体计算机使用,比如混合体。我认为显然我们正在触及消费者层面的表面。所以我目前的理论是,OpenClaw 是未来事物的一个愿景。OpenAI 与 OpenClaw 有关联是好事,但他们绝没有权利赢得它。我一直在追求的一般论点是,就像 2025 年是编码智能体的一年一样,2026 年是编码智能体突破遏制去做其他一切事情的一年。所以编码智能体继续获胜,但因为他们生成软件,而软件吞噬世界,这有点像传递性质:软件吞噬世界,编码智能体吞噬软件,因此编码智能体吞噬世界。这是一个有趣的……
I don't know. Something around consumer agent computer use, like hybrid. I think obviously we're scratching the surface on the consumer side. So my current theory is that OpenClaw is a vision of things to come. And it's good that OpenAI has the association with OpenClaw, but by no means do they have the rights to win it. The general thesis I've been pursuing now is that the same way 2025 was the year of coding agents, 2026 is coding agents breaking containment to do everything else. So coding agents continue to win, but because they generate software and software eats the world, it's kind of like the transitive property: software eats the world, coding agents eat software, therefore coding agents eat the world. Which is an interesting...
是的,但突破遏制在消费者环境中总是比企业环境更容易的阶段。你看到人们在个人生活中运行这些非常酷的实验。我认为弄清楚如何……显然现在每个人都专注于企业方面,如何创造这些体验。我觉得感觉上——人们喜欢这些一切都完全转变的叙事。就像,实际上,OpenAI 在组织上,抛开波动不谈,是伟大的产品、伟大的团队、伟大的模型。世界上其他所有人都被激励着希望再多两三个。每个人都希望有更多伟大的模型公司。所以我觉得当任何一家公司过于成为焦点时,世界的自然力量会反抗,对吧?生态系统中有太多人被激励着不让这种情况发生。所以如果我们在未来 6-12 个月的某个时候没有看到感觉的回摆,我会感到震惊。也许不会完全反过来,但至少会更平等一些。
Yeah, but breaking containment is always an easier phase in the consumer context than the enterprise one. You've seen people run these really cool experiments in their own personal lives. I think figuring out how you... obviously everyone's focused on the enterprise side now around how you create these experiences. I feel like the vibes—people love to have these narratives of everything completely shifting. It's like, actually, OpenAI organizationally, volatility aside, is great products, great team, great models. Everyone else in the world is incentivized for there to be two, three more. Everyone would love more great model companies. So I feel like the natural forces of the world revolt when any one company is too much the star of the show, right? There are so many people in the ecosystem that are incentivized for that not to happen. So I'd be shocked if we don't have a reversion of vibes. Not maybe completely the other way, but at least a little more equal at some point over the next 6-12 months.
我认为只是不同的阶段。当你谈到世界想要更多模型公司时,我想到了 Neo Labs。我的意思是,我不知道,说他们中没有一个在过去一年真正突破,这样说公平吗?
I think there's just different stages. When you talk about the world wanting more model companies, I think about the Neo Labs. And I mean, I don't know, is it fair to say none of them have really broken through in the past year?
我认为这完全公平。这很艰难。
I think that's totally fair. Which is rough.
嗯,那么,我们如何增加选择上的多样性呢?就是这样。
Um, and well, how are we going to grow that diversity in choice? Like, that's it.
是的,看看最终会发生什么会非常有趣。你看到像 Nvidia 这样的公司,非常有动力确保有一个更广泛的其他模型提供商的平台。我认为……
Yeah, it'll be really interesting to see what ends up happening with that. And you've seen folks like Nvidia, very incentivized to make sure there's a broader platform of other model providers. I think...
我不知道。人们这么说,但我认为他们没有努力尝试。Nvidia 更努力地构建新云,而不是新实验室。嗯,他们确实非常努力地构建新云。所以是的。但是,比如,让我们说像 CoreWeave 这样的公司,比建立在它们之上的任何新实验室都更快乐。
I don't know. People say this, but I don't think they tried it hard. Nvidia tries harder to build new clouds than new labs. Well, they try pretty damn hard to build new clouds. So that's yeah. But like, let's call it the CoreWeaves of the world, much happier place than any new lab built on top of them.
是的,尽管有人可能会争辩说,让一个新云成功比……你不能像对待新云那样凭空变出一个新实验室。
Yeah, though one might argue it's easier to enable a new cloud to be successful than it is—you can't will a new lab into existence the same way you can with a new cloud.
是的,所以 Nvidia 对此有更直接的控制,当然。今天在初创企业方面还有什么吸引你的眼球?你是否担心——显然有这种叙事,基础模型宣布一个产品,所有股票都下跌 15%。你担心基础模型会蚕食一堆这些初创企业类别吗?
Yeah, so Nvidia has more direct control over it, for sure. What else is catching your eye today on the startup side? And are you worried—there's obviously this whole narrative of the foundation models announcing a product and every stock goes down 15%. Do you worry about the foundation models just eating into a bunch of these startup categories?
不太担心。我认为实际上,因为有一个……好吧,有一个作为初创企业投资者的观点,还有一个是你是否想创业的观点。我认为老实说,所有这些的下行风险都非常小,因为最坏的情况就是你最终被收购进这些实验室之一。
Not really. I think actually, as there's this... Okay, there's the point of view of being an investor in startups, and there's the point of view of do you want to start something. And I think honestly, the downside for all these is so minimal in the sense that the worst you do is you just get acquired into one of these labs anyway.
所以,我认为对于那些只是做事、尝试、并以称职方式执行的人来说,市场是存在的,即使商业上不成功,即使结果并不那么好,但这也是你进入这些领域的面试。所以,从非常小的初创公司角度来看,我没有这种感觉。中型初创公司,是的。我会说有很多死掉的 LM 基础设施,很多 LM 基础设施整合,比如 Langfuse 这类被 ClickHouse 吸收。我认为人们可能已经弄清楚了特定领域的玩法。我觉得这没问题。是的,我不太担心这个。
So, I think the market for people who just do things and try things and try to execute in a competent way, even if it doesn't work out commercially, even if it just wasn't that great anyway, but that's your job interview to go into one of these things anyway. So, I don't feel that from a very small startup's perspective. Mid-size startups, yes. I would say there's been a lot of dead LM infra, a lot of LM infra consolidation, like the Langfuses of the world getting absorbed into ClickHouse. I think people have maybe worked out the domain-specific playbook. And I think that's okay. Yeah, I'm not that worried about that.
好的。所以我会说,我更担心传统的 SaaS,比如低 NPS 的 SaaS。这就是一直在进行的 AI 与 SaaS 之争。而且说实话,我在我的公司里就正在经历这件事。所以我是在非常切身的层面上思考这个问题,对吧?
Okay. So I would say I'd be more worried about traditional SaaS, like low NPS SaaS. This is the whole AI versus SaaS debate that has been going on. And literally I'm going through that exact thing in my company. So I'm kind of thinking through this on a very visceral level, right?
一方面,有人说你们这些 vibe 编码者不欣赏 CRM 背后的大量工作。是的,你以为你能取代 Salesforce。你之前的 30 个创业者也是这么想的,对吧?你通常会低估你不深入了解的东西,而且你的目标受众并不是你。同时,我们从未能如此轻松地构建和定制软件。而且,是的,你不会用到 Salesforce 中 90% 的功能。所以,是的,你内部做了什么?
On one hand, you have the people who say you vibe coders don't appreciate the amount of work that goes into the CRM. And yeah, you think you can rip out Salesforce. So did the 30 entrepreneurs before you, right? You classically underestimate the things that you don't deeply know, and your target audience is not you. At the same time, we have never been able to build software so easily and customize software so easily. And yeah, you're not going to use 90% of the things in Salesforce. So yeah, what's the what have you done internally?
所以我们有用于活动管理和赞助管理的主要 SaaS,我们每年为此支付 20 万。不算巨大,但对我来说相当可观。而且,是的,我可能花 2000 就能构建一个定制版本。难点在于处理我团队其他成员的问题,让他们跟上,因为我是团队里最愤世嫉俗的人,但我不能自己做决定,对吧?
So we have the main SaaS that we do for event management and sponsor management, and we pay 200k a year for that. Not huge but chunky for my scale. And yeah, I could probably spend 2000 and build a custom version of that. The trick has been dealing with the rest of my team and getting them on board, because I'm the most cynical person on my team, but I can't make that decision myself, right?
我认为,就像我一直对其他 CEO 和团队领导说的那样,你可以超级云端化,你可以超级 LM 精神病,认为那没问题。但你必须带上你的团队。我认为公司中 LM 精神病日益扩大的差距正在造成真正的裂痕,因为一方面,那些不太 AI 原生的人没有跟上形势。他们实际上落后了。他们没有意识到,你认为必要的一切实际上并不那么必要。事实上,如果你捏着鼻子走进去,从另一边出来,只与自然语言的智能体对话,你的生活反而会更好,而你只是思想封闭。这是那个视角。另一个视角是,"哦,你这个 vibe 编码者,你周末就搞定了,得到了 80% 的解决方案,现在你其余的雇员不得不收拾你剩下的烂摊子,你以为自己很厉害,但实际上你没搞明白,而且实际上他们在这上面仍然毫无用处,等等等等。" 所以我认为现在每家公司都在进行这场巨大的辩论。我有一个小的缩影,但这让我犹豫是否要扣动扳机,但我迟早会的。也许我会推迟一年,但不会推迟五年。
I think in the same way I've been telling other CEOs and team leaders as well, you can be super cloud-pilled. You can be super LM psychosis and think that's okay. But you have to bring your team with you. And I think the widening disparity in LM psychosis in companies is causing real rifts, because on one hand, the people who are less AI native are not getting with the picture. They're actually behind. They're not waking up to the fact that everything you think is necessary is not actually that necessary. And in fact, it would be better for you if you just held your nose and went in and came out the other side only talking to agents in natural language. And your life would actually be better, and you're just close-minded. There's that perspective. The other perspective is, "Oh, you vibe coder, you did this in a weekend and you got the 80% solution, and now the rest of your employees have to pick up the rest of your mess that you thought you were so hot at, but actually you didn't figure it out, and actually all of them are still useless at this, and blah blah blah." So I think there's this huge debate going on in every company right now. And I have a small microcosm of it, but it's making me hesitate to pull the trigger, but I will at some point. Maybe I'll put it off for one year, but not five.
但所以 SaaS 肯定在受到挤压。这确实让我想知道,我确实认为存在一个更 AI 原生的记录系统的机会,不仅仅是 Postgres 或 MongoDB。虽然两者都非常好。也许像 Convex,或者人们经常提到 Convex。我不知道。我只是觉得,所谓的 AI 应用的 Firebase 还不存在,除了我们已有的。这没问题。只是我们可能可以先从一个更快速的迭代周期开始,然后再扩展到像 Postgres 或 MongoDB 这样的旧技术。
But so SaaS is definitely getting squeezed. It does make me wonder, I do think that there's an opportunity for a more AI-native system of record thing that is not just Postgres or not just MongoDB. Although both are very good. Maybe it's like a Convex or people bring up Convex a lot. I don't know. I just feel like the quote-unquote Firebase of AI apps isn't really a thing yet beyond what we have. Which is fine. It's just we could probably start in a more rapid iteration cycle first before scaling up to like a Postgres or a MongoDB, which are more old tech.
我和 Anthropic 的 CPO Mike Krieger 共进晚餐。我们当时在房间里轮流说,"你最担心什么?" 是的。对我来说,我没有提安全,而是提了生物安全。
I was at a dinner with Mike Krieger, the CPO of Anthropic. And we were kind of going around the room going like, "What are you most worried about?" Yeah. For me, instead of security, I brought up bio-safety.
是的,不过很经典。实际上,我说过这很老套也很经典。桌上其他人说,"你是什么意思,有人坐在家里就能制造出消灭一半人类的病毒?" 这就像 Geoffrey Hinton 的原话,这就是你应该害怕的原因。
Yeah, classic though. Actually, like I said it was cliché and classic. And the rest of the table were like, "What do you mean someone sitting at home can manufacture a virus that wipes out half of humanity?" That was like the OG Geoffrey Hinton, like this is why you should be scared.
我说,"是的,读读风险报告。这就是那件事。" 我觉得 Michael 就坐在那里,知道自己掌握着神话,然后说,"实际上,是安全。"
I'm like, "Yeah, like read the risk reports. This is like the thing." I think Michael was just sitting there knowing he was sitting on the mythics and going like, "Actually it's security."
我认为其中一部分是非常好的营销,太好了。我实际上会建议 Anthropic 调低营销力度,因为这也是一个非常好的模型,你不必围绕它做那么多营销宣传。同时,如果你把它给 40 家公司,每家有 1 万名员工之类的,那它就不是真正的私有模型,对吧?它不是私有的。
I think there's part of it that is very good marketing, too good. I would actually advise Anthropic to tune down the marketing, because also it's just a very good model and you don't have to make so many marketing claims around it. At the same time, it is not really a private model if you give it to 40 companies, each of whom have like 10,000 employees or whatever, right? It's not private.
就像里面有坏人。是的,希望不像广泛发布那么糟糕,但不,我的意思是,这是一个有趣的案例研究,关于从现在开始,有多少模型发布可能会像这样,对吧?
It's like there are bad actors in there. Yeah, hopefully not as bad as releasing it widely, but no, I mean it's an interesting case study for how many model releases might look from now on, right?
Anthropic 有一个整体产品策略,比如捆绑、限制访问、将产品与模型捆绑,而 OpenAI 在哲学上肯定更倾向于我们将在各处启用访问,我们不知道会从中产生什么,对吧?
There's an overall product strategy for Anthropic of like bundle, restrict access, bundle product with model maybe, whereas OpenAI has definitely been a lot more philosophically aligned on like we will just enable access everywhere and we don't know what will come out of it, right?
对。不过,我的意思是,当前这个时刻,显然愤世嫉俗的看法也仅仅与两家公司运行的算力量有关。
Right. Though, I mean this current moment obviously the cynical take is also just ties to the amount of compute that both companies are running.
是的,我认为这是真的。我确实认为大于 10 万亿参数模型的黎明非常有趣。我不认为这是一个暂时现象,因为未来 3 到 5 年,每个人都会有更大的算力集群上线。这已经是注定的了。所以,就我们是否会在 2 年内对超过 10 万亿的模型进行配给而言?我不这么认为。我认为每个人都能获得。
Yeah, I think that's true. I do think the dawn of larger than 10 trillion parameter models is very interesting. I don't think it's a temporary phenomenon, because we have much larger compute clusters coming online for everyone over the next 3 to 5 years. And this is already written in the cards. So, to the extent that will we have rationing of models above 10 trillion in 2 years? I don't think so. I think everyone will have access.
对下一阶段进行配给。对,对。
To have rationing of the next phase. Right, right.
但这几乎就是应该的样子。比如我的经典例子,这只是我自己的推测,不是谷歌确认的。谷歌发布 Gemini 时,其实宣布了三个尺寸:flash、pro、ultra。他们从未发布 ultra,只有 pro 和 flash。所以我的理论是,他们把 ultra 放在地下室里,不断从它蒸馏出 flash 和 pro。我觉得任何实验室都应该这么做,因为那些才是人们最终想用的模型,而且成本太高了。
But like that's as it should be almost. Like my classic example, which I this is just me theorizing not anything confirmed by Google. When Google announced Gemini, they actually announced three sizes, which was flash, pro, ultra. They never released ultra. They only have pro and flash. So my theory is they have ultra sitting in a basement and they just keep distilling from it for flash and pro. Which like yeah, I mean I actually think that's as it should be for any lab that they do that. Yeah, just because of those are the models that people actually want to end up using or and it's just like cost prohibitive.
更多,是的,是成本问题。不是需求问题,就是成本。
More Yeah, it's cost. It's not the wants, it's just the cost.
嗯,我确实觉得有趣的是,有一段时间我在考虑模型参数上限是两万亿的理论,我认为这被证明是错的。那么,如果我错了,我错得有多离谱?我们会做到 200 万亿吗?还是两千万亿?我不认为我们有确切的答案,但有趣的是,我们继续扩大参数数量,尽管大家都看到我们不会从这个范式获得下一个千倍或百万倍的提升。所以像世界上其他的实验室都在研究其他模型架构的改进。我们需要不同的缩放定律,因为我觉得人们已经觉得我们在这个方向上已经到头了。这个范式的最终状态是我们把世界大部分变成数据中心。
Um I do think like uh it is interesting that uh for a while I was considering the theory that models capped out at two trillion and I think that's proving to be wrong. And well, then if I'm wrong, how wrong am I? Do we do 200 trillion? Do we do two quadrillion or whatever? And I don't think we have the straight answer to that, but like uh it's interesting that we are continuing to scale number of params when everyone kind of like can see that we're not going to get like the next thousand or 1 million X from this paradigm. So like the others like the alias of the world are working on other um model architecture improvements. We need a different scaling law, I guess, because like where I feel like people are already feel like we're tapped out on this. Like the end state of this is we turn most of the world into data centers.
而且我不知道,我不知道我们是否想要那样。
And like I don't know. I don't know if we want that.
是的,但我的意思是,如果智能的回报在那里,也许没那么糟。我认为现在有大量的不可扩展性在困扰着人们的感受,尤其是在上下文长度方面。我的经典说法是,上下文长度是大语言模型中最慢的扩展因素。
Yeah, but I mean if the return of intelligence are there, maybe maybe not so bad. I think there's just a sheer amount of like unscalability that is wrangling people's sensibilities right now. It's especially in terms of like context lengths. My classic quote is like context length is like the slowest scaling factor in LLMs.
是的。
Yeah.
嗯,我们大概花了 3 年时间从 4000 上下文长度发展到一百万。就这么多。是的,Gemini 已经有一百万 token 上下文长度两年了,但没人用它。所以,是的,内存可能是所有这些事情上最大的限制约束。
Um we like we took maybe 3 years to go from a 4,000 context length to a million. And that's about it. Yeah. Like Gemini has had a million token context length for 2 years now. Um and no one's using it. So like yeah, it's memory is probably going to be the biggest limiting constraint on all these things.
是的,看起来确实如此。我很好奇,自从我们上次录制以来,这一年里你改变想法的一件事是什么?
Yeah, certainly seems that way. I guess I'm curious over the last year since we recorded last like what's one thing you've changed your mind on?
我觉得去年我对开放模型有点悲观。因为我和贝恩资本的 Ankur Goyal 做过播客,他对所有顶级 AI 公司有很好的了解。他说开源的市场份额是 5% 并且在下降。我认为这已经改变了,我认为它在上升。即使能力差距似乎在扩大,也很难说。
I feel like I was kind of bearish on open models like last year. Um in the sense of like I had just done the podcast with Ankur Goyal of Bain Capital where he and he I mean, you know, he has a good cross-section of all the top AI companies. And he says market share of open source is 5% and going down. Um I think that's changed. I think it's going up. Um and even if capability gap does seem to be increasing. It's hard to tell.
取决于……很难说。
Depending on the It's hard to tell.
是的,真的很难说。因为,对听众来说,能力差距扩大是在公共基准上。假设你比较 Mythos 和 GPT-OSS 或 GLM-5.1,真的很难说,因为即使它们在缩小差距,你也不会相信它们缩小了很多,因为操纵基准很容易。所以你真的不知道。你只知道的是,在自由市场中,OpenRouter 的统计有些客观地反映了人们的选择。人们确实大量选择一些开放模型,但很多都有大幅折扣。所以你需要对这些进行价格调整。所以,即使那是真的,我也不确定,我觉得这个数字现在是上升而不是下降。我认为顶级智能体实验室与普通 AI 初创公司或普通 GPT 包装器之间的差距足够大,你不应该担心行业平均数字,而应该把事物分成中位数、底部 80% 和顶部 20%。顶部 20% 的行为与底部 80% 非常不同。顶部 20%,也就是我关心的,肯定在走向更开放的模型。Fireworks 和 Together 正在崛起。所以所有微调者,对吧?所以,我想也许上次我们甚至说过微调即服务行不通。现在它会行得通。这是开放模型市场的衍生品。
Yeah, it's really hard to tell. Cuz like okay, for listeners, capability gap increasing is like on public benchmarks. And let's say you're comparing Mythos versus like I don't know, GPT-OSS or like GLM-5.1. And it's really hard to tell because even if they were closing, you would also not believe that they were closing that much because it's very easy to game the benchmarks. So you just don't really know. Um all you know is like there's somewhat objective open router stats on like what people choose in a free market. And people do choose some of these open models in significant volume. Except that a lot of them are heavily discounted. So, you need to kind of like price adjust these things. So, even if that were true, which I'm not sure like I feel like the number is just up now instead of down. Uh I think the separation between what the top-tier agent labs are doing versus the average startup in AI or the average GPT wrapper is significant enough that you should not worry about the sort of mean industry number and you should cohort things into like here's the median here's like the bottom 80% and here's the top 20%. And top 20% acts very differently than the bottom 80%. And so, top 20% is which is what I all I care about um is definitely going towards more open models. Um the fireworks and the together is a crushing. Um and uh so all the fine-tuners, right? So, like um I think maybe last time we even said things like fine-tuning as a service doesn't work. Well, now it's going to work. It's a derivative of the open market uh open models market.
嗯,而且在工作负载扩展到人们越来越关心成本和速度的程度。然后从纯粹的使用案例发现,比如这些模型能做什么,到好的,我们知道它们在规模上能做什么。现在,让我们把它们做得更便宜、更快。
Well, also in the workload scaling to the point where people care about cost and speed, you know, more and more. And then like the you know, moving from just pure use case discovery of like what can these models do to okay, we know what they can do at scale. Now, let's do them cheaper and faster.
是的,是的。嗯,所以这个变化我认为可能是最重要的。我总是喜欢做心算,比如这是我关于调度学习率的想法。当你错了一次,你还错在什么地方?我还在思考。对我来说,另一件事是编码,显然我现在已经完全转变了。但我认为人们没有充分重视暗工厂,我不知道你们在播客里讨论过没有。
Yeah, yeah. Um so like uh that change I I think is probably the most significant in in my mind. And like I I always like to do the mental math of like uh this is what I think about uh scheduling a learning rate. Like when you've been wrong once, what else were you wrong on? Um and I I'm kind of working through it. I I think to me the the the other thing was the coding one, um which obviously I I have now come full 360 on. But I think like people are not appreciating dark factories enough, which I don't know if you've discussed in the pod yet.
没有,是的。
No, yeah.
嗯,这是一个 Strong DM / Simon Willison 的术语。大致意思是,你可以有不同级别的 AI 编码狂热。第一个级别,顺便说一句,我在五个月前在 Cognition 第一次遇到,是零人类编写代码。对吧?这在现在看来是合理的,但在五个月前不太合理。下一个前沿,今天听起来和过去的零编码一样疯狂,是零人类审查。就像你直接提交代码而不审查。很少有人这样做,但 Open Eyes 在探索这个,我觉得这绝对是唯一可扩展的方式,这意味着你必须翻转 SDLC 或大量改变你通常做的事情。这可能是你本来就应该做的事情:更多测试,更多自动化验证等等。但这是一个前沿,当你在公司解锁它时,你将产生比以往更多的软件数量。它会如此之多、如此可丢弃、如此便宜,以至于你可能在质量上也能大量创新。数量帮助你达到质量。
Um Uh and so, this is a kind of a strong DM / Simon Willison term. Uh the general idea is okay, there's different levels of AI coding psychosis you can have. Um the very first level which I by the way I can encounter first in cognition five months ago was zero human written code. Right? Which like seems like a reasonable thing now was less reasonable five months ago. The next frontier that sounds as crazy today as it as as zero coding was in in the past is zero human review. Like you just just check it in without even reviewing it. And very few people are doing that but open eyes is is like exploring this and like I feel like it's it's definitely only scalable way to do this which it just means like you have to just kind of like flip the SDLC or change large amounts of what what you normally do. Which is probably things you should have done anyway. More testing, you know, more automated verification or whatever. But like that is a frontier at which like when you have it unlocked that in your companies, you are just going to produce much more quantity of software than than you've ever had. It's going to be like so much so disposable so cheap that you can probably innovate in quality a lot as well. Like that that quantity helps you get to quality.
是的。我认为人们对此非常不舒服,因为人们把更多数量与垃圾内容联系在一起。对吧。
Yeah. Which I think people are very uncomfortable with because like people associate more quantity with slop. Right.
现在回到我们正在讨论的话题,就是对这些 token 最大化排行榜的反应,以及这样一种观点:今天这或许不是产品效率的最佳指标,但未来你仍然会因此获得回报。所以,随它去吧。但我认为,2026 年表现出色的人,不会是那些愤世嫉俗者,他们觉得‘哦,那不过是垃圾,我不参与’。他们会说:‘好吧,不管有没有我,这都在发生。让我们把它引向正确的方向。’
Now it's back to exactly the discussion we're having on the reaction to these token-maxing scoreboards and the idea that today maybe that's not the best sign of product efficiency, but going forward, you still get rewarded for it. So it's whatever. But I think the people who do well in 2026 are not the cynics who go, 'Oh, that's just slop. I'm not going to participate in that.' They're like, 'Okay, this is happening with or without me. Let's bend this the right way.'
是的,我喜欢这个观点。对我来说,在输出优先模型方面,有一个相关的事情:很长一段时间里,我真的认为做任何形式的强化学习、后训练、预训练——任何能提高整体质量的事情——都没有意义。当然,对于延迟和成本来说,这对我来说一直是有意义的,但对于整体质量来说,就像,你会在 3 到 6 个月后从模型那里免费获得这些。我开始稍微改变想法的一点是,听到所有这些应用公司说,我们构建东西,然后 3 个月后因为模型改进就把它扔掉。你会想,‘好吧,那你为能力提升所做的事情,只是另一种版本,对吧?’我仍然不认为你的强化学习或后训练会让你在未来的很多年里拥有更好的模型。但也许你仍然需要非常严格地问自己,这是解决客户问题的最佳方法吗?很多时候,答案就是‘不,添加更多数据,喂更多数据,成为这些模型的对手,或者在后端做一些巧妙的工程。’但如果在这 3 个月的时间里,改善客户结果的最佳方法是以某种方式做后训练,真正提高模型的输出,即使 3 个月后因为通用模型赶上来而把它扔掉,它可能仍然是值得做的。所以,我想我对这个想法更加开放了。
Yeah, I love that. And I think for me, a related thing on the output-first model side is that for so long I really didn't think it made any sense to do any sort of RL, post-training, pre-training—anything you could do to improve overall quality. Certainly for latency and cost it always made sense to me, but for overall quality, like, you just get that for free in the models 3 to 6 months later. I think what I'm starting to change my tune on a little bit is, hearing all these app companies talk about how we build stuff and then we throw it out 3 months later as the models improve. You're like, 'Okay, well, then what you're doing for capability improvement is just another version of that, right?' I still don't think that your RL or post-training is going to make you have a better model for years and years to come. But maybe you still have to be pretty rigorous in asking, is that the single best thing you can do to solve a customer problem? Often times it's literally just, 'No, add more data, feed more data, be a counter to these models, or do some clever engineering on the back end.' But if the single best thing you can do for that 3-month time period to improve your customer's outcomes is post-training in some way that really improves the output of the model, even if you throw it out 3 months later because the general models get up there, it still might have been worth doing. So I think I'm more open to that.
你扔掉了结果,但你没有扔掉原始数据。
You throw out the results, but you don't throw out the raw data.
完全正确。所以,就像……
Totally. And so, like...
对,然后你只需再运行一次。所以,基本上有一个成本水平——显然在 1000 万美元的水平,也许那太高了,但有一个成本水平,在那里……
Right, then you just run it again. And so, basically there's some level of cost—obviously at the level of $10 million, maybe that's too much, but there's some level of cost where...
不,甚至不到 1000 万美元。
No, it's not even $10 million.
不,当然不是。你知道,显然有一个投资水平,相当于雇佣四名工程师去花 3 个月构建东西。
No, of course it's not. You know, there's obviously some level of investment at which it's the equivalent of just staffing four engineers to go build something for 3 months.
所以,另一件事,我真的——对于听众,我就留一些信息碎片。研究一下长期轨迹,人们正在做的合成规则工作非常重要,包括一个叫 Dr. GRPO 的东西。我就把这些关键搜索词留在这里。我认为这意味着强化学习将比人们想象的更加多轮。这意味着你可以在比传统的、我们称之为 SFT 或一年前做的浅层强化学习更具体的维度上定制模型。所以,就像数百轮。是的。我认为这会引导你走向完全领域专业化的道路。
And so, the other thing I really—for listeners, I'm just going to leave some droplets of info. Look into the long trajectory, the synthetic rubrics work that people are doing is very important, including something that's called Dr. GRPO. I'll just leave those key search terms in there. I think what it means is that RL is going much more multi-turn than people think. And that means that you can customize the models in way more specific dimensions than traditional, let's call it SFT or shallow RL that was done a year ago. So, like hundreds of turns. Yeah. And I think that leads you down a path of complete domain specificity.
明年你还在寻找什么?在当今 AI 的这些未解问题中,你密切关注的是什么?
What else are you looking for in the next year? Of these unanswered questions in AI today, are you paying close attention to?
对于下一个前沿是什么,我有几个论点。一个是记忆——我们谈到的记忆和个性化。另一个是真正的世界模型,我们做了一个小系列,从李飞飞一直到月之暗面,还有通用智能。关于这个的相对重要性有很多争论。我认为很多表现为 3D 静态世界,你可以在里面待一会儿,走来走去,它们很酷,但这对我的 B2B SaaS 有什么帮助?
I have a few theses for what is the sort of next frontier. One is memory—memory and personalization we talked about. The other is really world models, which we've done a small little series on, from Fei-Fei Li all the way to even MoonLake, and general intelligence. And there's a lot of debate as to the relative importance of this. I think a lot of it manifests as 3D static worlds that you kind of inhabit for a little bit and you walk around and they're like cool, but how does this help me with my B2B SaaS?
现在所有的炒作都是机器人,对吧?
All the hype now is robotics, right?
是的。嗯,世界模型与具身视觉和体验之间显然存在相关性,这导致了机器人技术。但我认为世界模型在提高智能本身方面非常有趣,从下一个词预测范式来看。所以我认为人们正在围绕这一点测试他们的边界。我们今年迄今为止的顶级文章之一是关于对抗性世界模型的。我确实认为,如果你什么都不做,就去读李飞飞关于空间智能的文章,关于为什么 LLM 没有它。她可能还没有解决方案,但她有正确的问题陈述。所以其他人都在试图以自己的方式解决这个问题陈述。让我们看看谁会赢。但我认为把世界模型等同于机器人技术,或世界模型等同于游戏,或某种当前的表现形式,对你没有好处,因为关键在于一个比仅仅回答问题重要得多的智能概念。它是:AI 是否理解桌子是什么?物质是什么?物理是什么?这几乎就像,对于那些电影迷来说,就像《心灵捕手》里马特·达蒙知道一切,因为他从书里读过,但他从未亲身经历过。
Yeah. Um, and there's obviously a correlation between world models and embodied vision and experiences which leads to robotics. But I think world models is very interesting in just improving intelligence itself, from the next-token prediction paradigm. And so I think people are kind of testing their edges around that. One of our top articles this year so far has been on adversarial world models. I do think if you don't do anything else, just read Fei-Fei Li's essay on spatial intelligence, on why LLMs don't have it. And she may not have the solution yet, but she has the right problem statements. And so everyone else is trying to solve that problem statement in their own way. And let's see who wins. But I don't think it does you any favor to equate world models to robotics or world models to gaming or some kind of current manifestations, because what is at stake is a much more important conception of intelligence than just answering questions. It is: does the AI understand what a table is? What matter is? What physics is? It's almost like, for those who are movie fans, it's like Good Will Hunting where Matt Damon knows everything because he read it in a book, but he's never lived it.
和罗宾·威廉姆斯的那场戏很棒。
Great scene with Robin Williams.
罗宾·威廉姆斯。我看着那个场景,我想,这正是非常聪明的 LLM 与知道一切但从未经历过任何事之间的区别。
Robin Williams. And I look at that scene and I go, that's exactly the difference between a very intelligent LLM who knows everything but hasn't experienced anything.
哇。这是一个很好的结尾。这是一个很棒——你以前用过这个类比吗?太棒了。
Wow. That's an awesome note to end on. That's a great—have you used that analogy before? That's great.
是的,这是一个——所以我对 Len's space 做的一件事是,我开始添加每日总结。所以有一次我在写每日总结时,我写了一个……
Yeah, it's a—so one thing I've done with Len's space is I moved to adding daily write-ups. And so one of the times I was doing this daily write-up I wrote a...
这是一个很棒的故事。我喜欢。嗯,非常有趣。非常感谢你来做客。
That's a great one. I love that. Well, it's been a ton of fun. Thanks so much for coming on.
加油,伙计。
Up, man.
我是 Jacob Effron,这里是 Unsupervised Learning,一个播客,在这里我可以和 AI 领域最聪明的人交谈,问他们大量关于模型正在发生什么以及这对企业和世界意味着什么的问题。正如我希望的那样,我从中获得了很多乐趣。这是一个夜晚和周末的项目,除了我在 Redpoint 做投资人的日常工作之外,但我们能请到这些了不起的嘉宾,真的来自于像你这样的人订阅播客、与朋友分享。这最终是让这一切运作起来的原因。所以请考虑这样做,非常感谢你的支持和收听。
I'm Jacob Effron, and this has been Unsupervised Learning, a podcast where I get to talk to the smartest people in AI and ask them tons of questions about what's happening with models and what it means for businesses and the world. As I hope is clear, I have a ton of fun doing this. It's a nights and weekends project in addition to my day job as an investor at Redpoint, but our ability to get these incredible guests on really comes from folks like you subscribing to the podcast, sharing it with friends. It's really what ultimately makes this whole thing work. And so please consider doing that, and thank you so much for your support and listening.
我们下期再见。
We'll see you next episode.