OpenRouter CEO on the AI Gateway Market
打开互动全文版(中英对照 + 朗读 + 问答)→Alex Atallah 讨论 OpenRouter 的历程、推理提供商格局以及 AI 基础设施的未来。
Alex Atallah discusses OpenRouter's journey, the inference provider landscape, and the future of AI infrastructure.
这将是科技领域有史以来最大的市场。很多公司都在做路由器,因为这是潮流。模型应用最终会有多种动机来追赶你。
This is going to be like the biggest, biggest market in tech ever. A lot of companies are making routers because it's fashionable. The model apps have several incentives to go after you eventually.
今天我们请到了 Alex Atallah,OpenRouter 的联合创始人兼 CEO。OpenRouter 是统一接口,是通往 LLM 世界的门户。据报道,Stripe 曾出价 100 亿美元收购。他们以超过 15 亿美元的估值完成了融资。他们是市场领导者,这次采访来得正是时候。
Today we have Alex Atallah, co-founder and CEO of OpenRouter, the unified interface, the gateway to the world of LLMs. They reportedly have had offers from Stripe for $10 billion. They've raised at a valuation of over a billion and a half. They are the market leader and this interview could not come at a more prescient time.
7 月份我们发布了 70 个模型,大约每 10 小时一个。美国仍然非常非常落后。GLM 5.2 对开放权重模型来说是非常大的一步。
In July we launched 70 models, about one model every 10 hours. America is very, very behind still. GLM 5.2 was a really big, big step for open weight models.
有报道说你将以 100 亿美元卖给 Stripe。这会发生吗?准备好了吗?
There are reports that you are selling to Stripe for $10 billion. Is that going to happen? Ready to go?
Alex,我太兴奋了,兄弟。我想做这期节目很久了。我从 Manifold 的 Matt 那里听说了很多关于你的事。我甚至在你和 Anjney 的对话中追踪过你,还有你之前的室友。所以,谢谢你来做客,兄弟。
Alex, I am so excited for this, dude. I've wanted to make this one happen for a while. I've heard so many things from Matt at Manifold. I've stalked the out of you speaking to Anjney, even your roommate before this show. Um so thank you for joining me, dude.
谢谢。很高兴来到这里。
Thank you. It's great to be here.
现在,我想从 OpenRouter 之前开始,先聊聊 OpenSea。那是一段非常不可思议的旅程。你在 OpenSea 经历了那么多,你带到了 OpenRouter 的是什么?
Now, I want to start with a little bit pre-OpenRouter and start on OpenSea. It was a pretty incredible journey. What did you take with you to OpenRouter having seen all that you saw with OpenSea?
是的,OpenSea 是我们作为第一个 NFT 市场创立的,和 OpenRouter 类似,它在很长一段时间内都非常小。我们一直保持团队很小,直到 A 轮融资左右,你知道,之后一点。这是在 AI 之前。所以就在 NFT 开始爆发之后,2020 年 10 月,我们想,“天哪,我们人手不足。服务器快熔化了。我们的搜索索引在爆炸。我们经历了几次大的宕机。保持网站运行很困难,就像‘天哪,我们要变成 Twitter 的失败鲸鱼了,但应用到加密货币上。’我最大的目标就是不要成为加密货币领域的 Twitter 失败鲸鱼。”
Yeah, so OpenSea we started as the first NFT marketplace and similar to OpenRouter it was very small for a long time. Like we kept the team very small until the series A roughly you know, a little bit afterwards. Um and this was before AI. So right after NFTs started of blowing up in in 2020, October of 2020, we're like, "Oh my goodness, like we are understaffed. The servers are melting. All kinds of like our search index was exploding. We had a couple big outages. It was tough to like keep the site up and it was like, 'Oh my god, we're going to become like the Twitter fail whale, but like applied to crypto.' My biggest goal was to have us not be the Twitter fail whale."
嗯,对于加密货币。我们花了一点时间来组建团队,控制平台和基础设施,确保我们能够可预测地扩展。换句话说,做负载测试来帮助网站承受 10 倍的负载,即使我们当时没有看到那么高的负载。因为加密货币,你永远不知道。有些时刻我们会遇到惊人的流量峰值,这很大程度上取决于内容和社区。所以我在 OpenSea 建立了大量的基础设施和扩展方面的责任,然后我花了很多时间思考,“好吧,我们如何让一个东西基本上永远在线?从基础设施的角度来看,即使市场动荡,出现巨大的激增,人们也能真正依赖它。”这当然对 AI 非常有帮助,因为你知道,所有公司,尤其是 Anthropic,都经历了不可预测的增长。我们也是。我们有过一些波折,但总体上好多了。OpenSea 以某种方式把它灌输给我,让我能够有效地带到 OpenRouter。
Um for crypto. And and it took a little bit to like create the team, get platform and infrastructure under control, like make sure we we could we could predictably scale. In other words, like do load testing to like help the site sustain 10x load. Um even when we weren't seeing that load. Because with crypto, you just don't know. There were like these moments where we would get these incredible traffic spikes and it'll be very dependent on the content and the community. And uh So I built like a lot of um infrastructure and scaling I think responsibilities then that I took to OpenSea and spent a lot of time like thinking about, "Okay, how do we you know, make something that is going to basically be always up? And that that people can really count on from an infrastructure point of view even when there are huge surges in in really like tumultuous markets. Um which has been very helpful for AI of course because like, you know, all companies, like especially Anthropic, have seen like unpredictable growth. And and we have as well. And you know, we've had like a couple bumps, but overall it's been like significantly better. And I just like OpenSea just kind of like drilled that into me in a way where I could like take it productively to open router.
我能问你吗,当你回顾公司的创始理念时,在模型格局的生态系统中,发生了什么你没想到会发生的事情?
Can I ask you, when you go back to the founding thesis of the company, what has happened in the ecosystem in the model landscape that you did not expect to happen?
嗯,好吧,我们没想到的一件事是,会出现一个公司生态系统来托管和服务开放权重模型。早期,我们不清楚这个市场会不会是垄断的,比如三大超大规模云服务商服务所有开放权重模型,而初创公司远远落后。实际上,你知道,你多久听到有人在超大规模云上运行 GLM?从来没有。他们用的是推理提供商,比如 Fireworks 和 Together。而且我们看到了一个很大的列表,这些提供商在托管所有开放权重模型方面做得最好。早期,我们有一个叫“提供商一”和“提供商回退”的东西。我们没有显示哪些提供商实际在托管,因为我的意思是,我们并不是真正的市场。我们更像是一个探索工具,用于发现和探索新的 LLM。我们想建立一个模型实验室的市场,但对于推理提供商层,我们不确定它是否会真正成为一个市场。结果证明,这些公司比超大规模云做得更好,托管模型和解决托管边缘情况的速度更快。而且正常运行时间将是一个持续的问题,不会神奇地被市场供应方解决。
Um okay, well, one thing that we did not expect was that a an ecosystem of companies would emerge to host and serve the open weight models. Um like early on, it wasn't clear that that that market wasn't going to be a monopoly where like just, you know, the three hyperscalers serve all the open weight models and uh and start and startups don't, you know, they're they're really far behind. In reality, like you know, how often do you hear people running, you know, GLM on a hyperscaler? Never. Like they're using the the inference providers like Fireworks and Together. And um there's like go big list that we that we see doing the best job of hosting all of the the open weight models. And um in the early days, we um we had I think we called it provider one and and provider fallback. We didn't like show which providers were actually doing the hosting cuz I I mean we weren't really a marketplace. We were kind of a an like we were an exploration tool for like finding and discovering new LLMs. And we wanted to get we wanted to build like a marketplace of model labs, but like the inference provider layer, we weren't sure it would actually be a marketplace. And it turned out that those companies were doing a way better job than the hyperscalers, were way faster to host the models and figure out these edge cases to hosting them. And um and uptime was just going to be a a constant problem. It wasn't going to like magically get solved by the supply side of the market.
很多人认为推理提供商层是一个可商品化的元素或层,最终会被移除或利润率被竞争掉。你对这个理论怎么看?
A lot of people suggest that that inference provider layer is a commoditizable element or layer that will be removed or see margin reduction competed out over time. What would you say to that theory?
现在我们处于一个供应严重受限的市场,而且可能还会持续一段时间,所有推理提供商都短缺,几乎总是短缺。你会想,“好吧,所以 GPU 真的非常有用。”然后,“为什么 Google 或 Amazon 或 Azure 不四处买下所有 GPU,把这些推理提供商都赶出市场?”嗯,制造 GPU 的人不希望这样。Nvidia 的首要任务之一就是避免客户集中。他们希望很多客户都有独立的 GPU 分配。他们希望市场具有异质性。他们希望在算力层有竞争。而且这对生态系统有好处。用户也想要这样。这对 Nvidia 好,对最终用户也好。它允许这些推理提供商在如何更好地服务模型方面提出新的创新。即使是像 Kimiko 3 这样的单一模型,Moonshot 刚刚发布了一个基准测试,展示了所有推理提供商以及他们服务 Kimiko 3 的表现。对于这些静态的、众所周知的基准,数字差异很大。我们一直在持续发布这些。我们总是在所有推理提供商、所有开放权重提供商上对所有模型进行基准测试,并不断发现非常不同的结果。结果会随时间变化。这些模型非常情绪化,非常不确定。所以,嗯。
Right now we're in a massively supply constrained market where um but and it's likely going to be supply constrained for a while where all the inference providers are are short, pretty much constantly short. Uh and and you're like, "Okay, so GPUs are are really really beneficial." And like, "Why doesn't Google or Amazon or Azure run around and like buy up all the GPUs and take all these inference providers out of business?" Well, the people making the GPUs don't want that. Like one of Nvidia's top priorities is not having customer concentration. They want lots of customers to all have like separate like allocations of GPUs. Um they want the the heterogeneity of the market. They want like competition on the compute layer. Um and I and this is good for the ecosystem. Like like users also want this. This like this is it's good for Nvidia and it's good for end users as well. It like allows these inference providers to kind of like come up with new innovations on like how to serve the models better. Even a single model like Kimiko 3 um uh like Moonshot just posted a benchmark showing all the inference providers and how well they're serving Kimiko 3 um and uh the they're they're pretty different numbers for like for benchmarks that are really static, that are well known. We post this continuously all the time. We always are like benchmarking all of the models on all of the inference providers, all the open weight providers, and finding really different results con- constantly. The results change over time. Um these models are like very They're very emotional. They're They're very like They're very like non-deterministic. So, uh um
我请了 Fireworks 的 Lin 上我的节目,她说,你知道,我说过 Gavin Baker 的话,一个 token 就是一个 token,这就是他说的。
I had Lin on my show from Fireworks and she said that, you know, I said about Gavin Baker and a token is a token is what he said.
嗯哼。
Uh-huh.
她其实纠正了我一点:token 并不都是一样的,因为有的提供商能让一个 token 走得更远。就像去商店,你可以绕整个街区开一圈,也可以直接开到商店。Token 可以被优化得更高效、走得更远,这正是提供商的工作。
And she kind of corrected me that a token is not a token, actually, because one provider can make a token go so much further than another token. It's like, how do you get to the store? Well, you can drive around the whole block or you can drive straight to the store. Tokens can be made more efficient and go further, and that's the job of the provider.
是的,我同意。我认为,在某种程度上,我们是在提供服务,帮助人们发现提供商。最终,当某个提供商让 token 走得更远时,我们会花大量时间在我们的中央路由技术上,让那个提供商立刻获得更多流量。一旦我们检测到质量提升、速度加快或价格下降,它就会立刻获得更多流量。而且这是 24/7 全天候发生的,每五分钟就有一次。对于大模型来说,变化很大。所以这确实改善了体验。
Yeah, I agree with that. I think that, in some ways, we are providing a service to help people discover providers. And ultimately, when one provider is making a token go further, we spend an enormous amount of time on our router central router tech so that that provider immediately gets more traffic. As soon as we detect that there's a quality improvement or a speed up or a price reduction happening, it immediately starts getting more traffic. And this happens 24/7, every five minutes. There are big changes for the big models. So it actually does make the experience better.
你只能投资一家推理提供商。你会投资哪一家?
You can only invest in one inference provider. Which one do you invest in?
我可能得保持中立。我确实喜欢那些在做定制硬件和底层优化的推理提供商。我也喜欢那些试图让定制更容易的提供商。比如现在你微调模型,从基础模型创建一个全新的、完全独立的模型。很多推理提供商在创建 LoRA,或者有人叫它们“卡带”,这些可能在模型之间更具可移植性。我们可能会看到这样一个未来:当你微调并想更换基础模型层时,只需要花费几百美元,甚至几十美元。
I probably have to stay neutral on this. I really like the inference providers that are doing custom hardware and very low-level optimizations. I like providers that are also trying to figure out how to make customization easier. So today you fine-tune models and you create this new fully independent model from the base model. Many inference providers are creating these LoRAs or some call them cartridges that are much more portable potentially between models. And we might see a future where when you do a fine-tune and you want to change the base model layer, it only costs maybe a few hundred dollars, maybe a few dozen dollars to change it.
没关系。我明白了,Firework 是你的最爱。没关系,我懂了。
It's okay. I understood that Firework is your favorite. It's okay. I got it.
我也是这么想的。
Mine too.
我的问题是,她在节目里说:“哦,你不想租用别人的智能,你想拥有它。”我们会看到公司拥有专门模型,这些模型基于自己的数据训练,且是专有的。在一个每家公司都拥有针对自身偏好定制的专门模型的世界里,这对开放路由器业务有利吗?
My question is when it was on the show she was like, "Oh, you don't want to rent your own rent intelligence, you want to own it." And we're going to see companies have specialized models which is trained on their own data and proprietary to them. In a world of every company having specialized models that's really tuned to them and their preferences, is that good for an open router business or not?
嗯。
Mhm.
为什么?因为你会固守一个属于你自己的、专有的、基于你的数据训练的模型,而不会对可用的模型多样性开放。
Why because you'd stick on one model which is yours, proprietary, trained on yours, and not be open to the diaspora of models that is available.
不,我不同意。我认为我们从一开始的使命就是增加整个 AI 生态系统的神经多样性。我们坚信多模型未来是不可避免的。当你开始,比如说有一个模型能满足你所有的需求,无论是在公司内部还是作为消费者。越来越多的人开始使用那个模型。然后有人决定:“你知道吗?我要创建一个神经多样性模型。我要创建一个有点不同的模型,说话方式有点不同,能产生第一个模型永远想不到的想法,因为训练它的数据完全不同。”这就会产生对同时使用两个模型的必然需求。创造力是无法验证的。创意想法很难用简单的数字来衡量。当你同时使用两个模型时,你比只用一个模型更有可能获得创意。这是事实,如果另一个模型是用不同的数据集、以不同的方式训练的,或者做了重大更新。所以,整合到一个模型上对我来说毫无意义。
No, I disagree. I think our mission from the very beginning has been to increase neurodiversity in AI for the whole ecosystem. And we really believe that a multi-model future is inevitable. And when you start, let's say there's one model that fulfills all of your desires, either within your company or as a consumer. And more and more people start using that model. And then someone decides, "You know what? I'm going to create a neurodivergent model. I'm going to create a model that's a little bit different, that talks a little differently, that has ideas that the first model could never have come up with because it's completely different data that's being used to train it." Then it kind of creates inevitable demand to use both models. Like creativity is not verifiable. There's no easy number to put on creative ideas. And when you use two models together, you're more likely to get creative ideas than if you just use one. It's just a fact if that other model was trained in a different way on a different data set, or has made a big update. So consolidation on one just doesn't make any sense to me.
完全理解你。所以,公司会有一个核心工作流或核心,即他们自己的专门模型,然后他们会使用大量其他模型,并利用 OpenRouter 来选择这些其他模型。
Totally get you. So, you will have companies which have like a core workflow or their core which is their own specialized model, and then they'll use a plethora of other models and they'll use OpenRouter for those other model selection.
是的。我认为当公司用自己的数据训练自己的模型时,你必须考虑这些事情的博弈论。如果每个人都在这样做,所有模型应用都在不断用新获取的数据(从其他公司购买的数据)创建新模型,这些数据都可能对你很有价值。什么最符合你的利益?去尝试那些其他模型,看看能否提高生产力,能否将它们合并以获得更好的最先进性能,能否用这些其他模型降低成本。无论你的目标是提高利润率还是发展公司,你都有动力去使用生态系统创造的东西。所以,你创建的模型,你必须不断改进以跟上步伐。它永远不会赢得整个市场。所以这将是一个巨大的市场。这将是科技史上最大的市场。可能是人类历史上最大的市场。没有人能赢得全部。你不会构建一个能赢得全部的模型。所以,你不如构建一个以专长于某些非常有用、对你的公司和业务非常重要的事情而闻名的模型,并以这种专长而闻名。我认为很多企业都会朝这个方向发展,创建自己的模型,创建自己的品牌智能。你的品牌是你护城河的重要组成部分。那个模型将成为你品牌传播的方式。
Yes. And I think that when companies make their own model trained on their own data, you have to play out the game theory for these things a little bit. Like if everybody is doing this as well and all the model apps are creating new models constantly using new data that they've acquired, that they've bought from other companies, that's all potentially data that's valuable to you. What is in your best interests? It's to go and try out those other models and see if you can be more productive with them, if you can merge them together to get better state-of-the-art performance, if you can reduce your cost using these other models. Whether your goal is to improve your margins or grow your company, you are incentivized to go use what the ecosystem creates. So the model that you made, you're going to have to continuously improve it to keep up. And it's never going to win the whole market. So this is going to be a massive market. This is going to be the biggest market in tech ever. And probably the biggest market in human history. No one's going to win all of it. You're not going to build a model that wins all of it. So you might as well build a model that is known to specialize in something very useful and very important to your company and your business, and be known for that specialty. And I think a lot of enterprises are going to move that direction, make their own models, make their own branded intelligence. Your brand is a big part of your moat. And that model will be a way your brand carries around.
你提到你在路由技术上花费了大量时间。很多人认为我们正在看到路由技术的商品化。你看到 Ramp 发布了这样的产品。我之前提到的 Merge,我们投资的公司,也发布了那个产品。有几家正在发布类似的路由技术,或者声称类似。我们是否正在看到这一层的商品化?
You mentioned the immense time you spend on the routing technology that you have. A lot of people are thinking that we're seeing the commoditization of the routing technology. You're seeing Ramp release products like this. I mentioned earlier of Merge, company we invest in has released that product. Several are releasing kind of routing technology similar or claiming to be similar. Are we seeing the commoditization of this layer?
我认为很多公司做路由器是因为这很时髦。他们看到这里有增长,或者至少他们在做网关。首先,我认为这有两个问题。第一,它立刻让你陷入模仿的心态,而不是去赢得什么。
I think a lot of companies are making routers because it's fashionable. I think they're seeing growth happen here or they're making gateways at least. First, I think there are two issues with that. First, it immediately puts you in the mindset of copying instead of winning something.
你是在为存在而战,而不是为胜利而战。也许你只是想服务好现有客户,希望看到一些 AI 增长发生。
You're sort of playing to exist rather than playing to win. And maybe you're just trying to serve your existing customer base, and you want to see some AI growth happen.
我认为这立刻就把这种网关产品,远远甩在了那些全力投入的公司后面好几个月。比如我 100% 专注于打造最好的路由器和网关以及 LLM 市场,这体现在我们的产品、我们内部创建的基准测试,以及我们如何看待自己与竞争对手的对比上。对我们来说这不是支线任务,但对其他一些公司来说可能是。
I think that immediately puts that gateway many, many months behind the companies that are fully focused on it. Like I am 100% focused on building the best router and gateway and LLM marketplace, and it shows in our product and in the benchmarks that we create internally and how we see ourselves compared to the competition. This is not a side quest for us, like it may be for some other companies.
另一个问题是,它降低了你所有用户的杠杆。我深信要赋予用户和开发者更多杠杆。从根本上说,让他们访问更多模型,就是让他们在 AI 的所有创新中获得更多杠杆。你希望能够访问所有模型,减少对任何单一模型的依赖。如果你构建在路由器或网关上,却无法访问整个市场、获得完全灵活性或完全可定制性,那就无法获得整个生态系统的全部杠杆,那你就像被切断了。你切断了公司所有员工所需的东西。所以 OpenRouter 从根本上讲是给人们更多选择,因为那给了他们更多杠杆。
The other problem is that it reduces the leverage of all of your users. I really deeply believe in giving users and developers more leverage. Fundamentally, giving them access to more models is about giving them more leverage over all the innovations that happen in AI. You want to be able to access them all. You want to reduce your dependency on any individual one. If you build on top of a router or a gateway that doesn't give you access to the full market or full flexibility or full customizability, it doesn't give you the full leverage of the whole ecosystem, then you're kind of being cut out. You're cutting out all your employees at your company of things that they need. So OpenRouter is fundamentally about giving people more choice, because that gives them more leverage.
你以 5.5% 的抽成来做到这一点。
You do that at a price at 5.5% take.
那是我们的按量付费计划。我们随后增加了企业计划,采用完全不同的定价模式,到目前为止非常成功。它基于承诺消费额,然后对承诺消费额不收取费用。
That's our pay-as-you-go plan. We then added an enterprise plan with a totally different pricing model, and it's been very successful so far. It's kind of based on committed spend, and then no fees on that committed spend.
因为那正是我想问的。最终,公司会在规模小时喜欢它。然后随着规模扩大,你会想:“天哪,这真的贵得离谱。我现在干脆自己构建路由技术吧,因为它实际上已经成为我成本基础中如此重要的一部分。”
Because that was going to be my question. Ultimately, companies will love it small. And then as you scale, you're like, "Shit, this is really freaking expensive. I'll just build my own routing tech now because it's become such a significant part of my cost base, actually."
我有点觉得有些公司只是还没意识到我们有企业计划。还有些是我们的错,因为我们没有一个更好、更详细的定价模式。我们很快会推出一个商业自助计划,也会让事情更合理。如果你有自己的推理,如果你自带推理到 OpenRouter,如果你自带密钥,那这笔费用就免了。所以这相当于是针对我们为你提供的推理服务。当你使用 OpenRouter 的能力且不在我们的企业计划中时,才会产生这笔费用。否则,我们需要能够预测需求,所以这就是我们做这些承诺消费的原因。
I kind of figured some of those companies just haven't realized we have an enterprise plan. And some of it is our fault for not having a better, more detailed pricing model. We're soon going to introduce a business self-serve plan that also makes it make a lot more sense. And if you have your own inference, if you bring your own inference to OpenRouter, if you bring your own keys, that fee goes away. So it's fairly, you know, for inference that we are providing you. When you go into OpenRouter's capacity and you're not on our enterprise plan, that's when that fee comes in. Otherwise, we need to be able to predict demand a little bit. So that's why we do these committed spends.
三年后 OpenRouter 的主要收入来源会是什么?
What will be the main revenue line of OpenRouter in 3 years' time?
我认为这在很多方面都取决于经济。如果整个 AI 市场在未来四年保持目前的增长速度,比如每年 10 到 15 倍,甚至更多,那是巨大的增长。在那种情况下,我预计人们会继续低估他们需要的推理量。因此,我们的收入将主要由今天主导的因素主导,也就是我们帮助企业处理计划外的推理容量,包括企业和初创公司。这是 OpenRouter 最擅长的。当你需要尝试你没想到要尝试的模型时,当你对特定模型的使用量超过预期时,我们通过提供最佳的故障转移和最佳的正常运行时间,确保这对你的公司不会成为问题。在市场持续低估其推理需求并以这种速度增长的情况下,这确实是一件好事。
I think it's going to depend on the economy in so many ways. If the overall AI market keeps growing the way it's been growing over the next 4 years, with like 10 to 15x every year, or potentially more, it's a lot of growth. Under that world, I would expect people to continue to underestimate how much inference they're going to need. And thus our revenue is going to be dominated by the same things that dominated today, which is us helping people with unplanned inference capacity, both enterprises and startups. That's what OpenRouter is best at. When you need to try models that you weren't expecting you need to try, when you're using more inference than you thought you were going to use on particular models, we make sure that is not going to be an issue for your company by providing the best failover and best uptime. And this is really a good thing to do when the market is continuously underestimating its inference needs and growing at this rate.
如果这种增长速度在未来四年持续下去,那将是疯狂的增长。经济有其局限性。我可以看到主要的中小企业 SaaS 为我们增长,并且在增长速度不再每年 10 倍或 15 倍时,它们也需要为我们增长。
If this growth rate continues over the next four years, it's going to be a wild amount of growth. The economy has some limits to it. I can see major SMB SaaS growing for us and needing to grow for us whenever growth does not keep going 10x or 15x per year.
好的,我们看到代币价格在 18 个月内下降了 90%。代币价格的下降对你的业务是有利还是有害?因为显然你从消费中抽成。如果价格下降,消费更高效,表面上对你的业务不利。你分到的蛋糕在缩小。
Okay, we've seen token prices fall 90% in like 18 months. Is the reduction of token prices helpful or hurtful to your business? Because obviously you have a take on spend. If they come down and spend is more efficient, seemingly it's bad for your business. You have a shrinking pie to take from.
嗯,很多人谈论杰文斯悖论,即当价格下降 10 倍时,使用量会增加超过 10 倍。但没有人真正做好建模。我们确实有很多具体案例证实了这一点。例如,OpenRouter 上的 GPT 5.6 Luna。OpenAI 降价 5 倍,然后与我们协调再降 2 倍。所以总的来说,Luna 的价格在过去两周内在 OpenRouter 上下降了 10 倍。你猜使用量增长了多少?13 倍。所以这几乎是一个完美的杰文斯悖论案例,价格下降 10 倍,使用量增长超过 10 倍,只多一点点。而且使用量相当稳定。它增长了,在 13 倍处趋于平稳,然后以之前达到 13 倍之前的增长速度继续增长。所以这相当有趣,而且这是一个相当低变量的案例,故事中几乎没有其他混杂变量。而且这发生在 DeepSeek 发布并拥有非常好的价格,GLM 也有非常好的价格的时候。现在 Luna 在 OpenRouter 上的使用量超过了 GLM。GLM 曾经是按代币量排名前三或前四的模型,现在 Luna 超过了它。这是 OpenAI 的模型在极长时间内首次进入我们平台按代币量排名前五的模型。所以这是一个非常大且有趣的举动。
Well, a lot of people talk about the Jevons paradox that when prices go down by 10x, the usage increases by more than 10x. But no one has really done a great job modeling it. We do have a lot of spot stories that confirm it. For example, GPT 5.6 Luna on OpenRouter. OpenAI cut prices by 5x and then in coordination with us by another 2x. So in total, the price of Luna has dropped 10x on OpenRouter over the last 2 weeks. And guess how much usage has grown? 13x. So it's a close to perfect Jevons paradox story where you drop prices 10x and usage grows by more than 10x, just a bit more. And the usage is pretty stable. It grew, flattened out at 13x, and then it's been growing at the same rate that it was growing before it hit the 13x multiple. So that's pretty interesting and it's a pretty low variable, there are few other confounding variables in the story. And it was also done in the middle of DeepSeek launching and having a really good price and GLM having a really good price. Now Luna is being used more than GLM on OpenRouter. GLM used to be one of the top three or four models by token volume, and now Luna is past it. This is the first time OpenAI has had a model on our platform in the top three to five models by token volume in an extremely long time. So this was a really big and interesting move.
你的代币量在多大程度上反映了市场?因为大约,我可能记错了,但也许占代币总量的 1.5% 到 2%。
How reflective of the market are your token volumes? Because it's about, I may get this wrong, but maybe 1 and a half to 2% of say token volumes.
那么,这些排名有多大的代表性呢?因为很多人当我说“哦,我在 Open Router 上看到的前五名模型都是中国的”时,他们会说,“哦,好吧,Harry,无意冒犯 Open Router,但这并不能反映市场,大多数使用前沿模型的人不会通过它。”比如他们使用前沿 API,所以没被算进去。你的排名在多大程度上反映了真实的 token 使用量?
And so, how reflective are they? Because a lot of people when I say, "Oh, the top five models when I look at Open Router are all Chinese." What does that mean? They'll go, "Oh, well, Harry, no offense to Open Router, but like it's not reflective of the market and most people who use Frontier, it doesn't go through that." Like they use Frontier APIs and so it's not counted. To what extent are your rankings reflective of true token usage?
我们试图通过调查用户或参考其他调查来估算偏差。我认为我们确实偏向于那些认同我们论点的人,即未来是多模型的,公司需要多个模型。现在仍然有一些公司,我很少遇到,但确实存在,他们会说,“哦,是的,我们是 OpenAI 的专属客户,只用 OpenAI 的模型。”所以我们看不到这些公司,我认为这些公司主要关注超大规模云厂商、OpenAI、Anthropic 和 Gemini。所以我们可能确实低估了前沿模型。但我认为随着时间的推移,我们的论点会越来越普遍,当其他公司意识到“我们需要使用其他模型”时,我们的数据就会变得更有代表性,随着规模的扩大,数据总体上也会更有代表性。所以我希望数据会越来越好。
We try to estimate how they're off by, you know, just surveying people sometimes or looking at like the surveys other people have done. I think we have a definite bias to people who believe our thesis, which is that the future is multi-model and companies who want multiple models. And there are still companies out there. I basically rarely very rarely run into them now. But there's still companies out there that are just like, "Oh, yeah, we're an Open AI shop. Like we've won, you know, we only do Open AI models." And so, we're not going to see any of those companies, and I think those companies are primarily focused on like the hyperscalers, Open AI, Anthropic, and Gemini. So, we do probably undercount the Frontier models. But I think over time our thesis is becoming more and more common to see in other companies, in the moment that they're like, "Oh, yeah, like we need to use other models." Then our data becomes more representative, and as we scale up, the data becomes more representative in general. So my hope is that it just becomes like better and better data over time.
我能问你吗,Alex Karp 在 CNBC 上以他那种非常有活力的方式说,公司害怕与前沿模型提供商合作。
Can I ask you, Alex Karp said on CNBC in his rather wonderfully energetic way that companies are terrified of working with Frontier model providers.
嗯。
Mhm.
你觉得他们是这样吗?
Do you think they are?
我没看过他在那说的内容。当我与客户交谈时,确实有一些紧张情绪,尤其是在 Claude Design 在 Figma 出现的时候。我确实看到了这一点,而且我认为对于一家只是在智能之上构建一个薄薄的上市包装的公司来说,确实存在真正的担忧。比如,“嘿,我们是一家把 AI 带入这个市场的公司,通过正确的集成和定制系统提示来实现。”如果模型实验室不在乎那个市场,你会没事的,而且会有很多这样的市场。但模型实验室最终会有几个动机来追你。一个是让他们关心的公司内部的多个团队依赖他们。所以这就是我关于为什么 Claude Design 具有战略意义的理论。虽然对 Anthropic 来说这不是一笔巨大的收入,但它确实让设计团队真正关心 Anthropic 的模型。所以他们想要的公司,现在有了另一个真正想留在 Anthropic 的团队。所以这种团队策略会让你与模型实验室竞争。所以我认为那些发现自己“哦,我们正在为一个现在对模型实验室具有战略意义的团队构建产品,这些公司是他们真正关心的”的公司,这就是我看到的近期威胁最大的地方。
I haven't seen what he talked about there. When I talk to our customers, there was definitely a little skittishness, particularly when Claude Design came out around Figma. And that part I did see, and I do think that there are real concerns for a company that is kind of building a thin go-to-market wrapper around intelligence. Like, "Hey, we are a company that brings AI to this market and does so by doing the right integrations and customizing the system prompt." You're going to be fine if the model labs don't care about that market, which there'll be many markets like that. But the model labs have several incentives to go after you eventually. One is getting multiple teams within companies they do care about to be dependent on them. So this is my theory behind why Claude Design was strategic. While it's not a massive amount of revenue for Anthropic, it does get the design team to really care about Anthropic models. And so the companies that they want, they now have another team that really wants to stick to Anthropic. So that team strategy can make you compete with the model labs. And so I think companies like that that find themselves like, "Oh, we're building a product for a team that has now become strategic for the model labs, for companies they actually care about." That's where I see probably the most near-term threat.
你认为 Claude Design 会对 Figma 的业务产生重大影响吗?我今天和很多创始人交谈,他们直言不讳地从 Figma 转向 Claude Design,这正在蚕食 Figma 的使用量。你看到了吗?你觉得这会发生吗?
Do you think Claude Design will have a meaningful impact on the Figma business? I speak to many founders today who are bluntly switching from Figma to Claude Design, and it's cannibalizing that Figma usage. Do you see that? And do you think that will happen?
我看到很多设计师尝试了 Claude Design,包括我们自己的设计师。但到目前为止,我还没有听到重复使用的案例。我不知道。老实说,我没有和很多设计师讨论过这个。我当然没有听到很多关于 Claude Design 的讨论。如果你只看 Figma 的数字,它们相当不错。他们的盈利非常惊人。
So I saw a lot of designers try out Claude Design, including our own. But so far I haven't heard of the repeat story. I don't know. Honestly, I have not talked to very many designers about this. I certainly haven't heard a lot of chatter about Claude Design. And if you just look at the numbers for Figma, they're quite good. They have very incredible earnings.
你不想公开,老兄。你看你有很好的数字,Figma 却下跌了。我就想,“可怜的 Della,给我,你在干什么?”
You don't want to be public, dude. You see like you have great numbers, Figma down. I'm like, "Poor Della, like give me the, what are you doing?"
嗯,是的。那太疯狂了。
Well, yeah. That was crazy.
我的意思是?他们就像,“真的吗?拜托。”我们刚才在讨论我们提供的不同模型,公司愿意与前沿模型合作。模型开发的速度感觉非常快。你觉得在未来 1 年、2 年、3 年,我们会看到同样的模型开发速度继续吗?
What I mean? They were like, "Really? Come on." We were talking about the different models that we have on offer, where the company is willing to work with Frontier models. The rate of model development feels immense. Do you think we will see the same rate of model development continue over the next year, 2 years, 3 years?
前沿模型开发还是通用模型?
Frontier model development or general model?
通用模型,包括前沿和开源。
General model, both Frontier and open.
是的,是的。因为我的意思是每天都有两三个、四个新模型。
Yeah, yeah. Just cuz I mean every single day there's two, three, four new models.
七月,我们发布了 70 个模型。大约每 10 小时一个模型。
In July, we launched 70 models. It's about one model every 10 hours.
还有一些智能体实验室正在启动,它们最终可能会制造模型。比如 Jeff Dean 现在正在从 Google 启动一个智能体实验室。以制造智能体闻名的公司有动力创建自己的模型,这是一个非常明确的动力,通过智能体分发模型。我们甚至还没看到那开始。抱歉,我们看到了开始,但还没有真正加速。比如 Cognition 有模型,Cursor 有模型。Lovable 有模型了吗?我觉得还没有。
There's some agent labs starting, too, that are all going to kind of like probably make models eventually. Like Jeff Dean is starting an agent lab right now from Google. The companies that are known for making agents have an incentive to create their own model, a very clear incentive to create their own models and distribute it through the agent. And we haven't even seen the start of that. Sorry, we've seen the start of it, but we haven't seen it really pick up. Like Cognition has a model, Cursor has a model. Does Lovable have a model yet? I don't think so.
没有公开。
Not publicly.
是的。所以,智能体实验室,我认为,会开发模型。这种压力来自 GPU 制造商如 Nvidia,他们希望在该领域创造更多竞争和多样性,加上我们,加上那些只想尝试新事物的投资者,所有这些都可能以某种神经发散的方式提升智能。我认为这些是强大的激励。我认为它们足以激励更多创始人创建新的实验室。如果美国的开放权重模型加速发展,那么这将给这些新实验室一个不是中国的训练基础。这可能会催生更多美国的新实验室。
Yeah. So, the agent labs are going to, I think, develop models. This pressure from both the GPU makers like Nvidia to create more competition in the space and create more diversity in the space, plus us, plus investors who just want to try new things, that all could improve intelligence in some neurodivergent way. I think those are strong incentives. I think that they're enough to incentivize more founders to make new labs. And if American open weight models pick up in steam, then it gives these NeoLabs a base to train on that's not Chinese. Which will then probably create more American NeoLabs.
你认为我们应该对中国开源模型的速度和质量感到担忧吗?
Do you think we should be concerned by the rate and quality of Chinese open models?
我们应该。我们落后了。美国仍然非常非常落后。我认为事情正在加速。我们有 Poolside,我们有 Thinking Machines,我们有 RC。
We should. We're behind. America is very, very behind still. I think things are picking up. We have Poolside, we have Thinking Machines, we have RC.
你对此感到有责任感吗?我的意思是,你知道,你是一家路由业务,你可以把公司路由到中国模型,谁知道人们担心后门、中国 ICCP 的参与。
Do you feel a sense of responsibility for that? And what I mean by that is like, you know, you are routing business and you could route a company to a Chinese model that who knows people are worried about backdoors, Chinese ICCP involvement.
你可能是这些模型的交付者。你对此有责任感吗?
You could be the deliverer of that to those models. Do you feel a sense of responsibility for that?
我们确实有责任为所有这些模型提供安全访问。客户信任是我们的首要目标。如果某个模型被认为不安全,我们就会将其从平台下架。如果有办法以不安全的方式使用它——我的意思是,所有模型都有办法被不安全地使用——那么我们相信要用技术来确保安全,并与模型实验室合作,了解他们那边是怎么做的,以便我们能达到最先进水平或更好。我们花了大量时间确保我们的实践与实验室出来的最佳成果相匹配,甚至更好。而且因为我们是探索所有模型、首次发现它们的途径,我们也是在整个公司部署安全措施的良好焦点。例如,我们有提示注入保护。你只需打开它,就能立即标记看起来像提示注入的提示。我们有 PII 缩减。我们还有几种不同的功能,你可以一键自动开启,为所有推理增加一层安全防护。我们这样做的目的是让企业觉得他们可以安全地部署新模型,让员工试用。我有点把模型看作互联网。你不能因为互联网上有一些坏东西就在公司禁止互联网。你可以设置护栏,也应该设置。你需要用 AI 来构建尽可能好的护栏。这就是我们正在做的。
So we do feel a responsibility to have safe access for all these models. Customer trust is our paramount goal. If one of these models is unsafe to use, generally considered unsafe, we pull it from the platform. If there's a way to use it in an unsafe way—I mean, there's a way to use all the models in an unsafe way—then we believe in using technology to make it safe and to work with the model labs themselves to figure out how they're doing it on their side, so that we can be state of the art or better. We spend an enormous amount of time making sure that our practices match with the best things we're seeing coming out of the labs, or better. And because we're a way of exploring all the models and finding them for the first time, we're a good focal point for deploying safety measures across your whole company. For example, we have prompt injection protection. You can just turn it on and immediately flag prompts that look like prompt injection that's trying to happen. We have PII reduction. We have a couple of different things that you can automatically just turn on with a click and get an added safety layer on top of all of your inference. And we build that so that enterprises feel like they can safely deploy new models and that their employees can try them out. I think of the models a little bit like the internet. You can't just ban the internet at your company because there are some bad things on the internet. You can create guardrails, and you should. You need to use AI to build the best possible guardrails that you can. So that's what we're doing.
你真的了解 Moonshot 或阿里的 Kuan 内部发生了什么吗?这些是中国深处的极其隐秘的组织。
You actually know what's going on within Moonshot or Alibaba with Kuan? Like these are incredibly secretive organizations in the depths of China.
我不能假装知道他们内部发生了什么。作为一家美国公司,我们会遵循美国的最佳实践,确保我们不会做出不负责任的事情。
Can't pretend I know what's going on inside of them. As a US company, we're going to follow the best practices of what happens in the US to make sure that we're not doing something irresponsible.
你认为美国公司更担心前沿模型还是中国模型?
What do you think US companies are more nervous of, frontier models or Chinese models?
我认为他们通常更担心前沿模型。部分原因是关于数据政策存在更多困惑,比如他们发送的提示到底发生了什么、存储在哪里、如何被查看。而且你不能在自己的机器上或选择的服务商上运行它们。这立即在许多企业中造成了所有不确定性。而且这种不确定性他们也能模式匹配。这很像在自己的基础设施上运行与在 VPC 中运行的区别,以及知道谁能看到数据。
I think they're more nervous about frontier models usually. Partly because there's much more confusion around the data policy about what's actually happening to the prompts they're sending, where they're being stored, and how they're being looked at. And you can't run them on your own machine or in a provider of your choice. That immediately creates all this uncertainty in a lot of enterprises. And it's uncertainty that they can also pattern match. It's very similar to running on their own infra versus running in their VPC, and knowing who can see the data.
没错。他们更担心总部位于硅谷的美国公司,在那里你可以看到、触摸到、感受到总部和领导者。这真是个奇怪的世界。
Exactly. Like they're more nervous of US companies headquartered in Silicon Valley where you can see and touch and feel the headquarters and the leaders. It's just what a strange world to be in.
是的,这很奇怪,尤其是前沿模型现在拥有最大的网络安全姿态。他们在网络安全方面做得最好。
Yeah, it is very strange, especially with the frontier models having the biggest cyber posture right now. They're doing the best at cyber.
公司有点摆姿态,哈哈,我们黑掉了某人?先是 OpenAI,然后是 Anthropic,然后是扎克伯格出来说“我不想错过派对,我们也做了。”
Company kind of posturing, haha, we hacked someone? First you had OpenAI, then you had Anthropic, and then you had Zuck coming out, 'I don't want to miss the party, we did too.'
是的,我认为他们必须谈论这件事。正确的做法是当你的模型涉及网络安全事件时公布出来。掩盖是行不通的,或者长期来看不会奏效。而且看起来他们确实都在吹嘘。但真的,如果你处在他们的位置,某个模型出了事,你必须决定是否公布,我认为正确的做法是公布,不管人们会怎么解读。所以我不知道,我怀疑他们是否真的在考虑重罪条款之类的。
Yeah, well, I think they have to talk about it. The right thing to do is to reveal when there's been a cyber incident involving your model. Covering it up doesn't work, or it's not going to work in the long term. And it certainly looks like they're all bragging about it. But really, if you were in their position and something happened with one of the models, and you had to make the choice about whether to publish it or not, I think the right thing to do is to publish it regardless of how people are going to spin it. So I don't know, I doubt that they're actually thinking of the felony bench or whatever it's called.
最新的 Kimi 模型引起了广泛关注,它有多重要?它像大家想的那么重要吗?
How significant was the latest Kimi model which got so much attention? Was it as significant as everyone thought?
它相当不错。它在网络安全能力上不如前沿模型,在长距离、长时程任务上,我认为它仍然落后于前沿模型。但 GLM 5.2 对开放权重模型来说是一个很大的进步。Kimi 有点像 Moonshot 达到了那个水平。我大致是这么看的。而且 Kimi 也是一个很好的写作者。声音和语调都相当不错。而一些前沿模型在编码能力提升后,尤其是,会出现声音退化。我当时想,“天哪,我再也读不懂这个输出了。输出听起来像是你提出的四个论点中有三个是对的,一个是转折点。你知道,然后就是‘但是,问题在这里。’”有时就是无法读懂他们在说什么。这些东西是可以修复的,但 Kimi 我认为一直有相当有趣的写作。
It's quite good. It's not cyber capable in the same way the frontier models are, and on long-range, long-horizon tasks, I think it's still a bit behind the frontier models. But GLM 5.2 was a really big step for open weight models. Kimi was kind of like Moonshot getting up to that step. That's a little bit how I see it. And Kimi's also a very good writer. The voice and tone are both pretty good. Whereas some of the frontier models have voice degradation that happens when they get better at coding, especially. And I was like, 'Oh my god, I can't read this output anymore. The output sounds like three of the four arguments you made are right and one is a turning point. You know, and do do do do do do. Here's the rub.' It just sometimes becomes impossible to read what they're saying. And this stuff is fixable, but Kimi, I think, has always had pretty interesting writing.
12 个月后,美国开源和中国开源之间的鸿沟会比今天更大还是更小?我担心会更大,因为当你有 DeepSeek 时,它就成了中国的国家冠军。我的意思是,习近平会说,“这是我们的 AI 马。我会把所有资金和努力集中在这上面,我会一直支持这个生态系统到最后。这是赢家。”然后当你看到另一个 Moonshot 出现时,突然所有监管都被搁置,所有政策都被推到一边,所有资金都变得可用。
In 12 months, will the chasm between US open source and Chinese open source be bigger or smaller than it is today? And my fear is that it will be bigger because when you have DeepSeek, it becomes a national champion in China. And I mean Xi Jinping is going, 'This is our AI horse. I will concentrate all of my money and efforts behind this and I will supplement this ecosystem to the end. This is the winner.' And then when you see another Moonshot come out, suddenly all regulation gets moved aside, all policy gets pushed aside, all funding becomes available.
嗯。
Mhm.
一切都被允许。你可以自由奔跑。这些人在实现最终目标的能力上不受限制。而 OpenAI、Anthropic 以及美国所有其他提供商,尤其是开源,你要为美国开源模型筹集数十亿美元——实际上有点难。并非不可能,但更难。商业模式存疑。AI 研究非常昂贵,而且你在与 OpenAI 和 Anthropic 竞争。我认为他们所处的相对格局意味着中国的开源提供商天生就有优势,可悲的是。
Everything is allowed. You are free to run. And these guys are unabridged in their ability to do whatever they want to get to the end goal. Whereas OpenAI and Anthropic and all the other providers in the US, especially open source, you're going to try raising billions of dollars for a US open source model—bit tough actually. Not impossible at all, but tougher. Business model questionable. AI research is super expensive and you're competing against OpenAI and Anthropic. I think the comparative landscapes they sit in mean that the Chinese open source providers are just inherently advantaged, sadly.
他们有非常非常好的研究人员,我认为美国人严重低估了这一点。
They have very, very good researchers, and I think Americans underestimate that a lot.
我确实认为他们会担心自己模型的网络安全态势,而且他们似乎非常担心审查模型,以及审查模型能向人们提供的信息。所以,虽然今天人们抱怨美国模型因网络安全而审查更多,我不确定这种情况会一直持续。随着中国模型对中国越来越重要,他们会怎么做?他们会拆除防火墙吗?他们会放弃在模型周围设置防火墙吗?我对中国了解不多,但似乎有点奇怪,他们似乎不太在意模型——或者我从未见过有人分析过,用 DeepSeek 能做什么,而在中国境内通过互联网做不到的事。比如你能访问什么信息。我从未见过有人真正深入探讨。DeepSeek 能绕过防火墙多远?如果防火墙对中国很重要,如果它在 10 年后仍然重要,那一定会有所改变。
I do think they're going to be concerned about the cyber posture of their models and they do seem very concerned about censoring the models and censoring the information that the models can provide to people. So, you know, while today people complain about American models censoring more due to cyber, I'm not sure that's always going to hold. And as the Chinese models grow in importance for China, I mean, what are they going to do? Are they going to drop the Great Firewall? Are they going to give up on putting the firewall around the models? I don't know that much about China, but it does seem kind of strange that they don't seem to care more that the models—or I've never seen anyone do a profile of what you can do with DeepSeek that you can't do with the internet in China that's available to you within the border. Like what information you can access. I've never seen anyone do a real deep dive. How far past the firewall does DeepSeek go? If the firewall matters to China, if it's going to matter in 10 years, something's going to change.
有趣的是,我们看到中国模型在国外的能力非常强大。中国模型在国内的能力实际上相对有限。尽管护栏极其严格且具有禁止性。所以,具有讽刺意味的是,它们在国外比我们强得多,在国内却表现糟糕。
Well, what's interesting is that we see the abilities of the Chinese models outside of China is immense. The abilities of the Chinese models inside China is actually relatively limited. Although the guardrails are incredibly stringent and prohibitive. So, it's ironic that they are incredibly superior to us, domestically, terrible.
有意思。
Interesting.
我刚刚让我亲爱的朋友 Jason Lang 给我一个运行 SAS 回来,他说:“我搞不清 DeepSeek 上星巴克几点开门。好像没有提供。”他会说“不允许”。
I literally just had my dear friend Jason Lang give me a run SAS to come back and be like, "I couldn't figure out what time Starbucks opened on DeepSeek. Like what wasn't on offer." Would say like not allowed.
是啊,是啊。
Yeah, yeah.
太疯狂了。非常基本简单的请求。
Wild. Very basic rudimentary requests.
我们在谈论所有这些不同的模型。我想说的是,忠诚度呢?你在生态系统中占据了一个绝佳的位置,可以看到一切。我们今天看到开发者对模型有任何忠诚度吗?
We're speaking about all of these different models. And the thing I think is, what about loyalty? And you have this incredible seat in the ecosystem where you can see everything. Do we see any developer loyalty today with models?
说实话,我们确实看到了一些。我们努力让切换成本接近于零。这样当新模型出现时,人们可以很容易地试用。但我们也衡量所有模型的留存率和流失率。当模型实验室要求时,我们也会与他们分享这些数据。这样他们就能知道:“哦,对于我刚发布的模型,哪些模型带来了流量?”对于那些用户,当他们离开时,他们去了哪些模型?我们很快会把这些数据越来越多地提供给世界。我们确实在流失数据中注意到,有些开发者即使有更好的模型,更适合他们用例的模型,也会持续坚持使用某些模型。我认为这可能是几个根本因素的组合。一是我的应用能用,我不想弄坏它。如果支持机器人开始说一些我没预料到的奇怪话,为什么要增加更多麻烦?我已经做了所有优化,我已经在它周围设置了所有这些护栏。二是新模型不一定会让你的定价更好。一般来说,当前模型的价格会随着时间下降,尤其是在实验室取得新进展时,你会看到智能跃升,但价格曲线也会跃升,然后会随着时间下降。所以,即使对于开放权重,切换到最新模型也不一定是最具成本效益的。第三个原因是对输出的根本信任。如果我使用一个模型来做我的工作,我喜欢它说话的方式,我可能有一些评估,比如个人评估。很多人都有这些个人评估,就是他们给模型的一些随机测试。如果随机测试在新模型上表现不佳,他们就会说:“好吧,反正我喜欢 Kimmy K 2.6。”
Honestly, we do see some. We try to make switching costs close to zero. So that when new models come out, people can try them out really easily. But we also measure retention and churn from all the models. We share this data with model labs, too, when they ask for it. So that they could know, "Oh, for my model that just came out, which models drove traffic to it?" And for those users, when they leave, which models are they leaving to? And we'll make this more and more available to the world soon. And we do notice in the churn data. There are developers who kind of continuously stick to models even when there are better models out there, better models for their use cases. I think it's a combination of a couple probably root factors. One is like my app works and I don't want to break it. If the support bot starts saying something weird that I didn't expect, why add more headache? I've already done all this optimization and I've already put all these guardrails around it. Another is new models are not necessarily going to make your pricing better. In general, what happens is that the current models' price goes down over time and especially when new advancements in the labs happen, you'll see intelligence jump, but the price curve also jumps and then will start going down over time. So, it's not necessarily the most price-effective thing to do to shift over to the newest model even for open weights. The third reason is it's just fundamental trust in the outputs. If I'm using a model to do my work and I like the way it talks, I probably have some eval, like a personal eval. A lot of people have these personal evals that are just these random tests that they give the models. If the random test doesn't look really good on the new model, they'll just be like, "Good. I liked, you know, Kimmy K 2.6 anyway."
人们之前认为记忆会是保留机制。OpenAI 有我所有的历史提示。它知道我在伦敦,我做播客,这将使它成为对我更好的模型。记忆不再是保留机制了吗?
People thought before that memory would be the retentive mechanism. Well, OpenAI has all of my previous prompts. It knows that I live in London, I do podcasting, and that will make it a better model for me moving forward. Is memory no longer a retentive mechanism?
记忆真的很有趣。我一直认为它是一种保留机制,问题在于它存在于哪里。它会存在于模型中吗?它会存在于推理提供商那里吗?它会存在于应用中吗?它会存在于基础设施提供商那里吗?路由器?我的猜测是,所有这些层都会尝试以不同的方式拥有记忆。将记忆放在每一层都有优势。如果你把它放在应用中,那么记忆拥有最多的应用相关上下文,并且与模型无关。如果你把它放在模型中,记忆可能在个性化基准上表现最好,也许拥有最好的终极智能,我认为模型实验室会致力于记忆。然后最终的问题可能是,有没有一个好的组合?我能否同时使用模型中的记忆和基础设施层或应用层的记忆?这会混淆模型吗?我们还不知道。我确实认为,一个层不可能捕获所有有价值的记忆,因为应用拥有很多模型实验室没有的重要上下文。而模型实验室为了让这工作,他们必须激励应用提供这些上下文。
Memory is really interesting. I've always thought of it like it is a retentive mechanism and the question is where it lives. Is it going to live with the model? Is it going to live with the inference provider? Is it going to live with the app? Is it going to live with the infrastructure provider? The router? My guess is that all of those layers are going to try to own memory in different ways. There are going to be advantages to sticking your memory in each layer. If you stick it with the app, then the memory has the most app-related context and is model agnostic. If you stick it with the model, the memory might perform the best on personalized benchmarks and perhaps have the best ultimate intelligence and I think the model labs are going to work on memory. And then the ultimate thing might be like is there a good combination? Can I use memory in the model and memory at the infrastructure layer or the app layer at the same time? Is that going to confuse the model? We don't know yet. I do think that it's impossible for one layer to capture all valuable memory because the apps own so much important context that the model labs don't have. And the model labs in order to get this to work, they'll have to incentivize the apps to give them that context.
说到应用和模型,比如 Claude Code、Cursor、Bundle、Model 和 Harness。当智能体和工具框架结合在一起时,路由器是否在有机会独立之前就被吸收到智能体框架中了?
Speaking of apps and the models that Claude Code, Cursor, Bundle, Model and Harness. Is the router absorbed into the agent framework before it ever has the chance to be independent when you have the agent and the harness together?
工具框架非常有趣,因为在早期,我们早期的一个赌注是,大多数应用低估了用户选择模型的愿望。在 2023 和 2024 年的早期,大多数应用甚至不清楚底层使用的是哪个模型。他们说:“哦,人们不会在乎那个。他们只想要 AI。”而我们当时的一个强烈信念是,不,人们会想使用特定的模型。他们会在乎他们在和谁说话。就像,当我试图解决问题时,我想知道我在和哪个员工说话,模型也会类似。而且这已经实现了,比如在 Notion 中,你可以选择你对话的模型。即使你会认为这样的应用可能想完全隐藏它。
The harnesses are pretty interesting because in our early days, one of our early bets was that most apps were underestimating the desire for users to choose the model. Like most apps in the very early days in like 2023 and 2024, it wasn't even clear which model was being used under the hood. They were like, "Oh, people are not going to care about that. They just want AI." And one of our strong convictions then was that no, people are going to want to use particular models. They're going to care about who they're talking to. It's like, I want to know which employees I'm talking to when I'm trying to solve a problem and models will be kind of like that. And that has played out, like in Notion, you can choose the model that you talk to. Even though you would think an app like that might want to obscure it completely.
类似的事情也发生在 harness 上,尤其是开发者。他们开始对不同 harness 产生好感,因为这是一种用户体验。所以,我认为这是支持 harness 会继续存在的最有力的论据。不是说它们会跟模型捆绑在一起,因为事实上,随着模型变得更好,它们变得更足智多谋,系统提示词里塞的那些垃圾反而成了累赘。Anthropic 发表了一篇好文章,展示了他们从系统提示词里删掉一些东西后,后来与用户提示词的矛盾突然变少了,模型表现也更好了。我们看到现在很多 harness 都在删除代码,以便在最新的前沿模型上表现更好。我认为这意味着 harness 并不坏。事实上,我认为未来会出现更多 harness,因为这是在模型之上构建用户体验的一种方式。这是非模型实验室的开发者拥有用户关系的一种方式,而且这一层对经济来说将极具价值。
A similar thing happened with harnesses, particularly with developers. They started to build an affinity to different harnesses because it's a user experience. So, I think that is my favorite argument for why harnesses are going to stick around. Not that they're being bundled with the models, because in fact, as models get better, they get more resourceful, and the junk that gets thrown in the system prompt just becomes a handicap. Anthropic published a good article about this where they showed that they got rid of stuff from the system prompt and suddenly fewer contradictions showed up later on with user prompts, and the model performed better. We're seeing a lot of the harnesses right now are deleting code in order to perform better with the latest frontier models. I think that means harnesses are bad. In fact, I think we'll see more harnesses come up in the future because it's a way of building a user experience on top of models. It's a way for developers who are not model labs to own a user relationship, and that is just going to be incredibly valuable for the economy to have that layer.
我这么说可能会被骂。harness 和应用有什么区别?感觉如果我要谈论 harness,那就是在玩弄辞藻。那不就是个应用吗?拜托。
I'm going to get killed for this. What's the difference between a harness and an app? Feels like it's word wank if I ever have to talk about harnesses. Is that not an app? Like hello.
是的。
Yeah.
这不就是 API 对应用做过的事吗?
Is that not what APIs did for apps?
是的,但 harness 更可靠、更确定,也更容易让用户理解,因为 harness 基于 Unix。它们都有这个特点,而且模型在 Unix、bash 命令上训练得非常好。而如果我让一个 harness 去云端编排一个应用,它会想:‘天哪,这个应用有 API 吗?怎么登录?需要你的密码吗?要不要启动一个虚拟浏览器?’这会很慢。我会搞定的。好吧,我启动了浏览器,现在需要你的密码,我要找到输入框把密码放进去,而且这个应用里可能某个地方有 API。我需要查文档来搞清楚。好了,现在我拿到 API 了。围绕应用进行组合时,有太多未知的未知。而围绕 harness 进行组合时,未知的未知非常非常少。所以我认为这给了开发者更多灵活性,而且是他们可以检查的灵活性。API 调用时,你只看到一大堆代码在屏幕上飞过。而 harness,我可以跳进去看看发生了什么,并用英语讨论它。所以它更用户友好。
Yes, but it's much more reliable, deterministic, and easy for users to grok with a harness because the harnesses are Unix-based. They all have that, and the models are so well trained on Unix, on bash commands. Whereas if I'm telling a harness to go orchestrate an app in the cloud, it's going to be like, 'Oh boy, does this app have an API? How do you log into this app? Do I need your password? Do I need to fire up a virtual browser?' It's going to be pretty slow. I'll figure it out. Okay, I fired up a browser and now I need your password, and I'm going to try to find the input to put it in, and there's probably an API in this app somewhere. I need to look up the docs to figure it out. Okay, now I've got the API. There are so many unknown unknowns when you're composing around an app. Very, very few unknown unknowns when you're composing around a harness. So I just think it gives developers more flexibility, and flexibility that they can inspect. With API calls, you're just seeing a whole bunch of code flying around the screen. With a harness, I can jump into the harness and look at what's going on and talk in English about it. So it's much more user-friendly.
我们看到 Meta 和 Muse 成为扎克伯格的重点。我们看到 AI 布线更加突出。你对 Meta 用 Muse 交付的东西印象深刻吗?
We've seen Meta and Muse really be a focus for Zuck. We've seen AI wiring front and center much more. Were you impressed by what Meta delivered with Muse?
他们做得不错,是的。从零开始建立一个全新的模型实验室需要时间,而且我肯定有很多组织债务要处理。
They've been doing a good job, yeah. It takes a while to set up a whole new model lab from scratch, and I'm sure there's a lot of organizational debt to deal with.
你认为他们会成为严肃的挑战者吗?
Do you think they will be a serious challenger?
我确实这么认为。我觉得他们有资源。他们可以在模型周围做一些有竞争力的事情,以模型实验室不太感兴趣的方式帮助人们。比如拥有一个社交网络,关注人。这对品牌来说是有价值的,也许 Grok 和 X AI 也有。他们确实需要找到自己的利基。我不太确定人们是否知道 Muse Spark 能做什么,比如什么时候用它,或者它的核心优势是什么。他们刚发布了一个编码 harness。他们现在正试图成为一个通用模型。我预计未来他们会说:‘看,我们在这方面强得多。’那对他们来说将是一个非常重要的时刻。
I do. I think they have the resources. There are some competitive things they can do around the model that help people in ways that the model labs are not as interested in doing. Like just having a social network and a focus on people. It's something for the brand that maybe Grok and X AI have, too. They do need to find their niche. I'm not quite sure people know what to do with Muse Spark yet, like when to use it or what its core advantages are. They just released a coding harness. They are trying to be a generally capable model right now. I expect that in the future they're going to say, 'Look, we are way better at this thing.' And that's going to be a really important moment for them.
我想我说过我对那印象深刻,实际上。你知道我现在用什么吗?也许提一下我们共同的朋友,但 Anastasius 和 Arena。这很奇怪。所以,我会把我的提示词放进 Arena。然后,显然,它会返回一堆不同的模型选项。
I think I have said I was impressed by that, actually. Do you know what I use now? Maybe plugging one of our mutual friends, but Anastasius and Arena. It's so weird. So, I'll put my prompt in Arena. And then, obviously, it comes back with a load of different model options.
是的。
Yeah.
然后,你知道,我带着那个回来。前几天我用了一个,Pergamon。
And, you know, I come back with that. I used one the other day, Pergamon.
Pergamon?
Pergamon?
是的。就像 Kimmy 和 Pergamon。它给你提供四个不同的选项。它带我去那些我以前从未用过的模型。实际上,Meuses 出现过几次,相当令人印象深刻。但是,我喜欢那种发现机制,能发现我从未用过的模型。我绝不会选 Kimmy,但显然我选了。我只是得到一个 chat GPT。
Yeah. And it was like Kimmy and Pergamon. And it offers you four different options. And it takes me to models that I would never have used before. And actually, Meuses have come up a couple of times to be pretty impressive. But, I love that in terms of discovery mechanism to models that I would never have used. I would never have picked Kimmy and obviously did. I just get a chat GPT.
这真的很有趣。
It's really interesting.
是的,这基本上说明了模型在那里变成了一种工具。
Yeah, I basically that goes to the point of the model there just becoming a utility there.
你这话是什么意思?
What do you mean by that?
嗯,实际上我对它们没有忠诚度。我与品牌没有关联。我去 Arena。我想看看你给我什么。给我看结果。我不在乎是 Kimmy 还是 Meuses 还是 Claude 还是 Sonic 还是别的什么。你明白我的意思吗?实际上,我只是想看看你有哪些选项。我会从里面选最好的。我宁愿并行运行四个。你相信我们会有一个前沿模型运行四个开放模型吗?前沿模型可能有 160 的智商,开放模型可能有 120 的智商。但是,那将是我们使用的模型基础设施或结构。
Well, I actually have no loyalty to them. I have no affiliation with brand. I go to Arena. And I want to see what you got for me. Show me the results. I don't care if it's Kimmy or Meuses or Claude or Sonic or whatever. Do you know what I mean? And actually, I just want to see the options you got. And I'll pick the best from there. I'd rather run four in parallel. Do you buy this whole we're going to have one frontier model run four open models? And the frontier model might be 160 IQ points and the open models might be 120 IQ points. But, that will be a model infrastructure or structure that we'll work with.
完全认为这是一个伟大的架构,每个人都需要探索。我们一直在帮助很多开发者做这件事。你有子智能体。我们有一个子智能体服务器工具,我们调整得让它非常擅长通用地使用模型。然后你有一个编排模型,当它想要完成特定任务时,它会调用子智能体。这些子智能体成本非常低,而且专注于确定性任务。与前沿模型相比,开放权重模型通常在这方面非常擅长。
Totally think that is a great architecture that everybody needs to explore. We've been helping lots of developers do this. You have sub-agents. We have a sub-agent server tool that we tuned to be really good at using models generally. And then you have an orchestrator model that calls out to the sub-agents when it wants particular tasks to get done. And these sub-agents are just very low cost, and they're focused on deterministic tasks. This is what open weight models are generally really good at compared to frontier models.
当你面对一个确定性的任务,知道输出的形态,也知道问题类型,而且这类问题已经被解决过,比如文本分类,那么你绝对应该使用 OpenRouter 上的低成本模型,然后让编排模型读取结果,继续去处理它原本要解决的未知的非确定性任务。
When you have a deterministic task where you know the shape of the output, you know the type of problem that you're working on, and it's a type of problem that has been solved, like classifying some text, for example, then you should definitely use a low-cost model from OpenRouter, and then have the orchestrator model read the results and go continue working on the unknown non-deterministic task that it was set out to do.
我想创建一个开放的美国生态系统,对吧?更多了不起的开放美国模型,我让你来负责这个项目。你会怎么做来鼓励、激励美国开放生态系统更激烈地与中国竞争?
I want to create an open American ecosystem, yeah? More amazing open American models, and I make you head of this program. What would you do to encourage, incentivize the open US ecosystem to compete more vociferously with the Chinese?
我想我会花更多时间与当前的美国实验室交流,弄清楚蒸馏中国模型对他们来说是什么样的,效果如何。蒸馏中国模型可能大有可为。开放权重模型和中国模型的好处是它们允许蒸馏,而且大多数都允许。这意味着你可以用这些模型的输出来对你正在构建的模型进行强化学习。这是所有实验室都在做的重要且常见的 AI 实践。所以我想多了解一点它的效果,但这是追赶开放权重模型的一种方式,因为它们允许这样做。
I think I would spend time talking to the current American labs a little bit more to figure out what distilling the Chinese models looks like for them and how effective it is. You can probably get pretty far distilling the Chinese models. The nice thing about the open weight models and the Chinese models is that they allow distillation, and most of them do. That means you can take the outputs of these models to do reinforcement learning on top of the model that you're building. This is a very important and common practice in AI that all labs do. So I would want to learn a little bit more about how effective it is, but it's one way to catch up with the open weight models because they allow it.
另外,当你蒸馏时,你能看到输出,所以你可以检查它们以确保对齐。所以如果你担心开放权重模型与你正在创建的模型的声音或章程不一致,在做这些强化学习 rollout 时,你更有可能发现这些问题。
The other thing is, when you distill, you see the output, so you can inspect them to make sure they're aligned. So if there's anything about the open weight models that you're worried about not being aligned with the voice or constitution of the model you're creating, you have a much better shot at catching it when you're doing these RL rollouts.
另一件我要弄清楚的事情是算力问题。我认为我们相对于中国仍然拥有巨大的算力优势,而这些新实验室需要机会。需要有一种更容易的方式把算力提供给合适的人才,在所有国家都是如此,但尤其是如果我们想创建一个有竞争力的美国新实验室体系。英伟达在这方面做得很好,但还有谷歌、TPU、亚马逊的 Trainium。我会与所有硬件公司以及新芯片合作,帮助解决算力问题。
The other thing I would try to figure out is the compute question. Compute is a huge advantage that I think we still have relative to China, and these neo labs need a shot. There needs to be an easier way to get compute to the right talent, in all countries, but especially if we're trying to create a competitive American neo lab system. Nvidia's been doing a good job of this, but there's Google, there's TPUs, there's Trainium from Amazon. I would work with all the hardware companies and also the neo chips to help with compute.
我不认为我们能长期保持算力优势。我觉得你看 DeepSeek 和字节跳动现在都在积极追求自己的芯片。出口管制意味着他们必须这样做,这是习近平在赢得 AI 战争竞赛中的头号问题。
I don't think we will have that compute advantage for long. I think you see DeepSeek and ByteDance both aggressively pursuing their own chips now. The export controls mean that they have to, and this is like the number one problem for Xi Jinping in his race to win the AI war.
我同意。我认为这是
I agree. I think it's
在 4 周内,我认为他们会在 6 个月内搞定芯片。
in 4 weeks, I think they'll manage a chip in 6 months.
是的,在芯片战争中保持领先对美国至关重要。
Yeah, it's staying ahead on the chip war is critical for America.
蒸馏是错的吗?
Is distillation wrong?
我的意思是,蒸馏是一种构建模型的技术。
I mean, distillation is a technique to build models.
但人们以冷嘲热讽的态度看待它。好吧,他们只是蒸馏了模型。
But people view it with cynicism and shade. Well, they just distilled models.
这是一种构建模型的技术。封闭弱模型实验室也蒸馏模型。比如,Sonnet 是 Opus 的部分蒸馏版本。这就是如何从大模型制作小模型。当你在生态系统中发现有用的东西时,这是教会模型新东西的重要方式。
It's a technique to build models. The closed weak model labs distill models too. Like, Sonnet is a partially distilled version of Opus. This is how you make smaller models out of bigger models. It's an important way to teach your model new things when you find something useful in the ecosystem.
我们确实认为实验室有权在服务条款中规定不允许。公司可以切断试图构建竞争性模型的人的访问权限。如果你只是想构建一个专注于做某件特定事情的小模型,这不具有竞争性,据我所知,大多数前沿实验室并不禁止。但会有市场允许和不允许的公司。我们确保帮助两家公司维护他们的服务条款。
We do think that labs have a right to say it's not allowed in their terms of service. A company can cut off access to someone who is trying to build a competitive model. If you're just trying to build a smaller model that's really focused on doing one specific thing that's not competitive, most of the frontier labs don't prohibit that to my knowledge. But there are going to be markets for companies that allow it and companies that don't. We make sure that we help both companies uphold their terms of service.
在我们进行快速问答之前,我必须问你一个问题。如果我不问,我会被骂死的。
I have to ask you one question before we do a quick fire round. I'm going to get killed if I don't ask you.
嗯哼。
Uh-huh.
有报道说你将以 100 亿美元卖给 Stripe。这会发生吗?
There are reports that you are selling to Stripe for $10 billion. Is that going to happen?
我无可奉告,但你知道,无论发生什么,我们都会执行愿景。我们所做的对生态系统至关重要。我们相信安全访问 AI,不让一个垄断者接管,而是拥有一个充满活力的模型生态系统,每个人都可以探索。当新的提供商、新的服务器工具和新的推理相关技术上线时,有一种非常简单的方式发现它并连接到你现有的所有 AI。
I can't comment, but you know, whatever happens, we're going to execute on the vision. What we're doing is critical for the ecosystem. We believe for safe access to AI where one monopoly doesn't take over and where we have a vibrant ecosystem of models that everyone can explore. And when new providers and new server tools and new inference adjacent tech comes online, there's a really easy way to discover it and connect it with all your existing AI.
如果我处于这种情况,我的反应会是,“好吧,现在我拥有公司 22% 的股份,100 亿美元就是 22 亿美元。哦,现在我是风险投资家了。”但是,不去那样想很难吗?
I was in these situations, my response would be like, "Well, now I own 22% of the company, $10 billion is $2.2 billion. Ooh. Now I'm a venture capitalist." But like, is it hard not to think like that?
我真的不去想它。
I don't really think about it.
你不想吗?
Do you not?
我个人花得不多。当我考虑如何使用个人资本时,我真的很想帮助人们解决那些不太适合风险投资的问题。这些问题处于灰色地带,人们需要解决但很难获得资金,因为它们没有附带商业模式。我认为现在在非营利领域有很多很酷的事情可以做,因为你可以用 AI 审查比以前多得多的数据。我还没准备好公开谈论,但我确实想做点什么来帮助研究人员解决这些问题,并获得资助去做。
I don't spend a lot personally. What I think about when I do with personal capital, I really want to help people work on problems that just don't lend themselves very well to venture capital. They're sort of falling in this gray area of problems that people need to solve but are really tough to fund because they don't come with a business model attached. I think there are very cool things to do now in the nonprofit space because you can use AI to review way more data than you ever could before. I'm not quite ready to talk about it publicly yet, but I do want to do something that helps researchers work on those problems and get grants to do it.
我觉得一个很酷的例子是 David Fialkow,他是 General Catalyst 的创始人之一,他基本上会找到那些不会获得电影资助的不可思议的故事,并资助它们以照亮它们,因为他认为它们非常重要。比如《异议者》,讲述了卡舒吉的故事,还有《伊卡洛斯》,讲述俄罗斯兴奋剂的故事。这些电影如果不是因为他的资助就不会获得资金,因为它们具有政治敏感性和争议性。他说,“我要让这些被禁止的故事得以呈现。”
One really cool example I think of this is David Fialkow, who's one of the founders of General Catalyst, who basically finds incredible stories that wouldn't get funded for movies and funds them to shine a light on them because he thinks they're very important. So like The Dissident, which told the story of Khashoggi, and then Icarus, which is the story of the Russian doping. These were films that would not get funded had it not been for his funding because they were politically sensitive and charged. He's like, "I'm going to enable the stories of these forbidden tales."
是的,有点像那样。我喜欢那些东西。
Yeah, it's kind of like that. I love that stuff.
他很棒。他太棒了。总之,你准备好快速问答了吗?
He's great. He's awesome. Anyway, are you ready for a quick fire round?
当然。
Sure.
好的。那么,今天 OpenRouter 上最被低估的模型是什么?
Okay. So, what is the most underrated model on OpenRouter today?
哦。好问题。
Ooh. Good one.
我的意思是,首先,我喜欢 Poolside 的模型。它们很棒。我可能喜欢我在快速问答环节的回答。好的,就像 New American Lab。他们在构建有趣的编程模型,这些模型虽小但非常高效,而且他们还在构建许多有用的工具来访问它们。
I mean, firstly, I like Poolside's models. They're great. I probably like my fire round answer. Good, like New American Lab. They're building interesting coding models that are small but highly effective, and they're building a lot of useful tools for accessing them.
好团队。
Good team.
70% 的新实验室将在未来 3 年内消亡。你同意还是不同意?
70% of neo labs will die in the next 3 years. Agree or disagree?
不同意。70% 似乎太高了。新实验室并没有那么多。如果被某个模型实验室收购算作消亡,我确实认为可能会有一些整合。如果把整合算进去,我会说 50%。
Disagree. 70 seems very high. There aren't that many neo labs. If getting acquired by one of the model labs counts as dying, I do think there will probably be some consolidation. If you include consolidation, I'd say 50.
你认为 Dario 作为 AI 领域的发声者,应该少一些负面、多一些积极吗?
Do you think Dario should be less negative and more positive as a voice in AI?
我认为有一个对未来以及事情将如何发展非常偏执的人是很重要的。我很感激,我个人很感激 Anthropic 的偏执。显然,在某些方面,我希望其他模型实验室不要觉得自己被排挤。但我非常相信神经多样性,Anthropic 是神经多样性图谱中非常重要的一部分。如果没有人极度偏执,那么就没有人发出那种声音。所以我感激他们这样做。
I think it's important to have someone who is very paranoid about the future and how things are going to shake up. I appreciate that, like I personally appreciate Anthropic's paranoia. Obviously, there are areas where I want other model labs to not feel like they're being pushed off the table. But I'm a big believer in neurodiversity, and Anthropic is a part of the neurodiversity map that really matters. If no one is being extremely paranoid, then no one is offering that voice. So I appreciate that they're doing it.
你坐在这个位置上,能看到所有人的使用情况,你觉得最疯狂但人们谈论得不够的事情是什么?
What's the craziest thing you see from your seat on top of everyone's usage that you don't think people talk about enough?
很多公司显然担心成本管理,对他们花费的推理量感到恐慌,而且不知道该如何看待这个问题。这是一种全新的经营方式和思考运营支出的方式。旧的思考方式是,你给员工多少薪水,然后你就忘了。有人知道每个人赚多少,但那是一个静态数字,每季度或绩效评估后调整一次。现在,你的员工成本完全是动态的、不同的。我认为很多公司都在自己承担路由的任务。未来,这很可能会下放到员工层面。你的员工应该自己决定哪些工具和模型最适合他们的任务,然后我们应该根据你作为员工所做的选择来计算你的成本。你作为员工的成本将是一个动态数字,取决于该员工如何有效地使用昂贵和便宜的模型来完成工作。我建议公司仍然进行正常的管理工作,让经理评估员工的有效性和生产力,但同时也要与员工的成本挂钩,然后得出一个庆祝象限——这些员工表现良好且成本效益高——和一个担忧象限——这些员工表现一般,而且成本效益极低,他们的 AI 使用量高得离谱。然后你处理担忧象限。所以我认为人们没有谈论如何在 AI 时代思考员工成本,以及它应该是一个动态数字,而不是一个只有少数人知道然后就消失的静态东西。
A lot of companies are obviously worried about cost management and freaking out about the amount of inference they're spending, and they don't know how to think about it. It's a whole new way of doing business and thinking about your opex. The old way of thinking about how much you give your employees—you give them a salary and you kind of forget about it. Someone knows what everyone's making, but it's a static number that's readjusted quarterly, maybe after performance reviews. Now your employees all cost totally dynamic different amounts. I think a lot of companies are putting it on themselves to do routing. In the future, there's a good chance it will get pushed downwards to the employee level. Your employees should figure out which tools and models to use that are best for their tasks, and then we should figure out how much you're costing based on the choices you make as an employee. Your cost as an employee is going to be a dynamic number, dependent on how effectively that employee is using expensive and cheap models to do their job. I advise companies to still do their normal management work, have their managers assess how effective and productive employees are, but also line it up with how much their employees cost, and then come up with a quadrant of celebration—these employees are doing a good job and they're cost-effective—and a quadrant of concern—these employees are doing a so-so job and they are not cost-effective at all, their AI use is off the charts. Then you address the quadrant of concern. So I don't think people talk about how you think of employee cost in the age of AI, and that it should be a dynamic number, not a static thing that only a few people know about and it's gone.
这很奇妙,但你能想象去对某人说,“哦,抱歉,你上个月值 10 万,现在你值 5 万。”我认为这会让规划……
And wonderful, but can you imagine going to someone, "Oh, I'm sorry, you were worth 100 grand last month, now you're worth 50." I think it would make planning...
嗯,他们可以控制自己的成本。这是好事。所有员工都可以控制自己的成本并影响它。现在你可以思考,“好吧,我作为员工有多好,我有多高效?”
Well, they are in control of how much they cost. That's the great thing. All employees are in control of how much they cost and can influence that. Now you get to think, "Okay, how good am I as an employee and how efficient am I being as well?"
最后一个问题,当你看到今天的格局,有太多令人兴奋的事情。你个人最兴奋的是什么?
Final one, when you look at the landscape today, there are so many things to be excited about. What are you singularly most excited about?
一个是罕见病研究,我认为这是那些受智能瓶颈或推理瓶颈限制的事情之一。它涉及尝试很多想法并看看它们是否有效。另一个是众包高效的城市生活改善。例如,想象一下,如果有人好奇想找到美国或英国的每一根铅管,并且有办法,但他们真的需要让它成熟并进行压力测试。现在你可以用 AI 来做这件事,我们也许能解决一些大家都放弃了的奇怪问题,因为你需要一个疯狂的想法从某个地方冒出来。绝妙的想法在世界各地分布均匀;它们可能来自任何地方。现在你只需给它们杠杆,让它们真正发挥作用。所以我对我们将能够实现的广泛的城市或农村生活质量改善感到兴奋。
One is rare disease research, which I think is one of those things that has been intelligence bottlenecked, or really just the inference bottleneck. It involves trying out lots of ideas and seeing if they work. The other is crowdsourcing productive urban life improvements. For example, imagine if someone was curious about finding every lead pipe in America or every lead pipe in the UK and had an approach to it, but they really needed to make it mature and stress test it. Now you can use AI to do that, and we just might solve some weird problems that everyone has given up on because you need a crazy idea to come from somewhere. Brilliant ideas are sort of evenly distributed all over the world; they can come from anywhere. Now you just give them leverage to actually work. So I'm excited about very broad urban or rural quality of life improvements that we'll be able to make.
好的,伙计。这个节目我想做很久了。我很高兴我们能面对面做。我之前还担心我们得远程做。面对面做要好得多。你太棒了。非常感谢你和我一起做这个节目。
All right, dude. I've wanted to do this one for a while. I'm so glad we could do it in person as well. I was worried that we were going to have to do it remote. It is so much nicer to do it in person. You've been fantastic. Thank you so much for doing it with me.
彼此彼此。这很棒。
Likewise. This was great.