The Case for Specialized Intelligence
打开互动全文版(中英对照 + 朗读 + 问答)→Fireworks 创始人 Lynn Quo 认为,AI 的未来在于基于专有数据的专用私有智能,而非仅仅通用模型。
Fireworks founder Lynn Quo argues that the future of AI lies in specialized, private intelligence derived from proprietary data, not just generalized models.
我不希望看到只有一家公司拥有智能,这对我来说没有意义。我认为去年是编程之年,今年是协作之年。今天的热座嘉宾是一位我仅用 15 分钟会面就开出 1000 万美元支票的创始人——Fireworks 创始人 Lynn Quo。这是我 10 年投资生涯中最轻松的投资决定之一。
What I don't want to see is there's only one company owns intelligence. That doesn't make sense to me. I think last year is the year of coding and this year is the year of co-work. And in the hot seat today, a founder who I wrote a $10 million check for after just a 15-minute meeting. Lynn Quo, founder at Fireworks. This was one of the easiest investment decisions that I've made in a 10-year investing career.
我确实认为 token 成本将大幅下降。未来三年成本降低 10 倍,而这 10 倍的降幅将驱动 100 倍的使用量。我们绝对不会进入应用层。至于是否向下进入数据中心等,这始终在讨论范围内,但问题是……
I do think the cost of token will go down drastically. 10x cost reduction in the next three years and this 10x cost reduction will drive a 100x usage. We absolutely are not going to move into application layer. very clear to us whether we will move down into data centers and so on that could be always be on the table but the question is
准备好了,Lynn。我对此非常兴奋。我听说了很多很棒的事情。我刚和你的联合创始人 Dimma 通了电话,还和 Alfred、Lynn Sonia、Matt Miller 等很多人聊过。所以,谢谢你接受我的邀请。
Ready to go Lynn. I am so excited for this. I heard so many great things. I just got off the phone with your co-founder Dimma, I spoke to Alfred, Lynn Sonia, Matt Miller, many more. So, thank you for joining me.
哦,谢谢你的邀请。
Oh, thanks for having me.
我听说 Eric Vishrier 有一条规则:不投资大型科技公司的高管,但他为你打破了这条规则,这非常特别。
Now, I heard that Eric Vishrier has a rule, don't invest in big tech directors, but he broke that rule with you, which is very special.
我也这么觉得。有个有趣的故事。我们握手决定合作后,他打电话给我,说他跟一位顾问聊了,顾问问他:你见过多少大型科技公司的高管成功创业?很少。他告诉了我这件事。我很惊讶,心想:我们现在要毁约吗?但没有,从那以后我们合作得非常紧密。
I think so too. So, a funny story. After we decide to handshake, he did call me and said he talked with one of his advisers and his advisor questioned him: hey, how many big tech executives have you seen being successful in starting a company? Very few. And he told me that. I was surprised, like, are we breaking our handshake now? But no, but we since then we work very closely with each other.
Eric 是最棒的之一。你创办公司时已经 48 岁了。
Eric is one of the best. You also started the company when you were 48.
哦,是的。
Oh, yeah.
这相当晚了。我能问问你,当如今我们推崇 15 岁左右创业时,你如何看待自己 48 岁成为创始人这件事?
That's quite late. Can I ask you, how do you reflect on being a 48-year-old founder when we glorify starting a company when you're pretty much 15 these days?
我没有深思过这个问题。我一直想自己创办一家科技公司。实际上,我在 2015 年就想创业了。因为我是一代移民,2000 年来到美国。我获得了分布式系统方向的计算机科学博士学位,尤其专注于数据库。数据库是一个非常复杂的系统,需要优化很多不同的目标。加入研究实验室后,我几乎接触了数据处理的所有方面。然后我去了 LinkedIn,进一步构建系统和产品,以驱动实际影响。那时我觉得自己准备好创业了。我了解所有技术,知道要构建什么产品。我有一份商业计划书,还有一份想一起创业的名单。我花时间思考,但暂停了,因为我认为自己不具备组建公司所需的人员技能。这不仅仅是产品和技术的问题,实际上是人的问题。我决定去一个能学到最多关于人的地方。当时最好的公司是 Facebook,它是硅谷的一颗新星。我暗自计划学习一两年就离开,回去做自己的事业。结果我在那里待了 7 年。
I didn't think deeply about that. I always want to have a tech business myself. I actually want to start a business in 2015. Because I'm a first generation immigrant and came to US in 2000. I did my PhD in distributed system, computer science, especially focused on databases. Database is a very complex system to build, with a lot of different objectives to optimize for. And pretty much touched every single aspect of processing data after I joined a research lab. Then I moved to LinkedIn to further build systems and products to be used to drive real impact. At that time I felt I'm ready to start a company. I know all the tech, I know what product to build. I have a business proposal. I have a list of people I want to start a company with. And I spent time thinking about it and I paused because I don't think I have the skill set on people to build a company. It's not just about product. It's not just about tech. It's actually about people. And I decided I want to go to a place I can learn the most about people. And the best company at that time is Facebook. It's a rising star in Silicon Valley. And secretly I was planning to learn for one year or two and leave and go back to do my own business. I stayed there for 7 years.
那么,通过 Fireworks,你在推理方面看到了世界尚未关注的东西。世界关注的是训练,我认为让人们理解这个堆栈会很有帮助,因为在你之下显然是芯片提供商,比如英伟达这样的公司,而在你之上是模型提供商,你处于中间位置。为什么这是堆栈中有价值的部分,而不是一种商品?
So with Fireworks you saw something in inference that the world was not focused on. The world was focused on training and I think it's helpful for people to understand kind of the stack because beneath you there's obviously kind of chip providers and your Nvidias of the world and then you've got above you the model providers and you sit in between. Why is that a valuable part of the stack and not a commodity?
这是个非常好的问题。但为什么要费心搞专门智能呢?为什么不直接用通用智能,这样你操心的事情更少,对吧?你只需在前沿模型提供的 API 之上构建,那不是更容易吗?所以论点如下:如果你认为智能是数据的衍生品,那么大部分数据实际上并未用于训练通用智能模型。训练数据来自公共互联网和标注数据。与全球数据相比,公共互联网是一个非常小的语料库。全球大部分数据实际上是私有的,锁定在应用程序和企业内部,永远不会与任何人共享,因为这是公司的专有知识产权。所以这就很有趣了。如果你审视这个领域,它会变得非常有趣,因为大部分数据没有被激活以产生任何智能。而这正是我们所相信的:激活这些数据。我们相信智能的未来前沿实际上是私有智能,或者说专门智能。所以这就是 Fireworks 从一开始就专注的方向,我们一直在推动价值。
That's a really good question. But why bother specialized intelligence? Why not just use generalized intelligence and you worry less things, right? You just kind of build on top of an API provided by frontier models, wouldn't that be much easier? So the argument is the following. If you think intelligence is a derivative of data, then majority of the data is actually not used for training a general intelligence model. The training data is coming from public internet and the labeled data. Public internet is a very small corpus of data compared with world's data. Majority of world's data are actually private, locked inside applications, locked inside enterprise, it will never get shared with anyone else because this is company's proprietary IP. So then it's interesting. If you look at the space, it becomes very interesting because majority of data is not being activated to derive any intelligence. And that's where we believe in: to activate that data. And we believe the future frontier of intelligence is actually private intelligence, or specialized intelligence. So that's kind of where Fireworks was from the beginning, we have been focusing on driving the value.
我有好多问题想问你。我完全理解你关于一些大公司内部私有数据价值的观点。但这不是 Anthropic 企业业务的前提吗?比如 Claude 的协作功能以及他们正在构建的许多周边产品?Dario 难道不会说这正是他们追求的目标吗?
I have so many questions to ask you. I totally understand you in terms of the values in private data within some of these largest companies. Is that not the premise of what Anthropic's enterprise business is though with Claude co-work and with a lot of their adjacencies that they're building? Would Dario not say that that's exactly what we're going after?
这很有趣,因为我认为 Anthropic 是一家完全相信 AGI 的公司。AGI 的定义是存在一个模型能以最佳方式解决所有问题。对我来说,这就是 AGI 的定义。这意味着你不需要专门化,那个模型应该能解决所有问题。它如此智能,对业务和工作的每个部分都有如此多的知识,可以胜任。那你为什么还要费心专门化呢?所以这本身就验证了我们生活在一个并非由单一原则统治的世界。我们生活在一个完全多样化的世界。给你举个例子,对吧?不同地区有不同的价值体系,有不同的政策。我们有不同的商业方式,有不同的生活方式。这全是品味、选择和判断的结合。我认为这定义了我们是人类。我们不是机器人。如果我们未来的世界被单一标准、被一家公司决定的品味所统治,我们就变成了机器人军队。这对我来说非常令人沮丧。我认为智人区别于其他物种的是创造力,是追求新事物、发现新生活方式的深层渴望。这定义了我们是人类,而这一部分是无法复制的。这是我的基本信念。这就是为什么硅谷和世界各地有如此多的创造力。有如此多构建新业务的创造力。什么是新业务?我在 Jensen 的 GTC 主题演讲后和他进行了一次有趣又愉快的对话。我们实际上录了下来,很有意思。
That's interesting because I view Anthropic as a company fully believing in AGI. The definition of AGI is there's this one model that can solve all the problems in the best way. That to me is the definition of AGI. To me that means you do not need to specialize, and that one model should be able to solve all the problems. It is so intelligent, has so much knowledge of every part of the businesses, every part of the jobs, it can fulfill. Then why do you need to bother specialize? So that itself is a validation that we're living in a world that's not ruled by one principle. We are living a fully diversified world. Give you one example, right? Different regions will have different value systems, will have different policies. We'll have different ways of conducting business. We'll have different lifestyles. It's all taste, choices, judgment combined. I think that's what defines us as human. We are not robots. If our future world is going to be ruled by one standard, a taste dictated by one company, we turn ourselves into an army of robots. And that's very depressing to me. And I think what separates Homo sapiens from other species is the creativity, is the deep desire of pursuing new things, of discovering new ways of living. That defines us as a human being, and that part cannot be copied. That's my fundamental belief. That's why in Silicon Valley there's so much creativity across the world. There's so much creativity of building new businesses. What is new business? I had this fun interesting conversation with Jensen after his GTC keynote. We actually recorded it and it's interesting.
我看了,非常棒。
I watched it. It was great.
是的。和 Jensen 录制其实不算是录制,他只是开始和我聊天。
Yeah. Recording with Jensen is not really recording. He just started having conversation with me.
我不知道他的团队已经开始录制了,我们就一直聊,太自然了。我们聊到了专门化智能。他对我说了一句话:“林,你说得对。没有所谓的专门化通用公司,因为每家公司都建立在一种独特的做事信念之上,否则它们就没有存在的理由。”这听起来很合理。但后来我回想他的话,觉得非常深刻,因为每家公司都在做独特的事情来证明其存在的合理性。这种独特性深深植根于它们的产品设计、软件设计、系统构建、数据、与用户的互动、对用户意图的深刻理解、用户参与度等等。所有这些都是公司存在的基础,是外部其他公司无法学习或保证的。
I didn't know his crew already started recording and we just kept talking, it's so easy. We talked about specialized intelligence. He said one thing to me: 'Lin, you're right. There's no specialized general company as in every company is built on a special belief of doing things, otherwise there's no reason they should exist.' It feels logical. But then I started to think back about what he said, it's profound because every single company is doing something unique that justifies their existence. And this uniqueness is deeply baked into their product design, software design, system building, data, interaction with users, deep understanding of user intent, engagement, and so on. All of that is the fundamental base of why a company should exist, which is not learnable or assured by another company sitting outside.
那你能帮我理解一下吗?作为一个播客主持人,我擅长问基础问题,所以请见谅。但为什么像 Dario、Sam、Larry 和 Sergey 这样的人,会以那种方式谈论 AGI,认为它是不可避免的?
Can you help me understand then? As a podcaster I specialize in asking basic questions, so forgive me. But why then do people like Dario, like Sam, like Larry and Sergey talk about AGI in the way that they do, as inevitable?
我认为他们构建的东西非常棒,因为他们基本上是在建造电力线路,来分配一个非常棒的智能来源,其他人可以在此基础上构建。这就是我对他们贡献的看法。如果我们没有这个基础基础设施,我们就不会有各种家用电器。我喜欢我的咖啡机,它是特别品牌的。但没有电力,我们就无法做那些有趣、独特、特别、体现我们品味的事情。所以我认为这非常重要。但问题是,这条电力线路会取代我们做的一切吗?我不这么认为。
I think what they build is fantastic because they are basically building power lines to distribute a really great source of intelligence that everyone else can build on top of. That's how I view their contribution. If we don't have this fundamental infrastructure, then we won't have all kinds of appliances living in our home. I love my coffee machine, it's specially branded. But without that power, we don't get to do the things that are fun, unique, special, that encode our taste. So I do think that's very, very important. But the question is, is this power line going to replace everything we do? I don't think so.
作为投资者,我的问题是:电力线路是好生意吗?你提到了 PyTorch 和开放生态系统。过去三个月我们都意识到,开源正在飞速加速,能力提升到了无法比拟的程度,但它已经达到了 90% 的效率,成本效益却高出 15 倍。在开源的世界里,电力线路还是好生意吗?
The question for me as an investor is, are power lines good businesses? You said about PyTorch and the open ecosystem. Open source in the last three months we've all realized is accelerating so fast and capabilities have increased to such an extent that it's not comparable, but it's getting 90% as efficient with 15 times more cost effective. Are power lines good businesses in a world of open source?
我是这样看待开源的。早期我们创立公司时,联合创始人之间有过一场深刻的辩论:我们该怎么做?是构建自己的模型,还是在开源模型之上构建?那时开源模型几乎处于起步阶段。如果我们走那条路,那是一个巨大的赌注。但凭借我们在 PyTorch 的经验,我们相信开源社区。我们相信开放性。这是我们运营的基本原则,因为开放性赋予了控制权。开放性把控制权交给了用户。想想开源模型:一旦模型发布,你就完全控制了权重。你可以随意更改它。它是你的。然后你可以在它之上构建。所以这是一个根本不同的运营原则,我们相信它,因为我们之前有开源的根基。所以我们下了那个赌注,它确实得到了回报。过去两年,开源模型和闭源模型的质量都显著提高,以至于这两条流都跨过了质量门槛,能够解决很多问题。在 Fireworks 内部,我们使用开源模型来驱动招聘流程、候选人搜寻、反馈收集,甚至一些内部财务流程,对于编码,我们使用模型来帮助我们调试。我们在 Fireworks 内部有大量的智能体,而且我们注重成本。所以两种模型类别都跨过了门槛,能够解决各种问题。其次,开源模型跨过了门槛,变得更容易微调。能够引导模型是模型智能的一部分。模型智能已经传递,引导起来更容易,尤其是用少量数据。一家公司拥有的少量独特数据,然后我们可以朝着你的评估指标爬山。通常爬山的结果是用你的数据解决你的独特问题,你比通用模型做得更好。
Here's how I view open source. Early on when we founded the company, we had a pretty deep debate among the co-founders: what do we do? Do we build our own models or build on top of open models? At that time, open models were almost in their infancy. It was a big bet if we were going to take that direction. But with our PyTorch experience, we believe in the open community. We believe in openness. That's a fundamental principle we operate with, because openness gave control. Openness gave control to the user. Think about open models: once the model is released, you have full control of the weights. You can change it however you want. It's yours. And then you can build on top of it. So that is a fundamentally different operating principle that we believe in because of our roots in open source before. So we took that bet and it did pay off in the sense that both open model and closed model quality significantly improved over the past two years, to the point that both streams crossed a quality threshold, they can solve so many problems. Within Fireworks, we use open models to drive our recruiting process, candidate sourcing, feedback collection, even some internal finance processes, and for coding we use models to help us debug. We have tons of agents within Fireworks ourselves, and we are cost conscious. So both model categories crossed a threshold to solve a variety of problems. Second, open models crossed a threshold where they are so much easier to tune. Being able to steer a model is part of the model intelligence. And the model intelligence has passed through, it's much easier to steer, especially with a small amount of data. A small amount of unique data a particular company has, and then we can hill-climb towards your eval. Often the end result of hill-climbing is to solve your unique problem with your data, you are better than a general purpose model.
当 90% 的企业工作流都可以像你说的那样用开源模型完成时,前沿模型的使用量就不会像需要它做所有事情时那么大了。那么,如果大多数工作都可以通过开源完成,这些公司是否实际上被严重高估和过度估计了?
When 90% of enterprise workflows can be done as you said with open models, the usage for frontier models will not be as large as it was if it was needed for everything. So are these companies actually dramatically overvalued and overestimated if the majority can just go through open?
我认为人们开始意识到这一点了。
I think people start to realize it.
我记得两年前,我去不同的地方谈论一个有趣的现象,这在过去的 SaaS 时代是不存在的。在 SaaS 时代,产品市场契合度和持久业务几乎是等价的。最难的是找到产品市场契合度,一旦找到,就尽可能快地规模化,因为 CPU 是商品,你构建的基础设施也几乎是商品,甚至不用担心。现在产品市场契合度和持久业务是两个不同的概念。对于初创公司,我们有很棒的公司,它们有产品市场契合度,客户愿意付费,非常看重它们的产品,但它们无法规模化,因为一旦规模化,它们可能会规模到破产。你听说过“规模到破产”吗?这是一个真实的问题。对于现有企业来说,问题更大。对于大型数字原生公司,因为它们有巨大的流量,它们从十年前还是初创公司时就赢得了胜利。一旦它们推出那些 AI 功能,它们会覆盖所有客户群,但它们负担不起,因为它们的 CFO 看着成本提案和预测,根本无法证明其合理性。所以这对所有创新者来说都成了一个真正的问题:我们真的很想接入这项新技术,但我们负担不起。我们需要找到一种替代方案来负担得起。而替代方案就是控制你的开源权重模型,并推出你自己的模型。
I remember from two years ago, I went to different places and talked about an interesting phenomenon that didn't exist in the past in the SaaS era. During SaaS, product-market fit and a durable business were almost equivalent. The hardest thing is finding product-market fit, and once you find it, just scale as fast as you can, because CPU is a commodity, the infrastructure you build on top is almost like a commodity, don't even worry about that. Now product-market fit and durable business are two separate concepts. For startups, we have great companies that have product-market fit, customers want to pay them and they really value their product, but they cannot scale because once they scale, they could scale into bankruptcy. Have you heard about scaling to bankruptcy? That's a real problem. It's an even bigger problem for incumbents. For big digital native companies, because they have huge amounts of traffic, they get a winner from a decade ago when they were startups. Once they roll out those AI features, they reach all their customer base, and they cannot afford to do it because their CFO looks at their cost proposal and forecasting, and there's no way to justify it. So it becomes a real problem for all those innovators: we really want to plug into this new technology, but we cannot afford it. We need to find an alternative to be able to afford it. And the alternative is to have control over your open weights model and roll out your own model.
或者你看看 Sam Altman 最近几天发布的,就是成本大幅降低的模型。我记不清具体数字了,但我觉得大概是便宜了一半,或者便宜了三分之二。
Or you see what Sam Altman released in the last few days, which is just dramatically lower cost models. I can't remember the amount, but I think it's like half as expensive or maybe three times cheaper.
下一步是否就是前沿模型价格大幅下降?
Is the next step actually we just see a massive reduction in price from the frontier models?
有可能,但我认为同时这也是一个非常不同的运作原理。对于开源权重模型,获取模型本身没有成本。显然,有些公司训练了这些模型并愿意开源。在美国,有多家公司这么做,包括英伟达正在训练 Nemotron。我们与他们合作非常紧密。所以一旦模型存在,谁使用它都没有成本,但前沿实验室需要投入资金来投资这些模型并收回研发成本。第二,你无法定制那些通用模型。你只能按原样通过 API 使用,没有控制权。而开源模型你可以完全控制,可以按需微调,按需使用。尤其是在 Fireworks,我们是一个专门的智能平台,提供各种工具让你轻松为特定用例定制模型。模型经过高质量微调后,我们进一步优化推理部署。在 Fireworks,我们认为每个模型部署都是“一码一配”,只针对你的工作负载,从质量、速度和成本角度进行优化。我们相信这绝对必要,因为一旦达到生产规模,面对数百万、数千万甚至数十亿用户,即使 5% 的成本降低也意义重大,更不用说我们过去看到的是五到十倍的降幅。
It could be, but I think at the same time it's just a very different operating principle. For open-weight models, model acquisition has no cost. Obviously, some companies train those models and are willing to open them up. In the US, there are multiple companies doing that, including Nvidia, which is training Nemotron. We're working very closely with them. So once the model is there, whoever uses those models, there's literally no cost, but there's a fundamental cost for the frontier labs to invest in those models and recoup R&D cost. Second, you just cannot customize those general-purpose models. You use them as is on top of an API, you have no control. With an open model, you have full control. You can tune it however you want, use it however you want. Especially at Fireworks, we are a specialized intelligence platform. We offer all sorts of tools for you to easily customize the model for one specific use case. After that model is tuned with high quality, we further optimize for inference deployment. At Fireworks, we think about every single model deployment as one size fits one. It's unique for your workload only, optimized for your workload only from quality, speed, and cost point of view. We believe that's absolutely needed because once you think about a production scale reaching millions, tens of millions, billions of users, even 5% cost reduction means a lot. It's a massive amount, let alone what we have seen in the past: five to ten times cost reduction.
我必须问的一个问题是企业对国家安全的担忧。看看 OpenRouter,今天排名前六的模型都是中国模型,质量惊人,开发速度也惊人,但它们是中国的。分析中国开源的力量时,我们是否有严重的国家安全担忧?
The one question that I do have to ask is the concern that enterprises have is national security concerns. When you look at OpenRouter, I think the top six models today are Chinese models and they're incredible quality. The speed of development is incredible, but they are Chinese models. Do we have serious national security concerns when analyzing the power of Chinese open source?
我认为这是整个行业正在激烈辩论的问题。一旦模型开源,你可以围绕它设置各种针对你业务的护栏。我想对所有模型说,无论开源还是闭源,你都应该设置自己的护栏。根本原因如下:模型提供商会将自己的判断和品味注入模型训练过程,你无法保证它与你的匹配。记住,这又回到了 Jensen 的评论:没有专门化的通用公司。每家公司都是特殊的,都有特殊的设计原则、特殊的品味、特殊的目标受众。正因为这种特殊性,一家公司的判断、品味和设计原则必然与你的公司错位、不对齐,而你的公司也是一个特殊的问题。所以你需要微调这些模型来匹配你的。我真的相信未来不会是少数几个 AGI 模型主宰世界。我相信未来——可能听起来吓人,但我觉得是真的——会有数百万个专用模型,每个应用、每个用例一个。
I think it's a huge debate happening right now across the industry. Once the model is open, you can put all kinds of guardrails specialized to your business around it. I would say to all models, doesn't matter if open or closed, you should put your own guardrails around it. The fundamental reason is the following: a model provider will infuse their own judgment, their own taste into the model training process. You cannot guarantee it matches yours. Remember, it goes back to Jensen's comment: there's no specialized general company. Every company is special. Every company will have a special design principle, a special taste, a special target audience to serve. Because of that specialty, it's guaranteed that the judgment, the taste, the design principle from one company would mismatch, would misalign with your company, which is a special problem. So that is the reason you need to tune those models to match yours. I really believe the future will not be a few small number of AGI models dominating the world. I really believe the future will be—it may be scary but I think that's true—millions of specialized models, one per application, per use case.
上周我们看到报道说,中国正在考虑限制对其开源模型的访问,因为他们看到发展太快太好了。考虑到美国缺乏开源模型,如果中国真的开始限制对其开源模型的访问,会发生什么?
We saw in the last week actually reports that China were looking at actually restricting access to their open models because they were seeing the development being so fast and so good. What would happen in a world where China actually started restricting access to their open models given the lack of open models we have in the US?
我认为短期内会有很大影响。但开源生态的美妙之处在于它不是一个供应商——这就是为什么它叫开源。它通常会吸引众多感兴趣的参与者。我相信,在人才密度和资源方面,美国有能力自己建立那个开源系统,我们也应该这么做。我已经看到这种情况在发生。再次强调,在许多开源系统中,百花齐放,这就是它的美妙之处。
I think it will be a big impact in the short term. But the beauty of an open ecosystem is it's not one provider—that's why it's open. It usually attracts many, many interested parties to participate. I do believe in terms of talent density and resources, I do believe the US will be able to build that open system by ourselves, and we should. I've seen this happening. Again, in many open systems, there are a thousand flowers blossoming, and that's the beauty of it.
当我们谈到企业内智能的专门化时,就像你刚才说的,如果举一个非常典型的例子——虽然我不太想举这个例子,因为我是 Lagora 的投资者,而且我知道你会站在哪一边——但法律领域有两家竞争公司:Harvey 和 Lagora。Harvey 承诺构建自己的模型,而 Lagora 没有。一年前,看起来没有承诺自建模型的公司是对的,因为前沿模型的能力增长非常快。现在看起来他们错了。像 Harvey 和 Lagora 这样的公司应该自建模型吗?如果不自建,会发生什么?
When we talk about the specialization of intelligence within enterprises as you have done just there, if we take a very prime example which I don't particularly want to take because I'm an investor in Lagora and I think I know which side you're going to fall on here, but you have two companies that compete in the legal space: Harvey and Lagora. Harvey have committed to building their own model and Lagora have not. A year ago, it looked like companies that didn't commit to their own model were right because frontier models were increasing so fast in terms of capability. Now it looks like they're wrong. Should companies like Harvey and Lagora be building their own model, and actually if you don't, what happens?
我有一个观察,很多人也有:软件开发,尤其是 SaaS 领域,已经被编码的通用智能显著颠覆了。应用开发生命周期在时间和资源需求上大幅压缩。过去,需要几十个非常强的产品工程师和产品经理,从想法到实现再到生产规模,需要多个季度甚至多年的投资。那是深度模式。今天,一个人几周内就可能把想法变成产品并快速规模化。这是前所未有的。这也重新定义了竞争格局,因为仅仅靠应用本身的想法来竞争非常困难,因为很多人有类似的想法。现在实现不再是那么大的障碍。
So here's one observation I had, and many people have: software development, especially the SaaS space, has been significantly disrupted because of the general intelligence of coding. The application development lifecycle has significantly collapsed in terms of timeline and resources needed. In the past, it required tens of very strong product engineers and PMs to convert from idea to implementation to production scale, multiple quarters or years of investment. That's a deep mode. Today, one person in a few weeks can possibly launch their ideas into a product and scale quickly. This is unprecedented. And that also creates interesting dynamics in redefining where the competition is, because it's really hard to compete just on the idea of an application by itself, because many people have similar ideas. Now implementation is no longer such a big barrier.
但当你考虑企业部署和企业推广时,这真的成立吗?如果你和世界上最大的律师事务所合作,企业销售周期至少需要多年,需要建立关系,这非常困难,然后部署也非常定制化。不像 11 Labs 那样拿来就用。这不一样。
Is that actually true though when you're looking at enterprise deployment, enterprise rollout? If you're working with some of the biggest law firms in the world, I mean the enterprise sales cycle is at least multi-year with relationship building, that's very tough, and then you have deployment that's very customized. It's not like 11 Labs where you pick it up and go. It's different.
而且我认为法律领域尤其具有挑战性,因为律师通常更保守。法律也完全不容忍错误,对吧?这就是律师赚钱的原因。你需要构建一个非常坚实的案例。如果模型产生幻觉并生成错误判断,那你就麻烦了。所以我确实认为法律领域是一个很有趣的渗透领域,这些公司都做得很好。
And also I think the legal space is particularly challenging because lawyers are usually more conservative. Legal is also not tolerant at all of errors, right? That's why lawyers get paid. You need to build a very rock-solid case. If something hallucinates and generates wrong judgment, then you're in trouble. So I do think the legal space is a very interesting space to penetrate, and these companies are both doing a great job.
但另一方面,我确实认为这两家公司都拥有专有知识和信息,知道如何构建那些助手,进行案例研究,深入推动法律研究等等。我对法律的理解非常浅薄,但案例有那么多不同的版本和类型。所以我确实认为他们处于独特的位置,能够转化那种深刻的理解,而且他们都有数据。这不仅仅是关于他们的业务有多强的防御性。而是关于,当他们构建那些助手时,通常有一个整合框架,用于决定和编排使用哪个 AI 工具、调用哪些工具,这是定制化的,调用那些工具的准确性以及调用哪种工具很重要。甚至那个框架也需要与驱动它的模型共同训练,对吧。所以有很多例子表明,通过在 workflow 层拥有自己的智能,可以将业务推向卓越。所以也许这是一个时机问题。以编程为例,我认为在编程领域,Cursor 可能是最早开始微调自己模型的公司之一,现在几乎所有编程公司都在微调自己的模型。模型开发的速度会放缓吗?因为似乎每天我们都有一个具有新能力的新模型,然后就像,哦天哪,Cursor 的最新模型太棒了。接下来,另一个人的最新模型太棒了。Gemini 的最新模型太棒了。三年后,模型开发的速度还会这么快吗?模型的优势还会如此短暂,一天是这个,第二天又是另一个吗?
But on the flip side, I do think both companies are owning proprietary knowledge and information how to build those assistants, to do case studies, to go deep in driving legal research and all this, right. And my understanding of legal is so shallow, but there's so many different versions of flavors of cases. So I do think they are in a unique position to convert that deep understanding and they all have data. It's not just about how defensive their business is. It's about, often when they build those assistants, there's a harness integrating and deciding orchestrating which AI tool to use, which tools calling to, and this is bespoke, this is customized, and the accuracy of calling those tools and calling to what kind of tools is important. And even that harness needs to be co-trained with the model powering it, right. So there are ample examples of driving that business to excellence by owning their own intelligence of how to do that in the workflow layer. So maybe it's a timing. Coding for example, I think in coding space, Cursor probably is one of the pioneer starting to tune their model, and now almost all coding companies tune their own models. Does that pace of model development slow down? Because every single day it seems like we have a new model with a new capability and it's like, oh my gosh, Cursor's newest model is amazing. Next, we have someone else's newest model is amazing. Gemini's newest model is amazing. In three years time, will the pace of model development still be so fast and model superiority be so transient where one day it's one and the next day it's another?
所以,模型进步有几个层面。有基础通用 IQ 的进步,这些会呈现阶梯式跳跃。这就是为什么他们发布时,总有主要版本或次要版本,对吧。主要版本的阶梯式跳跃,正如你记得的,去年年初,整个思考过程是新的,对吧。模型不再立即吐出答案;模型会自己思考然后给出答案,这样好得多。所以那是一个阶梯式跳跃。我们见过很多阶梯式跳跃,但我认为每一年或每三个季度就会有一个重大飞跃。但与此同时,在这些基础模型之上,我可以看到专业化开始加速。因为正如我所说,这真的像一棵树,对吧。有很多树枝和树叶可以挂在树干上。随着基础模型质量开始出现阶梯式飞跃,我们可以做更多的事情来实现专业化。所以我确实看到,在世界上,专业化的速度将比通用智能部分快得多。
So, there are a few layers of model advancement. There's base general IQ advancement, so those will take step functions. That's why when they release, there's always major releases or minor releases, right. The major release of step functions, as you remember, beginning of last year, there's this whole thinking process is new, right. The model just doesn't spit out answer immediately; the model will think by itself and spit out answer, and it's much better that way. So that's one step function. And there are many step functions we have seen through, but I see those as every year or every three quarters there's a major leap. But at the same time, built on top of those base models, I can see the specialization start to accelerate. Because as I said, it's really like a tree, right. There are so many branches and leaves that can possibly hang on the trunk. And as the base model quality starts to have step function leaps, there's so much more we can do to specialize. So I do see in the world, specialization is going to accelerate much faster than the general intelligence part.
当我们考虑通用智能部分时,就在我们进一步深入多模型堆栈之前,Sam Altman 提出了 OpenAI 和其他公司向政府赠送 5% 股份的想法。你认为我们是否已经达到了这样一个阶段:模型开发如此先进、对社会如此重要,以至于它们将部分由政府或行政机构拥有?
When we think about the general intelligence part, just before we move further into the stack of like multimodel, Sam Altman proffered the 5% kind of gifting of OpenAI and others to the administration. Do you think we've reached a stage where model development is so advanced and so important to society that they will in part be government or administration-owned?
这是一个非常有趣的问题。我认为有先例。如果我们把那些通用智能模型的基础层视为大经济运作的基础设施,那么有先例,比如 PG&E 拥有电力和天然气,等等,对吧。所以我实际上不知道,但我不希望看到只有一家公司拥有智能。我认为这对我来说没有意义,因为智能有不同的类型。有一种通用的共同智能,惠及所有人。还有一种专门的智能,实际上帮助我们在历史上进步,以不同的方式思考,创造新的生活范式或新的商业和塑造行业的范式。我不希望那因为只有一家公司能做到而消亡。我认为这没有意义。
That's a very interesting question. I think there were precedents of that. If we think about the foundation tier of those general intelligence models as fundamentally a base infrastructure for the big economy to operate around, there have been precedents like PG&E owns electricity and gas, and so on, right. So I actually don't know, but I don't want to see is there's only one company owns intelligence. I think that doesn't make sense to me because there are different flavors of intelligence. There's this general common intelligence that benefits everyone. And then there's a specialized intelligence that actually helps us advance in history to think differently, to create new paradigm of living or new paradigm of doing business and shaping industry. I don't want that to die because there's only one company can do that. I don't think that makes sense.
随着多模型繁荣理论,有一种想法是,你将根据模型的专业领域将任务路由到不同的模型。
With the many models blooming theory, there's the idea that you will route tasks to different models dependent on what they specialize in.
我认为是这样。
I think so.
考虑到这一点,你不会构建自己的开放路由系统来迎合这一点吗?
With that in mind, will you not build your own open router of the world to cater to that?
是的,你可以说他们是最好的构建者,因为他们深刻理解自己的用例,并且有评估标准。所以再说一次,我对前沿的理解不仅仅是这一个模型。前沿可能是你为业务定制的特殊路由机制。你根据以下方式分解:为了完成这个任务,你通常需要一个高度智能的层,可能是最昂贵的开放或封闭模型,来处理最高复杂度的判断。通常人们还会构建子智能体来解决较小的问题,然后那些可以交给较小的开放模型,这些模型还可以进一步定制以适应你的特殊设计。所以我看到很多人今天已经在这样做了,我们也认为存在构建一个可以自我学习的自动路由系统的空间。结合自动调优系统,最终我们认为这一切都应该自动化,然后你会看到一个基于产品流程自我进化的系统,而你的产品也在不断进化。你的产品是活的,对吧?所以你不断部署和发布新功能与用户互动,那将完全是一个自我进化的自动化系统。
Yes, you can argue they're the best builder because they deeply understand their use case and they have the evals. So again, my thinking of what is the frontier is not just this one model. The frontier could be your special routing mechanism for your business. And you decompose that based on, in order to fulfill this task, you usually need a highly intelligent layer, maybe the most expensive open or closed models to judge at the highest complexity. And usually people will also build sub-agents to solve smaller problems, then those can go to smaller open models, and those can also further be customized to fit into your special design. So I've seen a lot of people already doing that today, and we also think there's a space to build an automatic routing system that can learn by itself. And that compound with automatic tuning system, eventually we think it should all be automated, and then you can see a self-evolving system based on what flows through your product, and your product keeps evolving. Your product is alive, right? So you keep deploying and launching new features to interact with your users, and that just kind of will be totally self-evolving automated system.
那么你认为,如果路由层可以自动化或独立构建,它有价值吗?这是一个有价值的层吗?
Do you think then that routing layer of the stack is valuable if it can be automated or it can be built on its own? Is that a valuable layer to have?
我绝对这么认为。
I definitely think so.
你真的这么认为?
You do think so?
我确实这么认为。
I do think so.
如果它可以自动化,或者公司可以自己构建,为什么还需要 Requesty 或 OpenRouter?
If it can be automated or companies can build it themselves, why would you need a Requesty or an OpenRouter?
你可能不需要。
You probably don't.
是的,我们还没到那一步。但我确实认为这可能是创新的领域。
Yeah, we're not there yet. But I do think this could be an area of innovation.
你说 Cursor 在创新方面是领跑者。我完全同意,但我听说你的 CTO Dema 在 Cursor 驻场了几个月,构建强化学习基础设施。这是必须这样做吗?这可以规模化吗?
You said Cursor being the front runners in terms of how innovative they've been. I completely agree with you, but I heard that your CTO Dema was embedded at Cursor for months building the RL infrastructure. Is that how it has to be done and is that scalable?
所以情况是,通常在新技术的早期采用曲线上,早期采用者都是黑客。黑客不是贬义;它没有负面含义。他们在某个领域有深厚的专业知识,并且想要控制很多事情。而在新技术采用曲线的后期,它开始变得更容易被更广泛的用户群体接受,这些用户没有深厚的专业知识,他们需要更少的控制。
So what's happening is, usually in the early adoption curve of new technology, the early adopters are all hackers. Hacker is not in a bad way; it doesn't have negative connotation. They have deep expertise in certain area and they want to control a lot of things. Versus in the late stage of a new tech adoption curve, it starts to get more accessible to a much bigger cohort of users who don't have deep expertise and they need less control.
所以通常总是先深入控制,然后再逐步放松控制。我们最终的目标是后一个阶段,但理解达到那个阶段需要什么也极其有价值。这就是为什么我们与 Cursor 深度合作。他们是尝试这些想法的先驱,拥有来自前沿实验室的研究人员,他们想要控制每一个细节。同时,我们也在推动边界。我们在做从未存在过的事情,构建从未存在过的系统,因为我们在这个特定场景下推动独特的边界。这里独特之处是什么?通常,如果你考虑训练,训练是非常资本密集的,通常发生在大型公司。他们有很多钱,用来购买非常昂贵的训练集群,通过 InfiniBand 互联,超级昂贵。一旦你有了那个昂贵的大型集群,通常你不需要深入思考如何高效,只需专注于工作。Cursor 和我们一样是初创公司,对吧?我们都对资本非常敏感,希望在提高效率的同时不拖慢研究创新。所以我们一起想出了一个非常聪明的方法来推动他们的训练过程。他们进行大量的后训练,基于强化学习,我们把强化学习分成两部分:训练器调整模型权重,不断生成新模型版本;然后新模型部署到我们称为 RL rollout 的环节,与合成环境(比如编码环境或真实编码环境)交互,然后获得奖励来判断模型版本的好坏。这是一个粗略的过程。我们把这两部分解耦。过去,在大型超大规模云厂商中,他们把所有东西都运行在一起。如果你有 10,000 或 100,000 个芯片通过 InfiniBand 互联,那极其昂贵且难以获得,但他们需要快速运行。我们设计了一个完全分布式的系统,跨五到六个数据中心区域全球运行,利用分散的 GPU,他们能够运行大规模任务。但挑战在于我们需要跨所有这些不同区域同步模型权重。这有多难?这很重要,因为发送这些权重的延迟决定了奖励的新鲜度。如果太陈旧,结果就会偏差太大。所以这是一个平衡,但我们创新了一种方法,可以快速分发新鲜的模型权重,不会太陈旧。从数值上看仍然合理,同时我们不受限于昂贵的 GPU 集群部署。这些就是我们与 Cursor 合作推动边界的创新,促成了他们最近的模型发布。我们为他们感到非常自豪。
So it always goes into deep control first, usually, and then little control later. So we are definitely aiming towards the later stage as the ultimate ten want to target, but it's also extremely valuable to understand what is required to get there. So that's why we partner deeply with Cursor. They are the pioneer trying those ideas. They do have researchers from frontier labs and they want to control every single thing. And at the same time, we're also pushing to the boundary. We're doing things that never existed before. We're doing systems that never existed before because we push the boundary that is unique to this particular setting. Okay. What is unique here? Typically, if you think about training, training happens, training is very capital intensive. And it usually happens in big companies. They have a lot of money. They put that money to buy very expensive training clusters interconnected with each other. Super expensive. And then once you have that expensive large fleet, usually you don't need to think too deeply about how to be efficient; you just focus on doing your work. Cursor is like us, their startup, right? Both of us are very capital conscious, and we want to be efficient while we don't want to slow down the research innovation. So together we figure out a very smart way to drive their training process. They do massive post-training, which is reinforcement learning based, and the reinforcement learning we break into two pieces: the trainer that is tweaking the weights of the model and basically generates new model versions constantly, and that new model will deploy to what we call the RL rollout. It basically deploys that new version to interact with a synthetic environment, a synthetic like coding environment or real coding environment, and then get the reward back to judge if that model version is good or bad. So that's a rough process. And we decouple these two. In the past, in large hyperscalers, they run that all together. If you think about getting 10,000, 100,000 chips all interconnected together through InfiniBand, it's extremely expensive and really hard to find, but they need to go really quickly. And we designed a fully distributed system and we run across five or six data center regions globally, tapping into scattered GPUs, and they are able to run massive jobs. But the challenge there is we need to sync model weights across all these different regions. How hard can that be? It matters because the latency of sending these weights over is going to dictate how fresh the rewards are. If it's too stale, then you are too off. So it's a balance, but we innovated a way we can distribute fresh model weights quickly. It's not too off. Numerically it's still sound, while we are not limited by a very expensive deployment of GPU fleet. So those are the innovations we work together with Cursor to push the boundary, leading to their recent model launches. We're very proud of them.
我能直白地问一个问题吗?Cursor 是一个了不起的客户,他们与你们合作取得了惊人的进展,看到这种伙伴关系很棒。这对你来说是一个非常大的客户。你如何看待在 SpaceX 收购后客户流失的担忧?
Can I ask a question bluntly? Which is, incredible customer to have, amazing progress they've had with you, and it's wonderful to see that partnership. It's a very large customer for you. How do you think about the concern of churn in the wake of a SpaceX acquisition?
是的,每个人都担心。整个行业,就模型驱动的应用创新而言,只有少数公司非常非常成功。它们达到了逃逸速度,但只是少数。这就是整个行业的形态。去年,Cursor 是其中之一。我想说所有模型公司都集中在 Cursor 上。我们专注于同一批应用公司。但自那以后,情况确实发生了变化,对吧?所以我们拥有非常健康多元化的客户群。特别是,我认为去年是编码之年。所有主要的编码公司都在使用我们。而今年是协作之年,协作本身比编码更加多元化,因为有通用型协作,比如帮助你做各种研究的通用型协作。你会问:“嘿,两年后英伟达 GPU 的价格会是多少?”“Anthropic 的 IPO 后股价会是多少?”这些都是深度研究,通用型深度研究。或者有各种不同类别的专用型协作:法律,我们刚刚谈到了两家很棒的法律公司;金融;客户支持;招聘;销售;营销;医疗保健。所以创新应用公司在协作领域非常广泛。它们做得非常好,并且是我们的客户群。更有趣的是,我们开始看到面向消费者的公司也开始关注生成式 AI 技术,它们正在改变对传统业务(比如推荐系统)的思考方式。这对我来说非常有趣,因为我们显然在世界上最大的推荐系统之一工作过。我们非常渴望看到这如何转化为我们的新经济。
Yeah, everyone is concerned. The whole entire industry, in terms of application innovations by model, in the sense that there are few companies that are very, very successful. They escape velocity, but few of them. So that's the shape of the whole entire industry. And last year, Cursor is one of the few. I would say all model companies are concentrated on Cursor. We concentrate on the same group of app companies. And since then, it does change, right? So we do have a very healthy diversified customer base. Especially I think last year is the year of coding. I think all major coding companies are on us. And this year is the year of co-work, and co-work is much more diversified by itself than coding because there's general purpose co-work, for example, general purpose like co-work to help you do all kinds of research. You want to ask, "Hey, what will be the Nvidia GPU price two years later?" "What will be Anthropic stock price after IPO?" So those are deep research, general purpose deep research. Or there are so many different categories of special purpose co-work: legal, we just talked about two great legal companies; finance; customer support; recruiting; sales; marketing; healthcare. So there's a very broad set of co-work space of innovation app companies. They are doing really well, and we have them as our customer base. And then more interestingly, we start to see an uptick of consumer-facing companies all starting to look into GenAI technology, and they are changing how they think about their traditional business of doing recommendation, for example. And that's very interesting to me because we have obviously worked at a huge recommendation system in the world. And we are very eager to see how that transforms into a new economy for us.
抱歉我可能问得有点天真。人们在推理领域是只与像你们这样的一个提供商合作,还是同时与你们和 Together 或该领域的其他公司合作?
I'm sorry for being naive here. Do people work with just one provider in the inference space like you, or do they work with you and with Together or anyone else in the space?
我认为人们在这个领域更倾向于多供应商策略,因为他们不知道会发生什么。拥有多个提供商来平衡感觉更安全。但我们不把自己视为推理提供商。再说一次,我们把自己视为交付专门智能,帮助公司调整他们的模型。我给你一些数字。今天我们每天处理超过 40 万亿个 token。这些 token 中的大多数来自定制模型,而不是现成模型。它们来自定制模型。
I think people are more inclined to a multi-vendor strategy in this space because they don't know what's happening. It feels safe to have multiple providers to balance things out. But we don't view ourselves as an inference provider. Again, we view ourselves as delivering specialized intelligence where we help companies tune their model. Let me give you some numbers. Today we process more than 40 trillion tokens a day. So the majority of those tokens are coming from a customized model, not from off-the-shelf models. They are coming from customized models.
到明年年底,这个 token 数量会是多少?
What will that token count be at the end of next year?
可能在 20 到 100 倍之间。
Anywhere ranging from 20 to 100x could be possible.
20 到 100 倍。
20 to 100x.
是的,我们现在正处于 S 曲线爆发的非常早期阶段。
Yeah, we're at a very early stage of S-curve explosion right now.
20 到 100 倍。如果是 20 到 100 倍,那么认为我们处于资本支出泡沫的想法是荒谬的,我们迫切需要比以往建议的更多的算力资本支出。对吗?
20 to 100x. If it's 20 to 100x, the idea that we are in a capex bubble is ridiculous, and we are desperately needing far more capex than we are ever suggesting for compute. Is that right?
是的,没错。同时,我认为 Jensen 有一个五层蛋糕。五层 AI 蛋糕,从上到下:应用、模型、基础设施、芯片、能源。我们在供应链方面受限于 AI 蛋糕的下层。
So that is right. At the same time, I think Jensen has a five-layered cake. Five-layered AI cake, from top down: application, model, infrastructure, chips, energy. We are bottlenecked by the lower part of the AI cake in terms of supply chain.
在物理世界中,制造速度是一个瓶颈。历史上,所有这些行业都不是为大规模 Scaling(规模扩张)设计的——说到 100 倍 Scaling,没有人为之设计过。我和很多制造商聊过;我们有点像被小零件、晶体管这些最微小的部件卡住了,它们拖住了整个服务器生产线的节奏,而这些服务器最终要部署到数据中心并用于生成 Token。
In the physical world, how fast we can manufacture is a bottleneck. In history, all these industries are not designed for massive scaling—speaking about 100x scaling, no one was designed for that. I talk with many manufacturers; it's kind of like we're bottlenecked by small parts, transistors, the smallest tiny parts that hold off the whole manufacturing line of servers that can deploy to data centers and be used to generate tokens.
你必须全栈覆盖 Jensen 的五层 AI 蛋糕吗?你必须全栈才能赢或减少依赖吗?我们看到 OpenAI 推出了……Jalapino,名字真难听。Anthropic 在和三星谈自研芯片。DeepSeek 也在自研芯片。扎克伯格宣布 Meta 自研芯片。你必须什么都做吗?
Do you have to be full-stack going to Jensen's five-layered AI cake? Do you have to be full stack to win or to reduce dependencies? We've seen OpenAI come out with... Jalapino, terrible name. Anthropic talking to Samsung about building their own chips. DeepSeek building their own chips. Zuck came out with Meta building their own chips. Do you have to be all of it?
这确实取决于公司的理念。对我们来说,敏捷就是一切,我们需要赢得构建任何东西的权利。所以专注就是一切,我们希望专注于我们基于自身优势能创造最大价值的地方,并且我们希望借助他人的优势来构建。具体来说,我们希望在全球所有可能的 AI 芯片上运行,而不受限于我们自己能带入数据中心的芯片数量,无论这些数据中心是我们自建还是租用的。但随着时间推移,当业务变得非常大时——我还记得 Meta 小的时候并没有什么都自己造,当他们变大后,自己造就有意义了:你赢得了为自己巨大的基础设施而建造的权利,如果它能节省五倍以上的成本,那你当然会去做。但我认为在早期阶段——这就是为什么我要给你讲一个有趣的故事:在编程领域,Cursor 是最早决定与我们合作的公司之一。我记得他们和我们合作时,营收还是个位数百万美元……
It really depends on the company philosophy. To us, agility is everything and we need to earn the rights of building anything. So focus is everything for us and we want to focus on where we add the biggest amount of value based on our strength, and we would like to leverage other people's strength to build on top of. In particular, we want to run everywhere on all possible AI chips in the world, and we don't want to limit it by how many chips we can bring into our data center, whether we constructed or rented. But over time, when the business grows very big, I still remember when Meta was young they didn't build everything, and when they're big it makes sense to build—you earn the right to build for your own giant infrastructure, and if it saves like five times more cost then you sure go do it. But I think at the early stage, that's why I give you an interesting story: in the coding space, I would say Cursor is the first company that decided to work with us early on. I remember when they worked with us they were single-digit million dollar...
哇。
Wow.
非常小。这只是两年前的事。他们在两年内增长了 100 倍到 1000 倍,差不多这样。但他们很早就决定与我们合作,因为他们认识到自己只想专注于产品创新和后来的研究。他们不想专注于平台创新。他们知道我们把所有研发都投入在那里,他们想找到最好的合作伙伴来赢得大市场。所以我确实认为这种专业化的心态是对的,我们也想专业化。我们不想拥有整个全栈。这不是我们公司的目标。
Very small. This was only two years ago. They grew by 100x to 1000x over two years, something like that. But they decided to work with us early on because they recognized they only want to focus on product innovation and later on research. They do not want to focus on platform innovation. They know we are putting all our R&D in there and they want to find the best partner to win big. So I do think that's the right mentality to specialize, and we want to specialize. We do not want to kind of own the whole entire stack. That's not our goal as a company.
抱歉我插一句。为什么 Jensen 跳过了你们那一层蛋糕?因为他正在用模型做 Neotron。他为什么不想也蚕食你们的业务?
I'm sorry to be hopping on. Why does Jensen skip your layer of the cake? Because he's doing Neotron with models. Why does he not want to cannibalize your business too?
嗯,Jensen 也没有在构建云,对吧?你可以说:‘嘿,Jensen,你完全有权利建一个 Nvidia 云。’但他并没有构建云基础设施。我想他也提到过这一点,很多人问过他这个问题。他还提到他想专注于他们有权做的事情。为什么做模型?我认为纯粹是一个供应链问题:如果美国没有一个美国原生的开放模型,那是个问题——这是个供应链问题。所以他纯粹是为了解决供应链问题。但如果不存在供应链问题,因为像我们这样的公司提供了这个专门的智能平台层,那他就不需要担心了。所以他只是希望整个五层 AI 蛋糕都能顺畅流动。没有阻塞,如果有阻塞,他有兴趣解决那些问题。
Well, Jensen is not building a cloud either, right? You can say, 'Hey, Jensen, probably you have all the rights to build an Nvidia cloud.' So he's not building a cloud infrastructure. I think he mentioned that as well and many people ask him that question. And he also mentioned he wants to specialize in what they have the rights to do. Why models? I think it's purely a supply chain question: if the US doesn't have a US-native open model, it's a problem—it's a supply chain problem. So he is solely there to solve the supply chain problem. But if there's no supply chain problem because companies like us are providing this specialized intelligence platform layer, then he doesn't need to worry about it. So he just wants to make sure the whole entire five layers of AI cake is flowing. There's no blockage, and if there's a blockage, he's interested in solving those problems.
Mark Benioff,你们本轮投资方之一——显然这会在本轮融资公布后出来——说他将 Salesforce 开发者薪资的约 3.8% 花在了 Anthropic 和 Claude Code 上。我认为这是一个有用的类比,因为如果你假设这就是花在 Claude Code 编码工具上的比例,那说明了市场的一方面;但如果达到 20%——哇,我们低估了这些公司能有多大。当你展望一两年后,你认为我们会花开发者薪资的百分之多少?是更少,因为这些工具会变得更便宜,还是更多,因为它们会越来越好?
Mark Benioff, one of your investors, I think in the new round, which obviously this will come out after the round is announced, said that he spends about 3.8% of developer salaries at Salesforce on Anthropic and Claude Code. And I think it's a useful analogy because if you assume that that is what's spent on Claude Code coding tools, that says one side of the market, but if it's 20%—wow, we're underestimating how big these companies can be. When you think forward a year or two, what percent of developer salaries do you think we'll spend? Is it less because these tools will get cheaper, or is it more because they'll get better and better?
我确实认为 Token 的成本会大幅下降,因为再次……
I do think the cost of tokens will go down drastically, because again...
到目前为止还没有。
It hasn't so far.
到目前为止还没有,因为供应链限制,但我们生活在一个自由经济中。所以想想:每当出现短缺,价格就会高;价格高会吸引很多人来解决问题,并带来竞争;竞争会降低成本,最终导致一个非常经济的解决方案。但实际上这对每个人都有好处,因为更实惠的基础设施会吸引更多使用。所以我的预测是,随着基础设施成本的下降——这就是我对明年情况预测的依据——基础设施成本会下降,使用量会因此爆发。当你不再把它当作一个问题时,你就只是使用它——如果它是一种公用事业,你就只是使用它。
It hasn't so far because of supply chain constraint, but we are living in a free economy. So think about: whenever there's shortage, price is high; price is high will invite a lot of people coming to solve the problem, and it will invite competition; competition will bring down the cost, and then eventually it will lead to a very economical solution. But actually that's good for everyone because much more affordable infrastructure will invite more usage. So my prediction is with the decrease of the infrastructure... that's where it comes to my prediction of how far next year will look like, because the infrastructure costs will go down and usage will explode because of that. The moment you don't think about that as a problem for you, you just use it—if it's a utility, you just use it.
Token 成本会下降多少?是减半吗?是降到百分之一吗?
How much will token cost come down? Is it like a halving? Is it like a hundredth of the cost?
有不同的思考方式。并非所有 Token 都是平等的。我认为我们应该建立最佳实践来评估每个任务的 Token 经济,因为不同模型吐出 Token 的方式不同。有些模型比其他模型冗长得多。所以你可以想象,一个模型比另一个便宜 2 倍,但解决同一任务需要多 2 倍的 Token,那么它们的成本是一样的。
There are different ways to think about this. It's not all tokens are equal. I think we should establish best practice to evaluate the token economy per task, because different models have different ways of spitting out tokens. Some are much more verbose than others. So you can imagine one model is 2x cheaper than the other but it's 2x more tokens to solve the same task, and then they're at the same cost.
是的。
Yeah.
对。所以总体而言,我认为随着模型质量的提升,精确性将成为优化的一部分。所以这是一个层面的优化:解决一个任务,我们应该需要更少的 Token。第二个层面是针对一个 Token——如何做到这一点,你需要定制模型来更好地、更精确地解决你的问题,这涉及到模型调优。其次,对于这些模型吐出并处理的每个 Token,我们也通过我们的平台专门优化单位经济性。
Right. So overall, I think as the model quality improves, being precise is going to be part of the optimization. So that's one level of optimization: to solve one task, we should need less tokens. And the second is for one token—how to do that is you need to customize the model to solve your problem especially better and more precise, that goes into model tuning. And second, for each token spit out from those models and processed by those models, we also specialize in making the unit of economics much better through our platform.
第三是底层基础设施,比如 GPU、周边内存等等。目前供应链严重受限,情况会好转。我认为未来一年到一年半内可能不会有变化,但长期来看,两到三年内应该会改变,成本会压缩。总体而言,我可以想象未来三年成本降低 10 倍,而 10 倍的成本降低将带来 100 倍的使用量增长。
And third is underlying infrastructure, like the GPUs, the surrounding memories, and all these. Today it's under stark supply chain constraint. It's going to get much better. The situation will get much better. I don't think probably in the next one year or a year and a half the situation will not change, but in the long term, two to three years, it should change, and that cost will compress. So overall, I can imagine a 10x cost reduction in the next three years, and this 10x cost reduction will drive a 100x usage.
你刚才提到了 token 效率,以及如何让客户更高效。但有了这种效率,你们收费更高。我调研时对比了竞争对手,发现 Together 是价格之王。我不是贬低他们,但他们更便宜。想要便宜就去那里;尊重地说,想要更高质量的产品就来找你们。但你们确实更贵。你觉得这个评价和类比公平吗?
You said there about kind of token efficiency and how you enable your customers to be much more efficient. With that efficiency, you do charge more. You know, when I did the research, when I compared to competitors, I got like Together's price king. And I don't mean this disparagingly, but like they're cheaper. If you want cheap, you go there. And respectfully, if you want better quality product, you go to you. But it is more expensive. Do you think that's a fair assessment and a fair analogy?
我认为我们可能不是在同类比较,因为这又回到了我们的业务。我们大部分流量是定制化模型,我们优化质量第一,永远是质量。质量是指模型质量,针对你的应用、你的具体业务、你的用例等等。第二,我们在推理时交付这些模型,也是质量。我们非常重视质量,以至于做了一些极端的事情。例如,在训练期间,有一个很难实现的目标叫做零 KLD。这有点技术性。KLD 是质量的衡量指标。它的意思是,在训练系统和推理系统之间,当模型迁移时,我们实现比特等价,即数值完全一致,不损失任何精度。这很难实现。但我们推动并交付了它,因为我们知道我们的主要业务是模型定制和定制模型的推理。我们希望客户在训练上投入的每一美元都能最大化。如果训练到推理的边界不是比特等价的,质量就会下降,那就等于你为训练投资支付了折扣后的质量。为什么要这么做?所以质量第一,质量确实带来额外价值,这就是为什么我们对商品化的一刀切模式不感兴趣——那种现成模型以相同方式部署给所有人。我们的业务总是定制模型,以独特方式为你的特定工作负载部署。
I think we're probably not comparing apples to apples, in the sense that again it goes back to our business. Majority of our traffic is customized model, and we optimize for quality number one. Always quality. Quality as in model quality towards your applications, your specific business, your use case, and so on. The second is when we deliver those models in inference, it's also quality. And we care quality so much we do extreme things. For example, during training time, there's a very hard thing to achieve called zero KLD. It's a little bit technical. The idea here is zero KLD. KLD is a measure of quality. And what it means is between the training system and the inference system, when model moves over, we have bit equivalence, so the numeric are fully the same. We do not lose a bit of accuracy. That's really hard to achieve. But the reason we push that, we deliver that. And the reason we push that is because we know our primary business is in model customization and inference of customized model. And we want our customers every single dollar investing in training to maximize it. And then if the cross training-inference boundary is not bitwise equivalent, they just drop the quality down, and then it's like you pay your training investment by discounted quality. Why do you do that? So quality first, and quality does bring additional value, and that's why we are not interested in commoditized one-size-fits-all, you know, this off-the-shelf model deployed in the same way for everyone. That kind of business will always customize model, deploy in a unique way for your particular workload.
两个问题。你们必须有一个 FDE 模型才能使定制化模型高效吗?
Two questions. Do you have to have an FDE model to make the customized model efficient?
事实上,我们确实有一个 FDE 团队,叫做应用机器学习工程团队。他们的主要工作是加速这种定制化部署,实际上也构建智能体来自动化许多部署。
We, as a matter of fact, do have an FDE team. It's called Applied Machine Learning Engineering team. So their primary job is to accelerate this customized deployment. As a matter of fact, to also build the agent to automate a lot of deployments.
考虑到我们在技术栈中的位置,我们有很多复杂性,我们的利润率结构与传统 SaaS 的 80% 有点不同。我不确切知道这里的利润率,但传统上我们在 30% 到 40% 的范围内。这是我们现在的新常态吗?
Given where we are in the stack, a lot of the complexity that we have, we have a margin structure that's a little bit different to like traditional SaaS being 80%. I don't know the margins precisely here, but they traditionally in the 30 to 40% range for where we are. Is that the new normal for where we are?
我不认为这是新常态。我认为这反映的是,至少对我们来说——我不知道其他公司——对我们来说,这反映了我们处于超高速增长阶段。在超高速增长阶段,你有选择,对吧?你要么优化……对我来说,利润率优化是一个约束问题。比如,嘿,我们想要 70% 的利润率,想要 80% 的利润率,然后我们倒退一步,施加这些约束来保证那些利润率。而约束通常会减缓创新。例如,在系统开发和高速扩张阶段,我们不想过度建设,因为我们处于高度实验阶段。我们在测试什么会留下,什么不会留下。优化没有意义。一旦我们知道这是我们想要 100% 构建的系统,然后我们要将其扩展一千倍,那时我们才会全力优化。我想你对业务的看法也一样。在超高速增长期,如果我们只专注于优化毛利率,我们绝对可以做到。但我们也牺牲了增长速度,因为我们想去所有地方。我们想进入不同的地理区域,应对不同的用例,不断创造不同的产品线。这些都不是优化的时机。这是我的观点。
I don't think that's the new normal. I think that is a reflection, at least for us, I don't know other companies, for us it is a reflection of we are in a hypergrowth phase. During hypergrowth phase, you have the choice, right? You either optimize to... to me, margin optimization is a constraint problem. As in, hey, we want to go to 70% margin, we want to go to 80% margin, and then we are going to go backwards and impose those constraints to guarantee those margins. And usually constraints slow down innovation. So for example, during system development and in a high velocity system expanding phase, we don't want to overbuild because we are in kind of high experimentation. We're testing what will stay, what will not stay. Optimization doesn't make any sense. Once we know this is a system that we want to build 100%, and then we are going to scale this a thousand times bigger, then we go optimize the heck out of it. I think you think about business the same way. Well, in the hypergrowth, if our focus is only to optimize gross margin, we absolutely can do that. But we are sacrificing the speed of growth as well because we want to go everywhere. We want to go into different geo regions. We want to tackle different use cases. We want to constantly create different product lines. And those are not the time for optimization. That's my opinion.
那么,我们可以在不进入技术栈不同层的情况下提高利润率。
So we will be able to increase margin without moving into different layers of the stack.
不是进入……我们绝对不会进入应用层,这对我们来说非常清楚。至于我们是否会向下移动,比如你提到的建设数据中心等,这总是可能的,但问题是时机。
Not into... we absolutely are not going to move into application layer, very clear to us. And whether we will move down into, for example, you mentioned build data centers and so on, that could be always be on the table, but the question is timing.
是不是有句话叫要么死,要么活得足够久去建自己的数据中心?就像 Elon 或 Zuck 现在花 100 亿在加拿大建最新数据中心。你想建数据中心吗?
Isn't the statement you either die or you live long enough to build your own data centers, as Elon or Zuck now spending I think 10 billion on the latest data center in Canada? Would you like to build data centers?
我建过重要的数据中心,那里也有很多创新可能。也没有一刀切的方案。构建 GPU 原生数据中心也很有趣,尤其我认为有一个潜在方向……所以这是一个权衡,对吧?从运营角度来看,构建异构部署要好得多。所有芯片相同、所有 SKU 相同,尽可能大,运行多个工作负载。这样是可互换的,对吧?管理起来非常容易。坏节点……你建立一个原则、一个流程来进行维护操作。但这又回到了优化。一旦规模如此之大,任何优化都会带来巨大的经济回报。例如,我们谈到 Nvidia 最近收购了一家叫 Guac with Q 的公司,它是一个基于大容量 SRAM 的 ASIC 加速器。我在节目前和 Jonathan 聊过。
So I have built data centers that matter, and also lots of innovation possible there. There's no one size fits all as well. And building a GPU native data center is also interesting, especially I think there is a potential direction of building... so it's a trade-off, right? From operation point of view, it's much better to build a heterogeneous deployment. It's all the same chips, all the same SKU, as big as possible, and run multiple workloads. So it's fungible, right? It's very easy to manage it. Bad nodes... you build one principle, one process to do maintenance operation. But again, it goes to optimization. But once it's so big, then any optimization is going to drive a lot of economical return. For example, we're talking about Nvidia recently acquired a company also called Guac with Q. It's a large SRAM-based ASIC accelerator. I spoke to Jonathan before this show.
Jonathan 非常出色。
Jonathan is excellent.
哦,我也是他的粉丝。所以,这是计算密集型 GPU 和 SRAM 密集型 ASIC 的绝佳组合,因为计算密集型非常适合语言模型处理的前半部分,即预填充,处理提示等;而 SRAM 密集型非常适合生成。这只是模型架构的本质。将这两者结合,而不是在同一个芯片上同质化运行,这很棒,对吧?但这需要非常独特的系统设计和数据中心部署。
Oh, also a fan of his. So, but it's a great combination between a flops intense GPU and SRAM intense ASIC, because flops intense is really good for first half of LM processing. It's prefill, called prefill processing the prompts and so on, and SRAM intense is really good for generation. That's just the nature of the model architecture. It's great to combine these two instead of running homogeneously on the same chip, right? And that requires a very unique system design and deployment into data center.
而且它实际上是异构的。之前,我真的认为同构设计对运营要好得多。这是异构的,那么如何运营这种异构设计就需要在数据中心部署等方面进行独特的创新。所以数据中心并非商品化。你可以专注于数据中心部署,一个数据中心可能比另一个更好,数据中心部署可以做得很好,也可能很糟糕。数据科学非常复杂。如果你从头到尾考虑,从建设到电力部署,你需要有合适的电力供应、光纤通道、合适的冷却,尤其是新芯片需要液冷。要把所有这些都做好,部件可能会出问题,以及如何更换它们。这都是非常深厚的专业知识。这不是开玩笑。不是明天我就能成为数据中心运营商。我不能。
And it is heterogeneous actually. Before, I really mean homogeneous design is much better for operation. This is heterogeneous, and then how to operate this heterogeneous design requires unique innovation in data center deployment and so on. So data centers aren't commoditized. You can specialize in data center deployment, and one data center is better than another, and data center deployment can be done well and badly. Data science is so complicated. If you think about the beginning all the way from construction to power deployment, you need to have the right power to come in, fiber channel, the right cooling, especially new chips requires liquid cooling. To get all this right, the parts can fall apart and how to replace them. It is all very deep expertise. It's no joke. It's not tomorrow I can be a data center operator. I cannot.
这难道不是你会长期押注中国的地方吗,恕我直言?尤其是在美国,数据中心部署的最大障碍之一是政策以及当地的法律基础设施,这些都会阻碍部署。在中国,你没有这些问题,数据中心部署要快得多。我认为总体而言,基础设施,即中国的实体基础设施建设进展非常快。我真的看到某种立交桥在一周内建成。那里的速度非常非常高。而我家附近的一条高速公路,一年后还没完工。所以这也是一个对比。
Is that not where you would bet long on China with the greatest of respects? Especially in the US, one of the biggest barriers to data center deployment is policy and is kind of local legal infrastructure that prevents it. In China, you don't have any of that and data center deployment is much much faster. I think in general infrastructure, the base physical infrastructure construction in China is going really fast. I literally see some kind of crossover bridge being built within a week. The velocity is very very high there. And there's a highway close to my home after one year. It's not done yet. So this is also a crossover.
所以我确实认为有一个独特的优势,可能是因为人口密度,而且他们专门从事这类建设工作。但我确实认为这里我们也有那些专业人才。只是我甚至听说连电工都严重短缺。
So I do think there's a unique strength probably because of the population density and they are specializing in those kind of construction related work. But I do think here we also have those specialty people. It's just even I heard even electrician is under severe shortage.
是的,我们这里也受到全球供应链的限制。
Yeah, we are under global supply chain constraint here.
进入数据中心层会对利润率带来什么变化?会从 30% 提高到 50% 吗?还是说没那么有意义?比如,那会对利润率产生什么变化?
What change would moving into the data center layer cause to margins? Would that take it from 30 to 50? Would it be not that meaningful? Like what would that change due to margins?
如今我们如何计算毛利率很有意思。因为硬件折旧的时间已经发生了显著变化。过去是六年,整整六年。而硬件发布通常每三年一次,这已经很快了。现在,仅一家供应商一年内就有三个 SKU。更新的模型通常在最新的硬件上运行最佳。模型折旧也非常快。我们每周都在发布新模型。模型的价值在下一个模型出来之前达到顶峰。而新模型喜欢最新的硬件。想象一下两年后的这种节奏。哪个模型会在两年前的硬件上运行?那将是两年前的模型。那些模型还有价值吗?所以我认为这就是我们现在面临的真实动态。硬件仍然可以使用六年,但是……
How we calculate gross margin is interesting these days. Because how long does hardware depreciate has significantly changed. In the past it's six years. Solid six years. And hardware release is usually three years. That's fast. And now within a year from one vendor alone we have three SKUs. And the newer model usually runs the best on the newest hardware. Model depreciation is also very fast. Every week we are launching a new model. And the model is kind of peak in its value before the next model comes out. And the new model likes the newest hardware. Imagine this cadence after two years. Which model runs on the two years old hardware? It'll be two year old model. And are those models still valuable? So I think that's kind of the real dynamics we are facing right now. The hardware will last for six years still, but...
但你的意思是模型开发的速度远远超过了芯片和硬件折旧的速度。
But what you're saying is the speed of model development far outstrips the speed of chip and hardware depreciation.
模型的速度绝对是最快的,但即使是硬件创新本身也是最快的。所以三年后,如果每年有三个硬件 SKU,三年间就有九个硬件 SKU。你还想回到九代前的旧硬件上运行三年前的模型吗?这值得怀疑。也许会有战争。它仍然有价值,但以这样的创新速度,值得怀疑。现在,不同的折旧周期改变了自建与自持、自建与购买之间的动态。这又回到了我最初的论点:你是优化增长还是优化毛利率?一切都关乎时机。
The speed of model definitely is the fastest, but even the hardware innovation itself is the fastest. So after three years, if every year there's three hardware SKUs, after three years there are nine hardware SKUs in between. Do you still want to go back to nine generation older hardware running three years old model on that? That's questionable. Maybe there's a war. It's still valuable, but with this pace of innovation, it's questionable. Now with a different depreciation cycle, it changed the dynamics of build versus own, build versus buy. And again, it goes back to my original thesis: do you optimize for growth or do you optimize for gross margin? It's all about timing.
你自己是怎么考虑这个问题的?当你周日下午坐在扶手椅上想,嗯,我们在优化增长?那么,什么时候是优化毛利率的时机?
How do you think about that question for yourself when you're sitting there in an armchair on a Sunday afternoon thinking, hm, we're optimizing for growth? Now, when is that time to optimize for gross margin?
嗯,我会说我们优化。我们想要两者都优化。所以我是这样想的。优化增长需要大量的业务规划,假设存在产品市场契合。优化毛利率就是优化差异化。我认为我想避免过度优化毛利率,但我们应该持续优化毛利率,就像我们应该持续优化产品差异化一样。这毫无疑问。而且我认为我们想要持续优化到一个健康的毛利率,这能让我们快速增长,这是一个权衡,我们不想妥协。妥协是指我们过度优化毛利率导致增长非常缓慢。优化毛利率的一种可能方式:我们根本不增长。我们只是拼命优化它。我知道我们可以达到一个很高的数字,但那绝对是灾难性的结果。
Well, I would say we optimize. We want to optimize for both. So here's how I think about it. Optimize for growth requires a lot of business planning assuming there's product market fit. Optimize for gross margin is optimized for differentiation. I think I want to avoid overoptimizing for gross margin, but we should optimize for gross margin continuously as in we should optimize for product differentiation continuously. There's no question about it. And I think we want to continuously optimize towards a healthy gross margin which allows us to grow really fast, and it's a trade-off and we don't want to take compromises. The compromise as in we overoptimize gross margin resulting in very slow growth. One possible way to optimize gross margin: we do not grow at all. We just optimize a heck out of it. I know we can climb to a high number, but that's absolute disaster outcome.
好的。有意思。如果我们只是说,嘿,那么毛利率,我们打算从 30% 降到 10%。这是一个赢家通吃的市场吗?我们可以吃掉所有人的午餐,然后再优化毛利率?
Okay. Interesting. If we just said, hey, so gross margin, we're going to take it from 30% to 10%. Is it a winner take all market where we could eat up everyone else's lunch and then optimize gross margin later?
我认为赢家可能不是某个时间点的快照。这将是一个长期的情况。我们确实看到某个行业会波动,然后开始稳定在少数几家好的公司。以法律行业为例。我曾在一次晚宴上,有趣的是,大约两年前似乎有很多这样的公司,但现在基本上只剩下两家了。
I think winner is probably not a snapshot in time. It's going to be a long-term situation. We do see a particular industry will oscillate and start to settle with a few good ones. Take legal for example. I was on a dinner table and interestingly seems like there were a lot of those companies around two years ago, but now it's pretty much two.
两家。
Two.
两家。所以我认为是的,这是一场漫长的游戏。
Two. So I think yeah, I think it's a long long game.
你如何看待你所在市场更成熟的状态?是像云市场那样,显然有 Azure、AWS、GCP,还是像 Uber 和 Lyft 那样,一家占 90%,其他家争抢残羹?
How do you see the more mature state of your market? Is it like a cloud market where you have obviously Azure, AWS, GCP or is it an Uber and a Lyft where one takes 90% and the others kind of fight for scraps?
我们正处于采用曲线上,越来越多的 AI 领域的公司开始认真考虑转向专用智能,开始认真考虑拥有自己的智能比租用更好。因为回到这个优化问题,什么时候是好时机?所以这是我们自己在回答的问题:自建与购买。而我们的客户也在考虑自建与购买,或者自建与租用,或者拥有与租用。我认为 AI 之旅或 AI 采用之旅已经走得更远了。很多公司有可观的流量,很多公司正在将 AI 部署到生产中。很多公司处于规模化阶段。而这正是优化发挥作用的时候。当优化开始时,你需要有控制权才能优化。如果你没有控制权,你就没有优化的空间。而为了拥有控制权,你必须在某个开放模型之上构建。你必须将你的数据转化为你的智能。
We're in the adoption curve where a lot more companies in the AI space start to seriously think about moving to specialized intelligence, to start to seriously think about owning their intelligence is better than renting. Because going back to this optimization, when is the good timing? So it's the same question we're answering for ourselves: build versus buy. And our customers are also thinking about build versus buy or build versus rent or own versus rent. I think AI journey or AI adoption journey has gone further along. A lot of companies have meaningful traffic, a lot of companies are deploying AI into production. A lot of companies are at the phase of scaling. And that's where optimization kicks in. When optimization kicks in, you need to have control to optimize. If you don't have control, you just don't have the range to optimize. And for you to have the control, then you have to build on top of some open model. You have to kind of turn your data into your intelligence.
说到拥有自己的智能与租用智能,这确实适用于国家层面。我们看到 Fable 在某些情况下被政府短暂封禁了 19 天,尤其是在欧洲,我们突然意识到:‘天哪,我们不能受制于 OpenAI 和 Anthropic,我们的医疗服务可能被禁用,基础设施可能被某个政府关停。’我们是否会看到主权模型的未来,即大国或国家集团拥有自己的主权模型?
Speaking of owning your own intelligence versus renting it, that does apply to a national layer. When we've seen Fable be banned in some cases by the administration briefly for 19 days, especially in Europe, we suddenly went, 'Oh my gosh, we cannot be at the hands of OpenAI and Anthropic where we can just be banned in our health services, sit on the infrastructure of something that an administration can turn off.' Do we see a future of sovereign models where large nations or nation blocks own sovereign models?
我绝对看到这种可能性。我也看到,如果我们把通用智能模型看作电力层、看作输电线,那么每个国家都应该拥有自己的输电线,对吧?所以我认为这是一个非常可怕的时刻:我的输电线被切断,我所有基本的日常活动都无法进行。我自己在家停电时感到非常沮丧。当我无法连接 Wi-Fi 时,我感到非常沮丧和焦虑。显然,国家运作极其重要,它建立在这个基础之上。对于每家公司来说也是如此。不仅仅是国家应该拥有独特的主权独立性,每家公司也应该拥有自己的独立性。你不想让任何一个人切断你的连接。那是一个非常可怕的时刻。
I definitely see that possibility. I also see, if we think about the general intelligence model as the electricity layer, as a power line, every country should own their own power line, right? So I do think that is a very scary moment: my power line is going to be cut off and all my fundamental day-to-day is not going to work. I feel so frustrated whenever there's a power outage in my home alone. I feel so frustrated when I cannot access my Wi-Fi. I feel so anxious. Obviously, operating the country is extremely important, built on top of this fundamental baseline. And for every single company, the same thing. It's not just whether a country should have their unique sovereign independence, but every single company should have their independence. You don't want any single person to cut you off. That's an extremely scary moment.
为什么你会进入数据中心领域,却不进入芯片领域?
Why would you move into the data center space but you wouldn't move into the chip space?
因为我知道造芯片极其困难。
Because I know building a chip is extremely hard.
我以前也这么想。好吧,我承认我……这也是为什么我觉得这个节目还算成功。我以前也这么想。但为什么每个人似乎都在做这件事,好像这只是另一个产品?正如我所说,OpenAI、Anthropic、DeepSeek、Meta,我们都在造自己的芯片。
I thought so too. Okay, again, I admit to being a... which is why I think the show is a little bit successful. I thought so too. But then how come everyone is seemingly doing it as if it's like just another product? As I said, OpenAI, Anthropic, DeepSeek, Meta, we're building our own chips now.
我认为 Meta 造芯片已经超过五年了,远远超过五年。MTIA 项目从 2018 年甚至更早就开始了。因为 Meta 在 AI 上投资了很长时间,在 AGI 之前,他们就有巨大的 AI 工作负载,专注于排名和推荐,而且 Meta 过去也建造过其他硬件。所以当使用量超过某个阈值时,建造底层供应在经济上就合理了,对吧?而且你可以针对你的工作负载进行专门化。这是另一种专门化:将你的逻辑固化到硬件中,这个硬件是为你的特定工作负载量身定制的。你最好确保这个工作负载不会改变,因为一旦硬件流片,就很难回头修改。虽然可能,但成本非常高。所以一旦你的工作负载稳定了,业务稳定了,不会频繁变化,那时才是考虑造芯片的时候。我仍然认为整个 AI 世界,尤其是模型定制,非常动态,工作负载模式非常非常动态。想想应用领域有多少能量,人们在尝试各种东西。你不知道哪个会起飞,而且它们会迅速起飞。哪个起飞后能持续?只有少数能持续。那时才是:‘哦,现在我们知道了这个模式,我们应该把这个模式编码到硬件中,然后把硬件带到数据中心等等。’这是一个级联的过程,然后它会一直级联下去……这是一个根本问题:我们处于工作负载成熟度的哪个阶段?我的意思是,我们仍处于工作负载成熟度的早期阶段,还不足以支撑一个耐用的芯片。所以现在你回过头来看:‘哦,我们有很多加速器,它们很成功,有些非常成功。’但请记住,那些公司是在 AGI 之前起步的。它们从优化某个工作负载开始,然后转向 AI,试图适配 AI 工作负载。这几乎就像你在 AI 工作负载出现之前下注,现在变成了一个机缘巧合的问题。你是否足够幸运,让它刚好奏效?有些确实奏效了。一些基本设计,比如在芯片上放置大量 SRAM,对 AI 模型很好,因为它们对内存很渴求,这确实加速了推理执行等等。所以那些成功了,有些则没有。
I think Meta has been building their chips for more than five years, way more than five years. And MTIA has been a project since 2018, maybe earlier. So because Meta has been investing in AI for a long time, pre-AGI, and they have a huge AI workload focused on ranking and recommendation, and Meta has been building other hardware as well in the past. So whenever the usage has passed a certain threshold, it makes economic sense for you to build the underlying supply, right? And you can specialize towards your workload. That's another form of specialization: specialized to bake your logic into hardware, and this hardware is purpose-built for your particular workload. And you better make sure this workload doesn't change, because once the hardware is taped out, it's really hard to go back and change it. It's possible, but very costly. So once your workload stabilizes, once your business stabilizes, it doesn't change too often. Then that's the time to consider building a chip. I still see the whole AI world, especially model customization, is very dynamic, very very dynamic workload patterns, very dynamic. So think about how much energy in the application space, people experimenting all kinds of things. You don't know which one is going to take off, and they will take off quickly. And which one, once they take off, is going to sustain? A few ones will sustain. Then that's the time: 'Oh, now we know this is the pattern, and now we should probably encode this pattern into hardware, and bring this hardware into a data center, and so on.' It's all cascading, and then it's going to cascade down to... It's a fundamental question: where are we in the stage of workload maturity? I mean, we're still in the early stage of workload maturity to warrant a chip that will be durable. So now you go back to: 'Oh, we have so many accelerators, they are successful, some are really successful.' But remember, those companies started before AGI. They started from some optimized workload, and they operate to AI and try to fit AI workload. It's almost like you bet before this AI workload emerges, and now it becomes a serendipity question. Are you lucky enough that this just works? And some really worked. Some fundamental design of putting a lot of SRAM on the chip is great for AI models because they are memory hungry, and this really accelerates the execution of inference, and so on. So those works, and some didn't work.
你认为当前最大的瓶颈是什么?我记得我请 Groq 的 Jonathan 上节目时,他说 HBM 是最大的瓶颈,这就是为什么价格涨了 5 倍。你认为人们谈论不够的最大瓶颈是什么?
What do you see as the greatest bottleneck today? You know, I think it was when I had Jonathan from Groq on the show who said like HBM was the greatest bottleneck and that's why you've seen the 5x increase in price. What do you see as the greatest bottleneck that people don't talk about enough?
我仍然认为我们没有为非常大的模型设计出好的系统。我真的相信底层基础设施成本会下降。所以解决一个任务,我们应该需要更少的 token,这会增加。所以总体而言,成本会显著降低。因此,未来我们可以更普遍地运行最高智能的模型。但我们没有为此设计的系统。例如,我们今天没有为 10 万亿参数模型设计出好的系统,这需要从模型到定制服务平台层再到芯片层的非常聪明的工程协同设计。芯片不是单个芯片,而是系统中芯片的集合,以及整个封装。我认为我们还有很多创新可以做。
I still think we don't have a great system for very large models. I really believe the fundamental lower-level infrastructure cost will go down. So for solving a task, we should need fewer tokens, that will increase. So collectively, the cost will significantly reduce. Therefore, we can run the highest intelligence model much more ubiquitously in the future. But we don't have a system designed for that. For example, we don't have a great system designed for 10 trillion parameter models today, and that will require very smart engineering co-design from the model to the customization serving platform layer all the way to the chip layer. Chip is not the individual chip, but the system collection of chips in a system, and all as a total package. I think there's still a lot of innovation we can do.
你最近宣布年收入达到 8 亿美元,真是不可思议的成就,增长如此之快。到今年年底会是多少?
I think recently you announced that you were at 800 million in AR, incredible feat and scaled so fast. What is that at the end of this year?
我们认为至少能翻倍。
We think we can at least double.
到年底。哇。作为风险投资人,我觉得特别有意思的是:我投资了 10 年。过去 Slack 是金童,18 个月内收入从 100 万到 1000 万就很了不起了。现在我们有 Fireworks 这样的公司,几年内收入就达到 8 亿美元。你提到 Cursor 几年内收入达到数十亿。公司收入增长的速度真是无与伦比。
By the end of the year. Wow. You know what's so interesting for me as a venture investor? I've been in investing for 10 years. We used to be in the day where Slack was the golden child where like 1 to 10 million in revenue in 18 months was like amazing. And now we have companies like Fireworks where you scale to 800 million in revenue in a matter of years. And you mentioned Cursor scaling to billions in revenue in a matter of years. The speed of company revenue growth is just unparalleled.
我认为这是因为这项技术带来了根本性的颠覆,它赋能一切。赋能一切的意思是,它触及我们每一个人,让我们发挥创造力,释放了我们原本无法获得的巨大创造力。这就是为什么我们看到这种极速增长的现象,因为需求在那里。
I think it's because there's a fundamental disruption in this technology that is all empowering. And all empowering in the sense it reaches out to every individual one of us to be creative, and it unleashes a lot of creativity that we just don't have access to. And that's why we're seeing this phenomenon of extremely fast growth because of a demand.
在快速问答之前最后一个问题。你雇了 George Hu,他之前是 Salesforce 的总裁。他非常出色。
Final one before we do a quick fire. You hired George Hu, who was president of Salesforce. He's exceptional.
他是我见过最直接、不废话的运营者之一。但你一两年或一年前见过他,当时你说“哦,我们还没准备好请你”。为什么那么说?为什么现在觉得时机到了?
He's one of the most direct no BS operators I've ever met. But you met him a couple of years before or a year before and you were like, 'Oh, we're not ready for you yet.' Why did you say that? And why did you decide now was the time?
对。一年前,我们大概只有 50 人。现在我们有 200 人了,规模还是不算大。
Right. So, a year ago, I think we were probably just 50 people. Today we're at 200 people. We're still not that big.
哇。你们已经领先 400 万了。
Wow. You're four million ahead.
对。在 50 人规模时,我更优先考虑的是先把产品做大规模,然后再大规模扩展业务。我们聊过,我非常尊重他,我知道他是个传奇。他是红杉、阿里巴巴的传奇运营者。我只是觉得我们对他来说太小了。我告诉他:“嘿,我们可能对你来说太小了,但我希望以某种方式与你合作。”于是他帮我组建团队,面试了很多高管。他的反馈总是非常平衡、深思熟虑,我们以那种方式合作到去年年底。我们增长非常快,他也知道,于是我们开始认真谈,早期的关系得到了回报。他真的很酷。他过去做了很多事,成就斐然,但我发现他有一个独特的特质:他经验极其丰富,商业视野很高,但同时非常好奇。他不会假设“我都懂,我看过所有剧本,都是一样的,让我像过去一样指挥”。他没有那种态度。他知道 AI 发展速度极快,他在学习的同时也完全拥抱 AI。实际上,他的团队,我们的 GTM 团队,正在使用各种 AI 智能体,他们分享技能,从而最大化生产力。他知道我们有一条超线性的需求曲线,我们只能以一定速度建设 GTM 团队。为了抓住这条曲线,我们需要建设团队,但团队也需要不断提高生产力来匹配。这就是他正在解决的问题。我很幸运能与他合作。总的来说,我觉得在 AI 领域,独特之处在于人们需要非常特殊的特质,几乎是矛盾的。比如经验丰富但超级好奇、学习曲线快。或者 Demma,我们之前聊过:他聪明绝顶,智力超群,但极其谦逊。这是一个奇怪的组合。
Yeah. So at 50 people, I'm more thinking about scaling the product first, then getting you know massively scale the business. We talked and I have huge respect for him, but I know he is a legend. He's legendary. He's a legendary operator in Sequoia, Alibaba. And I just feel like we're too small for him. I told him, 'Hey, we're probably too small for you, but I would like to work with you in some capacity.' So he helped me actually build up the team, interview a lot of executives. His feedback is always well balanced, very thoughtful, and we started to work together in that capacity until I think end of last year. We're growing really fast and he knows, and we start talking seriously, and that early relationship paid off. So he's really cool. He's really cool in the sense that he did a lot of things, great accomplishments in the past, but I find a unique character about him: he's extremely experienced, has high aptitude of business vision, but he is also very curious. He doesn't make assumptions like 'I know it all, I've seen all the movies, it's the same movie, let me just direct this as I did in the past.' So he didn't come with that attitude. He knows AI moves at an insanely fast pace, and he's learning along the way but also fully embraces AI. Actually, his team, our GTM team, is using all kinds of AI agents; they're sharing skills. So they maximize their productivity. And he knows we have a super linear demand curve, and there's just a certain pace we can build our GTM team. In order for us to catch this curve, we need to build a team, but the team also needs to have increasing productivity to match. So that's a problem he's solving. I feel very fortunate to work with him. In general, I feel in the AI space, the unique part is people need to have very special traits, almost like contradictory characteristics. For example, very experienced but super curious and fast learning curve. Or Demma, we talked a little bit earlier: he is brilliant, high intellectual horsepower, but extremely humble. It's a weird combination.
他也非常棒。
He's amazing too.
他几乎有一种东欧式的愤世嫉俗,但同时非常谦逊。
And he's almost like cynical in an Eastern European way, but also at the same time very humble.
我能来一轮快问快答吗?
Can I do a quick fire round?
好,来吧。
Okay, let's do it.
过去 12 个月里,你改变最大想法的是什么?
What have you changed your mind on most in the last 12 months?
我认为是我们增长的速度。我改变了想法,因为我一直很担心团队过早变得太大。所以当我见到 Georgia 时,我告诉他我们对你来说太小了,因为我不打算在人员上增长太快。我非常担心速度变慢、失去敏捷性。但自那以后,我们非常积极地使用我们的工具。我们开发了自己独特的方法来招聘特定类型的人——我们知道他们会高速前进,有极强的 ownership 意识,非常善于沟通,从不接受“不”作为答案。我们也学会了如何找到这些人。现在我对快速前进感到舒服多了。
I think how fast we grow. I changed my mind because I have been quite worried about too big a team too early. So that's why when I met Georgia, I told him we're too small for you, because I don't intend to grow very fast in terms of people. I worry about slowing down and losing agility and velocity very deeply. But since then, we have been very aggressively using our tools. We have developed our own unique way of hiring certain type of people that we know will be charging forward with high velocity, extreme sense of ownership, very communicative, and never take no as an answer. So we also learned how to get those people. And now I feel much more comfortable skating really fast.
你招什么样的人?我知道这听起来奇怪,但我们招的人其实非常具体:几乎只招移民。英国人不怎么努力,抱歉。非常科学严谨,大多数事情用数据驱动。我实际上认为创造力往往来自数据,并由数据启发。坚定不移地负责和 ownership:没有任何事是别人的错,都是我的错,即使真是别人的错。这就是 20VC 的人。你的标准是什么?
What's your type of people? And I know that sounds weird, but like our type of people is actually really specific. Pretty much only hire immigrants. British people don't work very hard. Sorry. Very scientific and rigorous. Use data for most things. I actually think creativity often comes from data and is informed by data. And unwaveringly accountable and ownership: nothing is anyone else's fault, it's all my fault, even if it's someone else's fault. That's a 20VC person. What would you say yours is?
不是奇怪的能力,不是能力本身。我们需要自信的人,但更重要的是,他们能否在这一波浪潮中(尤其是在 Fireworks)做好的强烈指标是:他们是否真正具备 extreme ownership。Extreme ownership 意味着我们不把人放进任何框框里,然后把框堆成塔。人们会自动认领:“嘿,这个端到端的问题,我要从头到尾搞定,和一群人合作实现它,无论如何都要交付。”这类人跑得最远、最长,他们的成长曲线惊人。
It's not in a weird way, it's not competence. We need people with high confidence, but more importantly, the strong indicator of whether they will do well in this wave, especially in Fireworks, is whether they are really built for taking extreme ownership. Extreme ownership as in we are not putting people in any boxes and stacking the box together into a tower. People just automatically claim, 'Hey, this end-to-end problem, I'm going to see through the whole thing and work with a bunch of people to make it happen, and I'm going to deliver it no matter what.' Those kind of people have the highest longest mileage, and their growth curve is amazing.
你从与黄仁勋的合作中学到的最重要一课是什么?是什么让他如此特别?
What's your biggest lesson from working with Jensen Huang on what makes him so special?
他无处不在。我真的觉得他有几百个克隆体插在各个地方。比如我给他发邮件,他一分钟内就回复。我不明白他怎么能一直关注细节。但运营公司四年后,我理解了他为什么这么做。这定义了速度。因为领导力是什么?领导力就是判断。不是特权,是判断。你基本上需要掌握背景信息。你需要正确的背景信息来为团队做出正确判断。尤其是在高速领域,如果你不知道正在发生什么、什么有效、什么无效、差距在哪里,你就会做出错误决策。在慢速领域,你可以等待信息层层上传下达再做决策。但在快速演进领域,你等不了,因为信息在层层传递中必然会有损失。人总是会这样,不了解实际情况就做判断会导致糟糕的领导力。他在这个疯狂的 AI 热潮之前就以这种方式运作。以前我佩服他做这件事的巨大能力。现在我理解了背后的智慧,因为我也以这种方式运作。我需要知道一线发生了什么,才能为公司做判断。
He's everywhere. I seriously think he has a clone of like hundreds of Jensens somehow plugged in. For example, I sent him an email. He will reply in one minute. I just don't understand how he's constantly in details. But now I operate a company for four years, I understand why he's doing that. That defines velocity. Because what is leadership? Leadership is just judgment. It's not privilege. It's judgment. You basically have the context. You need to have the right context to make the right judgment for the team. Especially in a high velocity space, if you do not know what's happening, what works, what doesn't work, what are the gaps, you make the wrong call. In a slow-moving space, you can wait for the cascading information up and down and make those calls. But in a fast graduation space, you just cannot wait, because it's guaranteed there is information loss in transition layer after layers. People always happen, and not knowing what exactly is happening and making judgment from that position makes bad leadership. He demonstrated through his own example, even before this crazy AI thing, he's operating that way. Before I was admiring him for his sheer amount of volume of capability of doing that. Now I understand the wisdom behind that, because I also operate that way. I need to know what's happening on the ground to make the judgment for the company.
在 Fireworks 的旅程中,你等了什么,但希望自己当初没有等?
What did you wait on in the Fireworks journey that you wish you hadn't waited on?
营销,我们聊过这个。我们在这方面有点书呆子气:在旅程的最初阶段,我们没怎么讨论过,但我们觉得产品会自己说话。最终,产品立得住,我们想把所有精力和焦点放在打造产品、与客户合作、验证产品市场契合度上,然后从那里出发。我们根本没花多少时间在营销上。
Marketing, we talk about it. So, we are a little bit nerdy in this way that at the very beginning of our journey, we kind of didn't discuss it, but we felt product speaks for itself. At the end, product stands, and we want to devote all our effort and focus on building product, working with customers, validating product-market fit, and go from there. And we didn't spend much time marketing at all.
我们过去没有优先教育客户,让他们了解正确的趋势方向和价值。但现在我们认为,我认为这很重要。营销不是关于流量,更多的是教育,更多的是清晰度。我们正在努力改进。
We didn't prioritize educating our customer what's the right direction to think about the trend and the value. But we do think now, I do think it's important. Marketing is not about flows, it's more about education, it's more about clarity. And we are working on that.
在你看来,如今 AI 的哪个领域投资不足?你提到了冷却或服务器。哪些领域投资不足?
What area of AI is underinvested in today in your mind? You mentioned like cooling or servers. What are like underinvested in?
我认为 AI 的性感之处在于它是一项如此创新、有创造力的技术,而在此基础上构建东西是重点。但监控 ROI,我认为行业开始关注这一点。但最终重要的是:不是花了多少钱,而是回报是多少,成本是多少,归因是什么。所以我认为在未来几年,随着 AI 越来越多地投入生产,人们会非常关注获得这种清晰度和纪律性。所以 token 最大化只是暂时的,但我们会很快转向 ROI 最大化,这完全是关于经营业务。
I think AI has the sexy part of this is such an innovative creative technology and building something on top of it is the focus. But monitoring the ROI, I think the industry starts to kind of pay attention to it. But eventually that's what matters: not how much spend, but how much is return, and what is the cost and what is attribution. So I think in the next couple of years, as AI is getting more and more into production, there will be a lot of focus on getting that clarity and getting that discipline out. So the token maxing is just, I think, a thing in time, but we'll quickly move into ROI maxing, which is all about running a business.
你最希望拥有但目前还没有的大客户是哪家?
What large customer do you not have that you would most like to have?
我们还没有在传统企业领域投入太多时间。我认为这只是因为我们当时规模很小。现在随着我们公司的发展,我认为即使我们没有主动投资,我们也有像 Geico、Capital One、Mercury Insurance 和 RBI 这样的客户,所有这些公司,即使我们没有主动去争取传统企业客户,他们也来找我们并成为了客户。但我认为那是一个非常大的市场。
So we haven't spent too much time in traditional enterprise segment. I think that's just because we were very small. And now as we build out our company, I do think even without our us investing, we have customers like Geico, like Capital One, like Mercury Insurance and kind of RBI, all these companies even without us pursuing enterprise traditional enterprise, they come to us and they are customer. But I do think that's a very big market.
在年底之前,必须发生什么尚未发生的事情,你才会认为这是好的一年?
What has to happen before the end of the year that hasn't happened for you to consider it a good year?
我对我们推动业务的能力充满信心。对我来说,今年我想证明我们可以通过保持同样的速度快速扩张,这对我来说非常重要。如果我们达到那个点,达到那个验证点,明年我就更有信心继续极其激进地扩张。我想确保我们今年做对了。
I'm confident in our capability of driving the business. And to me, this is a year I want to prove we can scale quickly by keeping the same velocity and that's very important to me. If we reach that point, reach that proof point, and next year I have a lot more confidence to continue to scale extremely aggressively. I want to make sure we do it right this year.
最后一个问题。关于未来三年,没有人看到但你看得非常清楚会发生或不会发生的事情是什么?
Final one for you. What does no one see about the next three years that you see very clearly happening or not happening?
我真的看到人们会拥有——每家公司都会拥有自己的智能,这是必需品,不是可选项。这是我看到的趋势。因为类比软件:每家公司都构建自己的软件栈是有原因的。没有现成的标准化软件可以直接用来解决他们的问题,因为每家公司都在解决独特的问题,他们想要构建软件是因为他们想要完全控制。显然他们会选择自己构建栈的哪些部分,哪些部分是常识,没有构建的必要,但每家公司都拥有自己的软件栈——显然我们在 SaaS 时代就在讨论这个,对吧。所以同样,我认为在 AI 时代,每家公司都应该拥有自己的智能。
I really see people will own their — every single company will own their own intelligence as a must-have. It's not optional. That's a trend I'm seeing. Because there's an analogy to software: there's a reason why every company builds their own software stack. There's no standardized software you just use off the shelf to solve their problem because every single company is solving a unique problem and they want to build software because they want to have full control. And obviously they will pick and choose which part of the stack they want to build themselves, which part of stack is common knowledge there's no point of building, but every single company owns their own software stack — obviously we're talking about this in the SaaS time, right. So same, I think AI time, every single company should own their own intelligence.
Lynn,你知道是 Matt 最先介绍我们认识的。我很高兴认识了你和 George。非常感谢你亲自前来参加。能面对面交流真是太好了,你太棒了。
Lynn, you know it was Matt that introduced us first. I've had the joy of getting to know you and obviously George. I can't thank you enough for joining me, for coming in person. It is so wonderful to do it in person and you've been fantastic.
这个演播室太棒了。你问了很多有趣的问题。和你聊天非常开心。
That's an amazing studio. You asked a lot of interesting questions. I had a lot of fun talking with you.
我们之前做了很多研究,是吧?
We do a lot of research before, huh?
是的,你确实做了。
Yes, you did.
非常感谢,你太棒了。
Thank you so much for that and you're fantastic.