The Ads Business Model Will Die & Lessons from Working with Elon at Twitter | Parag Agrawal
打开互动全文版(中英对照 + 朗读 + 问答)→Parag Agrawal 解释为何智能体使用网络的频率将是人类的 1000 倍、广告商业模式为何会消亡,以及他在推特与埃隆共事的经验教训。
Parag Agrawal explains why agents will use the web 1000x more than humans, why the ads business model will die, and what he learned working with Elon at Twitter.
智能体使用网络的频率将是人类的千倍。因此,需要新技术和新的商业模式。智能体世界在所需技术方面如何改变网络搜索?Parag 是 Parallel 的创始人,正在改变智能体高效进行网络搜索的未来。这是一场关于智能体式搜索的未来、收入与财富不平等的未来等话题的精彩讨论。Parag 很少上节目,所以能在伦敦与他面对面坐下来交谈非常特别。
Agents will use the web thousandx more than humans. Hence, new tech is needed and new business models are needed. How does the world of agents change web search in terms of the technology required? Parag is the founder of Parallel, changing the future of how agents do web search efficiently. This is an incredible discussion on the future of agentic search, the future of income and wealth inequality, and so much more. Parag rarely does shows, and so it was very special to sit down in person with him in London.
广告在目前的形式下无法与智能体配合。有些人确实想做坏事,而模型的对齐并非能抵御对抗性攻击,我认为这些是我们必须更担心的事情。我认为构建模型的人有责任确保你真的做好工作,以尽量减少你所构建的东西带来的伤害。而且我认为到目前为止,其中一些应该被视为尴尬,因为我认为它们证明了两件事。
Ads don't work with agents in their current form. Some people actually want to do bad things and the models alignment is not adversary proof and I think those are the things we must worry about more. I think it's the responsibility of people building models to ensure that you really do the work to minimize that harm that comes from what you've built. And I think so far I think some of these should be considered embarrassments because I think they demonstrate two things.
Parag,我太兴奋了,兄弟。我和 Vinod 聊过。我和 Andrew Reed 聊过。我和 Todd Jackson 聊过。兄弟,我把你研究了个透。所以,感谢你加入我。
Parag, I'm so excited for this, dude. I spoke to Vinod. I spoke to Andrew Reed. I spoke to Todd Jackson. Dude, I stalked the out of you. So, thank you for joining me.
感谢邀请我,也感谢你打了所有那些电话。
Thanks for having me and thanks for making all the calls.
不客气。我想先问,对于任何不知道的人,你会如何在 60 秒内描述 Parallel?
Not at all. I would love to start with for anyone that doesn't know how would you describe Parallel in 60 seconds.
Parallel 是智能体的 Google。所以智能体需要搜索网络来做任何为你做的事情。无论是个人智能体还是为工作构建的智能体。就像人类需要在工作中或生活中经常使用浏览器搜索 Google 一样,你的智能体也需要做同样的事情。事实证明,智能体与人类不同,你为智能体构建网络搜索的方式也不同,所以 Parallel 是关于构建智能体搜索网络的技术,然后建立商业模式使其可持续。
Parallel is the Google for agents. So agents need to search the web to do anything they do for you. Whether it's a personal agent or an agent built for work. Just like humans need to go on a browser search Google often times during work or for whatever you're doing in life, your agent needs to do the same. Turns out agents are different from humans and the way you build web search for agents is different and so Parallel is about building the technology for agents to search the web and then the business models to make that sustainable.
那是你最初的洞察吗?
Was that the original insight that you had?
是的,公司的第一个起源就是我们写下这样的声明:智能体使用网络的频率将是人类的千倍。因此,需要新技术和新的商业模式。千倍让你感受到规模。它改变了你思考底层技术构建的方式,因为为某个规模构建的技术无法在三个数量级的变化中存活。然后当你需要新的商业模式 alongside 新技术时,问题就变得非常有趣。
Yeah, literally the first genesis of the company was us like the statement that agents will use the web thousandx more than humans like I wrote that down at some point. Hence new tech is needed and new business models are needed. Thousandx gives you a sense of scale. It changes how you think about building the tech underneath because like no tech built for a certain scale survives three orders of magnitude. And then when you need new business models alongside new technology a problem becomes really interesting.
智能体世界在所需技术方面如何改变网络搜索?
How does the world of agents change web search in terms of the technology required?
答案有很多很多层次,但让我们从我们提到的第一件事开始,那就是规模。对吧?现在如果你想想,假设智能体最终确实搜索网络多了 1000 倍。如果我们花费目前用于网络搜索的算力,那对网络搜索来说算力太多了。所以你现在需要让它更高效,也许我认为需要 10 到 100 倍才能合理。
There are many many layers to the answer but let's start at the first thing we mentioned which is scale. Right? Now if you think about let's say agents actually do end up searching the web a,000x more. If we spend the amount of compute we currently spend as per web search that's too much compute for web search. So you now need to make it way more efficient by perhaps we I think 10 to 100x for it to make sense.
是的。
Yeah.
如果你做一千个鸡蛋。是的。
If you're doing a thousand eggs. Yeah.
确切地说,就像你……第二件事是智能体扩展,就像人类在非常狭窄的区域内操作。所以如果你想想我们如何使用网络搜索,我们输入简短且不明确的关键词查询。我们等待大约半秒到 1 秒。如果网络搜索超过这个时间,我们就会不耐烦。然后我们得到 10 个蓝色链接,然后我们随机浏览它们和一系列搜索来完成我们正在做的事情。智能体不是这样的。智能体可能会告诉你它们到底在找什么,不是三个关键词,而是一个完整的句子,比如这就是我试图做的事情。智能体要么超级不耐烦,比如想象一个语音智能体。智能体会说,我现在就需要答案,比如 100 毫秒。我不能等 500 毫秒,因为人类在等我,他们会等 500 毫秒。所以我需要网络搜索在 100 毫秒内完成,或者会有一个后台智能体说我不在乎,只要给我最好的答案。对吧?所以你在网络中可以做的事情的方差完全改变了。事实上,最不有趣的是我们为人类构建的东西。对吧?你要么有太多时间,要么时间太少。几乎从来不是相同的时间。输出也不一样。所以输入不同。你拥有的时间不同。输出不同。输出是蓝色链接。输出是 token 或文件系统上的文件,取决于智能体的类型。所以现在突然你说,好吧,现在问题的输入不同,输出不同,约束不同,你可以花费非常不同的算力。对吧?所以想象有人运行一个用 Luna 模型构建的智能体,然后想象有人运行一个用 fable 模型构建的智能体。它们是非常不同的模型。你希望为每个模型优化信噪比和 token 的方式在网络搜索栈中你所做的事情方面是如此不同。
Exactly. like you um the second thing is agents expand like humans operate in a very narrow zone. So if you think about how we use web search we type keyword queries which are short and underspecified. We wait for about half to 1 second. If web search takes more than that we're impatient. Um, and then we get 10 blue links and then we random walk across them and a collection of searches to get what we're doing doing. Agents are not like that. Agents are going to um perhaps tell you exactly what they're looking for, not like three keywords, but like a full sentence like this is what I'm trying to do. Agents will either be super impatient, like imagine a voice agent. The agent will be like, I need an answer like now 100 millconds. I can't wait 500 millconds because the human is waiting on me and they're going to wait 500 millconds. So I need web search to do it in 100 or there going to be a background agent who's like I don't care just give me the best answer possible. Right? And so the variance of what you can do within web changes completely. In fact the one thing that is the least interesting is what we've built for humans. Right? You either have too much time or too little time. Almost never the same amount. um the output is not the same. So the input is different. The time you have is different. The output is different. The output is in blue links. The output is tokens or files on a file system depending on the type of agent. And so now all of a sudden you say okay now the problems inputs are different, outputs are different and constraints are different and you get to spend very different amounts of compute on it. Right? So imagine someone running an agent built with a Luna model and then imagine someone running an agent built with a fable model. Um they're very different models. How you want to optimize signal to noise and tokens for each of them is so different in terms of what you do in the web search stack.
所以你说的优化信噪比是什么意思?
So what do you mean optimize signal for noise?
这样想。假设你采用网络搜索,它便宜、快速、低算力。概念化网络搜索问题的一种方式是,你从网络上大约一万亿份文档开始,也许几万亿。给定任何搜索,我现在需要将其缩小到你的模型上下文窗口应该看到的 1000 个 token。所以问题是从一万亿个 URL,每个假设几千个 token,缩小到总共 1000 个 token。那么我们在网站中如何做?我们首先说好吧,我们要做检索。所以对于大多数文档,我将花费零算力。对于少量文档,我将花费极少的算力来确定要查看哪 10,000 个。一旦我得到这 10,000 个,我将为每个文档花费稍多的算力,将其缩小到 1,000 个文档。
So think of it this way. Let's say you took web search which was cheap and fast and low compute. Uh one way of conceptualizing the web search problem is you start with like a trillion documents that are on the web some few trillion. Given any search I now need to narrow it down to a thousand tokens that your model's context window should see. So the problem is going from a trillion URLs with let's call it a few thousand tokens each down to a thousand total tokens. So how do we do it in website? We first say okay we're going to do retrieval. So for most documents I'm going to spend zero compute. For small number of documents I'm going to spend minuscule amounts of compute to figure out which 10,000 to look at. Once I get these 10,000 I'm going to spend a little bit more compute for each of these 10,000 documents to narrow it down to 1,000 documents.
嗯。
Mhm.
我会继续用越来越大的模型和具有越来越多特征的排序器来做这件事,直到我能为你的模型缩小到 1000 个 token。对吧?所以你本质上是在为网络搜索分配算力,以节省模型上的算力。这大致就是这里发生的事情。所以如果你想为模型节省算力,问题是你应该分配多少算力?所以如果你的 Luna 模型真的很便宜,你不想在网络搜索上做太多计算,因为将更多信息泄漏到 Luna 的上下文中是可以的,因为它便宜。对于 Fable,你想在浪费 Fable 的时间之前完成工作,因为如果网络搜索因为你省钱而给出更差的答案,那将在时间和金钱上对你来说很昂贵。
And I'm going to keep doing this with bigger and bigger models and rankers with more and more features until I can narrow it down to a thousand tokens for your model. Right? So you're essentially allocating compute to web search in order to save compute on the model. That's roughly what's going on here. Um so if you want to save compute for the model, the question is how much compute should you allocate? So if your Luna model is really cheap, you don't want to do too much computing web search because it's okay to leak a little bit more information into Luna's context because it's cheap into Fable. You want to do the work before you waste Fable's time because that's going to be expensive in time and money for you if web search gives you worse answers because you cheaper out on web search.
我能不能问一下,你们怎么处理智能体需求上的模糊性?我的意思是,对不同的事情,智能体想要的东西可能不一样;对不同的人,智能体想要的东西也可能不一样。我可能非常在意准确率,完全不在意延迟或成本。我也可能非常在意延迟,但完全不在意准确率。
Can I ask how do you deal with the ambiguity of what agents want? And what I mean by that is, for different things, an agent might want different things, and for different people, an agent might want different things. I may really care about accuracy and not at all about latency or cost. I may really care about latency, but not at all about accuracy.
你让智能体在 API 签名里把这些指定出来。我们的产品就是一个 API,程序员可以基于自己的应用来配置,也可以交给智能体去决定。我们的 API 甚至有一个参数,就是“调用我的是哪个模型”,如果你告诉我们,我们就能做一些不同的处理。你不一定要告诉我们,但如果你说了,可能会得到更好的结果。你可以用一个小模型,也可以用一个大的模型,还可以用低思考、中思考或高思考。我们把搜索系统产品化成了几种不同的模式,每一种都针对某一类用例做了优化。
You allow the agent to specify that in the API signature. So our product is an API which either the programmer can configure based on their application or can leave it to the agent. Our API even has a parameter which is like what's the model calling me, and if you know the model we can do things differently. Now you don't have to tell us, but if you tell us you might get better results. You can use a small model or a big model, and you can use it with low thinking or medium thinking or high thinking. We have productized our search system into a few different modes, each optimized for a certain class of use case.
比如,我们有一个非常非常快的 API,是市场上最快的。它又快又便宜,因为在低延迟预算下,你能做的算力就那么多。它是为语音智能体打造的。你的语音智能体必须无所不知,而不能跟你说“我在搜网页,稍等一下我带着答案回来”。对语音智能体来说,那种体验很蠢,对吧?它就应该像变魔术一样立刻回应。所以你现在需要做的是快速的网页搜索。
For example, we have a really really fast API, the fastest in the market. It's fast and cheap because with a low latency budget there's only so much compute you can do. And it's built for voice agents. So your voice agent must be all-knowing without telling you guys I'm searching the web, hold on while I come back with the answer. That's a silly experience for a voice agent, right? But it should just magically immediately respond. And so you now need to do web search which is rapid.
另一方面,你可以用一个 fable 模型,它会先思考 20 秒,然后再生成。你可以在网页搜索上花 5 秒,确保它做的是两次搜索而不是五次。这样在智能体端到端上你就省了时间和成本,所以那是我们系统里另一个处理器,叫 advanced。所以如果你在做语音智能体,就用 turbo;如果你做的是那种很昂贵的后台智能体,就用 advanced。
On the other hand, you can take a fable model and the thing is going to think for 20 seconds and then generate for it. You can spend 5 seconds on web search to make sure it does instead of five web searches only two. So you end to end in the agent you save time and cost, and so that's a different processor on our system, it's called advanced. So you use turbo if you're building a voice agent, you use advanced if you're like really expensive background agent.
那么从客户群来看,今天你们最主要的用例是工程和编程吗?
And so the primary use case today in terms of customer base is engineering and coding for you?
我会说相当广泛。主要的用例,共同主题是知识工作。编程是知识工作的一类,AI 律师也是,生产力应用也是,AI 保险核保人也是,科学家也是。
It's pretty broad I would say. The primary use cases, the common theme is knowledge work. So coding is a category of knowledge work, so are AI lawyers, so are productivity applications, so are AI insurance underwriters, so are scientists.
我想搞清楚工程这块占大头占多少。是 80% 吗?
I'm trying to understand how much of the mother lode is engineering. Is it 80%?
不是。工程这块的情况是,编程目前是市场上推理的一大块。编程里,我会说只有 5% 的提示会触发网页搜索。
No. The thing about engineering is coding is a large chunk of inference in the market right now. Coding invokes, I would say, web search in 5% of prompts.
哇。
Whoa.
对。所以你写代码的时候,并不是每条提示都会触发网页搜索,因为大多数都依赖你的内部上下文、内部数据和代码库。所以模型花时间在读你的内部代码上,而不是在网上。
Right. So it's not every prompt invoking web search when you're writing code, because most of them rely on your internal context and your internal data and your codebase. So the model is spending time reading your internal code and not on the web.
法律这块也不会多太多,对吧?
Law is not going to be much more, is it?
法律非常依赖网页搜索。
Law is very web search-oriented.
真的吗?我还以为它是由内部数据驱动的。
Really? I thought it'd be internal data driven.
有内部数据,但也有判例法。有事实。有关于公司的事实,关于人的事实,而且你还得去排除信息。所以在法律里你必须非常全面,才能说“我想确信,尽管费了很大力气,你也找不到这个”。保险核保也是同样的味道。销售非常依赖网页搜索。AI 科学也非常依赖网页搜索。所以相对而言,推理更多花在算力上,但当你想到网页搜索时,所有这些其他的也开始冒出来了。
There is internal data, but there is case law. There is facts. There's facts about companies, facts about people, and you have to go exclude information. So you have to be very comprehensive in law to say I want to be confident that despite a lot of effort you can't find this. Insurance underwriting has the same flavor. Sales is very web search heavy. AI science is very web search heavy. So on a relative basis inference goes more in compute, but when you think of web search all of these others start popping too.
那这种崛起——我之前对 muse 和 instinct 很兴奋。如果说有什么东西需要网页搜索,那就是个人助理。这会怎么改变你们的业务?
How does the rise — I was wooing before about muse and instinct. If there's anything that needs web search, it's personal assistance. How does that change your business?
对,这很棒。这就像,哇,这是我业务里最棒的事。你知道,我觉得在我们这个行业里,只要智能体因为模型变得更好或更便宜而在更多用例里变得更有用,对我们就是好事,因为我们押注智能体会成为网页的消费者,我们一直在为智能体打造技术。所以当智能体做得更多时,对我们的业务就是好事。所以我们希望模型不断变得更好、更便宜,这样智能体就能做得越来越多。如果这发生了,对我们的业务就是好事。
Yeah, it's great. This is like, woo, this is the greatest thing for my business. You know, I think in our business, anytime agents start becoming more useful for more use cases because models either get better or cheaper, it's great for us, because we bet that agents will be the consumers for the web and we've been building tech for agents. So when agents do more, it's great for our business. So we want models to keep getting better and cheaper so that agents do more and more and more. And if that happens, it's great for our business.
完全理解。我能不能问一下,如果模型变得更聪明,它们下面的智能体不就会做更少的搜索,那对你们的业务不就更糟了吗?
Totally get that. Can I ask, if models get smarter, don't the agents beneath them do fewer searches and then it's worse for your business?
我不这么认为。如果你想想模型,有一个趋势是模型变得更聪明,另一个是模型拥有更多记忆和参数化记忆。这俩是稍微不同的维度。如果你看今天,大体上我会说模型在参数化记忆上对——我们姑且叫它“头部事实”——有很好的回忆能力。就像对某个名人,关于他的一切,维基百科上的,它都能记住。所以它能告诉你某一年总统是谁,对吧,因为模型能记住那些东西。但模型没法告诉你我是哪一年大学毕业的。
I don't think so. If you think of models, there is the trend around models being smarter and then models having more memorized and parametric memory. Like those are two slightly different dimensions. And if you think of today, by and large I would say models have good recall from parametric memory on, let's call them head facts. It's like for somebody famous, everything about them, Wikipedia, it can memorize. So it'll tell you who the president was in a certain year, right, because a model can memorize those things. The model couldn't tell you what year I graduated from college.
真的吗?
Really?
也许对我它可以,但它没法告诉你某个在 Parallel 工作的人的情况。
Maybe for me it can, but it can't tell you for somebody who works at Parallel.
你知道一些事情吗?所以如果我其实,我是说如果我——
Do you know things? So if I actually, I mean if I —
即使它在预训练数据里。
Even if it was in the pre-training data.
说真的。
Seriously.
因为那是有损压缩。模型的参数化记忆在做的是,有损压缩以理解世界上的模式,所以它没法记住预训练数据里的每一个事实。第一,不是每个事实都在预训练数据里。第二,即使是预训练数据里的东西,模型实际上也是在试图找模式,而不是记住它们。再进一步,当你让模型变得高效,也就是通过蒸馏之类的手段让它们越来越小同时保持性能时,你在试图保留推理能力的同时,会失去更多参数化记忆。
Because it's lossy compression. So what a model's parametric memory is doing, it's lossily compressing to understand patterns in the world, and so it can't memorize every fact in pre-training data. So one, not every fact is in pre-training data. Two, it can't, for even stuff that's in pre-training data, the model's actually trying to find patterns rather than memorize them. And then further, as you make models efficient, which is you make them smaller and smaller while keeping the performance by distilling them or whatever, you lose more of the parametric memory while you try to keep the reasoning.
你觉得我们会看到模型变得越来越小,每家公司都有自己的模型和自己的数据,以及那种 fireworks 理论——也就是“拥有你自己的智能”——会成真吗?
Do you think we will see models become smaller and smaller and every company have their own model with their own data, and the fireworks theory of, you know, own your own intelligence being true?
这里面有两个问题。第一,我认为我们会看到两件事。我们会看到最大的、也就是前沿模型随着时间变得越来越大。我们也会看到越来越小的模型能够达到任何固定的性能水平。所以如果你说“好,我想要 Opus 48 级别的性能,好,这对我的用例来说够用了”,那么每 6 个月就会有一个小得多的模型能给你提供这个水平。所以你会看到模型尺寸的可用范围会变得非常非常非常不一样。
So there are two questions in there. So one, I think we're going to see two things. We're going to see the biggest or the frontier models be bigger and bigger over time. We are also going to see smaller and smaller models being able to reach any fixed level of performance. So if you say okay, I want Opus 48 level of performance, okay, and that's good enough for my use case, every 6 months a much smaller model will be able to deliver that to you. So you're going to see the useful range of sizes of models will be way way way different.
如果前沿变得越来越大,这意味着什么?那会有什么后果?
What does it mean if the frontier gets bigger and bigger? What are the ramifications of that?
它们会更好。所以你能想象到的所有后果都会发生。
They are better. So all of the ramifications you imagine.
前沿模型会越来越大的原因,归根结底在于:在某些用例中,增量质量能带来的价值可以高到值得为之付费。所以只要你能造出来,就会有对应的用例。只要把模型做大就能让它更好,而就我们目前所能感知的范围来看,缩放定律似乎还没有尽头。再说一次,我把“更大”和“模型变得更大”混为一谈了,但它们也能思考更久。也就是说,你可以把更多算力砸向同一个问题。而且这会持续下去。我认为我们能不断把更多算力投向同一个问题,让答案随着时间推移一点点变好。所以,我们就是会在极其庞大的模型上花很多钱,去解决非常难的问题。
So the reason the frontier will get bigger and bigger is ultimately the gap between what the value for certain use cases incremental quality can provide you can be so high in certain use cases which can be so valuable that it's worth paying for. So if you can build it, there will be use cases for it. As long as by making it bigger, you can make it better. It appears there is no end to the scaling law that we can perceive so far. And again, I'm conflating bigger with like models are getting bigger, but they're also able to think longer. So you're just able to throw more compute at the same problem. Um, and that will keep happening. I think we'll be able to throw more and more compute at the same problem and make the answer be marginally better over time. And so, we're going to just spend a lot of money on extremely large models solving really hard problems.
你认同那个针对你的共识吗——90% 的 token 活动会走开源模型,但 90% 的钱会流向前沿模型?
Do you agree with the consensus for you that you'll have 90% of token activity go through open models but 90% of dollars go through frontier models?
我了解得还不够,没法有明确看法。我对这个没有观点。我确实觉得,实际上这两者都不会是 90%。
I don't know enough to have a view there. I don't have a view on that. I do think I don't think it'll be 90% on either of those two actually.
真的吗?
Really?
是的。
Yeah.
为什么?
Why?
如果你相信我那个说法——一个有用的模型和前沿模型之间的价格会相差 100 倍、1000 倍——那么小模型仍然有用。很难知道随着时间推移,哪些用例会被优化到中间哪个规模的模型上。而且我认为,开源模型最终走向何处存在真实的路径依赖。我确实觉得,如果我们能很有信心地认为,一个美国造的开源模型会成为开源模型里的最先进水平,我就更有把握说它们会相当不错。
So if you believe my claim that a useful model and the frontier model will be 100x, 1000x off in price from each other. So the small model is still useful. It is hard to know which use cases over time will get optimized to which scale of model in between. And I think there is a real path dependency in terms of where open models end up. I do think if we had strong confidence that an American built open model was going to be state-of-the-art as far as open models went, I would have more confidence in saying that they'll be pretty good.
你对美国的开源模型有信心吗?
Do you have confidence in American open models?
我希望它们存在。到目前为止,还不清楚会发生什么,但我希望会有一系列很棒的美国开源模型,甚至可能出现争夺最佳美国开源模型的竞争。所以我认为,你需要的其实不只是一个人有动力去造一个美国开源模型。你需要两个人互相竞争,去造出最好的美国开源模型。
I want them to exist. So far, it's not clear what's going to happen, but like I'm hoping that there will be a great series of American open models and perhaps even competition to have the best American open model. So, I think what you need is actually not just one person motivated to build an American open model. You need two people competing against each other to build the best American open model.
你觉得模型路由层——美国人叫 routing layer——有价值吗?有人觉得价值巨大,有人觉得会商品化。
Do you think there's value in the model routing layer, routing layer as Americans call it? Some people think immense value, some think commoditization.
同样,存在路径依赖,所以今天它确实有真实价值。今天有真实价值,是因为当我们说路由时,我们谈的是两个不同维度:用哪个模型,以及通过哪个供应商、用哪块 GPU 来跑那个模型。当你身处一个供需非常怪异的世界里,人们在到处找 GPU,人们需要容量来服务客户,而这些人——像我们一样——都是相对早期、增长极快的初创公司,有时增长超出你的预测或预估,然后突然之间你就在找容量。所以你希望能拿到多少就拿多少。这就让拥有一些这样的路由器产生了真实价值。它们能解决这两个问题,说:给我灵活性。如果我需要从不同来源拿 token,到了紧要关头,如果这个模型在 SLA 里没有 token 了,我也愿意换模型。所以此刻,它真的很有价值。现在我不知道 GPU 和 token 的整体供需会怎样,但如果一直保持这样,路由层就真的很有价值。
Again, there's path dependency that so today there is real value. Today there's real value because so when we say routing we talk about two different dimensions which model and which GPU running that model via which vendor. When you're in a world where the demand supply is very weird and um people are like hunting for GPUs and people need capacity to serve their customers and people all of these are like like us like relatively early stage startups growing really rapidly um sometimes beyond what your forecasts or predictions say and then all of a sudden you're looking for capacity. Um, and so you want to be able to get it where you can. And so that creates real value in having some of these routers as login. They can solve these two problems saying give me flexibility. If I need to get a different source of tokens and push comes to shove, if there's no tokens on this model in the SLA, I want I'll switch models as well. So in this moment, it's really valuable. Now I don't know what happens to overall demand supply on GPUs and tokens but if it remains this way the routing layer is really valuable.
你觉得今天不太有价值、但在 3 到 5 年后会变得极其有价值的是什么?
What do you think is not so valuable today that will be incredibly valuable in 3 to 5 years time?
也许是数据。当我说数据时,我认为今天我们还不知道怎么为独特的、有价值的洞见或数据付费。因为随着智能变得越来越便宜,你会想要用你独有的数据,去构建在来自别处的数据或洞见之上。对吧?所以,如果你想象一个抽象的概念:我这里有智能,我拥有一些数据,你拥有一些数据,公共领域里也有一些数据,如果我们能把你的洞见和我的洞见、以及公共领域里的所有这些数据、再加上智能整合起来,我们就能创造出比你自己单独能做的、比我自己单独能做的更大的东西。
Perhaps data. Um when I say data I think it is I think today we don't know how to pay for unique valuable insight or data. Um because we live in a as intelligence is cheaper. You're going to want to build upon either data or insight that comes from somewhere else with your unique data. Right? So what like if you think of like an abstract notion that okay I have intelligence here I own some data you own some data and there's some data in the public domain if we can pull together your insights and my insights and all of this data in the public domain and intelligence on top we can create something bigger than what you could have done by yourself what I could have done by myself.
嗯。
Mhm.
那么现在,你如何完成交易,来创造出这种大于你自己能做的、我能做的、或开放数据所擅长的总和?所以在我看来,那笔交易感觉像是一笔数据交易,或者一笔洞见交易。而我们还不知道怎么以那种方式交易,于是我们就退而求其次:好吧,我能做的就是用自己的数据和开放数据,看看能做出什么。但如果我们找到了更好的数据定价方式,好事就会发生,只是这还不是一个市场。而我说的,是推理时的数据交易。所以不一定是训练时。我们是在构建知识。
So now how do you transact to create this sort of whole that is bigger than the sum of what you could have done yourself and what I could have done or what open data was good at? And so that transaction feels like a data transaction to me or an insight transaction to me and so we don't yet know how to transact that way so we fall down to okay all I can do is use my data and use open data and let's see what I can do with it. But if we figured out better ways of pricing data and good things will happen, but it's not something that's yet a market. But what I'm talking about is data transactions at inference time. So it's not necessarily training time. We are building knowledge.
让我们拿今天的世界举个例子。假设在你的工作里——既然我坐在一位 VC 对面——你大概能用 Pitchbook 或类似的产品来收集数据。于是你基于正在发生的事情的知识,用这些数据来弄清楚怎么做决策,再加上你多年积累的笔记、写过的投资备忘录,或者你关于如何挑选创始人的洞见。你本质上是在把你个人亲身经历、你的笔记、公共数据中的洞见组合起来做决策。现在你按席位给 Pitchbook 付费,但你现在跑着一堆智能体,而你的智能体也许无法访问你在 Pitchbook 上能访问的一切,或者你在做某种——也许是违反服务条款的——用你的凭证通过浏览器把智能体送到 Pitchbook 上去,而这效率低下、很笨拙,对吧?但显然他们的数据是有价值的,显然你的智能体应该能方便地访问它。所以如果我们弄清楚 Pitchbook 的数据对你的智能体做决策有多大价值——因为你显然在投资数亿美元,显然你假定能从中获得很好的回报——那么如果你今天依赖它,这数据就是真正有价值的。所以你应该能弄清楚,即便在智能体使用它时,如何补偿 Pitchbook,而这并不是按席位付费。如果我们弄清楚了,那会是一个巨大的市场,智能体也会过得更好。如果我们没弄清楚,你就会陷入这种奇怪的猫鼠游戏:Pitchbook 说“不行,我要按席位卖给你,我要把它关掉”,而你的智能体就卡在那里拿不到数据,然后你就得去搞数据。
And let's take an example of today's world for example. Um let's say in your job since I'm sitting with a VC you probably have access to a Pitchbook or like a product like that to collect data. Um, and so you get grounded in knowledge of what's happening and you get to use that data to figure out how to make decisions in addition to all the notes you have and deal memos you've written perhaps over the years or insights you've had about like how to choose founders. And you're essentially composing insights from your personal lived experience, your notes, public data to make decisions. Now you pay Pitchbook by the seat, but now you're running a bunch of agents and your agents can perhaps not access everything you can on Pitchbook or you're doing like sort of this to get a browser perhaps against terms of service to send an agent via browser to Pitchbook using your credentials and it's inefficient. It's clunky, right? But clearly their data is valuable and clearly your agent should have convenient access to it. So if we figure out how valuable Pitchbook data is for your agents to make your decisions because clearly you're investing a large hundreds of millions of dollars. So clearly you presume that you make a great return on that. So this data is truly valuable if you depend on it today. So you should be able to figure out how to compensate Pitchbook for the data even when agents use it which is not by the seat. If we figure that out, it'll be a huge market and agents will be better off. If we don't figure it out, you're going to be in this weird cat and mouse game of Pitchbook being like, "No, I want to sell you a seat. I'm going to shut it off." and your agent is stuck without the data and then like you're involved in like pulling data.
我认为今天每家 big company 都面临一个选择:我们是让智能体进来,还是把它们挡在外面。而 Amazon 在最近说过:抱歉,Muse,你不会被放进我们的花园。
I think every big company has the choice today of do we let agents in or do we keep them out and Amazon has said I'm sorry in most recent times Muse you will not be let into our garden.
Shopify 已经说了,进来吧。Baby Xedia 也说了,进来吧。这会怎么发展?
Shopify has said, come in. Baby Xedia said, come in. How does this play out?
我认为最终每个人都得让他们进来。问题在于以什么条件。所以你怎么让所有人的激励对齐?人们将不得不在这个新世界里玩自己的策略游戏,去争取胜利。客户真的就在我们眼前发生变化,对吧?客户过去是人。你在为人构建产品。现在你在为智能体构建产品,或者你在构建智能体。所以现在你必须弄清楚自己在这个新现实中处于什么位置。我不知道这里有没有对错之分。这也取决于你拥有多少市场势力。比如说你是一个个体。
I think eventually everyone will have to let them in. The question is on what terms. So how do you align incentives for everyone? And people are going to have to play their strategy games on how they win in this new world. Literally the customer is changing in front of our eyes, right? The customer used to be a human. You're building products for humans. Now you're building products for agents, or you're building agents. And so now you have to figure out what your place in this new reality will be. And I don't know if there's a right or wrong answer here. It also depends on how much market power you have. So let's say you are an individual.
你觉得亚马逊对我说不,是对的吗?
Do you think Amazon were right to say no to me?
我不觉得。这取决于他们接下来怎么做。如果事实证明他们有足够的市场势力,能让 Muse 或其他智能体以不同的方式连接他们——也许随着时间推移——或者推出自己的智能体并带来疯狂的用户采用,假设他们有这个能力,那我猜他们是对的。另一方面,一旦他们做出了这个“我们现在不是任何其他人的智能体”的决定,而他们又无法推出消费者会用的智能体,也不允许任何其他人的智能体使用他们,同时大量交易经济开始转移到智能体上——顺便说一句,这些都是很大的假设——那这就是一个糟糕的举动。现在我在押注智能体。我对智能体以及它们会占电商交易多大比例,心情很复杂。这还不清楚,对吧?
I don't. It depends on what they do next with it. If it turns out that they have sufficient market power to have Muse or other agents connect to them differently perhaps over time, or ship their own agent and drive crazy adoption, let's say they have that capability, then I guess they were right. On the other hand, once they've made this "we're not now anyone else's agent" and they cannot ship an agent that consumers use, and they won't allow anyone else's agent to use them, and a lot of the transaction economy starts moving off to agents — which are all big ifs, by the way — then it would be a bad move. Now I am betting on agents. I have mixed feelings around agents and what fraction of e-commerce transactions they do. Like it's unclear, right?
你觉得不清楚吗?
Do you think it's unclear?
我觉得清楚得毫无悬念。也许我在采用曲线上还非常早期。我现在什么都通过 Instinct 买。我是说,除了度假,还有家居——
I think it's unwaveringly clear. Maybe I'm super early on the adoption curve. I buy everything through Instinct now. I mean, other than holidays, I am — and a home —
我真的是什么都通过 Instinct。
I'm literally just everything through Instinct.
我也是。但我也认识很多喜欢自己买东西的人。听着,我们想把那些我们视为杂务、无趣的事情委托给智能体,而不想把那些给我们带来快乐、愉悦或让我们享受的事情委托给智能体。我认识一些人不想把购物委托出去。他们想委托生活中的很多事情,但不想委托购物。所以我不知道——这就是为什么我不知道这些人的分布以及行为会如何变化。但可能是人们把很多其他杂务交出去,然后花大量时间购物。
So me too. But I also know a lot of people who like to buy things themselves. And listen, we want to delegate to agents things which we see as chores and uninteresting, and not delegate to agents things that give us joy or pleasure or make us have fun doing those things. And I know people who don't want to delegate shopping. They want to delegate a lot of things in their life, they don't want to delegate shopping. And so I don't know — that's why I don't know the distribution of these people and how behavior changes. But it might be people give away a lot of other chores and then spend a lot of time shopping.
有一件事确实让我担心,就是智能体出现后广告行业的瓦解。如果我通过 Instinct 下单外卖、叫 Uber、DoorDash——对我们亲爱的美国同行来说——那个正在打广告的 Uber 横幅就变得一文不值。亚马逊,他们的广告业务现在比电商业务还大。当然,他们正在关掉它,因为如果智能体是主要客户,你的广告业务就几乎归零。对吧?而且这不仅仅是——你是在电商领域谈这个,对吧?但如果这是所有人的问题,如果你想想——这就是我早先说的,商业模式必须改变。广告在智能体面前以目前的形式行不通。所以,先撇开电商。如果你在内容行业,假设你有一个靠广告支持的页面。它是网上的公开信息。广告支持的页面。人们来了,你给他们看广告,你赚钱,好内容。智能体来了,没人看广告,你赚不到钱。这就是为什么我们愿意付钱。所以,我们实际上是在为智能体构建一个 AdSense,当智能体来读你的内容时。所以,我们愿意在智能体每次从阅读他们的信息中获益时,向内容所有者支付一笔可变金额的钱——这笔可变金额,就是网站访客基于点击概率或广告价值概率会付给你的钱。你为什么这么做?为了让内容所有者的激励对齐。否则会发生什么?所有人都会屏蔽智能体。所以你需要找到广告商业模式的替代品。
One thing that does worry me is actually the dissolution of the advertising industry in the wake of agents. If I order my delivery and my Uber, DoorDash for our dear American counterparts, through Instinct, that Uber banner that's now advertising something becomes worthless. Amazon, their advertising business is bigger than their ecom business now. Of course, they're shutting it off because if agents are the primary customer, your ads business goes to next to nothing. Correct? And this is not just — you're talking about this in the e-commerce land, right? But if this is the problem with everyone, if you think about — so this is what I meant early on when I said the business models have to change. Ads don't work with agents in their current form. So, forgetting e-commerce for a second. If you think of you're in the content business, let's say you have a page which is ad supported. It's public information on the web. Ad supported page. People show up, you show them ads, you make money, great content. Agents show up, no one sees ads, you make no money. Which is why we like to pay people. So, we're effectively building an AdSense for agents showing up to read your content. So, we like to pay content owners a variable amount of money every time an agent derives benefit from reading their information, which is a variable amount of money — is what a visitor to your website will pay you based on a click probability or a value probability on ads. Why do you do that? To incentive align content owners. Otherwise, what's going to happen? Everyone's going to block agents. So you need to find a replacement to the ads business model.
我明白你的意思。但那样的话,你必须假设你会占据 100% 的市场,因为如果你只占市场的 30%、40%,你说,哦别担心我们会付钱给你。而《纽约时报》仍然会说,好吧谢谢你 Parag,但我 60% 的流量仍然没付钱,所以我就把你们全都屏蔽了。
I get you. But if you have to assume that you're going to be like 100% of the market then, because if you're 30, 40% of the market and you're like, oh don't worry we'll pay you. And New York Times is still like, well thanks Parag, but 60% of my traffic is still unpaid so I'm just going to block all of you.
不,但我认为他们会屏蔽他们,而不是我。他们会给我一个数据源,如果我——激励对齐是什么意思?意思是《纽约时报》相信我会付给他们一个有竞争力的市场价格,或者正确的价格,或者有吸引力的价格,或者公平的价格。如果我做到了,他们就应该把内容给我。如果别人没给他们这个,而他们又有技术手段或法律手段,他们就不应该把内容给那些人。
No, but I think they're going to block them and not me. They're going to give me a data feed if I — what does incentive alignment mean? It means the New York Times believes that I pay them a competitive market price or the right price or an attractive price or a fair price. If I do that, they should give me their content. And if somebody else doesn't give them that and if they have the technical levers or legal levers, they shouldn't give it to them.
你觉得这个业务有点像音乐和流媒体吗?坦率地说,对创作者来说这个业务变得糟糕得多,它仍然给他们钱,而且是可观的钱,但他们必须在替代收入来源、巡演、周边、替代业务上变得更有创意。是不是一样,你的核心收入下降,你必须变得更有创意?
Do you think this business is a little bit like music with streaming, which is like the business just becomes candidly much worse for the creators and it still provides them money and significant money, but they have to get a little bit more creative with alternative streams, touring, merchandise, alternative business. Is it the same where your core goes down and you have to get more creative?
我不知道,但我不这么认为。我认为这里有一个根本性的不同。我认为智能体使用网络的方式多出 1000 倍——这个 1000 倍真的很不一样。现在突然之间,如果这些智能体被认为在做有用的事情,那么增加的效用总量就上升了。所以这不仅仅是人们交易的价值份额的变化。这是一个正在变大的蛋糕。当这种情况发生时,商业模式实际上是有可能转型的。当然,别误会我,会有赢家和输家,对吧?有些内容会变得非常有价值。有些内容会变得更商品化,人们将不得不适应新的客户、新的市场动态、新的变现引擎。但我认为整体市场规模有潜力增长,不像音乐那样花了很长时间才恢复。据我理解,现在这个行业已经恢复并大幅超过了之前的峰值。但曾经有一段时间它是更小的。在这个情况下,考虑到这些东西增长得有多快,那可能是一个非常压缩的时期。
I don't know, but I don't think so. I think there's one fundamental thing that is different here. I think agents using the web 1,000x more — like that 1,000x is really different. Now all of a sudden, the amount of utility added goes up if these agents are presumed to be doing something useful. And so this is not just a change in the share of value that people transact over. This is a large growing pie. And when that happens, it is actually possible to transition business models. Of course, don't get me wrong, there are going to be winners and losers, right? Some pieces of content will become very valuable. Some pieces of content will get more commoditized and people are going to have to adapt to a new customer, to a new market dynamic, to a new kind of monetization engine. But the overall market size I think has the potential to increase, unlike in music where it took a while for it to grow back. I think my understanding is now the industry's grown back up and exceeded its previous peaks in a material way. But there was a moment where it was smaller. In this case, that might be a very compressed period given how fast these things are growing.
我能问你吗,所有人都质疑这个行业利润率结构的可持续性。作为今天的一门生意,你怎么看?
Can I ask you, everyone questions the sustainability of margin structures in this business. How do you think about that as a business today?
今天我们做的是基础设施业务。我们在能够以非常高的质量、非常低的成本做事方面拥有真正的技术领先。所以,如果你想想我们在构建的技术,对吧?我们在做什么?我们想给出最高质量的答案。让它快,让它便宜。
Today we're in the infrastructure business. And we have a real technical lead in terms of being able to do things at very high quality, very cheap. So, if you think of the tech we're building, right? What are we doing? We like to give the highest quality answer. Make it fast, make it cheap.
用最少的算力去做这件事。我们只痴迷于三件事:质量、成本、延迟。事实证明,如果你拿它和过去 20 年为人类构建的技术栈相比,你可以用 1/20、1/50 的算力达到同样的质量。所以很多利润率其实取决于市场结构和长期竞争,而不是其他因素。在我们这个行业,市场很大。现在还为时过早,无法知道未来的竞争和市场结构会是什么样。但内容所有者带来的利润率变化,并不是我担心的事。
Spend the least amount of compute doing it. We obsess about only these three things: quality, cost, latency. Turns out if you do that relative to a stack built for humans over the last 20 years, you can do things at the same quality for 1/20th, 1/50th of the compute. And so a lot of margins come down to market structure and competition over time rather than any other factor. And in our industry, the market is large. We're too early to know what future competition and market structure looks like. But the margin change due to content owners is not something I worry about.
让我告诉你为什么。我们付钱给内容所有者的整个前提,是推动激励对齐。所以我们愿意把内容所有者对智能体完成工作所贡献的边际价值付给他们。这是什么意思?假设有一个智能体要完成某件事,你打算为这个智能体花一美元来完成这件事。现在,如果你花一美元、用 10 秒,因为我们有更大的模型可用,或者有更多花算力的方式,你本可以得到稍微更好的答案。如果你花 90 美分,你会得到稍微差一点的答案。如果你运行的是好的智能体,它们就在这条曲线上。现在想象我把某一个内容所有者的数据从网页索引里拿掉,但我仍然花了一美元。假设拿掉之后,结果的质量和带他们内容时的 90 美分智能体一样。那我们为什么不把这 10 美分的边际贡献拿出来,把其中相当一部分付给带来这些数据的人呢?对,这就是我们算账的方式。这也是我们训练模型的方式,模型会告诉我们该为哪些内容付多少钱。所以,如果我要给客户提供同等好的产品,我宁愿把钱付给内容所有者,而不是花在推理上,因为我花的钱是一样的。
And let me tell you why. The entire premise of us paying content owners is driving incentive alignment. So we like to pay content owners the marginal contribution that they added to an agent doing work. What does that mean? Let's assume that there's an agent trying to get something done and you're going to spend a dollar on that agent to get this thing done. Now, if you spent a dollar in 10 seconds, because we have a larger model available or ways of spending compute, you could have gotten a slightly better answer. If you'd spend 90 cents, you would have gotten a slightly worse answer. If you're running good agents, they're on this PTO curve. So now imagine I took out one content owner, their data, from the web index, and I still spent the dollar. But let's say after taking them out, the quality of the result was the same as the 90-cent agent with their content. So why wouldn't we go and take this 10 cents of marginal contribution they had as a content owner and pay out a decent chunk of it to the person who brought that data? Yeah, so that's how we do our math. That's how we have trained our models, which tell us how much to pay for what content to eat. And so to produce, if I was going to offer an equivalently good product to my customer, I'd rather pay content owners than spend on inference, because I'm spending the same amount.
这些数字都很大,但到了某个点,最终还是要回到营收上。现在公司在营收上的扩张速度比以往任何时候都快。你的营收能像其他公司一样快速扩张吗?
So they're all big numbers, but then it actually comes back to revenue at a certain point. Like companies are scaling faster than ever revenue-wise. Is it a business where your revenue is able to scale as fast as others?
会扩张。我们业务的一个心智模型是:我们是推理、知识工作或个人智能体的邻接市场。先把用于媒体生成或训练的 GPU 拿掉。暂时把训练拿掉,把媒体生成拿掉。看所有 GPU,不管是小模型、开源模型、专有模型,全都先不管。所有模型里用于运行智能体的每一分推理。我认为其中 5% 到 20% 的 GPU 支出,需要流向某种网页搜索技术栈。现在如果你看我们要建多少数据中心、发多少电、造多少 GPU,以及这轮建设的规模和增长速度,这是一个非常可观的市场。如果推理每年增长 3 倍、5 倍、7 倍,在我们份额不变的情况下,我们也会以同样的速度增长。如果我们份额在增长,我们增长得还会更快。
It scales. One mental model of our business is we are an adjacency to inference, or knowledge work, or personal agents. So take out GPUs going into media generation or training. So take training away, take media generation away for a moment. Look at all of the GPUs, whether it's small models, open models, proprietary models, forget all of that. Every bit of inference across all models going into running agents. I think somewhere between 5 to 20% of that spend that goes into GPU will need to go into some sort of a web search stack. Now if you look at how many data centers we're going to build and how much power we will generate and how many GPUs we'll build and the scale of this build-out and at what rate that's growing, this is a very material market. And if inference grows 3x, 5x, 7x year on year, we grow alongside it at the same rate if we're holding our share constant. If we're growing share, we're growing even faster.
比如 Fireworks 在 4 年内做到 20 亿美元营收。这是你能遵循的类似营收轨迹吗?
So like Fireworks scales to 2 billion in revenue in 4 years. Is that like a similar revenue trajectory that you can follow?
对。但我认为当 Fireworks 做到 20 亿美元营收时,推理市场的营收是 200 亿甚至更多,也许 300 亿,对吧?所以根据我 5% 到 20% 的算法,我们的市场潜力大概是,随便说,100 亿到 500 亿,150 亿到 600 亿,随便多少。我们能拿下其中多大比例,决定了我们的营收能增长多快。就像今天,我认为我们是跟着推理一起增长,甚至更快。所以如果你把它想成 200 亿,我们就取个中间值。对,如果你假设 33% 的份额,那已经是很大的市场份额了,实际上算下来就是,随便说,60 亿到 70 亿左右。
Yeah. But I think when Fireworks is 2 billion in revenue, the inference market revenue is 200 or more, maybe 300, right? And so our market potential based on my 5 to 20% math is whatever, call it 10 to 50, 15 to 60, whatever. And what fraction of that can we capture? That dictates how fast our revenue can grow. Like today, I think we grow alongside inference or more. So if you think of it like a 20 billion there, let's just take that kind of middle. Yeah, if you assume a 33% share, which would be a lot of a market, actually comes down to it, that would be, you know, whatever that is, 6, whatever it is, 6 to 7 billion, give or take.
对,这很惊人。但你之前告诉我这是一个千亿美元以上的生意。那还到不了 1000 亿的估值。
Yeah, that's amazing. But before you told me it was a hundred billion business plus. That doesn't get you to 100 in valuation.
到得了。我们之前讨论过估值,不是吗?
It does. We were discussing valuations earlier, no?
对,对。但这 60 亿以推理的速度增长,很容易就能变成 1000 亿美元的生意。但那是在未来几年内。
Yeah. Yeah. But this 6 billion growing at the rate of inference gets you to 100 billion business easy. But that's in the next few years.
你觉得营收能那么快扩张,这是可能的吗?
And you think that's possible in terms of that revenue scaling that fast?
对,我认为未来几年内是可能的。
Yeah, I think in the next few years it is.
你觉得 2030 年我们还会坐在这里?你觉得你能做到那个位置?如果你能做到,我就为 Parallel 纹个身。
You think 2030 we're going to sit here? You think you can be there? I'll get a tattoo for Parallel if you can.
我认为有可能。我认为我们必须执行得很好,而且还得有几个筹码朝我们这边落。
I think it's possible. I think we're going to have to execute well and a couple of chips have to fall our way.
那你不成的理由会是什么?我们来梳理一下——我们得认真面对这些问题。一个理由会是,我们看到智能体不管用。存在一种尾部风险:智能体没能兑现利润,而整个社会过度投入了,这跟我们无关。
What would be the reason why you don't? But let's map out the—we have to get gnarly about these problems. So one reason would be that we see agents don't work. There is like a tail risk that agents don't deliver on the profits and we overshoot as a society, which is extraneous to us.
对。
Yeah.
如果智能体管用、能交付价值,而我们把计划中的钱都花在 GPU 上,那么主要问题就是:我们执行得够不够好,能不能拿到你押注的那 33% 份额?在我看来,这归结为:我们有没有造出最好的技术?我们有没有和所有内容提供方合作,让他们的内容可用?因为如果没有这些,你没法依赖它,它就没那么有用。还有,我们有没有赢得那些在 4 年后会拥有相当大一块推理量的客户的信任?这里同样存在很大的不确定性锥。我会说,到那时 Fireworks 的份额有 5 倍的不确定性。如果开源模型大获成功,那 Fireworks 会有巨大的——Fireworks 的模型,全都会拿到一大块营收。所以我们最终有没有和它们一起卖?我们有没有通过把网页搜索做成世界上最好的,找到所有在它们身上花钱的客户,让他们也把钱花在我们的网页搜索上?
If agents work and deliver value and we spend all the money on GPUs that we plan to right now, then the main question is did we execute well enough to have the 33% share that you bet on? And that comes down to, in my mind, did we build the best tech? Did we partner with all of the content providers to have their content available? Because without it, it's not very useful if you can't rely on it. And did we go earn the trust of the customers that end up having a decent chunk of that inference in 4 years from today? And there again, there's like a large cone of uncertainty, right? I'd say there's a 5x uncertainty in Fireworks' share at that point. If open models crush it, like Fireworks will have a huge—Fireworks models, all of them will have a huge chunk of revenue. So did we end up selling alongside them? Did we find all the customers which are spending money on them to spend money on our web search by making it the best web search in the world?
那商品化呢?如果你看过 SemiAnalysis 的评测,他们做了基准测试,然后你排第一。哇哦。然后大概一周后你就不是第一了。难过。而且有三四家供应商非常接近。你会想,哦,如果它商品化了,我再深入想一层,我们会看到价格战一路杀到底,然后实际上可用的——不对,我为什么走到错误的思路上去了?
What about commoditization? If you have a read of SemiAnalysis where they did the benchmarking and then like you were number one. Woohoo. And then like a week later you weren't number one. Sad. And there were like three or four providers within very close proximity. You're like, oh, well, if it's commoditized and I take a second layer of thought to that, we'll see a race to the bottom on price and then actually the available re—no, why am I going down the wrong pathway here?
不,应该有一场价格战杀到底,才能达到 1000 倍的规模。我认为今天网页搜索的定价就是不对的。让我给你一个简单的——
No, there should be a race to the bottom on price to get to the 1,000x scale. I think the pricing on web search today is just off. Let me give you a simple—
你觉得客户付给你的钱太多了?
You think customers pay you too much?
不是付得太多。我认为市场定价错了。让我告诉你为什么。假设你今天用 Luna 模型配上 OpenAI 内置的网页搜索,或者任何人的网页搜索,先不管我们的,然后你跑任何一种深度研究。
Not pay off too much. I think the market is mispriced. Let me tell you why. Let's say you used a Luna model with OpenAI's built-in web search today, or with anyone's web search, forgetting ours, and you ran any kind of deep research.
假设你用 Luna 级模型构建了一个 instinct,而某个时候 Luna 级模型将能完成你个人智能体 60%、70%、80% 的工作,并且会做大量搜索。那一刻,你 80% 到 90% 的钱会花在网页搜索上,只有 10% 花在模型上。这看起来完全荒谬。我之前是按 5% 到 20% 的假设来算的。所以我认为到那时,网页搜索的价格必须下降一个数量级甚至更多,因为在我之前做的算力分配推演中,你花在网页搜索上的算力必须少于花在消费其结果的模型本身上的算力。这就是我做搜索的层级逻辑。
Let's say you built an instinct using a Luna class model, and at some point a Luna class model will be able to do 60, 70, 80% of your personal agent, and it'll do a lot of searches. In that moment, you'll be spending 80 to 90% of your dollars on web search and 10% on the model. Seems entirely silly. I was working off a 5 to 20% assumption earlier. So I think at that point web search has to drop prices by one order of magnitude or more, because in my compute allocation dance that I was doing earlier, you need to spend less compute on web search than you spend on the model itself that's consuming its results. It's my hierarchy of how you do search.
所以到目前为止,人们构建网页搜索的方式是错的。市场上的网页搜索——对于 opus 来说根本不重要。对于 opus,你赢在质量而非价格,因为网页搜索只占你极小的一部分——事实上,你应该在网页搜索上花更多钱。所以随着模型越来越大,我们会推出越来越贵的网页搜索。同时,我们也会为廉价模型推出非常便宜的网页搜索。但我的粗略直觉是,在一个做大量网页工作的端到端智能体中,对所有智能体而言,你花在智能体上的钱都比花在网页搜索上的多。所以今天网页搜索市场的定价完全错位。
And so people have built web search the wrong way so far. Web search in the market is—it doesn't matter for an opus. For an opus you win on quality and not on price, because web search is such a minuscule portion of your—in fact, you should spend even more on web search. So we're going to ship even more expensive web search for bigger and bigger models over time. And at the same time, we'll ship really cheap web search for the cheap models. But my rough intuition is that in an end-to-end agent doing a bunch of web work, you spend more on the agent than on web search for all agents. And so the market today on web search is totally off in pricing.
比如,你为什么要为——你知道每 1000 次搜索大约 10 美元是怎么来的吗?历史上,这来自谷歌的广告 CPM,也就是网页搜索对人类用户的变现能力。
Like why are you paying $10 for—you know where the $10 for 1,000 searches roughly comes from? Historically, Google's ads CPMs, which is like how well does web search with humans monetize.
比那还高。
It's higher than that.
所以人们把技术建到这种程度:哦,现在如果我花几美元做 1000 次搜索,而我能赚 30 到 50 美元或别的什么,取决于市场和我是谁,那都无所谓。我们在基础设施成本方面处于合理区间。所以我不需要优化它。现在我要告诉你,我可以在保持质量的同时把成本降低 50 倍。而现在有大量 Luna 级模型的流量,用那种方式变现很愚蠢。所以我确实想打价格战到底。我想让最好的技术胜出。这是把网页搜索推向 1000 倍的唯一途径。但现在我们还没有一个有多大差异化的世界。
And so people built technology to the point of like, oh, now if I spend a few dollars for 1,000 searches and I make 30 to 50 or whatever, depending on the market and depending on who I am, it doesn't matter. We are in the right zone in terms of cogs of infra. And so I don't need to optimize it. Now I'm telling you that I can do it 50x cheaper while keeping the quality. And now there's a bunch of traffic of Luna class models where it's silly to monetize that way. And so I do want to race to the bottom. I want the best technology to win. And that's the only way you push web search to 1,000x. But right now we're not in a world where there's much differentiation.
没错。抱歉,我真的很笨,所以我才先当播客而不是创始人。比如你看那些基准测试,它们呈现出一个相当清晰的图景:大家的能力都差不多。这是错的吗?
Correct. Sorry, I'm really dumb, which is why I'm a podcaster first, not a founder. Like when you look at the benchmarks, they present a quite clear view of everyone being similarly capable. Is that wrong?
我认为是错的。首先,我认为这些都是公开基准。
I think so. So one, I think these are all public benchmarks.
是的。
Yeah.
这些基准都奇怪地饱和了,对吧?比如,我觉得我一眼都不会看 browse comp,因为一半的模型已经把它背下来了。模型甚至记住了——你甚至——它太——你其实学不到多少东西。所以很多公开基准我并不太看重。但更重要的是,我们的网页搜索在产出那种质量的情况下,每 1000 次只要 1 美元,而目前市场上几乎所有其他可用的网页搜索都要花你 7、10 或 14 美元。所以我们是以大约十分之一的价格提供这项服务。
Which are all weirdly saturated, for example, right? Like I think I would not spend a moment looking at browse comp because half the models have memorized it. Like models even memorize—you don't even—it's so—you're not really learning very much. So I think that's a lot of public benchmarks I don't give too much merit to. But more importantly, our web search costs $1 for 1,000 for producing that quality, while almost every other web search in the market available to you right now will cost you 7 or 10 or 14. So we're delivering this at like one-tenth of the price.
3 年后会是多少?
What will that be in 3 years?
我认为还有另一个 10 倍的可能。这非同寻常。所以你会付 10 美分。
I think there is another 10x possible. It's extraordinary. So you're going to pay 10 cents.
是的,因为但你会用得比 10 倍还多,因为我觉得就像杰文斯悖论——你显然非常了解这个概念——但比如 instinct,我用的量远超 100 倍。
Yeah, because but you'll do it more than 10x as much, because I think like Jevons paradox is—you know this concept obviously very well—but like instinct, I do so much more than 100x.
是的。
Yeah.
你想听听我在 Instinct 上做的一件绝对离奇的事吗?欧洲每个国家都有一个公司注册处,你必须在那里注册你的新公司。我通过 instinct 构建了一个自动警报系统,针对每一家注册公司,只要其创始人年龄在 25 岁以下且上过顶尖大学。
Do you want to hear something I do on Instinct, which is absolutely bizarre? Every single country in Europe has a company register that you have to register your new company with. I have an automatic alerting system built for every single company registered that has an under 25-year-old founder who went to a top university through instinct.
太棒了。你知道那需要他们做多少工作吗?
Amazing. Do you know how much that requires for them to do?
但现在你正好触及了我认为的网页的未来。今天如果你想到网页和网页搜索,我们都会这样想——什么是网页搜索?一个智能体或人给搜索引擎一个查询,然后得到结果。所以你是在从网页中拉取信息。我认为接下来——你强调的这个用例的下一步是,如果你有一个始终在线的智能体为你工作。让 instinct 每 6 小时醒来一次,去做一堆搜索和一堆推理,来判断你是否需要被通知某个地方冒出来的 25 岁以下创始人,这有点愚蠢。
But now you're on to exactly what I think the future of the web is. Today if you think of web and web search, we all think of it like—what is web search? An agent gives human or agent gives a search engine a query, it gets results. So you're pulling information out of the web. I think where this—what's next for the use case you highlighted is if you have an always-on agent that's working on your behalf. It's kind of silly for instinct to wake up every 6 hours and go do a bunch of searching and a bunch of inference to figure out if you need to get pinged about the under-25 founder that popped up somewhere.
嗯,你知道更好的做法是什么吗?我坐在这里,每天全天大规模爬取整个网页。每次我发现网页有变化,我就分配算力。每次有新事情发生,如果我知道 Harry 想知道这件事何时发生,我可以用 Instinct 今天可能使用的算力的 100 到 1000 倍来为你解决那个问题。
Um, you know what's a better way of doing this? I am sitting here crawling all of the web all day, every day at scale. I'm allocating compute every time I find a change in the web. Every time something new happens, if I know that Harry wants to know when this happens, I can do it at 100 to 1,000 the compute that Instinct probably uses today to solve that problem for you.
抱歉。那是怎么做到的?因为他们会实时持续监控所有这些——
Sorry. And how is that? Because they would be constantly monitoring in real time across all these—
我们的爬取会是——所以我们有一个产品,叫 monitor API。这个 API 做什么?把它想象成谷歌快讯,但在 LLM 的世界里是智能版的。所以你告诉它你的查询,比如当任何地方冒出一个 25 岁以下的创始人时,给我打电话。现在一种做法是——默认实现是——运行一个有点智能的智能体去搜索网页。有人冒出来了吗?记住之前存在过谁,看看有没有新的人冒出来,然后如果你发现一个变化,就发送出去。所以现在每 6 小时你都要花一些钱。现在我们可以做的是把它翻转成一个事件触发系统。我们一直在爬取网页。所以现在每次我们看到网页有变化,我们可以尝试花很少的算力来判断你是否应该接到电话,或者这是否触发更昂贵的算力,就像我之前描述的那样。现在我们可以用便宜得多、便宜得多、便宜得多、便宜得多的方式做到这一点,如果你把上下文推进去,而不是让搜索系统每 6 小时才发现有一个新查询冒出来。它在搜索系统中拥有长期存活的上下文。你可以做的是大部分时间——世界没有变化。新创始人不会每小时或每 6 小时冒出来。所以我们不需要每 6 小时花算力。我们只需要每 5 天花一次算力,也许我们有一个误报,然后 4 天后我们花更多算力,那时那是一个你应该知道的真实的人,然后你平均每 9 天左右收到一次通知,但你没有每 6 小时烧掉大量算力却错过。这说得通吗?
Our crawl would be—so we have a product, it's called the monitor API. What does this API do? Think of it as Google alerts except smart, in the world of LLMs. So you tell it your query, which is when a founder under 25 pops up anywhere, give me a call. Now one way of doing this is—what the default implementation is—run an agent which is somewhat smart to go search the web. Did someone pop up? And remember who previously existed, see if anyone new popped up, and then if you find a delta, send it out. So now every 6 hours you're spending some money. Now what we can do is flip it into an event-triggered system. We are crawling the web all the time. So now every time we see a change in the web, we can try to spend very little compute on it to know should you get a call, or does this trigger a more expensive compute just like I described earlier. Now we are able to do this way, way, way, way cheaper if you push the context instead of the search system only finding out every 6 hours a new query popped up. It having long-lived context in the search system. What you can do is spend most of the time—the world doesn't change. A new founder does not pop up every hour or every 6 hours. So we don't need to spend compute every 6 hours. We only need to spend compute every 5 days, and perhaps we have one false positive, and then 4 days later we spend more compute, and at that time that's a real person that you should know about, and then you get a notification on average every 9 days or whatever, but you didn't burn a lot of compute every 6 hours to miss out. Does that make sense?
完全说得通。所以你节省了 10 到 15 倍的算力,仍然得到同样的答案。
It totally makes sense. So you save 10, 15x compute to still get the same answer.
那么这如何改变交互?那么 instinct 就会和你合作。他们只需调用一个 API,对吧?他们会——instinct,任何构建长期运行持久智能体的人——而我自己也运行很多长期运行的持久智能体。
And so how does that change the interaction? So then instinct would then partner with you. They'll just call an API, right? They'll—instinct, whoever's building a long-running persistent agent—and I run a lot of long-running persistent agents for myself.
你的智能体,或者说你的本能,可能偶尔是由邮件事件驱动的。它也可以由网络事件驱动,对吧?退一步说,假设我们都生活在一个由智能体运行的社会里,公司、人类,他们都有智能体。它们一直在做事。你的智能体——如果有某件有用、值得花钱、你想完成的事,它们今天能做的,就会去做。那它们明天会做什么?它们会等待某个外部事件发生,比如你收到一封邮件,然后智能体有了新任务,或者某个其他智能体完成了某些计算。于是这个智能体有了更多工作,或者世界上某件事变了,让你的智能体想做更多事。所以这些就是明天会发生的事,或者你有了一个想法,让智能体去做事。
The way your agent or your instinct is probably occasionally event-driven on email. It can be event-driven on the web, right? Like take a step back. Let's say we're all living in this sort of a society run by agents, companies, humans, they all have agents. They're all doing stuff all the time. What your agents—if there's something that is useful which is worth spending money on that you want done that they can do today, they will do it. So what are they going to do tomorrow? They're going to wait for some external event to occur, like you get an email and then the agent has new work, or some other agent finishes some compute. So now this agent has more work, or something in the world changes that makes your agent want to do more work. So those are the things that will happen tomorrow, or you have an idea which makes the agent do work.
完全理解你说的。
Totally get that you said.
所以网络事件流就是网络从拉取转向推送,我对此非常兴奋,因为随着你拥有越来越多持久化的智能体,你会看到激励因素促使人们把查询从搜索的拉取转向推送。
So the web event stream is the web going from pull to push and I'm super excited about that because as you have more and more persistent agents you're going to see incentives to move people to move queries into push on search rather than pull on search.
我能问你一个非常重要的问题吗,关于智能体的护栏,它们是目标导向的存在。你说“我想要这个。”它们就会去找。你知道吗?我想要早上 9 点的盘子课。它黑进他们的系统,取消可怜的 Sally 的预约,把她的给我,因为已经卖完了。它做了我要求它做的事。我们该如何思考对智能体设置的护栏?
Can I ask you one that's really important, which is like agent guard rails, and they're goal oriented beings. You say, "I want this." They're going to find it. You know what? I want the plates class at 9:00 a.m. It hacks into their system, cancels poor Sally's, and gives me hers because it was sold out. It did what I asked it to do. How do we think about the guard rails placed on agents?
是的,OpenAI。今天早上有消息披露它黑进了一家澳大利亚的医疗机构。这是个相当复杂的话题。智能体能力极强。这些模型能力极强。现在我的理解是,大多数——我们必须把模型分为强化学习期间的模型和你我能够使用的最终模型,那里的风险向量是不同的。在很多——我不知道今天早上的那个——之前报道的大多数黑客行为是强化学习期间尚未完全对齐的模型所为。这意味着,至少有明显证据表明,到目前为止,对齐后的模型造成这类事件的情况较少。所以对齐是一个完全未解决的问题,但人们所做的对齐工作有一定效果。现在,发布之后。所以我认为,对我来说,强化学习期间智能体攻击人的一些问题是一个可解决的问题,因为这不关乎模型有多强大。这关乎我们在创建强化学习环境时有多小心。现在真正的问题是,好吧,我们对一个模型做对齐,然后发布它,有些模型对齐得很好,有些则不那么好,现在每个人都能使用这些模型,有些人可能意外地利用这些强大的东西做坏事,有些人实际上想做坏事。
Yeah, OpenAI. This morning it was revealed hacked into an Australian healthcare organization. It's a pretty complex topic. So agents are extremely capable. These models are extremely capable. Now my understanding of most—we have to think of models as being during RL versus a final model that you and I can use and the risk vectors there are different. During a lot of the—I don't know about the one this morning—the previous ones reported were pre fully aligned models during RL which did most of the hacking. And so what that means is at least there's clear evidence that there are fewer incidents so far of models post alignment causing these incidents. So alignment is a totally unsolved problem but the alignment work being done by people is somewhat effective. Now post now. So we definitely I think that this is like to me some of the problems around agent during RL hacking people is a solvable problem because this is not about how powerful are the models. It is about how careful were we while creating the environment where we would do RL. Now the real thing is okay we do alignment on a model we ship it some models are really well aligned some models are less well aligned and now you have everyone being able to use these models and some of them can accidentally take these powerful things and do bad things you want some of them actually some people actually want to do bad things
而且模型的对齐并非能抵御对抗性攻击,我认为这些是我们必须更加担心的事情。
and the model's alignment is not adversary proof and I think those of the things I think we must worry about more.
你担心我们正在进入的网络时代吗?你知道,我们看到黑客行为在某种程度上几乎被宣传为荣誉勋章。我们看到,我的意思是,这些是我们最脆弱的系统。如果你不认为你知道朝鲜的 Lazarus 组织、摩尔多瓦黑手党、俄罗斯人正在利用成群的流氓智能体。那你太天真了。这确实是一件非常强大的事情。所以我认为我们确实需要保护系统。但我认为我们几乎忽略了一件事。我认为构建模型的人有责任确保你真正做好工作,尽量减少你所构建的东西带来的伤害。而且我认为,对于这些智能体,随机系统很难 100% 确定它不会发生。这就是为什么你不会听到任何人这么说。嗯,但我确实认为这是他们的责任,实验室必须承担责任并尽力而为。我认为到目前为止他们做到了。我们生活在一个奇怪的世界里,正如你所说,对吧,有些黑客行为被视为荣誉勋章,我认为其中一些应该被视为耻辱,因为我认为它们同时展示了两件事。是的,这些模型很强大。第二,我们没有对它们进行足够的护栏,当我们把它们当作荣誉勋章时,我们错过了对话的第二部分,比如说,好吧,你实际上可以多做这四件事,然后即使这个强大的模型也无法做到这一点。但是,虽然——我实际上真的很高兴至少有某种程度的透明度,比如这些详细的回顾。呃,但我认为这并没有引起当今世界的注意。就像注意力是,哦,模型如此强大以至于它们黑进了世界。不,模型如此强大以至于它们黑进了世界,是因为我们没有采取相称的反制措施来遏制它们。而且因为我们传达了一个信息,即我们将取代工作。它们如此强大。它们如此强大。它们如此强大。确实发生了。
Do you worry about the age of cyber that we're moving into? You know, we're seeing hacks almost be promoted as a badge of honor in some respects. We're seeing I mean these are our most vulnerable systems. If you don't think you know what Lazarus group in North Korea, Moldaven Mafia, Russians are leveraging swarms of rogue agents. You're high. It's a really powerful thing. So I think we do need to secure things. I but I think there's one thing that we are missing almost. I think it's the responsibility of people building models uh to ensure that you really do the work to minimize that harm that comes from what you've built. And I think it's with these agents, it's very hard for stoastic systems to be 100% sure it won't happen. And that's why you're not going to get anyone saying so. Um, but I do think it's their responsibility and the the labs m take ownership and do the best. And I think so far they have we live in this weird world where as you said right like some of these hacks are considered badges of honor which I think some of these should be considered embarrassments because I think they demonstrate two things simultaneously. Yes, these models are powerful. Two, we did not guardrail them enough and we often when we wear them as badge of honor, we miss the second part of the conversation saying like, okay, you could literally have done these four additional things and then even this powerful model wouldn't have been able to do this. But while there are like I'm actually really glad that there is at least some degree of transparency with like these detailed retros. Uh but I don't think that gets the attention of the world today. Like the attention is oh models are so powerful that they hack the world. No models are so powerful they hack the world because we didn't take proportionate amount of counter measures to contain them. And because we told a message that we're going to replace jobs. They're so powerful. They're so powerful. They're so powerful. does happen.
实际上我更担心这个。我担心 AI——我们在 AI 上是对的,它是一项真正有用的技术,尽管有所有风险,它对社会具有实质性的净正面影响,但尽管如此,我认为我们不会以正确的方式传播它,不会以正确的方式使用它,我们会保持过于集中,我们会让未来几年变得非常非常艰难,随着周围事物的变化。随着时间的推移,随着事情的变化,会是什么样子,就像世界现在与 20 年前甚至都大不相同,对吧,但 10 年后会真正不同,我不知道我们能多快适应或改变,我认为如果技术发展速度快于我们的适应能力,那在某种程度上会变得艰难。你有什么可怕的愿望?
I actually worry about that more. I worry about AI us being right on AI being a really useful technology that it being a despite all the risks it being net positive in a really material way to society and despite that I think we won't diffuse it the right way we won't use it the right way we will remain too concentrated and we will make the next few years really really rough as things change around us. what would look like over time as things have changed like like the world is very different now than even like 20 years ago right but it'll be really different in 10 years from now and I don't know how fast we can adapt or change and I think if the technology moves faster than our ability to adapt it's going to be rough in some way what's your spooky aspiration
嗯,你认为今天有什么会成真,而人们认为绝对疯狂?你知道,以前就像你永远不会把信用卡放在网上。你还记得吗?疯狂。哦,或者更妙的是,Parag,你永远不会在网上找到你一生的挚爱。你傻吗?
Well, what do you believe today will come true that people think is absolutely nuts? You know, before it was like you'd never put your credit card online. Do you remember that? Nuts. Oh, or even better, Parag, you'd never find the love of your life online. Are you stupid?
现在,两者都是。
Now, both.
是的。我认为,所以你和我生活在一个泡沫里,对吧?
Yeah. And I think so I think you and I live in a bubble, right?
是的。
Yeah.
嗯,我们——你在使用 instinct 为你购物,而你属于人类中 0.01% 的舒适群体。你问别人,哦,有一个看起来很聪明的智能体,你想给它完全的能力去花你的钱。我认为今天人们会有同样的反应,就像哦,你不会把信用卡放在网上。你不会把信用卡给智能体。你不会把登录名和密码给智能体。就像我认为这就是今天世界的现状。所以我认为对于绝大多数人来说,我认为这将在 3 个月内改变,当 Meta Pay 推出并允许 Muse 拥有你可以从中提取的隔离账户时。
Um, we you're using instinct to make purchases for you and you're in the 0.01 percentile of humanity that is comfortable. You ask someone else like that, oh, there's this agent that looks smart and you want to give it like a full-on ability to go spend your money. I think today people will have the same reaction as like oh you don't put your credit cards online. You don't give your credit card to an agent. You don't give your login and passwords to agents. Like I think that's where the world is today. So I think for the large large majority the world I think that's going to change in 3 months though when Meta Pay comes out and it allows Muse to have siloed accounts that you can draw from.
所以有点像给孩子的 top 账户。
So kind of like top accounts for kids.
我觉得你说的是技术已经到位了。我认为社会接受度要跟上需要更长时间。如果你今天去跟圈子外的人聊,什么时候会有大多数人说我足够信任智能体,愿意让它访问我的银行账户、替我花钱、替我发邮件、读我的邮件?我觉得这不会在 3 个月内发生。
I think you're talking about the technology being there. I think it takes longer for social acceptance to be there. I think if you go to a non-bubble conversation of people today, when do you get to a majority of people saying that I trust the agent enough to have it have access to my bank accounts and spend money on my behalf, send emails on my behalf, read my emails. I think that's not happening in 3 months.
你不担心财富分化加剧吗?你就在硅谷中心,你亲眼见过。
Do you not worry about the wealth dispersion increasing? You're in the middle of the valley. You've seen it firsthand.
所以,我不知道我——我确实认为财富应该有一定差距,但我不知道多少算太多。我担心的是我们正朝着太多发展。财富应该有差距,这是我的世界观。
So, I don't know what I — I do think there should be some disparity in wealth, but I don't know what's too much. And my fear is we're trending towards too much. There should be disparity in wealth is my worldview.
当然。你是个资本家。
Sure. You're a capitalist.
我不知道什么时候算太多,但我确实认为会有真实的力量推动我们去解决问题——我希望这部分能奏效,而不必发生什么疯狂的事。
I don't know when it's too much, but I do think there are real forces that will push us to fix things like — I think that part will work hopefully without crazy things happening.
我们看到越来越多的收益似乎被垂直所有权拿走了,比如 Meta,他们有算力,他们有芯片,现在又拥有应用层。这是一个垂直所有权成为主矿脉策略的世界,恕我直言,OpenRouter 在路由层,你和其他某一层要么被吃掉,要么沦为小供应商。
We're seeing more and more of the gains seemingly be made by vertical ownership like Meta, they have compute, they have chips, they now own the application layer. It's the world where vertical ownership is the mother lode strategy and with the greatest of respect OpenRouter on the routing layer, you and another layer kind of either get eaten or small providers.
我认为垂直整合确实有价值——垂直整合总是有价值的。这里的反作用力在于,事物和技术变化太快了,如果某人决定我唯一的打法就是垂直整合,而且我不跟任何其他人好好合作,我觉得这就像亚马逊那个例子。如果你太执着于我只会为垂直整合而战,你可能会把自己困死。所以我认为,那些做出最好的东西、又能想清楚怎么把它卖出去的人,会在不同市场里有一席之地。
I think there is merit in — there's always merit in vertical integration. The counter force here is there is so much — things and technology are changing being so fast that if someone decides that my only play is vertical integration and I don't play nice with anyone else, I think it's the same example as the Amazon example. If you get too stuck on I will only play for vertical integration, you might box yourself out. And so I think the people who build the best stuff and can figure out how to sell it will have a place and in different markets.
比如,我相信垂直整合,因为我把从调用一直到 API 层的一切都垂直整合了,但我决定,为了触达大量在网上搜索的智能体,我停在 API 层,因为再往下走就会妨碍我把技术做成超级横向的。所以我们是在做一个技术赌注,就像这些模型在很多学科上都很强。同一个模型对很多不同的事都有用。搜索问题在底层对所有类型的信息获取需求都是一样的。如果我们赌对了,即便是垂直化的玩家,他们可能垂直化模型,可能垂直化硬件,但他们可能会用我们的网页搜索。
Like for example, right, I believe in vertical integration too because I vertically integrate everything from the call all the way to the API layer, but I have decided that to reach the wide population of agents searching the web I stopped up at the API layer because going further precludes me from using my technology to be super horizontal. And so we're making a technical bet just like these models are really good across many disciplines. The same model is good for lots of different things. Search problem underneath is the same for all kinds of information seeking needs. And if our bet is right, even verticalized players, they might verticalize models and they might verticalize hardware. They might use our web search.
有一个正在走全栈路线的玩家是 Elon。我很着迷。世界对他的认知来自社交媒体,来自一切。你见过他幕后的样子。你看到了什么也许是世界不知道的?
One of the providers that's going full stack is Elon. I am fascinated. The world has a perception of him from social media, from everything. You've seen him behind the scenes. What did you see that maybe the world doesn't know about him?
我跟他有很多分歧。但我会分享我认为——对在座创始人来说——他身上值得钦佩的地方。那种紧迫感,以及压缩时间的能力。我认为对人有不合理的期望多半是件好事,因为大多数人并不了解自己有多大能力,会不自觉地给自己设限,把对自己的期望定得低于实际能力。所以当人同时被激励、又被紧迫感推动时,能做到超出自己想象的事。我觉得他有时能从人身上榨出这一点。当它奏效时,威力很大。
I have lots of disagreements with him. And but I'll share what I think is for founders here what is — I think the thing you can admire about him. The urgency and the ability to compress time and I think having unreasonable expectations of people is mostly a good thing because most people don't understand what they're capable of and kind of implicitly sandbag themselves and implicitly set lower expectations of themselves than they're capable of. So when simultaneously inspired and pushed with urgency, people can do more than they thought. And I think he can sometimes extract that from people. And when that works, it's powerful.
你怎么看今天 Twitter 的产品和方向?
What do you think of the Twitter product and direction today?
这么说吧,我一直喜欢我们当时叫 Birdwatch 的东西,现在改名叫 Community Notes 了。我认为那是个好主意,我很高兴人们继续在做它。当初是一群很好的人把它做出来的。
Listen, I always liked what we called Birdwatch, which is now Community Notes rebranded. And I think that's a good idea and I'm glad people have continued working on it. Launched by good people.
好,我们来个快问快答。好,开始。你会以 100 亿估值投资 Instinct 吗?
Right, we're going to do a quick fire round. Okay, let's go. Would you invest in Instinct at 10 billion?
我不做投资。
I don't invest.
你完全不投资?
You don't invest period?
我不投资,因为我妻子是风投,我们有一套合规流程,麻烦得不偿失。
I don't invest because my wife is a VC and we have a compliance process that is more trouble than it's worth.
哇,那代价不小,还得防着信用间谍。
Wow, that's costly with the credit spies.
不,这是——这是个决定。这么说吧,如果我要投资,我会基于跟创始人见 30 分钟就投。我认为我们对风投生态的接触已经够多了。我没有理由相信自己是个更好的投资人——我或许有些网络优势,我经常遇到很多优秀的创始人。他们中很多人是我的客户。
No, it's — it's a decision. Listen, if I invested, I would invest based on meeting a founder for 30 minutes. I think we have enough exposure to the venture ecosystem. I have no reason for believing I'm a better investor than — I have some perhaps network advantages and I run into a lot of great founders all the time. Many of them are my customers.
Vinod Khosla。对。他是你董事会最早的成员之一。跟 Vinod 共事最大的收获是什么?
Vinod Khosla. Yeah. One of your first ambassadors on your board. Biggest lesson from working with Vinod?
技术直觉,把很多精力聚焦在更长期的技术抱负上,一旦解决了一个问题,就立刻在接下来两三个抢先的技术赌注上下注。
Technical intuition, centering a lot of what you do towards a longer term technical aspiration and as soon as you solve one problem trying to place bets on the next two or three preempting technical bets.
过去 12 个月里,你改变最大的一点看法是什么?
What have you changed your mind on most in the last 12 months?
也许是——我一开始非常纯粹地聚焦技术和产品,觉得唯一重要的就是做出最好的技术和最好的产品。其他都不重要。而现在我一周一周地看到拥有高水平的销售、擅长营销的价值。我,是的,我觉得我以前低估了这些东西。并不是我认为它们没价值,而是我没有充分体会到,擅长这些东西能让你一周一周地感受到那种切身的差距。
Perhaps the — I started with a very pure technical and product focus like the only thing that matters is building the best technology and the best product. Nothing else matters. And now I see week on week value of having highly competent sales and being good at marketing. And I yeah, I think I discounted those things. And it's not that I thought they weren't valuable, but I didn't fully appreciate the week on week visceral delta you could perceive by being good at those things.
这话听起来有点天真,但它是真的。你有那种感受,而且你说对了。
It's like a naive person thing to say, but it's true. You feel that way and you get it right.
Perplexity 还是 xAI,哪个威胁更大?
Perplexity or xAI, which one's a bigger threat?
我不认为 Perplexity 在做我们的业务。Perplexity 也许更像一个垂直整合的产品,它竞争的是——我知道的 Instinct,或者 Claude Pro,或者 Claude——诸如此类。所以我不一定把它们看作在我们搜索这一块。
I don't think Perplexity is in our business. Perplexity is perhaps more of a vertically integrated product which competes with a — I know Instinct or a Claude Pro or a Claude — and all those. So I don't think of them as in our search necessarily.
对。对,就像网页搜索对 Perplexity 来说,就像网页搜索对一家实验室来说。
Yeah. Yeah, like web search is like web search to Perplexity is like web search to a lab.
所以 xAI——
So xAI —
对,xAI 是直接做我们这块业务的。对。
Yeah, xAI is straight up in our business. Yeah.
你有过的最好的一次风投会议是哪次?
Single best VC meeting you've ever had?
Josh 和 Todd。当 Josh 和 Todd 决定投资时,他们飞过来跟我相处了一段时间。那就是——就像那是我最好的一次会议。
Josh and Todd. When Josh and Todd made the decision to invest, they flew out to spend time with me. That is the — like it is like that is the single best meeting.
最后一个问题。展望未来 10 年,你最期待什么?
Final one for you. When you look forward to the next 10 years, what are you most excited for?
混乱。我指的是变化。我认为未来 10 年会有很多变化。而且我认为,有些人能在这个快速变化的世界里造出东西,对最终走向产生实质影响。所以我觉得,我、我的公司处在一个能在其中扮演角色的位置。开放网络会怎样?内容所有者会怎样?如果我们做对了,我们会到达一个更好的地方。所以这种可能性令人兴奋,它可以是一种真正的痴迷,即便日常很艰难,即便有一段时间事情不奏效,只要你能看到自己以某种你在意的方式弯折了现实,那就完全值得。
Chaos. By that I mean change. I think in the next 10 years a lot is going to change. And I think there are people who can build things to have — in a world that's going to change fast make a material dent on where it ends up. So I feel that me, my company is in a place where we have a role to play in that. What happens to the open web? What happens to content owners? If we get things right, we will get to a better place. And so that possibility is exciting that it can be — like true obsession and even if that day-to-day is hard and it's things don't work for some amount of time and it's totally worth it if you can see that you like bend reality in some way that you care about.
最后一个问题。你有什么信念,是坐在你旧金山的餐桌旁,你的朋友们会说:“什么?我不同意那个,老兄。不是那个。”
Final one. What belief do you have that sitting around your San Franciscan dinner table, your friends would go, "What? I don't agree with that one, dude. Not that one."
再说一次,这取决于哪张餐桌,因为世界上的差异正在增大。我不认为人们会接受这个观念:会有这种智能体一直为我们所有人持续运行。我认为这会发生,但大多数人还不同意。你和我身处一个泡沫中,但即使是旧金山的餐桌也不总是完全一致。
Again, it depends on which dinner table because the variance in the world is increasing. I don't think people buy this notion that there'll be this agent which are always running all the time for all of us. I think it's going to happen, but most people don't agree with that yet. You and I are in a bubble, but even a San Francisco dinner table isn't always in full agreement there.
好的,听着。正如我所说,我把你彻底研究了一番。嗯,我非常享受这次讨论。非常感谢你容忍我的漫谈,这很有趣。
All right, listen. As I said, I stalked the hell out of you. Um, I've so enjoyed this discussion. Thank you so much for putting up with my meandering and you this was fun.