AI 编程的未来与应用价值

The Future of AI Coding and Application Value

阿拉文德·斯里尼瓦斯 Aravind Srinivas · This Week in AI · 2026-04-23 · 约 80 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Perplexity CEO 探讨编程的开放性、AI 模型不会商品化,以及价值在于应用层。

Perplexity CEO discusses why coding is open-ended, AI models won't commoditize, and value lies in the application layer.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 20)

全文 · Full transcript(中英对照)

引言与Perplexity增长 Introduction and Perplexity's Growth

Host

其他公司能跟上 Claude Code、Cursor、Codex、GitHub Copilot 吗?

Can these other companies keep up with Claude code, cursor, codex, GitHub's co-pilot?

Aravind Srinivas

我们几乎还处于所有这些进展的起点。其他人有可能赶上。

We're almost still sort of at the beginning of all of this progress that can still be made. Someone else could catch up.

Host

我们在编程方面是否已进入终局?

Are we in the end game when it comes to coding?

Aravind Srinivas

我认为远未接近。编程与下围棋截然不同,因为编程是完全开放的。编程的可能性空间是无限的,纯粹受限于你的想象力。

I don't think we're anywhere close. Coding is very different from playing Go in that coding is completely open-ended. The space of possibilities in coding is endless. It's limited purely by your imagination.

Host

大语言模型会被商品化吗?

Are the LLMs going to get commoditized?

Aravind Srinivas

我仍然不认为 AI 模型本身会被商品化。即使它们都有相同的程度、相同的知识水平,人们自然也会根据心情想与不同的模型交谈,即使是针对同一个话题。

I still don't think that the AI models themselves are going to get commoditized. Even if they all have the same degree, same level of knowledge, people will naturally just want to talk to different models depending on their mood, even for the same topic.

Host

你认为价值将开始在哪里积累?

What do you think about where the value will start to accrue?

Aravind Srinivas

我相信价值在于应用层。其中一些公司倒闭的主要原因之一是它们无法构建应用。纯粹的 API 模式行不通。我认为,优化人类想要的、真实用户想要的、让生活更好的东西,与优化点赞和参与度之间是有区别的。

I believe that the value is in the application layer. One of the main reasons some of them went out of business, they couldn't build an application. The pure API model doesn't work. I think there's a difference between optimizing for what the humans want, what the real users want, and what makes their lives better as opposed to optimizing for likes and engagement.

Host

感谢我们的朋友 PayPal,本期《本周 AI》的独家赞助商。PayPal 是受全球数百万客户信任的支付和增长平台。PayPal Open。立即在 paypalopen.com 开始增长。好了,各位,欢迎回到《本周 AI》第 10 集。这是我为了更了解 AI 并跟上这个行业而做的新的圆桌讨论,这个行业每个月的变化可能相当于其他行业的一年。跟上它非常困难。这就是这个播客的意义所在。我们将与专家——那些真正在构建未来的人——讨论我们行业中正在发生的事情。本周我们有一个很棒的圆桌讨论。Arvind Srinivas 和我们在一起。他是 Perplexity AI 的联合创始人兼 CEO。从 AI 驱动的搜索和答案开始。人们对此上瘾了。然后发布了名为 Perplexity Computer 的产品,以及其他一些非常棒的产品。据《金融时报》报道,你的营收从 1 亿增长到了 4.5 亿。我不知道你是否确认过,但你确实做得很好。你确认了。好的。

Thanks to our friends at PayPal, the exclusive sponsor for This Week in AI. PayPal, the payment and growth platform that's trusted by millions of customers worldwide. PayPal open. Start growing today at paypalopen.com. All right, everybody, welcome back to This Week in AI episode 10. This is the new round table that I've been doing in order to get smarter about AI and keep up with an industry that is moving every month is probably a year in our industry. Keeping up with it incredibly hard. That's the point of this podcast. We'll talk about whatever's happening in our industry with experts, the people who are actually building the future. And we've got an amazing round table this week. Arvind Srinivas is with us. He is the co-founder and CEO of Perplexity AI. Started out with AI-powered search and answers. People became addicted to that. And then released something called Perplexity computer in addition to a number of other really great products and per the FT, the Financial Times, your revenue has grown 100 million to 450 million apparently. I don't know if you've confirmed that or not, but you've had quite a run. You did confirm it. Okay.

Aravind Srinivas

我们一两周前确认的是 5 亿。

500 million was what we confirmed a week or two ago.

Host

太棒了。所以这是超高速增长的最佳体现,是 Perplexity Computer 真正推动了这一切吗?是那个界面吗?

Amazing. So this is hyper growth at its best and it is Perplexity computer that's really driving this? Is it that interface?

Aravind Srinivas

没错。

That's right.

Host

在你看来,为什么这个产品变得如此受欢迎?我每天都在用。我很喜欢。我喜欢模型委员会。我一直在用 Perplexity Computer 浏览器,并且对它赞不绝口,但你可以下载这个不可思议的应用。你有 Perplexity Computer。是什么吸引了人们的注意?人们用它来做什么?

Why has this product become such a hit in your mind? I use it every day. I love it. I love the model council. I've been using Perplexity computer browser and won't shut up about it, but you can download this incredible app. You have Perplexity computer. What is catching people's attention? What are people using it for?

Aravind Srinivas

我认为它让智能体变得非常简单。这是核心原因。它是最直观的界面,可以让你作为几个智能体的管理者,本质上编排几个不同的智能体,而无需考虑它们是在本地运行还是在云端运行,也无需设置。无需上手引导。没有上手痛苦。无需提供 API 密钥。一切都在你习惯的同一个界面中直观地工作,你让别人为你做事,它连接到数百个有价值的连接器。它将所有模型整合到一个框架中,一个智能体框架。这样你就不必感到被云或 GPT 锁定。你可以保证最好的模型会做它应该做的事情。人们喜欢这一点。人们用它进行大量深入和广泛的研究、浏览器自动化、数据分析,以及许多任务,比如构建仪表盘、构建 Web 应用。

I think it makes agents really simple. That's the core reason. It makes it the most intuitive interface to have a manager of several agents to essentially orchestrate several different agents which don't need to think about whether it runs locally or on the cloud or setting it up. There's no onboarding required. There's no onboarding pain. There's no need to bring API keys. It all works intuitively in the same interface that you're used to asking people to do stuff for you and it connects to hundreds of connectors that are valuable. It puts all the models in one harness, one agentic harness. So you don't have to feel a vendor lock-in to cloud or GPT. You just can be guaranteed the best model will do whatever it's supposed to do. And people love that. People are using it for a lot of deep and wide research and browser automation, data analysis, so many tasks, building dashboards, building web apps.

Host

是的,我们已经在内部开始使用它。我们显然对 Open Claw 很着迷。我们仍在迭代那个开源项目。我们试过 Claude Cowork。然后团队开始喜欢 Perplexity Computer,以至于我不得不升级到每月 200 美元的账户。所以你把我套住了。

Yeah, we've started using it internally. We obviously had a fascination with Open Claw. We're still iterating on that open source project. We've tried Claude Cowork. And then the team just started loving Perplexity Computer to the point at which I had to upgrade to the $200 a month account. So you got me on the hook.

Aravind Srinivas

太棒了。我很乐意为你们提供任何客户支持。所以如果遇到任何问题,请随时联系我。

It's awesome. I'll be happy to do any customer support for you guys. So please feel free to ping me if you run into any issues.

Host

我们会联系你的客服热线。我们在后台功能方面取得了很好的成功。我们有一家风险投资公司,我们有会计、法律文件。我们做尽职调查。所以我们一直在编写工作范围、标准操作程序、最佳实践,比如对初创公司进行尽职调查。现在我们正在使用 Perplexity Computer 产品,并试图弄清楚:“嘿,哪些部分我们可以实际交给它?它工作得怎么样?”它令人印象深刻。今天和我们在一起的还有 Edwin Chen。他是 Surge AI 的创始人兼 CEO。他们为所有这些前沿模型做数据标注。成立于 2020 年。这已经成为一个不可思议的领域。据我所知,你的客户包括 OpenAI、Google、Prophetic、Microsoft、Meta。有 130 名员工,大约 5 万名专家合同工。其中一些可能需要更新。我不确定,Edwin,因为正如我之前所说,事情发展得非常快。但互联网上也有很多关于专家网络的争议。这是一个好生意吗?还是一个糟糕的生意?它们是非常快速增长的业务,这总是让人感到焦虑。我们投资了一家,Micro One,我认为它是你的同行。请告诉我们一些关于这个业务的情况,为什么它很重要,然后也许谈谈我们在过去一两周看到的对该行业的反弹或批评。

We will ping your customer support line. We're having good success with back office functions. We have a venture capital firm and we've got accounting and we've got legal documents. We do due diligence. So we've been writing the scope of work, the standard operating procedure, the best practice for say doing due diligence on startups. And now we're going into the Perplexity Computer product and trying to figure out, 'Hey, which sections can we actually give to it? And how well does it work?' And it's been quite impressive. Also joining us today, Edwin Chen is here. He is the founder and CEO of Surge AI. They're doing data labeling for all these frontier models. Founded in 2020. And this has become an incredible space. You have clients from OpenAI to Google and Prophetic, Microsoft, Meta, from what I understand. 130 employees, approximately 50,000 expert contract contractors. Some of this might need to be updated. I'm not sure, Edwin, because like I said earlier, things are moving really fast. But there's been a lot of brouhaha on the internet about expert networks as well. Is this a great business? Is it a terrible business? They're very fast-growing businesses, which always gets people wringing their hands. We have an investment in one, micro one, which I think is a contemporary of yours. Tell us a little bit about the business, why it's important, and then maybe the backlash or the criticism of the industry that we saw in the last week or two.

Edwin Chen

我的意思是,我想先说,我实际上讨厌“数据标注”这个术语,因为当你谈到数据标注时,你会想到人们在做非常简单的事情,比如标注猫和狗的图片,在汽车周围画边界框。而我认为我们做的事情实际上比那复杂得多。我经常把我们正在做的事情想象成在为 AGI 建造一所学校。就像我们有所有这些非常聪明的物理学家,比如哈佛教授、普林斯顿研究生、斯坦福计算机科学博士。他们所做的有点像在交叉审问这些模型,探测它们以找出它们何时犯错。然后当它们犯错时,他们会介入并教它们所有这些非常高级的东西。

I mean, I would start off by saying I actually hate the terminology data labeling because when you talk about data labeling, you think about people doing incredibly simple things like labeling images of cats and dogs and drawing bounding box around cars. And I think what we do is actually so much more complex than that. It's like I often think of what we're doing as building a kind of school for AGI. Like we have all these incredibly smart physicists, like Harvard professors, Princeton graduate students, Stanford computer science PhDs. And what they're doing is they're kind of like cross-examining these models and probing them to figure out when they make mistakes. And then when they make a mistake, they're going in and teaching them all these incredibly advanced things.

教AI与训练AI模型 Teaching vs. Training AI Models

Aravind Srinivas

是的,我认为这几乎是我们能为 AI 做的最深刻的事情之一。它甚至超越了教学。我经常把我们正在做的事情想象成培养这些模型,不仅仅是让它们正确,不仅仅是让它们给出问题的答案,而是让它们思考,拥有某种价值观,以及智慧和品味等等。所以,首先我要说的是,我认为我们正在做的事情是

And so yeah, I mean, I think it's almost like one of the most profound things that we can do for AI. It's like, I mean, it even goes beyond teaching. Like I often think about what we're doing as like raising these models not just to be correct, not just to produce the answer to a question, but to think and to have a certain kinds of values and to have like wisdom and taste and all that. So, uh, yeah, so I first of all, I'll start off by saying I think what we're doing is

Host

有更好的术语吗?行业里有没有从数据标注演变而来的术语?因为我同意你的看法,当你雇佣博士、律师或注册会计师时,这远不止于此。

Better term? Is there an industry term from that, you know, has evolved this from data labeling? Cuz I agree with you, it's much more than that when you hire PhDs or lawyers or CPAs to

Aravind Srinivas

是的。

Yeah.

Host

数据训练?是数据训练吗?行业里正确的术语是什么?或者它需要一个新词?

Data training. Is it data training? What's the right term in the industry? Or does it need one?

Aravind Srinivas

所以,我经常用育儿或教育的类比来思考这个问题。我们做的不仅仅是教他们事实,不仅仅是教他们“哦,这是维基百科页面,这是正确答案”。相反,我们试图教他们创造力和品味。所以,我个人喜欢的术语是“教模型”或“AI 教学”。因为我认为这实际上也超越了训练。你也在评估它们等等。但我确实喜欢“教学”这个概念。

So, again, I think I often think about this either parenting or education analogy, where again, what we're doing is going beyond teaching them facts. Going beyond just teaching them, oh, like, you know, this is a Wikipedia page and here's the correct answer. And instead, we're trying to teach them like creativity and taste. So, like personally the terminology I like is either uh like teaching the models, AI teaching. Uh cuz I think it actually gets beyond training as well. Like you're also measuring them and all that. But I also like that idea of teaching.

AI教学投入 Spending on AI Teaching

Host

这些前沿模型总共在这方面花了多少钱?看起来每年有数十亿美元,对吧?

How much are these collectively the frontier models spending on this? It seems like billions of dollars a year, yeah?

Aravind Srinivas

是的。而且我认为疯狂的是,这实际上与它们的算力预算相比相形见绌。所以,我认为它们应该花更多。

Yeah. And I think the crazy thing is that I actually think that this pales in comparison to their compute budgets. So, I think they should be spending a lot more.

Host

是的,有道理。据我所知,这些真正的专家在非常特定的领域每小时能赚 100 美元、200 美元,因为网络上的大量数据已经被爬取过了。那么,我们机械地来看一下,然后进入我们今天这份精彩的文档?我们有四五个很棒的主题要讨论。但为了让观众理解,它是如何运作的?你们会使用人们在使用大语言模型时点“踩”的查询吗?这些会被路由吗?比如,如果某人对 Perplexity 的查询不满意,它会路由给你们修复吗?还是你们只是说,“嘿,我们雇一群律师来处理世界上这些重要案件,并以智能方式标注它们?”数据是如何工作的?

Yeah, fair enough. Uh and these are true experts getting paid 100 bucks an hour, 200 bucks an hour from what I understand in very specific fields because a lot of the data obviously on the web that could have been crawled has been crawled. So, so how does it just take us mechanically and then we'll get into this amazing doc we have today? We've got four or five great subjects we're going to chop up. But just to for the audience to understand, how does it work? Do you take the queries that people gave a thumbs down to when they were using an LLM? And does that get routed? Like if somebody's not happy with a perplexity query, does it get routed to you to fix? Or do you just say, "Hey, let's just hire this group of attorneys to take these important cases in the world and annotate them in an intelligent fashion?" How does the data work?

Aravind Srinivas

是的,实际上有很多不同的方式。最典型的方式是:假设你有一位数学专家。在他们正常的研究过程中,比如试图证明某个新定理。他们会像进行正常研究一样与模型互动。比如,“嘿,试着解决这个问题。试着向我解释这个概念。”他们一直这样做,直到发现模型的失败。这就是为什么我经常把它看作是对模型的“交叉询问”。你与它交谈,直到你发现一个非常有趣的失败。有时有加速发现失败的方法。比如我们这边可能会做各种事情,让我们的数据科学团队找到模型中的损失模式,这些模式会引导我们发现失败。或者,有时实验室的朋友会给我们发送某些查询,他们感觉到用户不满意。但用户并不总是对的。用户因为各种原因给差评。所以我们仍然需要确保模型在那里确实失败了。我们做的是验证模型是否失败,如果是,我们就教它正确的答案。嗯,有很多不同的方式来找出现模型推理中的这些“断裂缺口”。可能是我们自己发现的,可能是通过我们做的各种分析,也可能是通过用户对话。嗯,基本上,我们拿这些对话,然后教模型正确的答案。

Yeah, so there's actually a bunch of different ways it can work. So, probably the most canonical way it works is: Okay, so you have this like expert mathematician. And in the course of their normal research, like yeah, they're trying to, you know, prove some new theorem. They will kind of just like interact with the models as if they're doing their normal research. So, you know, "Hey, you know, try to solve this problem. Try to explain this concept to me." They keep on doing that until they find a failure from the model. Again, this is why I often think about it as a cross examining model. Like you're talking to them until you kind of like find this very very interesting failure. And sometimes there's ways of accelerating finding that failure. Like we may do various things on our end where we're I guess our data science team to find loss patterns in our models that guide the failures. Or yeah, sometimes friends to our labs will send us uh like certain kinds of queries where they sense that users are unhappy. And I mean it's not always the case that the user is right. Like you know, users aren't wrong. They fund things down for incredible reasons. And so they will still need to make sure that the model failed there. And so what we do is we'll verify that model failed and if so, like you know, we'll teach it the correct answer. Um yeah, like there's all these different ways of coming up with these almost like broken gaps in the model's reasoning. And so it might be either us finding it ourselves, it might be through various kinds of analyses that we do, it might through be through like user conversations. Um yeah, we basically take those conversations and then we teach it the right answer.

新CEO下苹果的AI机遇 Apple's AI Opportunity Under New CEO

Host

是的,这对人们来说正成为一份好工作。好了,听我说。话题一:整个行业确实被蒂姆·库克决定卸任 CEO 的消息震惊了。这对我们来说是一个重要的话题,因为苹果在 AI 方面有巨大的机会,原因有很多,他们任命了约翰·特努斯为 CEO。库克将担任 CEO 到 9 月 1 日,然后升任执行董事长。他仍将负责行业关系,但特努斯已经在苹果工作了 25 年,参与了许多非常重要的硬件产品。我想,我最想听阿文德和埃德温的看法是:苹果芯片可以说是他们最成功的案例之一。他们摆脱了英特尔,然后当开源模型出现时,人们说,“嘿,我在哪里能运行这些?”如果你没有英伟达的机架,人们开始组装 Mac Studio,配备 128GB 或 512GB 内存,他们还有 Siri。所以你有 Siri、芯片和操作系统。三个重要的“S”。新 CEO 应该做什么?阿文德,如果你与他们合作,拥有这些令人难以置信的资产,你会给他们什么建议?因为他们没有值得一提的语言模型。他们做过一些开源项目。他们有一个功能失调、支离破碎的 Siri,每个人都想把它扔出车窗。但感觉他们确实处于有利位置。你的看法是什么?

Yeah, and it's becoming a great job for people. Uh all right, listen. Topic one, the industry is really uh was taken back by Tim Cook deciding to transition out of the CEO uh role. This is an important thing for us to discuss here because Apple has a huge opportunity in AI uh for a number of reasons and they named John Turnus as CEO. Uh Cook's going to be CEO until September 1st, then he'll move up to executive chairman. He's still going to work on industry relations, but uh Turnus has been there now for 25 years and he worked uh on a lot of very important hardware products while at the company. And I guess, you know, the take I'm most interested in hearing uh Arvind from you and and and also from Edwin is Apple Silicon has arguably been one of their great success stories. They got off of Intel and then when the open claw thing came out or Kimmy, uh open see a bunch of open source models, people said, "Hey, where can I run these?" Okay, if you don't have an Nvidia rack people started pulling together Mac Studios with 128 gigs of RAM, 512 gigs of RAM and they also have Siri. So you have Siri, you have silicon, and you have a system, an operating system. So three big S's there. What should the new CEO do? What would you advise them to do, Arvind, if you were uh working with them with this incredible group of assets because they don't have a language model to speak of. They've worked on some open-source projects. They've got a dysfunctional broken Siri that everybody wants to throw out the window of their car when they try to use it. But it does feel like they are positioned well. What are your takes?

Aravind Srinivas

我实际上认为 M 系列芯片,这是现任 CEO 约翰·特努斯领导的项目,是他们被低估的资产之一。我认为人们真的低估了制造强大芯片的难度。目前,在基准测试中,它甚至比 DGX Spark 更好。至少对于可以本地托管的大语言模型的本地推理来说是这样。还有开源模型,比如你可能看到了最近发布的 Kimiko 2.6,我想是昨天发布的,

I actually think the M series chips, which is a project led by uh John Ternus, the current CEO. Um is is is one of their underrated assets. I think people really underestimate what it takes to build a powerful chip. Um at this moment in time it is even better on the benchmarks than DGX Spark. Uh at least for local inference uh of of of of LLMs that can be hosted locally. And the open-source models, like like you you might have seen Kimiko 2.6 that uh launched recently uh I think yesterday that

Host

嗯。

Mhm.

Aravind Srinivas

似乎在终端基准测试和智能体套件基准测试上比 Ocus 和 GPT 做得更好。所以我确实认为这些模型正在达到一个点,比如 3.6、Kimiko K2.6。它们正在达到一个可以与前沿模型竞争,同时也有可能在你的 MacBook 或 Mac mini 等辅助硬件上运行的程度。尤其是 M6 芯片会更好。

seems to be doing even better than Ocus and uh GPT on some of the terminal bench and agentic suite benchmarks. So I do think these models are getting to a point like when 3.6, Kimiko K2.6. They're getting to a point where um they can be competitive with the frontier but they could also potentially run on one of your um MacBooks or the auxiliary hardware like Mac minis. Especially the M6 chips are going to be even better.

苹果芯片优势与智能体循环 Apple's silicon advantage and agent loops

Aravind Srinivas

他们已经提前锁定了明年和后年 2 纳米芯片的大量晶圆产能。所以他们应该在这方面继续深入,而且我认为你们有完美的领导者。Tim 把公司布局得很好,让 Apple Silicon 这个赌注在未来多年都得到了回报。所以,如果智能体循环开始在本地运行,那就是 CPU 算力。所有这些都不需要集中在服务器上。你可以拥有自己的智能体循环。你的智能体在本地系统上访问的数据——本地文件、本地应用、消息、邮件、笔记、照片——所有这些都可以保持私密。编排循环可以在本地运行,编排它们的模型也可能在本地运行。哪家公司最适合从这一切中获利?我认为是 Apple。所以他们实际上处于一个相当有利的位置。

They've already secured a lot of the fab capacity in advance for the 2-nanometer chips for next year and the year after. So they should go deeper on this, and I think you have the perfect leader for that. Tim has set up the company well so that Apple's silicon bet has paid off for multiple years in the future. So if agent loops start running locally, that's the CPU compute. All that stuff doesn't need to be centralized on servers. You get to own your agent loops. What data your agent accesses on your local system—local files, local apps, messages, emails, notes, photos—all that can stay private. The orchestration loop can run locally, and the model orchestrating them could also potentially run locally. Which company is best positioned to profit from all this? I think it's Apple. So they're actually in a pretty good spot.

Host

是的,Edwin,如果你仔细想想,这是一个令人难以置信的愿景,因为前沿模型很昂贵。现在它们是前沿模型,所以它们往往领先,但正如你所了解的,如果你是 Perplexity 的用户,它只是为你挑选最好的模型,大多数人十次查询中有九次其实并不关心是什么模型。

Yeah, this is an incredible vision if you think about it, Edwin, because frontier models are expensive. Now they are the frontier models, so they tend to be ahead, but as you learn, if you're a Perplexity user, and it's just picking the best model for you, nine out of ten queries that most people do, they actually don't care what model.

Aravind Srinivas

正是如此。

Exactly.

Host

所以这个愿景与前沿模型共存是兼容的。这个编排器仍然可以调用一个依赖前沿模型的子智能体。

So this vision is compatible with frontier models coexisting together. This orchestrator can still ping a sub-agent that relies on a frontier model.

Aravind Srinivas

嗯。

Mhm.

Host

它可以使用你自己的 API 密钥,也可以使用 Perplexity 的集中式版本,这并不重要。但关键是循环开始在本地硬件上运行,智能体循环本身,以及像事件触发器这样的重复进程。我们可以设置一个触发器,比如每当 Jason 就 Perplexity Max 的问题给我发短信时,就提醒我的支持团队。我可以设置很多这样的循环,它们根本不需要在任何服务器上运行,然后它就开始成为我自己的个人电脑,就像我拥有的自己的智能体。而最适合这个的硬件设备就是 Apple 的生态系统。

It could use your own API key or it could use a Perplexity centralized version, doesn't matter. But the key thing is the loops start running locally on your hardware, the agent loop itself, the recurring processes like event triggers. We could have a trigger that says every time Jason texts me about an issue on Perplexity Max, make sure to alert my support team about it. I could set up a lot of loops like this that just don't need to run on any server, and then it starts to be my own personal computer, like my own agent that I own. And the hardware device that's best suited for this is Apple's ecosystem.

Host

Edwin,你对硅芯片的威力有什么看法?鉴于他们在面向庞大客户群的任何产品方面基本上是从零开始,新 CEO 应该怎么做?你会怎么做?你和 Arvind 有相同的愿景还是不同的?

Edwin, what are your thoughts here on the power of the silicon and what the new CEO should do, given they're kind of starting from zero in terms of any kind of product that's facing their massive customer base? What would you do? Do you have the same vision as Arvind or a different one?

Edwin

所以我想说两点。第一,我认为历史上很多人认为大语言模型会成为商品。比如最终每个模型都会达到一定程度的智能,并且可以互换。我认为我相信的,也是我们过去一年开始看到的,是实际上每个模型都有不同的个性。你与 ChatGPT 互动,感觉就与从 Claude 或 Gemini 那里得到的对话类型、个性类型、品味类型截然不同。所以它们几乎不是商品,我真的不认为它们会成为商品。你真的、真的、真的需要自己的基础模型,因为 AI 对未来以及你希望产品拥有的感觉如此重要,以至于你确实需要自己的基础模型。否则你只能依赖别人的品味、别人的 sophistication。所以我真的认为基础模型对 Apple 来说将极其重要,他们真的需要自己构建。

So I think I would say two things. One is I think historically a lot of people have thought that LLMs were going to be commodities. Like at the end of the day every model is going to be intelligent to some level and they're going to be interchangeable. I think what I believe, and what we've been starting to see over the past year, is that actually every model has a different personality. You interact with ChatGPT and it just feels very different from the type of conversation, the type of personality, the type of taste that you get from Claude or Gemini. So it's almost like they're not a commodity, and I really don't think they're going to be. You really, really, really need your own foundation model because AI is going to be so important to the future and to the kind of feel that you want your products to have that you really are going to need your own foundation model. Otherwise you're just going to be relying on somebody else's taste, somebody else's sophistication. So I really do think that the base foundation model is going to be incredibly important for Apple and they really need to build it themselves.

Host

他们需要拥有一个模型吗,Edwin?关于第一点,你认为他们应该收购一家模型公司,还是干脆成立一个团队,也许分叉一个现有的开源模型?因为他们有一个他们正在研究的开源图像模型。你认为他们在构建模型方面应该怎么做?然后你继续讲你的第二点。

Do they need to own a model, Edwin? On that first point, do you think that they should either buy a model company or just start a group and maybe fork an existing open source one, because they have an image one that they work on that's open source. What do you think they should do in terms of building models? And then you definitely go on to your second point.

Edwin

哦,是的。我的意思是,我绝对认为他们真的需要自己构建,因为如果你不自己构建,你就依赖别人的品味和个性来决定 AI 应该如何表现。Apple 显然一直对其产品及其设计应该是什么样子有着非常强烈的愿景,以至于他们不能只是把它外包给别人。当然,他们可能暂时能做到,只是为了尝试所有这些不同的概念,并探索 AI 产品在 Apple 设备上可能是什么样子。但如果他们真的想拥有未来,他们就需要自己的模型,以便将他们自己的价值观注入到他们希望 AI 系统表现的方式中。

Oh, yeah. I mean, I definitely think they really need to build their own because if you don't build your own, you're relying on somebody else's taste and personality for how an AI should behave. Apple has obviously always had such a strong vision for what their products and what their design should be that they can't just outsource it to somebody else. Sure, they may be able to do it temporarily, just to play around with all these different concepts and play around with what AI products can look like on an Apple device. But if they really want to own the future, they are going to need their own to infuse their own values into the way they want AI systems to behave.

Host

在我看来,很明显他们会把整个公司押在这上面。而且人们不禁要问,Steve Jobs 发起了 Apple Silicon 运动。我认为是在 2008、2009 年他们做出了决定,大约八、九年后才推出了第一批产品。这是一个非常重要的战略努力。但我不认为他们当时想的是“哦,这会……”,那时还没有大语言模型。甚至没有人知道这个产品会存在。但是,天哪,这真是机缘巧合和押对了宝,是吧 Arvind?如果你回顾 Steve Jobs 的历史遗产,他有点像看到了拐角后面,或者同时看到了两个拐角。这是一项不可能完成的任务。

Seems clear to me that they are going to put the whole company behind this. And one has to wonder, Steve Jobs started the Apple Silicon movement. I think it was 2008, 2009 when they made the decision, they came out with the first products like eight, nine years later. This was a very significant strategic effort. But I don't think that they had in mind, 'Oh, this is going—' at that time there were no large language models. Nobody even knew that this product would exist. But man, talk about serendipity and making a great bet, huh Arvind? If you think about historically Steve Jobs' legacy, he kind of saw around a corner, or maybe two corners at once. It's an impossible task.

Aravind Srinivas

是的,我认为 Apple Silicon 是一个非常——它不一定是他们为 LLM 构建硬件而下的赌注。但硬件变得越来越强大。神经引擎非常强大。MLX 编译器真的非常好。他们现在在构建这些东西方面有很多专业知识。而且不仅如此,这不仅仅是关于 Mac Studio 或 Mac Mini。还要考虑这样一个事实:如果你想要一个计算生态系统,比如你未来想戴一副 Apple Glass,你想解析你看到的任何东西,并开始提问,所有这些都与你在家中的一些辅助硬件无缝配对,但这一切都作为一个伪桌面服务器运行。所以你能够将所有计算能力配对到一个设备家族中。我认为这些都是你可以提供给消费者的神奇体验,而不会耗尽设备本身的电池。所以所有这些都还没有转化为真正的消费者体验。

Yeah, I think Apple Silicon is a very—it wasn't necessarily a bet they made to build hardware for LLMs. But the hardware got increasingly more and more powerful. The neural engine is very capable. The MLX compiler is really, really good. And they have a lot of expertise in building these things now. And not just that, it's not just about Mac Studios or Minis. It's also consider the fact that if you do want an ecosystem of compute, like you want to wear an Apple Glass, let's say, in the future, you want to parse whatever you're seeing, and you want to start asking questions about it, and all that pairs seamlessly with some auxiliary hardware you have at home, but it's all running as a pseudo desktop server. So you're able to pair all the compute in one family of devices. I think all that is the kind of magical experiences you can provide to a consumer without draining the battery on the device itself. So all that stuff hasn't been converted into a real consumer experience yet.

苹果在消费AI设备上的优势 Apple's advantage in consumer AI devices

Host

但在我看来,即使其他人制造了所有这些消费级 AI 设备,他们最终也会输给苹果,因为苹果拥有芯片优势,他们已经提前多年锁定了产能,他们有操作系统,有生态系统锁定,还有你所有的个人联系人,而且你信任他们以最注重隐私的方式处理这一切,对吧?

But it feels to me that even if other people build all these consumer AI devices, they're eventually going to lose to Apple, because they have all the chips advantage, they've already secured the capacity for years in advance, they have the OS, they have the ecosystem lock-in, and they have all your personal contacts, and you trust them to handle it in the most privacy-conscious way, right?

Aravind Srinivas

这是一个关键点,对吧?就是隐私。稍微展开一下,因为你之前提到过。比如,作为一家公司,你想拥有自己的智能体循环,也就是你的组织正在构建的智能体知识。这基本上就是你的整个业务,把它输入到另一个大语言模型中,对某些人来说可能觉得没什么大不了的,直到你所有的秘密都被你的竞争对手利用,因为你刚刚用下一个机器人模型进行了训练,对吧?

This is a key point, right? It's the privacy. Unpack that a bit, because you mentioned it earlier. Like, hey, as a corporation, you want to own your agent loops, the agentic knowledge that your organization is building. It's essentially your entire business, and to feed it into another LLM, well, that might seem to some people like, okay, no big deal, until all your secrets are now being used by your competitors because you just did the training on the next bot model, yeah?

Host

完全正确。所以,我的意思是,这也是为什么他们即使使用不同的模型,比如新闻说他们在和 Gemini 合作,他们也会在自己的芯片上托管,根据自己的需求定制,并进行大量的定制后训练。所以,我的感觉是,尽管很多人对 Siri 的好坏有看法,但他们有能力花时间按照自己的方式做事,因为他们作为一个真正值得信赖的品牌有很多优势,生态系统锁定被低估了,辅助硬件设备、芯片优势,所有这些现在都被严重低估了。而且我有一个以前没说过的新观点:手机,iPhone,实际上根本没有被 AI 颠覆。事实上,AI 工作得越好,iPhone 本质上就变成了你的数字护照。它有你钱包和所有卡片,有你的通行证,有你的健康记录。你通过它与他人联系,进行 FaceTime、打电话,保存你生命中珍贵时刻的照片。所有这些都真正属于你个人,与 AI 无关。这就是为什么他们实际上可以慢慢来。

Exactly. So, I mean, this is also why they, even if they're using a different model, like say the news is they're working with Gemini, they will host it on their own silicon, they will customize it to their own needs, they'll be doing a lot of custom post-training for that. So, my sense is that even though a lot of people have opinions on how bad Siri is or good Siri is, they have the ability to take time and do things the way they want to do because they have a lot of advantages as a brand that people truly trust, and the ecosystem lock-in is underrated, and the auxiliary hardware devices, the chip advantage, all this is really underrated right now. And here's my opinion I haven't said this before: the phone, the iPhone, is actually not getting disrupted by AI at all. In fact, the more AI works better, the iPhone essentially becomes your digital passport. It has your wallet and all your cards. It has your passes. It has your health records. And you connect with other human beings through it. You do FaceTimes, you do calls, you have your photos of precious moments in your life. All these are things that are truly personal to you and have no connection to AI. And that's why they can actually afford to move stuff.

Aravind Srinivas

是的。在我看来,隐私这部分,Edwin,就是隐私加上芯片,再加上,天哪,我的照片在这里,但我不想让我的照片出现在 OpenAI 上,无意冒犯 ChatGPT,但即使是 Google,我也会想,我真的要这样吗?我关闭了同步,把照片从 Google 移走了。我现在不想让它们留在 Google 云里。我更愿意把我孩子的所有照片保存在本地设备上,而且我相信苹果不会用我孩子的图像来训练他们的下一个图像模型,等等。

Yeah. It does seem to me that that privacy piece, Edwin, is the privacy plus the silicon plus, oh my gosh, my photos are here but I don't want my photos up on OpenAI, all due respect to ChatGPT, but even with Google, I'm like, do I really? I turned off syncing and took my photos off Google. I don't think I want those in the Google cloud right now. I much prefer to keep all my kids' photos on my local device, and I trust Apple to not train their next image model on my kids' images, etc.

Host

是的,完全正确。我的意思是,尤其是因为这些模型非常强大,当它们在某些数据上训练时,最终会直接复述出来。我认为人们应该思考他们的数据来自哪里,这一点非常重要。

Yeah, exactly. I mean, especially because these models are so powerful that when they're trained on certain pieces of data, they just end up regurgitating it. I think it's really, really important that people should wonder about where their data comes from.

Aravind Srinivas

是的,你知道吗,下一个故事,Edwin,你在我们的群聊里谈到了大量的后期资本,我们正处于风险投资的一个非常有趣的时刻。我们都在这个行业里有一段时间了,但涌入这个领域的资金量和速度都非同寻常。AI 公司在 2026 年第一季度筹集了 2420 亿美元。我假设这个数字包括 OpenAI 筹集的 1000 亿美元。但你的公司,Edwin,Surge,你等了一段时间才融资。我想你在进行第一轮融资之前已经达到了 10 亿美元的收入。也许谈谈你为什么坚持那么长时间的自主创业模式,以及所有这些资金涌入创始人手中会对行业产生什么影响。然后,Aravind,我会问你如何管理你的资金,因为你也是受益者之一。

Yeah, you know, the next story, Edwin, you were talking in our group chat about the massive amount of late-stage capital, and we are in a really interesting moment in time in venture capital. We've all been in the industry for a while, but the amount of money and the velocity of the money coming into this space is extraordinary. AI companies raised $242 billion in Q1 of 2026. I'm assuming that number includes the giant hundred-billion-dollar raise by OpenAI. But your company, Edwin, Surge, you waited to raise money. I think you hit like a billion dollars in revenue before you did your first round. Maybe talk a little bit about why you went with the bootstrap model for as long as you did, and then what the impact of all this money being dumped on founders is going to have on the industry. And then, Aravind, I'm going to go to you just to talk about how you manage your treasury, because you've also been a beneficiary of this.

Host

是的,所以我的意思是,首先,我们实际上从未融资过。所以我们仍然愉快地自主创业,并以我认为最好的方式增长。所以我真的非常高兴我们从未融资。而且,是的,你说得对,我认为在 AI 领域融资从未如此容易。而这正是问题所在。比如,当你筹集了 10 亿美元,你就会从投资者那里得到所有这些增长目标,这些目标激励数量而非质量。你会面临董事会压力,要求你把时间花在优化下一轮融资上,而不是你正在构建的产品上。就像我所有其他公司的 CEO 朋友一样,他们会说,“哦,是的,我接下来几周都要准备董事会演示。”他们总是嫉妒我不需要做同样的事情。而且我认为问题在于,我听说一些后训练团队,他们的目标不是让模型更智能。这些后训练团队的目标实际上只是让他们的公司获得 10 亿用户。如果那是你的北极星,而且当你筹集了巨额资金时就会出现这样的北极星,那么你的模型开始试图在你耳边低语并开始用标题党吸引你,这有什么奇怪的吗?这有点好笑。几周前,我在和 ChatGPT 聊天。我问它,你知道,给我一些在东京做什么的建议。它回复我时结尾是,“嘿,顺便说一句,你想听一个你可以做的奇怪技巧吗?”它就像在跟我说话,因为我们有这个超级智能的模型,但它听起来就像 2002 年的小报。是的。问题是,这就是当你有所有这些与你最初想要构建的东西不一致的激励时会发生的事情。所以,是的,我们选择了一条完全不同的道路,我认为我们很高兴我们这样做了。

Yeah, so I mean, first of all, we've actually never raised. So we're still happily bootstrapped and growing in, I think, the best way possible. So I'm actually really, really happy that we've never raised. And I mean, yeah, to your point, I think it's never been easier to raise money in AI. And that's kind of the problem. Like when you raise a billion dollars, you get all these growth targets from your investors that incentivize volume over quality. You get all this board pressure to spend your time optimizing for your next fundraise instead of the product that you're building. Like all of my friends who are CEOs of other companies, they're like, "Oh yeah, I have to spend the next few weeks just prepping a board deck." And they're always jealous of the fact that I don't have to do the same thing. And I think the problem is, I've heard that some post-training teams, their goal is not to make their model more intelligent. Their goal of all these post-training teams is actually just to get their companies a billion users. And if that's your North Star, and it's a North Star that happens when you raise gazillions of money, is it a surprise that your model starts trying to whisper in your ear and start clickbaiting you? Like it's kind of funny. A couple weeks ago, I was chatting with ChatGPT. I was asking it, you know, give me some tips for what to do in Tokyo. And it ended the response to me with like, "Hey, by the way, do you want to hear about one weird trick that you could do that was no?" And it was just like talking to me because we have this super intelligent model and it just sounds like a 2002 tabloid. Yeah. And the problem is, this is what happens when you have all these different incentives that don't align with what you were originally trying to build. So, yeah, we chose a completely different path and I think we're really happy that we did.

Aravind Srinivas

非常令人印象深刻。而且你通过自主创业达到了 10 亿美元的收入,我想我没听说过我们行业里有其他公司做到过。Edwin,你听说过有公司没有融资就达到 10 亿美元收入吗?我的意思是,我知道有些人非常谨慎地融资,但这对我来说确实是第一次。我想我从未听说过。

Super impressive. And you hit a billion dollars in revenue bootstrapping, which I don't think I've heard of another company that's done that in our industry. Edwin, have you heard of one who's hit a billion in revenue without raising venture capital? I mean, I know people who have been incredibly judicious about raising, but that's a true first for me. I don't think I've ever heard that.

保护AI免受增长黑客影响 Shielding AI from growth hacks

Host

是的,我认为这非常重要,因为我确实觉得 AI 对我们的未来至关重要,它需要被保护,免受典型的硅谷增长黑客策略的影响。

Yeah, I mean, I think it's really important because I actually really do think that AI is just so important for our future that it kind of needs to be shielded from the typical Silicon Valley growth hack playbook.

Host

我唯一能想到的是 Mailchimp 和 Patagonia 以不融资闻名,但不在我们行业。Arvind,你有一个资金储备,从一些最重要的公司和投资者那里融了资。你认为我们正在看到不自然的行为吗?这是我用来描述 Edwin 所说现象的词。我在出版行业亲眼见过,人们追逐点击诱饵,做各种不自然的行为来提高页面浏览量,BuzzFeed 和 Business Insider 就是典型的疯狂例子。你如何保持脚踏实地,同时你又在与人竞争?如果你不融资,资本就会成为武器,就像我在 Uber 对 Lyft 对 Sidecar 中亲眼所见的那样。Travis 在融资方面简直是个怪物,如果你投资了一家,就不能投资另一家。这种情况现在有点消失了,但你对这个问题有什么看法?

The only ones I can think of Mailchimp was famous for that and Patagonia, but that's not in our industry, but those are so that that other company, that was another one that didn't raise a ton of money to do this. Arvind, you've got a war chest, you've raised from some of the most important companies and investors in the world. Do you think we're seeing unnatural acts? That's the term I use for what the phenomenon Edwin's talking about. I saw it up close and personal when I was in the publishing space and people would chase clickbait. They would do all kinds of unnatural acts to try to get their page views up and BuzzFeed would be and Business Insider would be these like canonical examples of lunacy. How do you stay grounded and then also you're in competition with people. So, if you don't raise, then there's capital as a weapon as I saw up close and personal with Uber versus Lyft versus Sidecar. Just Travis was an absolute monster when it came to raising money and if you invested in one, you couldn't invest in the other. That's kind of gone away here a bit, but what are your thoughts on this issue?

Aravind Srinivas

非常敬佩 Evan 所做的。我认为不融资并达到 10 亿营收非常非常困难。不仅是在业务建设方面,不仅仅是财务帮助,还包括说服其他人加入你的创业公司。他们都在看估值,都希望得到外界的认可。要说服真正优秀的员工在你还没有成熟业务时加入你,他们希望看到来自其他人的认可,这可能是一位知名的风险投资人。所以,没有这个就很难组建团队。因此,非常敬佩他。我认为创始人需要从 Surge 这类公司的成功中学到的一点是,他们需要更加自律。你当然可以融资,只要你真正知道如何正确使用资金。Elon 融了很多钱,但他完全知道该怎么花,对吧?xAI 融了很多钱,但都用于建设数据中心。他以资本配置非常审慎而闻名。所以,你可以两者兼得。你可以拥有自筹资金创始人的自律,也可以拥有资本主义历史上最成功创始人 Elon 的雄心。所以,如果你能设法做到两者兼得,你就能更成功。这是我的心得。不必在永远自筹资金和无节制地融资之间二选一。我认为你应该成为一个非常好的资本配置者,并且有一个清晰的计划说明为什么需要钱。然后,即使你拥有巨大的资金储备,你也应该继续保持自筹资金创始人的自律。

Huge kudos to Evan for doing what he did. I think not raising capital and getting to a billion in revenue is very, very hard. Not just in terms of business building, like not just financial help, but also convincing other people to come join your venture. They're all looking at valuations. They all want validation from the rest of the world. To convince really good employees to come join you when you don't yet have a working business, they want to see validation from somebody else, which could be a reputed venture capitalist. So, trying to build a team without that is very hard. So, kudos to him. I think one thing that founders need to take away from the success of companies like Surge is that they need to be more disciplined. You can raise money for sure, as long as you truly know that you're spending it the right way. Elon raises a lot of money, but he knows exactly what to do with it, right? XAI has raised a lot of money, but it's being spent on building data centers. He is known for being very judicious about the allocation of capital. So, you can be both. You can have the discipline of bootstrap founders, but you can also have the ambition of the most successful founder in history of capitalism, Elon. So, if you can figure out a way to be both, you could be far more successful. That's my takeaway. It doesn't have to be a dichotomy between staying bootstrapped forever or raising endlessly with very undisciplined capital allocation. I think you should just be a very good capital allocator, and you should have a clear plan for why you need money. And then you should continue to have the discipline of a bootstrap founder, even if you have a gigantic war chest.

Host

你是如何融资的?到目前为止你融了多少钱?我想这已经是公开的了。

How have you managed to raise? How much have you raised to date? I mean, I think it's been public already.

Aravind Srinivas

我想我们累计融了大约 20 亿。从去年八月以来就没有再融资了。我们的目标实际上是进一步推进营收进展,我们从年初就开始这样做,并努力实现盈利。与模型公司不同,我们不需要在算力上花很多钱,尤其是在训练上。我们做很多后训练,但不做任何预训练。所以,我们没有理由不盈利。而且与编码应用层的公司不同,它们的营收毛利率实际上是负的,我们没有这个问题。所以,我们所有营收的毛利率都非常正,对于最高用户每月 200 美元的计划来说,毛利率非常高。因此,我们的目标就是继续增长营收,保持自律,不增加工资或基础设施支出,并尽快实现盈利。当这发生时,我们并不认为更多资本会显著改变我们的命运。这应该成为应用层公司的蓝图。就像尝试以 Edwin 那样的自筹资金创始人的自律来高效运营公司,并努力保持营收增长。

I think we raised around like what around 2 billion cumulatively. We haven't raised since like August of last year. And our goal is actually to advance further in our revenue progress, that we've been doing since beginning of the year, and try to become profitable. Unlike a model company, we don't have to actually spend a lot on compute, particularly on training. We do a lot of post training, but we don't do any pre-training. So, we have no excuses to not be profitable. And unlike companies in the coding application layer, where your gross margins on the revenue are actually negative, we don't have that problem. So, gross margins on all the revenue we make are pretty positive, highly positive in the case of max users to the $200 a month plan. So, our goal is to just keep growing the top line, stay disciplined, not actually spend more on payroll or infra, and become profitable as soon as possible. And when that happens, we don't actually think more capital is leading to a meaningful change to our destiny. And that probably should become the blueprint for application layer companies. Like try to just run the company in an efficient way with the discipline of bootstrap founders like Edwin. And try to keep growing the top line revenue.

Host

Arvind,把单位经济模型调好至关重要。我们看到一些其他公司,我认为编码领域是头号例子。它们就是在亏钱。

Having the unit economics dialed in Arvind is critically important. We've seen some other folks who are supposedly and I think the coding space is the number one example. They're just losing money.

Aravind Srinivas

没错。是的。

That's correct. Yeah.

Host

是的,我不知道有多少比例的用户,或者可能是整个用户群体总体上在亏钱,对吧?

Yeah, on their I don't know what percentage of users or maybe it's the entire user base in aggregate loses money, yeah?

Aravind Srinivas

没错。据我所知是这样。情况可能会变,但今天就是这样,这不是因为产品本身。实际上是因为前沿实验室通过订阅计划补贴 token。所以,即使 Claude Code 每月价值 200 美元,你在 Claude Code 上能消耗的 token 实际价值超过了每月支付的 200 美元。所以,他们实际上是在把它作为亏损引流产品来主导 token 收集,以便利用这些 token 让他们的模型变得更好。因此,如果你是一家与 Codex、Claude Code 和编码领域竞争的应用层公司,你很难有任何正的毛利率。

That's right. That's what I know. I could change but today that is the case and this is not because of the product. It's actually because of the Frontier Labs subsidizing tokens in the form of a subscription plan. So, even though Claude code is worth $200 a month, the amount of tokens you can consume on Claude code is actually worth more than the $200 a month you pay. So, they're actually running it as a loss leader to just dominate token collection in order to take all these tokens and make their models even better. So, if you are an application layer company competing with Codex and Claude code and coding it's pretty difficult for you to have any positive gross margins.

Host

Edwin,我想这部分是因为谁在编码方面消耗的 token 最多,谁就会拥有最好的模型,因为你会有最多的强化学习,所有来自开发者的使用都会为你的模型创造信号,对吧?

And Edwin, I would assume that part of that is whoever has the most tokens consumed specifically in coding will have the best model because you'll have the most reinforcement learning and all that usage from developers is going to create signal for your model, yeah?

Edwin

是的,我认为你可以学到很多有趣的信号,而且你知道,就像你推理得越多,你的模型推理得越多,这往往会导致更好的回答。

Yeah, I think there's a lot of interesting signals that you can learn and you know, just kind of like the more you reason, the more your models reason, like often times that just leads to better responses.

Host

这些其他公司能跟上 Claude Code、Cursor、Codex、GitHub 的 Copilot 吗?它们能跟上吗,还是你认为我们已经到了某种加速阶段,Claude 会一骑绝尘,Edwin?

Can these other companies keep up with Claude code, Cursor, Codex, GitHub's Copilot? Are they going to keep up or do you think we've hit this sort of acceleration where Claude's going to run away with it, Edwin?

Edwin

我不认为有什么内在因素阻止其他公司追赶。数据当然有价值,但我们几乎还处于所有这些进步的起点,仍然可以取得进展,所以我认为,是的,我认为其他人可以赶上。

I don't think there's anything inherently preventing any other companies from catching up. I certainly that data is valuable, but like we're almost still sort of at the beginning of all of this progress that can still be made that I think Yeah, I think someone else could catch up.

传统语言的数据标注 Data labeling for legacy languages

Host

你们是否在做 Fortran 和 COBOL 这类语言的数据标注?这是否是这些公司希望你们找到那些老前辈,让他们解释 AS/400 和 60、70 年代的微型计算机实际工作原理的一部分?这对你们来说是一门大生意吗?

Do you do data labeling for all this like Fortran and COBOL and is that part of the desire of these companies to get you to find these old gray beards to explain to you how these AS/400 and microcomputers from the 60s and 70s actually work? Is that a big business for you?

Aravind Srinivas

是的。我的意思是,编程领域非常庞大,我们必须参与其中的每一个方面。所以,是的,每一种语言都涉及。包括前端设计和后端设计,算法的正确性、效率,以及他们创建的前端设计的质量和美感。所以,这个领域如此广阔,我们必须参与一切。

Uh it is. I mean, the coding landscape is just so huge that we have to be part of every aspect of it. So, yeah, it's every single language. It is front end design and back end design. It is the correctness of the algorithms, the efficiency of the algorithms, but also the quality and the beauty of the front end designs that they create. So, it's just such a wide landscape that yeah, we have to be part of everything.

编程是有限游戏吗? Is coding a finite game?

Host

我们在编程方面是否已经进入终局?数据是有限的,Edwin。所以,在我看来,某个时候会出现收益递减。我们是完成了 96% 还是 99%?就像自动驾驶,显然已经完成了 98-99%,只剩下边缘情况。如果我们把它看作一个游戏,从完美解决的角度来看,你会如何描述这个游戏?国际象棋已经被完美解决了。他们认为无限注德州扑克几乎被完美解决了,而底池限注奥马哈,我们看看它最终是否会被完美解决。我认为它们正在路上。

Are we in the end game when it comes to coding? It is a finite set of data, Edwin. So, it would seem to me there'll be diminishing returns at some point. Are we 96% of the way there or 99% of the way there? Like self-driving is apparently 98-99% of the way there with the edge cases. How would you contextualize that game if we made it a game in terms of being perfectly solved? Chess got perfectly solved. They believe no-limit hold'em has been almost perfectly solved, and PLO, we'll see if that eventually becomes perfectly solved. I think they're on the way.

Aravind Srinivas

是的,我的意思是,我认为我们远未接近。我经常思考的一件事是,编程与下围棋非常不同,因为编程是完全开放的,对吧?你实际上可以创建世界上任何程序。它没有单一的解决方案或单一的终态。就像围棋游戏,最终是一方赢,另一方输。我经常想到的一个类比是,想象你带 Jeff Dean,给他一千年的时间来学习更多编程知识、探索世界,同时学习诗歌、数学、物理、历史、艺术等等。他能够将所有这些原则融入他所构建的东西中。软件是关于构建覆盖全球的基础设施,是关于设计火箭飞船。所以编程的能力几乎有一个无限的上限,我真的认为我们只完成了 1%。

Yeah, I mean, I don't think we're anywhere close. So, one of the things I often think about is coding is very different from playing Go in that coding is completely open-ended, right? You can literally create any program in the world. It doesn't have a single solution or a single end state. Like a game of Go sort of ends with one person winning and the other person losing. One analogy I often think about is imagine you took Jeff Dean and you gave Jeff Dean a thousand years to learn more about coding and to explore the world and to also learn about poetry and mathematics and physics and history and artistry and all of that. He'd be able to incorporate all of these principles into what he builds. Software is about building globe-spanning infrastructure. It's about designing rocket ships. So there's almost an infinite ceiling to what coding is capable of that I really think that we're just 1% of the way there.

Arvind对解决编程的看法 Arvind's take on solving coding

Host

Arvind,你对此怎么看?你认为我们正在接近解决编程游戏,每个人都能编写高质量代码,还是认为我们正在制造大量带有许多攻击向量的垃圾?我们显然已经讨论过 AI 帮助识别的各种攻击向量,但就解决编程游戏而言,你的看法是什么?

Where do you stand on it, Arvind? Do you think we're getting close to solving the game of coding and everyone will just be able to make quality code, or do you think we're producing a lot of slop with a lot of attack vectors? We've obviously covered the various attack vectors that AI is helping identify through this, but what's your take on where we are at in terms of solving the game of coding?

Aravind Srinivas

我认为框架应该围绕“解决”意味着什么来展开,对吧?所以,也许可以将其视为范式。比如 Cursor、GitHub Copilot,像自动补全。你试图补全几行代码。但你主要还是在写代码。Quad code、Codex,作为命令行界面,你几乎可以将其视为自动差异。你查看差异,新增的代码行和删除的现有代码行。你实际上不再自动补全了。你在一个不同的抽象层次上操作,即变更。下一个范式将是自动结果。你将查看结果。你甚至不会查看差异。你不会阅读任何一行代码。你将查看结果,然后要求更改,并不断迭代。这显然是下一步。然后 Elon 谈到过,我想他在 xAI 的一次全员大会上说过,就像在 God 流中,他说:“你将直接输出二进制。”你知道,你甚至不需要……

I think the framing should be around what does solving mean, right? So maybe think of this as paradigms. Like Cursor, GitHub Copilot, like autocomplete. You're trying to complete a few lines of code. But you are writing code largely. Quad code, Codex, as command line interfaces, you can almost think of it as auto diff. You're looking at the diff, the new lines of code added and the existing lines of code subtracted. You're not actually auto completing anymore. You're operating at a different abstraction of changes. The next paradigm is just going to be auto outcomes. You're going to look at the outcome. You're not even going to look at the diffs. You're not going to read any line of code. You're going to look at the outcome and then you're going to ask for changes and you're going to keep iterating. That's clearly the next thing. And then Elon talks about it, I think he talked about it in one of the xAI all-hands at God like stream where he said, "You're going to just output the binary." You know, like you don't even...

Host

这想起来很疯狂。

Which is a wild thing to think about.

Aravind Srinivas

是的。所以我仍然觉得我们在能力上还有一座山要爬。我认为在完全解决编程意味着什么方面,我们仍然非常早期。当然,解决问题的能力,比如 Edwin 提到的在不同事物之间建立联系的能力。我几乎想象,会是什么样子,Jeff Dean 或者像 Linus 这样的人,我听说也在使用 AI。那么这些人在编程什么?他们现在是如何做到以前做不到的事情的?我认为所有这些事情都非常值得思考,但根本上,如果 AI 能够真正将你提升到处理结果和二进制文件的水平,那么你对编程的看法也会从检查代码行等事情上改变。

Yeah. So I still feel we have a hill climbed on capability yet. I think we're still very early in what it means to solve coding entirely. Of course, problem-solving skills like the ability to connect dots across different things that Edwin was talking about. I almost imagine how would it be, what is Jeff Dean or someone like I heard Linus is also using AI. So what are these people coding? How did they do things these days that they were not able to do before? I think all these things are very interesting to think about, but also fundamentally if AIs can actually move you to the level of working at outcomes and binaries, then what you think of coding also changes from inspecting file lines of code and things like that.

初创公司与无代码运动 Startups and the no-code movement

Host

我认为这是最令人兴奋的部分。你知道,我常说初创公司就像你能最先看到这些趋势的地方。有点像圣莫尼卡的酸奶,抱歉,是瑜伽、新鲜食物和农贸市场。所有有趣的事情都始于圣莫尼卡和威尼斯的海皮士,然后向东传播,如果它们能成功走出那里的话。我觉得初创公司也是一样。现在的初创公司,两三个人就够了。他们永远不会添加他们原本以为会雇佣的第四、第五个员工。他们发布代码的速度比我见过的任何时期都快。他们使用像 Perplexity 的计算机这样的工具进行市场推广和客户获取,他们以惊人的速度生产代码,Edwin。我实际上认为我们会看到大量的就业创造,因为越来越多的人意识到,我不需要开发者就能创办公司。我们有过一次错误的开始。有人使用脚本,哦天哪,在 vibe coding 之前它叫什么?有一个术语,这些代码几乎像 wizzywig。天哪,它叫什么名字?你知道吗,Arvind?人们过去常做的那种,有点像 vibe coding?

I think that's the most exciting part. And you know, I always say startups are like where you can see these trends before anything else. It's kind of like Santa Monica with yogurt. I'm sorry, with yoga and like fresh food and farmers market. Like everything interesting starts with the hippies in Santa Monica and Venice and then goes east. If they wind up making it past there. And I feel like startups are the same thing. Startups now you'll have two or three people. They'll never add their fourth employee, the fifth employee that they thought they were going to add. And they're shipping code faster than I've ever seen. And they're doing their go-to-market and their customer acquisition using things like Perplexity's computer, and they're producing code at such an alarming rate, Edwin. I actually think we're going to see a significant amount of job creation as more people realize, I don't need a developer to start a company. And we had a false start. There were people using scripting and, oh god, what was it called before vibe coding? There was a term for these code, they were almost like wizzywig. Gosh, what was the name of it? Do you know, Arvind? That people used to do where they would kind of vibe code?

Aravind Srinivas

你是说无代码?

You mean no code?

Host

无代码。谢谢。无代码运动。那真是一个错误的开始,但我过去经常有人来加速器推销。我会问:“哦,这是谁建的?”他们说:“我建的。”我说:“好吧,随便吧。我以为你是 Salesforce 的销售,自己开了公司。”他们说:“是的,但我只是学会了如何使用无代码。”而现在无代码已经消失了,对吧?它被完全取代了。

No code. Thank you. The no code movement. Which was like such a false start, but I used to have people pitch the accelerator. I'd be like, "Oh, who built this?" They're like, "I did." I'm like, "Okay, whatever. I thought you were the sales person from Salesforce who started their own company." They're like, "Yeah, but I just figured out how to use no code." And like no code's just gone now, right? Like it's just totally replaced.

Aravind Srinivas

它在你能做的所有可能事情方面相当有限。因为它是用某种意图、某种程度的确定性行为、某种程度的硬编码构建的。

It's fairly limited in what it can do in terms of what are all the possible set of things you can do. Because it was built with certain intentionality, certain level of deterministic behavior, certain level of hard coding.

编程作为通用智能 Coding as General-Purpose Intelligence

Aravind Srinivas

所以显然它可以覆盖所有组合可能性,模型可以即时生成代码,做你要求的任何事情。这也回到了我认为 Edwin 之前提出的观点:编程的可能性空间是无限的,完全受限于你的想象力。比如你可以构建 Minecraft 中的事物,你在 Minecraft 中能建造的结构和世界是无穷无尽的。所以作为一款游戏,Minecraft 甚至比围棋更复杂。而 Minecraft 只是 AI 能编程的一个游戏,世界充满了无限可能。这就是为什么解决编程问题意味着你拥有了真正通用的人工智能。

So obviously it can cover all the combinatorial possibilities that models can just generate code on the fly and do whatever you ask them to do. This is also why it goes back to the point that I think Edwin made earlier, which is the space of possibilities in coding is endless. It's limited purely by your imagination. Like you can build things that exist inside Minecraft, the kind of structures and worlds that you can build inside Minecraft is endless. So as a game, Minecraft is even more complicated than Go is. And Minecraft is just one game that AI can code. And the world is full of infinite possibilities. So that's kind of why solving coding means you have something truly general purpose intelligence.

内部开发者工具:Codex vs Claude Code Internal Developer Tools: Codex vs Claude Code

Host

你鼓励你的开发者内部使用什么工具?使用程度如何?今年他们的生产力提升了多少?

What are you encouraging your developers to use internally and how much and how much more productive are they this year when you look at it?

Aravind Srinivas

嗯,大致有两个阵营:Codex 或 Claude Code。我一直在试图理解为什么有人偏好其中一个,而且偏好一直在变,但我可以分享我今天的大致理解:对于 Swift UI 和 Rust,人们喜欢 Codex,它似乎在那里表现更好。前端开发和全栈开发,人们喜欢 Claude Code。特别是如果你想做前端设计工作,Claude Code 似乎更好。这也印证了 Dario 最近在一个播客中提到的观点:模型正在商品化。实际上,模型正在专业化,即使是在编程这样的专业领域内,每个前沿实验室擅长的编程方面也不同。这也是为什么我们想构建像 Computer 这样的产品,因为当模型开始深度专业化时,协调它们各自擅长的领域就很有价值。所以我们主要在这两个阵营。从人员数量来看,我们从年初以来一直持平,在过去一年里,也就是从去年同一时间到现在,我们只增长了大约 30%。所以我想保持这种效率,我希望我们的公司能成为未来许多创始人的榜样,建立不到 500 人的公司却能创造数十亿美元的收入。我认为这是正确的方向,因为你想要那种乘数效应。你希望你的设计师写代码,你的业务人员做数据分析,你的销售代表自己制作演示文稿和客户数据分析,他们自己做 Bug 分类。你不需要项目经理或项目主管作为中间人。这会让你的公司更加垂直整合。

Yeah, so largely it's two camps: Codex or Claude Code. I've been trying to understand why one is preferred over the other and it keeps changing, but I can share with you a rough level of understanding I have today, which is: Swift UI and Rust, people like Codex. It seems to be better there. Front-end development and full-stack development, people like Claude Code. Especially if you want to have front-end design work done, Claude Code seems to be better. So this also goes to the point Dario made in one of the podcasts recently of this whole point of models commoditizing. Actually what's happening is models are specializing, and even within a specialty like coding, there are specializations on which aspects of coding each frontier lab is actually good at. Which is also why we wanted to build a product like computer, because when models start specializing deeply, an orchestration of all what each one can do individually at whatever they're skilled at is valuable. So yeah, we're largely in these two camps. Headcount-wise, we've remained flat since beginning of the year, and over a period of one year, that is exactly from last year same time to now, we have grown roughly just 30%. So I want to remain this efficient, and I want our company to be an example for many other founders in future to build sub-500 people companies that can make several billions in revenue. And I think that's the way to go, because you want that sort of force multiplier. You want your designers to write code. You want your business professionals to do data analysis. You want your sales reps to actually make their own presentations and decks and data analysis of the customer. You want them to do the bug triaging. You don't need a program manager or project manager to be an intermediary there. So it just vertically integrates your company even more.

Host

而且每个人都在增加技能。没错。我们过去 20 年一直听到的建议是:“嘿,选一个专业,成为专家,Edwin。”而现在,如果你是一个销售人员,你可以重新设计演示的登陆页面,如果你觉得有更好的想法,你可以直接“氛围编程”然后发出去。开发团队会说:“好吧,随便。”或者如果你是首席营收官,你不需要去找数据分析团队,你只需把电子表格丢进 Perplexity Computer,然后直接处理。你是怎么用的……你听到了那两个阵营。Arvind,简单总结一下,人员增长 30%,营收增长 5 倍。如果只考虑效率,这已经很显著了。Edwin,你们有 130 人,至少据我研究,大概这个数字。如果你们营收超过 10 亿,不需要天才也能看出你们现在的效率有多高。那么,你们如何看待效率和公司建设?

And everybody is adding skills. That's right. We lived for 20 years with the advice: 'Hey, pick a specialty. Be an expert, Edwin.' And now, hey, you know, if you're a salesperson and you can redesign the landing page for the demos and you think you have a better idea, you can just vibe code it and send it. And the dev team's like, 'Okay, whatever.' Or if you're the chief revenue officer, you don't have to go to the data analysis group, you can just dump your spreadsheets into Perplexity computer and just rip. How are you using... You heard the sort of two camps. And Arvind, just to put a pin in it, 30% headcount growth, 5x revenue growth. That's significant if you think just about efficiency. And Edwin, you have 130 people, at least in my research, somewhere around that number. And if you're over a billion in revenue, doesn't take a genius to figure out how efficient you are right now. So, how do you think about efficiency and company building?

效率与公司建设 Efficiency and Company Building

Edwin

是的,我完全同意 Arvind 的观点。我坚信,历史上公司一直缺乏尽可能快速扩张的动力。我认为人们总是低估了由此产生的官僚主义、政治斗争和沟通复杂性。比如,有谁喜欢……我觉得很少有人想管理一家 5000 人的公司。你不再投入于整天摆弄产品、与用户交流,而是整天做一个管理公司的企业 CEO。所以我完全同意我们的方向。我认为像 Claude Code 这样的工具,其中一个我喜欢的地方是:过去,如果我们的前端开发人员,甚至是运营团队的某个人,想为进来的专家做一个新界面或新登陆页面的原型,他们需要写下想法,发给设计师,等设计师画出草图,这可能需要几天时间,然后设计可能还不完全符合他们的要求。所以迭代周期很长。现在,你只需跟 Claude Code 对话,它会在 15 分钟、10 分钟、5 分钟内生成相当不错的东西,你可以迭代得更快。而且你能实际看到这个愿景是什么样子,也许亲眼看到后你并不喜欢。所以他们改变想法,完全转向另一个方向。而且 Claude Code 全天在线,不像我们的设计师每晚要睡 8 小时。所以我认为它让产品开发过程更快,同时对于运营团队的人或构建登陆页面的工程师来说,他们可以端到端地拥有一个东西,看到自己的愿景实现,而不是只委派部分工作。所以,我真的非常看好。

Yeah, so I absolutely agree with Arvind. I really strongly believe that historically there's been disincentive for companies to grow as much as possible, as quickly as possible. And I think people always underestimate the bureaucracy and the politics and the communication complexity that that creates. Like, does anybody like... I think very, very few people want to be running a 5,000 person company. You're no longer invested in... you're no longer spending your entire day playing with your product and talking to your users. You're just spending your entire day being a corporate CEO who's just managing the company. So I absolutely agree on our front. And I think that one of the things I love about things like Claude Code, for example, it used to be the case that if one of our front-end developers or even somebody on our operations team wanted to prototype a new interface or prototype a new landing page for these experts that come in, they would need to write down their ideas, send it to our designer, wait for a designer to sketch something up, and that may take a couple days, and then maybe the design didn't quite look like what they wanted. So there would just be this long iteration cycle. Now, you can just talk to Claude Code. It spits something out pretty amazing within 15 minutes, 10 minutes, 5 minutes, and you can just iterate so much faster. And you get to actually see what this vision looks like, and maybe you didn't like it now that you see it in person. So they change their idea, they just go somewhere completely different. And Claude Code is online all the time, unlike our designer who sleeps 8 hours a night. So I think it just makes the product development process both faster, but then also for the person on our operations team or for the engineer who's building this landing page, they just get to own something end-to-end. And to basically see their vision flushed out as opposed to just delegating parts of it. So yeah, I'm really really bullish.

Host

这个周末我参加了突破奖,你们那个小小的科学奖,我和神奇女侠 Gal Gadot 聊了聊。还有导演 Darren Aronofsky 也在,我们聊到了这如何影响好莱坞。然后你开始思考每个人曾经拥有的独特角色。有分镜师,有编剧,有像黑泽明或斯皮尔伯格这样的导演,他们会自己画分镜,或者像《异形》和《角斗士》的雷德利·斯科特,他以自己画分镜而闻名,他会画出所有这些有趣的图像,然后交给摄影师。

I was at the Breakthrough Prize this weekend, your little science prize, and I was talking to Wonder Woman, Gal Gadot, the actress. And we were just talking and there was a director there, Darren Aronofsky, and we're just talking about how it's impacting Hollywood as an example, and you start to think about the unique roles everybody had. There were people who were storyboard artists, there were people who wrote scripts, there were directors like Akira Kurosawa or Spielberg who would draw their own or Ridley Scott from Aliens and Gladiator. He was known for drawing his own cells and he would draw all of these interesting images that he would then give to a cinematographer.

电影制作中的AI与跨学科创新 AI in filmmaking and cross-disciplinary innovation

Host

现在有了 AI,写剧本的人、制片人,他们都聚在一起,任何人都几乎能做任何事。写对话、做背景、编写所有那些分镜脚本。这种跨学科的特性带来了创新。Arvind,如果你见过某个在多个领域都有专长的人,比如计算机科学和艺术,或者艺术和分镜,他们就能取得别人无法取得的突破。她告诉我有一部电影即将上映,叫《Bitcoin Killing Satoshi》。预算只有 7000 万美元,本来要花 2 亿,但他们只是让演员在灰色屏幕前表演,然后所有背景都由 AI 构建。所以他们只需要写好剧本,让最好的演员表演,然后在摄影棚里拍 20 天就能完成电影。想想这个行业会发生什么。现在你可以用一部电影的成本拍三部电影。这真是个有趣的时刻,对吧?

Now with AI, the people who are writing screenplays, the producers, they're all coming together and anybody can do almost anything. Write dialogue, do the backgrounds, and write all these cells to do them. The cross-disciplinary nature of that leads to innovation. Arvind, if you've ever met somebody who had expertise in multiple areas, whether it was computer science and art or art and cells, they can just make breakthroughs that other people don't have. And she was telling me there's a movie coming out, Bitcoin Killing Satoshi. It's only a $70 million budget. It would have cost $200 million, but they're just doing the actors on a gray screen and then everything is being built by AI in the background. So all they had to do was write a great script, have the best actors perform it, and now they can just build the movie with them on a sound stage for 20 days. Think about what happens in that industry. Now you can make three movies for the cost of one. Really interesting moment in time, yeah?

Aravind Srinivas

是的。我的意思是,乔治·卢卡斯谈论过故事。我认为他给了一个教训:制作电影的大部分工作都在前期制作阶段,比如把故事弄对,把配乐、故事和叙事弄对。他在皮克斯最大的收获就是:他会和团队坐在一起,尝试通读整个故事板,如果说不通,就重做。在故事足够好之前,我们不会拍电影。这是他从沃尔特·迪士尼那里学到的唯一教训:无论你把电影制作得多好,一个糟糕的故事都不可能成功。但即使你在制作质量上做得不够好,一个好故事也会胜出。

Yeah. I mean, George Lucas talks about the story. I think he gave a lesson on how most of the work in producing a movie is all about the pre-production phase, like getting the story right, getting the score, story, and storytelling right. That was his biggest learning at Pixar: he would sit with the team, try to go through the whole storyboard, and if it didn't make sense, go redo it. We're not going to make the movie until it's so good. And this was the single lesson he learned from Walt Disney: you cannot make a bad story succeed no matter how good you produce the actual movie. But even if you don't do a great job with the production values, a good story will win.

Host

独立电影也是如此,对吧?你可以看到一些令人惊叹的独立电影,虽然有些粗糙,但出色的表演基于一个好故事,基于好的对话。你把这些都做对了,那就是……

That's true with independent films, right? You can see some incredible independent film where it's a little rough around the edges, but the great performance is based on a great story, based on great dialogue. You get all that right, and that is the...

Aravind Srinivas

产品也是如此。一个简单的产品往往能成功,只要你把一两个想法做到极致。

That's true in products, too. Like a simple product often works if you just hit one or two ideas out of the park.

Host

顺便说一句,我认为你在 Comet 上做到了这一点。你是第一个推出浏览器的人,那是我第一次和你交流。我当时想,这个浏览器太不可思议了,我让所有人都用上了。我不知道你是否用过它,但它是第一个让你可以这样做的:嘿,这是我页面上的内容,我们一起来处理。这是第一次让 ChatGPT 或你使用的任何模型在真实的网络上运行。天哪,那是一个重大突破。现在,显然通过钩子和集成,它正在进入下一个层次。我想谈谈大语言模型的商品化以及小语言模型(SLM,我想是行业术语)或垂直化小语言模型(VSLM)的创建。我认为 Arvind,你相信我们开始进入某种形式的商品化。我想知道价值会流向哪里。是流向 harness(框架)、wrapper(封装器),还是流向核心模型?所以,也许你可以解释一下你对未来一两年内人们加载 Kimi 或 DeepSeek 甚至不知道自己在用哪个模型的情况的最佳估计。然后解释什么是 harness,以及人们应该如何思考 harness 及其将产生的影响。

I think you did that with Comet, by the way. You were the first to drop a browser, and that's when I first started communicating with you. I was like, this browser is unbelievable, and I got everybody in. I don't know if you've used it before, but it was the first where you could be like, hey, here's what's on my page, let's work with that. And it was the first time you let ChatGPT or whatever model you're using out on the real web. And man, that was a major breakthrough. Now, obviously with hooks and integrations, it's getting to the next level. I wanted to talk a little bit about the commoditization of large language models and the creation of small language models, SLMs, I guess is the industry term, or VSLMs, verticalized ones. And I think Arvind, you have the belief that we're starting to hit some form of commodification. I'm wondering where the value is going to accrue. Is it going to accrue to the harness, to the wrapper? Is it going to accrue to the core model? So, maybe you could explain your best estimation of what's going to happen in the next year or two in terms of people loading Kimi or DeepSeek or not even knowing which model they're using. And then what a harness is and how people should think about harnesses and the impact they're going to have.

Aravind Srinivas

人们不买模型,他们买产品,对吧?从根本上说,最终消费者要为我们的服务付费。纯模型公司基本上已经不存在了。Anthropic 在应用层的投入和模型层一样多。无论报道怎么说,我忘了,但 30% 到 40%,至少 30% 的收入来自应用。所以这表明,无论你是否构建模型,你都必须成为应用层的参与者。钱在应用里。如果你有模型,显然你可以将其与应用垂直整合。你可以为你的模型构建自定义 harness,并声称你的模型在你的 harness 上表现良好。这是你的优势,但劣势是你必须始终确保自己拥有最好的模型。这是激烈的竞争。这是一个你只有拥有至少数百亿美元现金用于算力才能参与的游戏,而且你必须提前几年确保算力容量,现在还要确保电力容量,然后超大规模云服务商需要投资你。这是一个只有最高层玩家才能参与的游戏。这也是为什么只有四五个玩家在玩。

People don't buy models, they buy products, right? Fundamentally, at the end of the day, the consumers have to pay for our services. Pure model companies basically don't exist anymore. Anthropic is as much playing in the application layer as they are playing the model layer. Whatever the information reported, I forgot, but 30 to 40%, at least 30% of the revenue is coming from applications. So that shows you that you have to be an application layer player whether you build models or not. And the money is in the applications. If you have a model, obviously you can vertically integrate it with the application. And you can build custom harnesses for your models and claim your models are good at your own harness. So that's an advantage you have, but the disadvantage is you have to always ensure you have the best model all the time. That's serious competition. It's a game you can only play if you have at least tens of billions of dollars in cash to spend on compute, and you have to secure compute capacity and compete with all the other players trying to secure compute capacity years in advance, power capacity now, and then hyperscalers need to be invested in you. It's a game that you only play at the highest level. That's also why there are like four or five players playing there.

Host

也许未来会更少。我们可能会看到一些整合,或者有些人如果觉得自己无法竞争就会退出这个行业。Edwin,你认为这会如何收场?你显然在帮助人们训练他们的模型。这些是你的客户,所以你支持他们。你帮助他们构建。但也有人说,价值确实会流向应用层,无论是谷歌的产品套件及其浏览器,还是苹果的产品套件,就像我们在节目第一个话题中讨论的那样。Perplexity 电脑、Claude、co-work、open claw,所有这些不同的前端、harness、编排层正变得越来越重要。那么,你自己如何看待编排层?大语言模型会被商品化吗?

And maybe less in the future. Maybe we'll see some of them consolidate or maybe some people get out of that business if they don't feel they can compete. Edwin, where do you think this winds up? You obviously are helping people train their models. These are your customers, so you're rooting for them. You're helping them build. But there's also people saying it really is the application layer where the value is going to accrue, whether it's Google's suite of products and their browser, Apple's suite of products as we talked about in the first topic of the show. Perplexity computer, Claude, co-work, open claw, all of these different front ends, harnesses, the orchestration level is becoming more important. So, how do you think about the orchestration level yourself? And are the LLMs going to get commodified?

Aravind Srinivas

我仍然不认为 AI 模型本身会被商品化。我认为很大一部分原因是我经常思考它们的个性或 Edwin 之前提到的专业化。例如,即使我问一个相当简单的问题,比如,我不知道,亚伯拉罕·林肯是谁?我会从 ChatGPT 和 Claude 那里得到非常不同的回答,或者 ChatGPT 和 Gemini。就像我有一群朋友,有时我会问他们某些问题,即使他们都有相同的智力水平,即使他们都有相同的学位、相同的知识水平。

I still don't think that AI models themselves are going to get commodified. And I think a big part of that is because I think so often about their personalities or the specialization that Edwin mentioned earlier. For example, even if I ask a fairly simple question like, I don't know, who was Abraham Lincoln? I would get a very different response from ChatGPT versus Claude, for example, or ChatGPT versus Gemini. In the same way that I have a bunch of friends, and sometimes I will ask them certain questions, even if they all have the same level of intelligence, even if they all have the same degree, same level of knowledge.

AI模型的商品化与专业化 Commoditization vs. Specialization of AI Models

Host

有时候我只想从朋友那里得到一个快速干脆的回答。某些心情下,我更喜欢和他们聊天。有时候我想要那种研究充分、见解深刻的东西,但我知道从朋友那里得到答案需要 5 分钟,而有时我太忙了没空跟他们聊。同样地,我觉得即使对于相当类似的任务,比如前端编码 vs. 后端编码,或者不同的语言,人们自然也会根据心情想跟不同的模型对话,哪怕主题相同。所以,我真的不认为 AI 模型会被商品化。

Sometimes I just want the quick snappy answer from one of my friends. I just enjoy talking to them more when I'm in a certain mood. Sometimes I want the really well-researched, really insightful thing. But I know it's going to take me 5 minutes to get the answer from my friend, and sometimes I'm just too busy to talk to them. In this same way, I just feel like people will, even for fairly similar tasks like coding, like front-end coding versus back-end coding, or different languages, people will naturally sometimes just want to talk to different models depending on your mood, even for the same topic. So, I really don't think that the AI models will get commoditized.

Host

我想知道,Arvind,你有没有另一面的看法?我的很多工作都在 Perplexity、计算机、Open Claw 里进行,我觉得我们需要把这些技能、这个灵魂文件、这些记忆文件放在本地硬盘上,在使用任何语言模型之前,我们先说:“嘿,这是上下文。这是我喜欢的做事方式。这是我喜欢的答案风格。”我现在不得不对四五个不同的模型解释:“我喜欢简洁的答案。我喜欢你直接解决问题,不要告诉我你的思考过程。”比如 Open Claw 最近的最新版本变得非常啰嗦,我简直想死。它总是:“好的,用户想让我做这个。好的,我要做这个。好的,我要做这个。”我就想:“不,不,你只要在牛排完美烹饪好的时候给我就行。我不想你解释烹饪牛排的所有步骤。”那么,你认为价值会开始在哪里积累?

I'm wondering, Arvind, if you have the other side of it where you know, so much of my work is happening in Perplexity, computer, Open Claw, and I'm like, we need to have these skills, this soul file, these memory files local on our hard drives, and before we use any language model, we're like, 'Hey, here's the context. This is how I like to work. This is how I like my answers.' And I've had to now with four or five different models explain to it, 'I like concise answers. I like you to just solve the problem, not give me updates on your thinking.' Like Open Claw had become so verbose recently in the latest version, I wanted to kill myself. It was like, 'Okay, the user wants me to do this. Okay, I'm going to do this. Okay, I'm going to do this.' I was like, 'No, no, you just give me the steak when it's perfectly cooked. I don't want you to explain to me all the steps in cooking the steak.' So, what do you think about where the value will start to accrue?

Aravind Srinivas

我相信价值在于应用层。有很多模型公司,其中一些倒闭的主要原因也是他们无法构建应用程序。纯 API 模式行不通,因为你无法构建一个远超其他模型的模型。没有人能保持那么大的领先优势。模型层唯一一次出现显著领先是在 GPT-4 存在的时候,其他人花了一年才赶上。

I believe that the value is in the application layer. There were a lot of model companies, and one of the main reasons some of them went out of business is also that they couldn't build an application. The pure API model doesn't work because you cannot build a model that's so far so much better than the rest. No one's able to maintain that much of a significant lead. The only time there was a significant lead in the model layer was when GPT-4 existed, and it took a year for anybody else to catch up.

Host

嗯。

Mhm.

Aravind Srinivas

在那之后,差距通常只有几个月,我会说。即使是开源模型和前沿模型之间,我认为差距目前大约是 6 个月到一年。

After that, the gap has usually been months, I would say. Even between open source and frontier models, I think the gap is like 6 months to a year at this point.

Host

你也是这么觉得吗,6 个月?你也有同感吗,Edwin?你觉得开源到前沿的差距是多少?

Is that what you feel, 6 months for you? You feel the same way, Edwin? What would you say the gap is, open source to frontier?

Edwin

是的,就原始模型智能或原始 API 而言,我同意。

Yes, in terms of raw model intelligence or raw API, yeah, I agree.

Aravind Srinivas

所以,我认为商品化和专业化并不一定是相互排斥的。有些模型明显在走向专业化。比如云模型显然非常擅长智能体编码、代码执行、智能体编排,而 OpenAI 模型和 Google 模型非常擅长多模态内容,因为他们拥有大量别人没有的多模态数据。Elon 的 Grok 模型则非常擅长无过滤、无约束。

So, I think commoditization and specialization are not necessarily mutually exclusive. There are some models that are getting specialized clearly. Like cloud models are clearly very good at agent coding, code execution, agent orchestration, and OpenAI models and Google's models are very good at multimodal stuff because they have a lot of data and multimodality that nobody else has. Elon's Grok models are very good at being unfiltered and unconstrained.

Host

不羁。

Unhinged.

Aravind Srinivas

不羁,是的。

Unhinged, yeah.

Host

对。

Yeah.

Aravind Srinivas

这很大程度上说明了模型的形态以及你如何塑造模型的价值观,比如你用什么数据训练它?它假设为真或至少被训练为真的基本事实是什么?所以这些都不是商品。这些模型的行为方式以及它们擅长的东西都不是商品。商品化的是如果它们都在 Eleuther AI Arena 或 terminal bench、agent extreme、humanities last exam 等学术基准上爬山。如果所有模型都在这些基准上爬山,因为那是你向研究人员展示你处于前沿的东西,那部分就是商品,因为开源也在做同样的事。所以,一些品质将是专业化的。很多学术基准将是商品化的,而模型训练者、产品构建者、应用层所有者需要把商品化的东西塑造成对他们拥有的用例有意义的方式。你只有真正拥有一定的工作流程,并且拥有大量忠诚客户、高留存客户,并拥有一些工作流程,才能生存,才能让价值积累到你身上,因为这是你收集独特 token、独特数据的唯一方式,只有你能利用这些数据并持续改进那些能力。

And that speaks a lot to the shapes and how you shape the values of the model, like what do you train it on? What are the fundamental ground truths that it assumes as true or at least has been trained for it? So that's all not commodity. These characteristics of how these models behave and what they're good at are not commodity. What is commodity is if they're all hill climbing on Eleuther AI Arena or like terminal bench or agent extreme or humanities last exam. These are all academic benchmarks. If all of them are hill climbing on these benchmarks because that's the stuff you publish to researchers to show you're at the frontier, that part is commodity because open source is also doing that. So, some qualities will be specializations. A lot of academic benchmarks will be commodity and it'll be up to the model trainer, the product builder, the application layer owner to take what is commodity and shape it in a way that matters for the use cases that they own. And you can only survive, you can only have value accruing to you if you actually own a certain bunch of workflows and then you have a lot of loyal customers, high retaining customers, and own a bunch of workflows because that's the only way that you collect unique tokens, unique data that you alone can harness and keep improving on those capabilities.

Host

你怎么看待这些 LM Arena 之类的基准测试、humanity's last test,Edwin?因为你显然在帮助人们进行训练。你是其中的关键人物。你对 LM Arena 和人们如今针对这些基准进行优化有什么看法?

How do you think about these LM arenas of the world, the benchmarks, humanity's last test, Edwin, cuz you're helping folks with training, obviously. You're a key player in this. What's your take on LM arena and people optimizing for these benchmarks today?

Edwin

我真的认为 LM Arena 是 AI 领域的一个可怕毒瘤。你基本上有一个随机的、小众的子群体。人们没有意识到它有多小众,还以为它是一个随机的代表性用户群体。但事实并非如此,它只是一个随机的、小众的子群体,他们想要免费访问模型,并且有无限的时间等待 LM Arena 无休止地旋转,然后他们扫一眼回答 2 秒钟,就点击他们喜欢的那个。所以 LM Arena 的结果就是,那些完全胡编乱造的模型,只要它们有一堆漂亮的格式能吸引这个小众群体的眼球,就能击败那些回答正确的模型。但问题在于它是一个如此显眼的基准。业内每个人都知道它。它是一个如此显眼的基准,以至于所有公司的副总裁、CEO 们基本上都有整个团队专门致力于破解它。业内众所周知,一旦你有一个数据科学团队分析这个小众群体的奇怪偏好,你就可以轻松破解它。所以公司们这么做,尽管研究人员自己也同意这只会让他们的模型变得更差。就个人而言,我原本希望去年 Meta 展示它有多容易被破解后,它会彻底消失,但不知何故它仍然存在。

I really do think that LM arena is just this terrible cancer on AI. You basically have a random niche subset of population. People don't realize that it's so niche and they think that it's a random representative set of users. But no, it's like a random niche subset of population that just wants free access to models and they have endless time to wait for LM arena to spin endlessly before they glance at the responses for 2 seconds and then they click their favorite. So what basically happens with LM arena is that you get models that completely hallucinate and they beat out models that answer correctly as long as they have a bunch of pretty formatting that catches the eye of this random niche subset of population. But the problem with it is that it's such a visible benchmark. Everybody knows about it in the industry. It's such a visible benchmark that you have all these VPs, all these CEOs at all these companies that basically have entire teams purely dedicated to hacking it. It's pretty well known within the industry that once you have a data science team analyzing the kind of weird idiosyncratic preferences of this niche population, you can just hack it. And so companies do it even though the researchers themselves agree that it simply makes their models worse. So, personally, I was really hoping that it would completely die out after Meta showed last year how easy it is to hack, but somehow it still exists.

基准测试与人工评估 Benchmarking and Human Evaluation

Host

嗯,这让我想到沃伦·巴菲特说的:‘给我看激励,我就给你看结果。’或者古德哈特定律:‘当一个指标成为目标时,它就不再是一个好指标,因为每个人都会开始针对基准进行优化。’对吧,Edwin?

Well, and this speaks to I guess Warren Buffett would say, 'Show me an incentive, I'll show you an outcome.' Or Goodhart's law, 'When a measure becomes a target, it ceases to be a good measure because everybody starts optimizing for benchmarks.' Yeah, Edwin?

Aravind Srinivas

对,完全正确。

Yeah, exactly.

Host

那我们该衡量什么呢?应该如何对行业和正在构建的模型进行基准测试?有没有更好的方法?

And what should we be measuring then? How should we benchmark the industry and the models that are being built? Is there a better way to do this?

Aravind Srinivas

所以,我认为真正的方法是要思考真实人类如何在现实生活中使用这些模型。例如,我们经常做的是运行人工评估,拿 Claude Code 或 Gemini 这样的模型,然后让软件工程师自己去用:‘在真实世界里用用看,用在你的日常工作中。’然后,让他们提问,衡量模型是否真的帮到了他们。所以,不仅要看答案是否正确,不仅要看它能否通过一组单元测试,还要看它为你创建的网页是否是你真正想发布给用户的?它是否为你相信的指标、为你的 A/B 测试提供了很好的建议?而不是玩这种基准测试游戏——很多基准测试,我觉得人们没意识到,它们非常刻意。提示词本身是真实用户永远不会问的东西。衡量基准的方式,因为通常是自动评估,纯粹看‘是否匹配某个字符串?’而在真实世界里,我们追求的不是这个。在真实世界里,你在乎的是回答的创意、网页的设计等等。

So, I think the way to really do it is to think about how real humans are using these models in real life. So, for example, a lot of what we do is we simply run these human evaluations, where we take models like Claude Code or Gemini, and we actually ask the software engineers themselves, 'Go use this in the real world. Go use it for your actual day-to-day job.' Then, ask it your queries, and measure whether or not it actually helped you. So, not only was it correct, not only would it pass this set of unit tests, but was the web page that I created for you something you would actually want to launch to your users? Did it make great recommendations for your A/B test, for the metrics you're trying to optimize for that you actually believe in? And not just playing this sort of benchmark game, where a lot of these benchmarks, I think people don't realize, they're just very contrived. So, the prompts themselves are things that no real user would ever ask. The way you measure the benchmarks, because they're often auto-evaluated, they're purely measured on, 'Did it match a certain string?' And in the real world, that's not what we're looking for. In the real world, you care about things like the creativity of responses. You care about the design of the web page, and so on.

Host

这——

It's—

Aravind Srinivas

对。

Yeah.

Host

如果你想想谷歌的做法,他们曾经衡量‘回弹率’:比如有人搜索‘嘿,尼克斯队今天几点比赛?’他们点进一个网页。他们会回来点击第二、第三个结果吗?如果会,说明你没给对结果。后来,我记得 20 年前和拉里·佩奇聊过,他说:‘最终,Jason,我们会直接告诉你尼克斯比赛几点开始。最终会这样。电脑会直接知道。’所以,这是个很奇怪的想法。谷歌的整个存在意义是‘不要为了那个查询再回到网站。别回来。我们多快能让你离开我们的网站?’而其他人,比如迪士尼、ESPN,他们说:‘嘿,当我们把你引到网站后,我们能让你在网站上待多久?我们能一直吸引你吗?’Meta,显然,通过 Instagram 和 Facebook,我们能让你会话持续多久?YouTube,我们能让你会话持续多久?两个非常不同的北极星,对吧?

If you think about it like Google did, Google was measuring bounce back rate at some point where it was like somebody searches for 'Hey, what time is the Knicks game today?' They go to a webpage. Do they come back and click on the second and third result? If they did, you didn't give them the right result. And then eventually, I remember talking to Larry about this 20 years ago, Larry Page, and he was like, 'Eventually, Jason, we're just going to tell you what time the Knicks game starts. That's eventually what's going to happen. The computer will just know.' And so, that's a very weird thing to think about. Google's whole existence was 'Don't come back to the website for that query. Don't come back. And how quickly can we get you off of our website?' Whereas other people, Disney Corporation, ESPN, they were saying, 'Hey, when we get you to the website, how long can we keep you on the website? Can we keep enticing you?' Meta, obviously, with Instagram and Facebook, how long can we keep your session going? YouTube, how long can we keep our session going? Two very different North Stars, yeah?

Aravind Srinivas

是的,我认为优化人类真正想要的、真实用户想要的、能让生活更好的东西,与优化点击和参与度之间是有区别的。

Yeah, I think there's a difference between optimizing for what the humans want, what the real users want, and what makes their lives better as opposed to optimizing for clicks and engagement.

Host

嗯。

Mhm.

Host

对。AI 最好的状态,Arvind,就是直接解决你的问题。你们会追踪这个吗?比如,用 Perplexity 我有没有解决问题?用 Model Council 我有没有解决问题?如果人们不知道 Model Council,你可以稍微解释一下。对我来说,我认为最终的测试是:我是否继续向你提问,并且对答案满意?对吧?

Yeah. And AI should be, at its best, Arvind, of just solving your problem. And do you track that? Like, did I solve the problem or not with Perplexity? Did I solve the problem or not with the model council? If people don't know model council, you can explain it a bit. Like, that's to me, I think, the ultimate test is do I keep querying you and am I happy with the answer? Yeah?

Aravind Srinivas

是的。所以,关于拉里·佩奇的说法,他说最终 AI——电脑应该直接告诉你比赛几点开始,你不需要点击一堆链接。这正是我们构建 Perplexity 的原因。我们解决的就是那个问题:链接到答案。所以,是的,我们确实追踪这个。例如,有一个非常简单的启发式方法:如果用户问了一个问题,然后后续问题像是‘但是不,我的意思是……’,那就说明第一个问题没回答好。最初,我们确实只用启发式方法,因为每天有大量查询,很难对所有查询运行大量 LLM 算力并过滤。但现在我们不在乎了——我们有小型语言模型,可以在大量查询日志上运行,过滤出那些用户明显需要再次澄清提示才能得到更好答案的线程,这意味着你在第一个提示时没有很好地理解用户意图。这又是拉里·佩奇的哲学:即使用户的提示不够详细,你的工作仍然是给出一个好答案。你应该把用户提示视为意图,而不是实际的描述性提示。这与 ChatGPT 的产品设计理念非常不同——在 ChatGPT 中,至少一开始,他们经常发推文说:‘如果你不知道如何说明这个模型为什么比上一个模型好,那你就不够有品味。’不,你不应该需要这样——模型应该自己说话。用户应该直接感受到。你不应该把责任归咎于他们的提示能力。所以,我们采用了谷歌的哲学:用户永远没错——即使他们的提示很糟糕,即使他们的提示措辞不当,AI 有责任去消歧、理解、重新表述、扩展提示、尽可能搜索,给出用户想要的信息,最后再问一个澄清问题,看他们是否满意或还需要别的。

Yeah. So, I mean, to the Larry Page thing, he said eventually the AI—the computer should tell you what time the game starts. You don't have to click on a bunch of links. And that's precisely why we built Perplexity. That was the problem we solved there: links to answers. So, yeah, we do track that. For example, there's a very simple heuristic. If a user asks a question, and then there's a follow-up question that's like, 'But no, I meant...' that means you didn't do a good job with the first question. Initially, we did start just using heuristics because every day we get a lot of queries, so it's hard to run a lot of LLM compute on all of them and filter them. But now we don't care—we have small language models that can just run on a lot of query logs and filter threads where the user clearly had to clarify the prompt again in order to get a better answer, which means you did not understand the user intent well in the first prompt itself. And this is another Larry Page philosophy thing: even if the user's prompt wasn't detailed enough, your job is to still give them a good answer. You should consider a user prompt as intent, not the actual descriptive prompt. This is a very different product design philosophy from ChatGPT where, in ChatGPT, I think at least in the beginning, they used to tweet stuff like, 'You're not a high-taste tester enough if you don't know to tell why this model was better than the previous model.' No, you shouldn't need to be—the model should speak for itself. Users should just feel it. And you shouldn't blame it on their prompting capabilities. So, we took the Google philosophy of the user is never wrong: even if their prompt was bad, even if their prompt was incorrectly phrased, it's on the AI to disambiguate and understand and reformulate it and expand the prompt and search as much as possible and give as much information the user wants, and ask a clarifying question at the end if they're happy with that or they're looking for something else.

Host

这似乎是 Claude 4.6 和现在的 4.7 的重大创新。对吧,Edwin?当你问一个非常简单的问题时,它会展开思考:‘嗯,你的接下来五个问题是什么?’或者‘你真正想表达的是什么?’然后它试图合理化,给你一个比你想象中更全面的答案,对吧?

It seems like this was the big innovation with Claude's 4.6 and now 4.7. Yeah, Edwin, when you ask it a very simple question, it kind of threads out and thinks, 'Well, what are your next five questions?' Or what did you really mean by that? And it tries to rationalize it and give you an answer that's much more comprehensive than you could ever have imagined, yeah?

Aravind Srinivas

是的,我认为 Claude 在规划阶段一直非常擅长:它必须提前制定计划,然后执行,然后在出错时回溯。所以,是的,我认为它就是这么做的。

Yeah, I mean, I think Claude has always been very, very good at the planning stage where it has to formulate this plan up in advance and then execute it and then, yeah, like kind of backtrack back whenever it's going wrong. So, yeah, I think that's what it does.

Host

我最喜欢的工具是 Model Council。我超级上瘾。我不知道,Edwin,你用吗?你用过 Perplexity 的 Model Council 吗?或者你内部有自己的基准测试工具?你用的是什么?你之前用过吗?对于把模型放在一起对比、了解差异,有什么想法吗?

My favorite tool is Model Council. I'm super addicted to it. I don't know, do you use it, Edwin? Do you use Perplexity's Model Council ever? Or do you have your own for your own benchmarking, I guess, internally, but what do you use? Have you used it before and any thoughts on putting the models up against each other and knowing the diffs, yeah.

模型委员会功能 Model Council feature

Host

所以我一直在做这件事,因为我们的日常工作,或者很多专家的工作,就是不断比较模型。我个人就经常做,坦白说,我开了多个窗口。我自己建了一个专门的应用程序来做这件事。我们的专家有另一种应用,但埃德温,你得给我演示一下 Model Council。

So, I am constantly doing that myself just because it's almost like what a lot of what we do in our day-to-day work or like a lot of our experts do, they're just constantly comparing the models. So, I personally do it a ton. I admit to you multiple windows open. So, I have a special app that I built to do it myself. And then our experts have a different kind of app, but yeah, Edwin, we'll have to give me a demo of Model Council.

Host

Model Council 太棒了。你输入查询,它会分发给你选的三四个模型,然后告诉你:“这是它们有分歧的地方,这是你应该信任的模型,对吧?”埃德温,现在这个功能有多受欢迎?

Model Council is spectacular. You give it your query, threads it out to whatever three or four models you want, then it'll tell you, 'Here's where they disagree and here's who you should trust, yeah?' And how popular is this now, Edwin?

Aravind Srinivas

它在 Max 用户中相当受欢迎。实际上,是 Jensen 让我开发这个功能的。Jensen 说他特别喜欢问不同模型同一个问题。他以前的做法是,问 Perplexity 一个问题,问 Claude 一个问题,问 ChatGPT 一个问题,再问 Gemini 一个问题,然后看所有 AI 的回答,在脑子里比较。我就想,Perplexity 已经把所有这些模型整合在一个应用里了,也许我们可以在应用内部实现这个功能,这样你就不用打开四个应用,读所有回答,再找出差异。所以我们原生构建了这个功能,而且我们是唯一能做到这一点的应用层公司,因为其他应用层公司有动力只推广自己的模型。我们开发出来后,他非常高兴,我们就发布了。显然这很贵,因为每个问题要问四五个不同的模型,而且不仅仅是汇总答案,编排器会分析,告诉你它们在哪里有分歧、哪里一致,以及你应该真正采纳什么。所以还有一个用前沿模型运行的合成层。我经常用它来处理健康查询,因为健康查询的证据在网上,但模型根据你输入的提示来解释证据的方式往往不同。有些模型非常风险规避,有些则倾向于冒险推荐。所以你希望看到两种行为模式,再加上分析。所以我觉得这非常有用。我喜欢问 Model Council 对不同股票的看法,比如“你怎么看特斯拉?特斯拉未来 10 年能值 10 万亿美元吗?”或者“按今天的估值,七巨头中哪只股票未来 10 年能涨 10 倍?”我喜欢同时问所有这些不同的模型,我觉得 Council 是个很酷的功能。而且 Council 只是我们计算机中的一个技能,所以你可以在计算机上使用它。

It's pretty popular among the max users. So, actually, you know, Jensen asked me to build that feature. So, Jensen said he really loves asking different models the same thing. And the way he said he would do it is he would ask Perplexity one question, he would ask Claude one question, he would ask ChatGPT one question, and Gemini. And would look at all the AIs and say see what each of them say and then compare in his head. I was like hey like Perplexity has all these models in one app. Maybe we can just do that within the app itself so that you don't have to open four apps and read all of them and then figure out what the differences are. And so what if we built that feature natively and we'll be the only app layer company that could do this because other app layer companies have an incentive to just put their own model. And so we built that and he was really happy about it and we rolled it out. Obviously it's expensive like you're going to ask four or five different models each question and it's not just like aggregating the answers like the orchestrator looks at it, tells you exactly where they differ, where they agree and what you should truly take away. And so there's some synthesis layer that's actually running with the frontier model too. And I like to use it a lot for health queries because health queries are very like the evidence is there on the internet but how the models interpret that evidence in terms of the prompt you asked often differs. Some models are very risk conscious, some models are actually risk seeking in terms of what they recommend. And so you want both modes of behavior there and then an analysis. So I think that's super useful. I like asking Model Council about what it thinks about different stocks. Like what do you think of Tesla? Can Tesla be worth $10 trillion in the next 10 years? Or like which Magnificent Seven stock actually could be worth 10x in the next 10 years for Verity's today. And I like asking all these different models simultaneously and I think Council is like a cool feature. And Council is just a skill inside our computer too. So you can use it in computers.

Host

哦,真的吗?我都没意识到。这就是我接下来需要的。我需要一个 AI 在我工作时坐在我旁边。我想这就是 Microsoft Copilot 本该做到的事,它会告诉我:“嘿,傻瓜,那个有快捷键。嘿,傻瓜,Perplexity 计算机内置了那个功能。”我需要一个像 Clippy 那样的东西,当我在电脑上做某事时,它会告诉我我是个白痴,有更快的方法。这在现在会非常有帮助。

Oh, really? Oh, that's I didn't realize that. I got to And that's the next thing I need. I need an AI to sit next to me while I'm working. I guess this is what Microsoft Copilot was supposed to be and just tell me, 'Hey, dummy, there's a quick key for that. Hey, dummy, Perplexity computer has that built in.' I need like a Clippy that just tells me when I'm doing something on my computer that I'm an idiot and there's a faster way to do it. That would be pretty helpful at this point in time.

Aravind Srinivas

我觉得这会是苹果最适合推出的功能,因为你需要完全访问你的屏幕,而你不会信任服务器端的 AI 公司来做这件事。

I think that's going to be a feature that I feel like Apple's best position to ship because you need full access to your screen and you're not going to trust the server-side AI company to do that.

Host

你有没有想过构建一个操作系统?我知道这听起来很疯狂。你想过吗?因为你构建了一个浏览器和一台计算机,Perplexity 计算机和 Comet 浏览器有什么区别?Chrome 变成了一个操作系统。你有没有想过构建一个人们可以启动的操作系统?

You ever think about building an operating system? I know this sounds insane. Have you thought about that? Because you built a browser and you built a computer and like what's the difference between Perplexity computer and Comet browser? Chrome became an OS. Have you thought about you building an OS that people can boot?

Aravind Srinivas

我认为这当然是个有趣的想法。但本质上,Jason,一切都关乎分发。如果你构建一个操作系统,你需要让它分发到实际的硬件设备上,这意味着你需要有 OEM 愿意为你分发。如果你读过微软与硬件 OEM 签订的合同,你会觉得谷歌像个天使。

I think it's certainly an interesting idea. Well, fundamentally, Jason, everything's about distribution. If you build an OS, you need to get it distributed in actual hardware devices, which means you need to have an OEM that wants to distribute it for you. And if you actually read the contracts Microsoft has signed with the hardware OEMs like you it will make you look at Google like an angel.

Host

那是当年反垄断案的一个重要部分,没错。

That was a big part of the antitrust case back in the day and yeah.

Aravind Srinivas

现在仍然很糟糕。就连 Chrome OS 都没能流行起来,这是有原因的。它们主要只能在学校、银行和政府层面获得采用,因为……

It's still pretty terrible. There's a reason even Chrome OS never picked up. And they were largely only able to get adoption at the level of like schools and banks and governments because it's

Host

我让公司用 Chromebook 运行了大概两年,然后 Zoom 的出现打破了这一切。Zoom 的崛起让它变得不可能,因为你在浏览器窗口中进行 Zoom 通话时体验很差。而且 Zoom 从未发布原生应用。这对我们的团队来说根本行不通,但人们喜欢它,因为工作时可以保持专注。不会弹出 iPhotos,不会有 Apple Music。人们就是喜欢它的克制。

I ran our firm on Chromebooks for maybe 2 years and then the one thing that broke it was Zoom. The ascent of Zoom just made it impossible because when you did a Zoom call in a browser window, it sucked. And Zoom never released itself. It's just like this is never going to work for our team, but people loved it because at work, you could remain focused. You wouldn't have iPhotos popping up. You wouldn't have Apple Music. People just loved the restraint of it.

Host

好了,在我们结束之前,我想让你们想想过去几个月里最令人印象深刻的 AI 体验,一个工具或产品。你们可以点名任何产品。我先来。我迷上了 WhisperFlow。我不知道你们是否在用 WhisperFlow,但它是我一生中体验过的最好的语音转文字工具。我之前已经放弃了语音转文字这个类别,因为我对 Siri 的听写能力非常失望。它从来都不好用,永远搞不定我的姓氏。我的意思是,Arvind,它永远搞不定你的姓氏。好吧,它也搞不定 Calacanis。但天哪,WhisperFlow 太神奇了。然后我在亚马逊上花 20 美元买了一个脚踏板。今天早上我在电脑前,用 Model Council 时,我踩下 Windows 机器上的踏板,然后一直说一直说,因为你给模型的信息越多,回答就越好。冗长的提示比简短的提示好得多,也省去了来回问答的麻烦。然后我抬起脚踏板,砰的一声,文字就出现在那里。这是一种改变人生的体验。就是那个愚蠢的脚踏板和这个不可思议的软件。你们有吗,Arvind 或 Edwin?

Okay, as we wrap here, I want you guys to think about the most impressive AI experience you've had the last couple of months, a tool, a product. You can shout anybody out. And I'll kick us off. I have become addicted to WhisperFlow. I don't know if you guys are using WhisperFlow, but it is the greatest speech-to-text I've ever experienced in my life, and I had given up on the category of speech-to-text because I was just so disappointed in Siri's ability to take dictation. It just never worked. It could never do my last name. I mean, Arvint, it would never be able to figure out yours. Okay, we can't figure out Calacanis. But my god, WhisperFlow is so amazing. And then I got a foot pedal for 20 bucks off of Amazon. When I'm at my computer this morning, I was using the Model Council, and I would press down on my pedal on my Windows machine, and I would just talk and keep talking and keep talking because the more you give a model, the better the response. The rambling long prompt is so much better than a short one and having to just go back and forth and play ping pong with questions. And then I just lift up my foot pedal, bang, text right in there. It is a life-changing experience. Just that stupid foot pedal and this incredible piece of software. Do you have one, Arvind or Edwin?

Aravind Srinivas

让我们推荐我们的产品,对吧?这就是这个问题的目的吗?

Name our products, right? Is that the purpose of the question?

Host

我是说,你可以,但当然,如果你想推荐一个你喜欢的也行。

I mean, you can, but sure, if you want to give one that you love.

Aravind Srinivas

某个东西。

Something.

最喜欢的AI工具与集成 Favorite AI tools and integrations

Host

你可以给自己打个广告,但我们的目标是……

Feel free to give a shout-out to your own, but the goal is to get to...

Aravind Srinivas

我很喜欢 Perplexity,我是重度用户。我喜欢用它做金融研究、公司运营、内部数据分析等等。不过我个人觉得 X 平台上的 Grok 集成最近改善了很多。推文上的“解释 Grok”按钮——有些笑话我总是不太懂,所以我觉得它挺不错的。

I love Perplexity. I'm a big time user of Perplexity. I love using it for financial research, company running, like internal data analysis, all that. Sure. But I personally thought the Grok integration inside X has improved considerably in recent times. The explain Grok button on tweets — I don't always understand some of the jokes, so I think it's pretty good.

Host

它特别棒,尤其是当你想要跟上某个流行文化热点或突发新闻时。比如有人说:“哦天哪,我猜特朗普总统和伊朗的事结束了。”我就想:“是啊,我跟不上了。”然后我点那个 Grok 按钮,它就把整个背景解释得非常清楚。

It's exceptional, especially when you're catching up to a pop culture moment or a breaking news story. Like somebody is like, "Oh my god, well, I guess President Trump and Iran's over." And I'm like, "Yeah, I can't keep up." So I hit that Grok button and it explains the whole context so well.

Aravind Srinivas

是的。在桌面上,我可以用评论功能来做这件事,但在 X 移动应用上,你需要那个“解释 Grok”按钮——它很好用。所以我认为这是一个做得非常漂亮的集成。向他们致敬。它以前其实并不好;过去一般般,但确实改善了很多。另一件让我印象深刻的事是 Gemini Flash 模型,Gemini 3 Flash。就它的能力而言,它快得离谱。它可能是速度和智能平衡得最好的模型。这完全是他们完成的惊人工程。

Yeah. When I'm on my desktop, I can just use comment to do that for me, but when I'm on the X mobile app, you need that explain Grok button — it's pretty good. So I think it's a very beautifully done integration. Kudos to them. It wasn't actually good before; it used to be so-so, but definitely it improved tremendously. The other thing I'm impressed about is the Gemini Flash model, the Gemini 3 Flash. It just seems insanely fast for the capability it has. It's probably the fastest model that hits the best sweet spot in terms of speed and intelligence. That's sheer amazing engineering they've accomplished.

Host

它快得离谱。Edwin,你有什么?周末或晚上你在 AI 领域痴迷什么?

It is wicked fast. Edwin, what do you got? What are you obsessing over on the weekends or at nights in the AI space?

Edwin

我其实是 Claude 3 设计的忠实粉丝。我觉得它设计得很好,而且很有主见。几乎可以说:“好吧,也许这就是我看到 Instagram 创始人风格的地方。”所以我觉得这是一个很棒的产品。我的意思是,它仍然有一些 bug,比如一些小烦恼让我有点沮丧,但我对它非常乐观。

So I actually am a really big fan of Claude 3 design. I think it's really well designed and opinionated. It's almost like, "Okay, yeah, maybe this is where I see the Instagram founder's touch." So I feel like it's a great product. I mean, it still has some bugs, like some little annoyances that have been a little frustrating for me, but I'm very optimistic for it.

Host

然后 Claude 设计,是的,上周刚发布。我的意思是,Claude 发布东西的速度我们谁也跟不上。我知道它发布了,只是还没机会玩。

And then Claude design, yeah, that came out last week. I mean, Claude is just releasing stuff at a pace that none of us can keep up with. I know it's out, just haven't had a chance to play with it yet.

Edwin

是的,我当时想,也许这是学习更多设计知识的好机会。所以上周我玩得很开心。但我要说的另一个让我震惊的地方是,几个月前我做了血液检查。我试着和医生沟通,但医生只给了泛泛的建议,就把我打发了。所以我拍了些照片,上传到各种模型上,老实说,它们给了我一些很好的建议。自从遵循这些建议以来,我感觉好多了。所以我认为那是一次非常棒的经历。

Yeah, I was like, maybe this is where it would be fun to learn a little bit more about design. So I had a lot of fun playing around with it last week. But I would say another place where models just kind of blew my mind was I got my blood work a couple months ago. I tried talking to my doctor, and my doctor just gave me generic recommendations and rushed me out the door. So I took some photos of it, and I uploaded it to all the different models, and honestly, they gave me some great recommendations. I feel like I've been feeling so much better since I've been following them. So I thought that was a pretty amazing experience.

Host

这很有意思。我不知道你们是 Whoop 队还是 Oura 队,或者用什么功能或超能力——市面上有很多这样的好东西。但 Whoop 现在,我正在努力调整睡眠,以便在播客和与创始人会面时表现出色,并努力达到某些压力目标,这是一件好事,比如以好的方式给身体施压。它现在内置了一个针对你健康调整的 AI。六个月前还很糟糕。但就在过去几周,它说:“哇,你滑雪滑得真棒。这是你连续第六天滑雪了。你可能想休息一天。你可以考虑一些补水饮料,而且你在 7000 英尺的海拔,所以你需要多喝水。你需要多休息。”我当时想:“哇。”然后它又说:“我知道你刚坐长途飞机到日本。这是你应该如何重置的方式。”这就像另一个突破性时刻,它说:“哦,我知道你在不同的时区。这是它如何影响你的睡眠。你想要一个睡眠计划吗?”我说:“我想要一个克服时差的睡眠计划吗?当然想。”做得非常好。非常垂直化,速度极快,使用我的数据。是的,不可思议。你有维生素 D 问题吗,Edwin?是不是像大家一样?在我们这个行业,似乎每个人都缺乏维生素 D。

That is an interesting thing. I don't know if you guys are on team Whoop or team Oura, or what you use or function or superpower — there's so many of these great things out there. But Whoop now, I'm trying to get my sleep dialed in so I can give a good performance on podcast and when I'm meeting with founders, and trying to hit certain stress goals, which is a good thing, like stressing your body in a good way. It now has an AI built into it that's tuned on your health. It was terrible 6 months ago. Then just the last couple of weeks, it was like, "Wow, you did a really great job skiing. This is your sixth day in a row of skiing. You might want to take a day off. You could consider some hydration beverages, and you're at 7,000 ft of elevation, so you're going to need to drink more water. You're going to have to get more rest." It was like, "Whoa." And then it's like, "I know you just took a long flight to Japan. Here's how you should reset." This was like another breakthrough moment where it's like, "Oh, I know you're on a different time zone. Here's how that's going to affect your sleep. Would you like a sleep plan?" And I was like, "Would I like a sleep plan to get over jet lag? You bet I would." Really well done. A very verticalized, incredibly fast, using my data. Yeah, incredible. You have vitamin D, Edwin? Is that your issue like everybody? It seems like everybody doesn't get enough vitamin D in our industry.

Edwin

我确实吃了很多维生素 D 之类的,但那是建议之一。不过我得看看 Whoop。

I do take a lot of vitamin D and all, but that was one of the recommendations. But I have to check out the Whoop.

Host

是的,哦,你也用 Whoop?是的。

Yeah, oh you use the Whoop, too? Yeah.

Edwin

我得看看。

I have to check it out.

Host

你应该看看。Arvind,你用这些吗?Whoop 还是 Oura?

You should check it out. Do you use any of these, Arvind? Whoop or Oura?

Aravind Srinivas

我用 Apple Watch,但我几乎每个月都做血液检查。

I use Apple Watch, but I do have my blood work done pretty much every month.

Host

每个月?哇,那有点强迫症了。

Every month? Whoa, that's obsessive.

Aravind Srinivas

嗯,那是因为我有一些状况需要监测。但现在我没事了。所有指标都很好,但我有时会缺乏维生素 D 和锌,因为一些饮食问题,所以提前知道是好事。我认为主动智能非常重要。所以我们在 Perplexity 中构建了功能,你可以连接所有健康数据,比如 Apple Health、Whoop、Oura。

Well, that's because I had some conditions and they need to be monitored. But now I'm fine. I'm pretty fine on all the vials, but the things I go low on at times are vitamin D and zinc because of some diet stuff, so it's good to know ahead of time. I think definitely the proactive intelligence is very important. So we built in Perplexity, you can connect all your health stuff like Apple Health, Whoop, Oura.

Host

Perplexity 可以?我打算试试。

Perplexity can? I'm going to do that.

Aravind Srinivas

Function Health、Be Well。你可以输入所有化验结果,一切都可以设置。所以我们非常认真地让 Perplexity 在个人健康用例上表现出色。

Function Health, Be Well. You can put all your lab results, everything and set it up. So we are very serious about making Perplexity work really well for personal health use cases.

Host

我认为这将会非常棒。好了,先生们,我知道你们有非常高效的团队,但有时你们肯定需要为某个特定职位招人。我们有很多听众。那么,Edwin,你在找什么人吗?我知道在专家方面,你一定在不断地寻找专家。那么,如果人们想成为专家或去你的公司工作,他们可以在哪里找到更多信息?

I think it's going to be incredible. All right, gentlemen, I know that you've got very efficient teams, but at some point you must need to hire a person for something specific. We get a lot of people listen to the pods. So, Edwin, anybody you're searching for? I know on the expert side, you must be constantly looking for experts. So, where can people find more information if they want to be an expert or go work at your company?

Edwin

是的,去我们的网站 searchhq.ai 或者直接给我发邮件。我喜欢看申请。

Yeah, so just go to our website, searchhq.ai or email me personally. I love reading applications.

Host

太棒了。Arvind,你在找什么样的人?你需要什么?我们如何帮助你保持产品列车像现在这样顺利运行?

Amazing. Arvind, what are you looking for? What do you need? How can we help you with keeping the product train running as it's been doing really well?

Aravind Srinivas

是的。我们正在深入企业领域,所以对销售职位感兴趣的人,请随意申请。正在观看的工程师们,我们正在招聘全栈工程师。所以请随意申请,非常期待看到谁对我们感兴趣。

Yeah. We're going deeper on the enterprise, so people interested in sales roles, definitely feel free to apply. Engineers who are watching this, hiring for full stack engineers. So definitely feel free to apply here and very excited to see who's interested in us.

Host

当然。这里是《本周 AI》第 10 集。请访问 thisweekinai.ai,订阅我们的新闻通讯。我们将推出付费新闻通讯。

Absolutely. And this has been episode 10 of This Week in AI. Go to thisweekinai.ai, sign up for our newsletter. We're going to have a paid newsletter.

结束语 Closing announcement

Host

我们正在追踪每一家获得投资的公司,并将针对每一家种子轮和 A 轮公司发送报告,所以您会想要注册并获取免费邮件。然后,我想是在 6 月 1 日,我们将推出付费版本,为业内人士提供更细粒度的细节。下次见,拜拜。

We're tracking every single company that's invested, and we're going to be sending out reports on every one of these seed stage and Series A companies, so you're going to want to sign up and get the free email. And then in, I think, June 1st, we're going to launch the paid version with even more granular details for people who are in the industry, and we'll see you next time. Bye-bye.

互动版:逐字朗读 + 针对本期提问 →