Box 的 Aaron Levie:在 AI 时代重塑自我与企业级扩散

Box's Aaron Levie on Reinventing Yourself in the AI Age and Enterprise Diffusion

亚伦·莱维 Aaron Levie · Training Data · 2026-09-15 · 约 65 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Box 创始人兼 CEO Aaron Levie 探讨为何应用公司成为新 Neolabs、AI 如何连接模型与企业工作流,以及创始人如何在 AI 浪潮中重塑自我。

Aaron Levie, founder and CEO of Box, discusses why application companies are the new Neolabs, how AI bridges models and enterprise workflows, and how founders can reinvent themselves in the AI wave.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 39)

全文 · Full transcript(中英对照)

给学生的职业建议 Career Advice for Students

Aaron

我想给所有大二大三的学生发个紧急警报:在 Twitter 上关注这 20 个账号,而且赶紧加入 Twitter,因为这会对你的职业生涯有帮助。仅仅根据你的信息流,你要么领先一年,要么落后一年,而我的信息流非常灵通。我还会和 20 岁的人聊天,他们说,是啊,我看到一些文章,我就想,你什么意思,你怎么看到文章的?我甚至不知道那是什么意思。难道你只是运气好有人给你发了篇文章?你必须得灵通。我很享受。这很有趣。我什么都玩。

I want to send an emergency alert to everybody who's a sophomore or junior in college and just be like: follow these 20 accounts on Twitter, and also join Twitter, because this will just help your career. You're either a year ahead or a year behind simply based on your feed, and my feed is so wired in. I'll still talk to 20-year-olds that are like, yeah, I see some articles, and I'm like, what do you mean, how do you see articles? I don't even know what that means. Do you just get lucky that somebody emailed you an article? You just have to be wired in. I enjoy it. It's a lot of fun. I play with everything.

介绍Aaron Levie Introducing Aaron Levie

Host

今天我很高兴欢迎 Box 的创始人兼 CEO Aaron Levie。Box 正在为企业构建协作平台和内容管理平台,现在随着 AI 的发展,它焕发了全新的生机。我很期待今天和你聊聊 Box 和 AI,聊聊你对 AI 的总体看法,因为你是如此有思想领导力,也聊聊创始人如何在这场 AI 浪潮中重塑自我。所以,感谢你的参与。

Today I'm excited to welcome Aaron Levie, founder and CEO of Box. Box is building a collaboration platform and content management platform for the enterprise that has now taken on really new life with AI. And I'm excited to chat with you today about Box and AI, about your general thoughts on AI because you are such a thought leader, and about how founders can reinvent themselves for this AI wave. So, thanks for joining us.

Aaron

感谢邀请。我是这个播客的忠实粉丝。我每集必看,做得很好。不过我在想,我知道这是另一个播客,但你们什么时候聊根管治疗?

Thanks for having me. I'm a big fan of the podcast. I religiously watch every episode and so great work. I am wondering though, I know this is a different podcast, but when do you talk about root canals in this?

Host

天哪。那是——我们叫 Doug Leone 来了吗?我们有没有想出怎么把那个融入进来?

Oh my gosh. Is that— Have we called in Doug Leone? Have we figured out how to weave that in yet or not?

Aaron

我是说,我很乐意聊聊我经历过的另一种痛苦和折磨。

I mean, I'm happy to talk about different kind of pain and suffering that I've been through.

Host

我们应该就做这期播客,现场做根管治疗。就这么干,看看效果如何。

We should just do this pod with a live root canal. Let's do it and see how it goes.

Aaron

实际上,那会像《Hot Ones》,但你要做根管治疗,还得在牙医椅上谈论你的战略,而 Doug 就在另一头。

Actually, that would be like Hot Ones, but you get a root canal and you have to talk about your strategy while you're in a dentist chair and Doug is just on the other end of it.

Host

太棒了。太棒了。顺便说一句,Doug 现在对自己非常满意。好的。他看过每条推文吗?

Amazing. Amazing. Doug is so pleased with himself, by the way, right now. Okay. Has he seen every tweet?

Aaron

哦,是的。好的。哦。哦,是的。他对自己非常满意。

Oh, yeah. Okay. Oh. Oh, yeah. He's very pleased with himself.

Host

很好。好的。让我们开始——我很好奇你对这个的看法。应用公司是最热门的 Neolabs。同意还是不同意。

Good. Okay. Let's start with— I'm curious your take on this. Application companies are the hottest Neolabs. Agree or disagree.

应用公司即新实验室 Application Companies as Neolabs

Aaron

两年前,我觉得这不太说得通,那是什么意思?但很明显,我认为这就是市场上正在发生的情况,而且这主要都是因为开源。但很神奇的是,现在作为 LLM 包装器或模型包装器的整个概念实际上正在奏效,因为我认为人们低估了的是,在现实世界中,在企业中,你需要的是从模型能力到企业实际工作流之间的某种桥梁,而那座桥梁基本上可能有上万亿美元押注在两种结果之一。要么那座桥梁非常有限,要么那座桥梁实际上非常广阔,需要能够在组织内部承担很多深度。而这个赌注基本上看起来像是:你是只做多模型本身和超级智能,还是做多这种应用层,或者你知道以前可能只是纯粹的新实验室,但我认为非常明显的是,实际上模型和工作流之间有很大的差距,随着你随着时间的推移弥合这个差距,你会到达一个点,你意识到哦,我也应该做模型,然后你有足够的数据,你有足够的领域专业知识,那就会成为它自己的飞轮。所以我对这个想法非常看好,很酷的是它为初创公司打开了多个层面的机会,因为你可以是实际的应用公司本身,也就是新实验室,或者你可以是新实验室的基础设施提供商,你有多个层面去攻击那个空间,但我认为这对市场和大家对它的看法是一个巨大的更新。

Two years ago, I think it would have not made that much sense as like what does that mean? But very clearly I think this is what's playing out in the market and it's all working out mostly because of open source. But what's pretty amazing is right now the whole concept of being a sort of LLM wrapper or model wrapper is actually working out because what I think people underappreciated was that in the real world in the enterprise what you need is some bridge from the model's capability to the actual workflow that the enterprise has and that bridge you basically like probably a trillion dollars has been bet on basically one of two outcomes. Either that bridge is very kind of limited or that bridge is actually very vast and needs to be able to take on a lot of depth within organizations. And that bet basically looks like are you only long the model itself and superintelligence or are you long this sort of application tier or you know maybe previously would have been just pure neolab but I think it's very clearly playing out that actually there's a lot of gap between the model and the workflow and as you bridge that gap over time you get to a point where you realize oh I should also do the model and then you have enough data you have enough sort of domain expertise where that becomes its own sort of flywheel. So I'm very bullish on this on this idea and what's cool is it's opening up like multiple layers of opportunity for startups because you could either be the actual applied company itself i.e. the Neolab or you could be the infrastructure provider to the Neolab and you have like multiple layers of going and attacking that space but I think a huge update for the market and everybody's kind of view of it.

Host

Box Labs,我们上吧。

Box Labs, let's go.

Aaron

我们已经有了。这有点秘密,但我们目前非常关注。主要焦点是让智能体在任何模型上都非常非常出色,但随着时间的推移,显然你会剥离某些用例,要么基于每个客户,要么跨越整个数据集。

We already have it. It's a little bit secret but we pay very close attention right now. The main focus is let's make the agent really really good on any model but over time obviously you would peel off certain use cases either on a per customer basis or across the whole data set.

Host

是的,我的意思是似乎有两种力量在起作用。一是人们不希望狐狸看守鸡舍,他们不希望代币的卖家同时也是计量和把关的人,你知道,为每个用例确定最佳代币。

Yeah, I mean it seems like there's two forces that are happening. One is people don't want the fox to be guarding the hen house, they don't want the seller of the token to be the one that is also metering and gating you know what is the best token for each use case.

Aaron

然后第二个是,你知道,在那座要跨越的桥梁上实际上有很多工作要做,而且涉及真正的研究。而且它非常定制化,针对确切的工​​作流和确切的最终客户。

And then the second is you know there's actually a lot of work to do on that bridge to cross and there's real research involved in it. And it's pretty bespoke to the exact workflow and the exact end customer that you have.

Host

是的。我认为我在硅谷经常看到的情况是,而且有很好的理由,这也是为什么这些公司如此成功,就像每个人都如此沉迷于研究,这再次完全很棒,我是超级粉丝。它带来了所有这些突破。但那种感觉是,好吧,模型和模型中的智能是唯一重要的形式因素。然后你进入现实世界,你会看到智能实际上是如何在人们的工作流中展开的,模型可能是世界上最智能的,你知道,超级智能,但那个工作流仍然需要你连接到其他数据系统,仍然需要这些有人类在环互动的时刻。工作流中有延迟,所以智能体不得不闲置。实际业务流程有变更管理。有尚未现代化的遗留系统。所以你会经历这五到十件事,它们远比模型的纯粹超级智能更偏向运营,更偏向于阻挡和擒抱。而我认为经典的研究组织最不想做的就是去攻击每一个这样的事情。这有点像,这不是历史上的一个一次性事件。就像我们一直都有基础设施和应用之间的这种关系。

Yeah. I think what I tend to see happen in the valley is and for very good reason and it's actually why these companies have been so successful is like everybody is so kind of research pilled which is again totally awesome big fan. It's led to all these breakthroughs. But the sort of sense that okay the model and the intelligence in the model is kind of the only form factor that matters. And then you go to the real world and you sort of see how intelligence actually gets rolled out in people's workflows and the model could be the most intelligent you know superintelligence in the world but that workflow still requires you to connect up to other data systems still requires these moments where there's a human in the loop interaction. There's delays in the workflow and so the sort of agent has to sit idle. There's change management of the actual business process. There's legacy systems that haven't been modernized. So you kind of go through these five or 10 things that are way sort of much more operational much more blocking and tackling than just the pure superintelligence of the model. And the last thing that I think a classic sort of research organization wants to go do is go attack every single one of those things. And this is sort of not a there's like this is not a sort of a one-off in history. Like we've always had this relationship between infrastructure and application.

基础设施与应用价值 Infrastructure vs. Application Value

Aaron

显然,AWS、GCP 或 Azure 创造了数万亿美元的价值,即基础设施的市值。但你猜怎么着?还有数万亿美元的软件价值,仅仅因为有了这些基础设施才存在。如果你回到 10 年前,看看数据领域发生了什么,看看 GCP 早期版本在构建什么,或者 AWS 在构建什么,我保证你无法预测到 Snowflake 或 Databricks 的存在。你当时会觉得,基础设施已经能做到这些了,为什么还要向这些其他产品支付另外 100 亿美元的收入,它们只是让你能处理数据而已?同样的事情也会发生在智能领域,即模型将极其有价值,但将这些模型带入银行、生命科学、医疗保健和政府等实际工作流程中的应用,那将只是大量的软件。

Obviously AWS, GCP, or Azure have created trillions of dollars in value, kind of market cap of infrastructure. But guess what? There's also trillions of dollars of value in software that only exists because of that infrastructure. If you were to go back 10 years ago and look at what was happening in the data space, and you looked at what GCP early versions of GCP was building or AWS was building, I guarantee you would not have predicted Snowflake or Databricks existing. You would have been like, the infrastructure just already does that. Why would you pay another $10 billion of revenue to all these other products that are just making it so you can work with your data? The same thing is going to be true for intelligence, which is the models will be insanely valuable, but the application of bringing those models into real workflows in banking and life sciences and healthcare and government, that's just going to be a lot of software.

Aaron

现在,未来两到五年的挑战是,模型提供商需要在多大程度上向上移动这个堆栈,并尝试在那个层面竞争,还是他们有意或无意地将整个空间留给应用生态系统?实际上,在某种程度上,这对他们来说是一个重大的战略问题,因为一方面你想更接近客户,另一方面你也希望能够拥有一个生态系统,让人们信任你。所以,未来几年这将是一个非常有趣的张力。

Now the sort of challenge for the next, let's say, two to five years is how much do the model providers need to move up that stack and also try and go and compete at that layer, or do they either leave open intentionally or accidentally that entire space for the application sort of ecosystem? And actually, to some extent, this is like a big question strategically for them, because on one hand you want to be closer to the customer, on the other hand you also want to be able to have an ecosystem so people trust you. So that's going to be like a really interesting tension over the coming years.

Host

你觉得这会如何发展?我觉得这是我们每天都在纠结的问题。

How do you think that'll play out? I feel like that's the question we spend every day wrestling with.

Aaron

是的。呃,你可以,是的。你知道,我不想代表红杉资本发言,但你能感受到现在风投的一些存在主义恐惧,比如,我是不是应该再向 Anthropic 投入 10 亿美元,或者,你知道,我是否应该尝试看看应用层会如何发展。所以,没人会羡慕你们的位置,必须弄清楚这一点。但但显然,对企业家来说更难。呃,但你知道,我我显然对应用层非常看好,就像我也同样非常有偏见,我我有一个非常集中的赌注,多样化非常有限,赌它会成功,你仍然想购买那种理解你的工作流程并能触及核心企业数据的技术。但我不知道是否有,我只是看不到与整个历史不同的事件发生。我的意思是,你知道,我们只有大约 50 或 60 年的计算机软件历史。呃,但基本上,当你去那家律师事务所,去那家制药公司,去那家银行,他们需要一些东西来桥接核心技术到他们的工作流程和业务流程。而 AI 并没有真正改变这种需求或形态。最好的体现就是这种聊天机器人与智能体工作流程的划分。所以聊天机器人可以完全通用,可以完全横向,但你知道,突然你看着它说:“嗯,我的工作流程需要一些东西在过程中的正确时间提醒我,或者需要访问某种数据,而聊天机器人无法原生访问。所以必须有人进入那个组织,设置好,必须有人为这个模型提供领域专业知识,以便它真正理解我们特定的业务流程。”呃,所以,除非你真的承保,你知道,大型实验室,老实说,我不是在夸张,像 10 万名员工,如果你不承保,那么扩散经济将非常庞大,因为每一个这样的公司,无论是 50 人的公司还是肯定有几千人的公司,都需要一支大军去帮助他们进行转型、变革管理。所以,你知道,无论是 FTE 现象还是再次理解领域专业知识,那都将是一件大事。

Yes. Uh you can Yeah. You know, I don't want to speak for Sequoia, but you can sense some of the existential dread of a VC right now is like, should I just put another billion dollars into Anthropic or, you know, should I attempt to sort of see what at the applied layer is going to play out. So, nobody's, you know, that envious of your guys' position having to figure that out. But but obviously, it's even harder for the entrepreneur. Um uh but you know I I am you know I'm pretty long obviously the application layer like I'm also equally very biased like I'm I have a very concentrated bet with very limited diversification on it working out that you still want to buy technology that sort of understands your workflow and can get to the core enterprise data. But I don't know that there's I just don't see a different event happening than all of history. I mean, you know, we only have like 50 or 60 years of computer software history. Um, but basically, when you go to that law firm and you go to that, you know, pharma company and you go to that bank, they need something that bridges the core technology to their workflow in their business process. And AI has not meaningfully sort of changed the need or the shape of what that looks like. And the best sort of manifestation of this is sort of this chatbot versus kind of agentic workflow, you know, demarcation. So the chatbot can be totally universal and can be totally horizontal, but you know, all of a sudden you kind of look at that and you say, "Well, my workflow kind of needs something to kind of ping me at the right time in the process or it needs access to a certain kind of data that just, you know, the chatbot can't natively get access to. So somebody has to go into that organization like get it set up and somebody has to go and provide domain expertise to this model so it really understands our particular business process." Um, and and so then unless you really just underwrite, you know, the big labs at at honestly like I'm not exaggerating, like a 100,000 employees, if you don't underwrite that, then the diffusion kind of economy is going to be massive because every single one of those companies, whether that's a 50 person firm or certainly a multi-thousand person firm, is going to need an army of people to go in and help them with that transformation, that change management. And so this is, you know, like whether it's the FTE phenomenon or just again understanding that domain expertise that that's going to be a very big deal.

Aaron

然后你提到了这一点,呃,狐狸 guarding 鸡舍,呃,这可能实际上是这件事必须发生的唯一最大原因,即即使在完全仁慈的情况下,没有人实际上以任何狡猾的方式做事,但道理很简单,如果我要将任务交给一个智能体系统,我只希望该任务在保持准确性的同时成本优化。那么谁能做到呢?是那家不在乎 10 种不同模型中哪个模型执行任务的公司。所以从定义上讲,你会希望那是一家没有偏好的公司。

And then you kind of alluded to this point, um uh Fox kind of guarding the hen house and um uh and that that might actually be singularly the biggest sort of reason this has to happen, which is even under all sort of like complete benevolence and like nobody's actually doing something um uh in any kind of like you know sneaky way, it just stands to reason that if I'm going to sort of give a task to an agentic system, I just want that task to be cost optimized like with accuracy as kind of holding constant. And so who can do that? It's the company that doesn't care among you know 10 different models which model is performing that task. So like definitionally you would want that to be the company that does not have a sort of a preference.

Host

然后你面临的阻力是,模型公司可以补贴他们的模型或以不同的价格提供它们。

And then the force you have going against that is that the model companies can subsidize their models or offer them at different rates.

Aaron

不过我不知道这能持续多久,因为当这些公司上市时,我认为它们将基本上受到与其他人相同的资本主义法则的约束。所以补贴在一定的支出阈值内效果很好,我们可能已经超过了那个支出,当你达到像数百亿美元的资本时,事实上,如果有什么不同,它甚至可能更糟,因为呃最终你必须进去为你的训练运行付费。所以就像他们喜欢代币的补贴,我认为这只是一个暂时现象。

I don't know how long that lasts though because when these companies become public, I think they will be held to basically the same laws of capitalism that everybody else is. So the subsidization is working very well up to a certain threshold of spend and we might have we may have exceeded that spend when you're at like tens of billions of dollars of kind of capital and in fact if anything it might even be worse because uh eventually you have to go in and sort of pay for your training runs as well. So like they like the subsidization of tokens is a I think it just has to be a temporary phenomenon.

Host

你是从毛利率的角度还是从反垄断的角度?

You mean from a gross margin perspective or from like an antitrust perspective?

Aaron

完全是毛利率。就像如果我

Entirely gross margin. Like if I

Host

但他们现在在推理上有如此高的毛利率。

But they have such high gross margins on inference right now.

Aaron

但那么他们在补贴什么?那么实际上他们收取的是相当不错的费率。

But then what are they subsidizing? Then it's actually then they're charging actually like decent rates.

Host

是的。但他们有能力将 API 定价高于他们自己的第一方产品。对吧。

Yeah. But they can afford to price the API higher than their own first party products. Right.

Aaron

所以那部分完全公平。你知道有一个有趣的维度,就像如果 API 是高利润的东西,用来支付补贴,但你把所有客户都转移到应用产品上,没有 API 收入。所以你必须在这方面取得平衡。

So that part to totally fair. You know there's an interesting kind of you know dimension which is well like if API is the high margin thing that is sort of paying for the subsidization but you've moved all your customers over to the applied product and there's no API revenue. So like there is like an equilibrium you have to strike with this.

Host

完全正确。

Totally.

AI栈中的价值分配 Value Distribution Across the AI Stack

Aaron

与此同时,如果你有一些非经济行为体,或者只是有着完全不同博弈论的人——你知道,Meta 算一个,SpaceX 可能算一个,中国肯定是个巨头,甚至 Nvidia 也算一个——那也会改变整个盘算。这四类玩家不一定需要靠推理赚钱,至少不需要像 Anthropic 或 OpenAI 那样的利润率结构。所以他们可能愿意把推理利润率压到 10%,因为他们只是想为庞大的算力集群买单。只要这种情况发生,只要没有疯狂的专有秘密,那么无论如何,单位 token 成本都会在同等条件下下降。所有这些都意味着更多价值会流向应用层。这绕了一大圈,其实就是说我认为整个堆栈里的每一层都有价值。我只是不确定——我唯一不会押注的,就是只有一两个实验室拿走 95% 的价值创造。我认为环境会动态得多。老实说,如果我是那两三个最大的实验室之一,我也会更喜欢这个结果,因为回到你提到的反垄断问题,如果你成了唯一存在的智能,某个时候你就会被国有化。所以在这个生态里,你总归希望有一点健康的竞争。

All the while, if you have some sort of either non-economic actors or people just with a totally different game theory in this — you know, Meta being one, maybe SpaceX being one, China certainly being a giant one, even Nvidia being one — that changes the calculus as well. Those four cohorts don't necessarily need to make money on inference, at least at the same margin structure as what Anthropic or OpenAI needs. So they might be fine to bring down inference to 10% margin because they just want to basically pay for massive compute clusters. And as long as that happens, and as long as there aren't insane proprietary closely held secrets, then no matter what, you're going to have cost per token go down on a like-for-like basis. All of which means more value accrues to the application layer. Which is a very long-winded way of saying I think there's just value in everybody in the stack. I just don't know that — the only thing I probably wouldn't bet on is that one or two labs get 95% of the value creation. I think there's just going to be a much more dynamic environment. And honestly, if I were one of the two or three biggest labs, I think I'd prefer this outcome too, because back to your antitrust point, at some point you'll just be nationalized if you're the only thing that exists as intelligence. So you kind of want a little bit of healthy competition in this ecosystem anyway.

Host

对,完全同意。我们转到 Box 吧。我之后还会回来聊 token、杰文斯悖论、中国这些,但先聊聊 Box。

Yep. Totally. Let's transition to talking about Box. I'm going to come back to talking about tokens and Jevons paradox and China and all this stuff, but let's talk about Box for a second.

Aaron

我也很喜欢这个话题。

I love that topic, too.

Host

给大家——我猜很多听这个播客的人都在用 Box——但给大家简单讲讲 Box 的历史,以及你们如何用 AI 重塑自己。

Give folks — I'm guessing a lot of people that listen to this podcast use Box — but give people a brief explanation of the history of Box and how you're reinventing yourself with AI.

Aaron

我们创办这家公司,是为了在云端安全地存储和共享数据。想法很简单,但我们算是敲开了一个坚果,或者说触动了一根神经。对 Doug 来说,我们能够快速扩张。我们在企业市场迅速转向,当时的想法是,企业会从本地部署系统迁移到云端,他们需要一种更好的方式来安全地存储、共享、协作和管理所有这些非结构化数据——公司文件、财务文件、营销素材、研究资料。这就是这家公司。我们其实从 2015 年左右就开始尝试 AI 产品和体验了。如果你还记得第一次——至少在现代——2015 到 2018 年那波 AI 寒冬,当时大家觉得要来了,结果没来。那段时间我们觉得,这些非常早期的 AI 模型给了我们一些迹象:如果你看一张图片,能很好地分类,那对企业来说就很有用,因为也许你可以给所有图片数据打标签,或者做 OCR 把里面的文字提取出来。这非常有帮助。问题是它贵得离谱,而且每个用例都得有一个模型——你想做的每个工作流都需要一个专门训练的模型。所以我们暂时搁置了。几年后,我们开始关注 GPT。我们办了一些黑客松,有人说,哦,我们可以在笔记产品里做输入联想。那是早期版本。我们在文本检测和分类上做了一些早期工作,帮助了安全用例。然后显然 ChatGPT 时刻来了,那是让人脑洞大开的时刻。别的不说,他们找到了一种形态,让所有人都豁然开朗:哦,这些可以是交互式系统,你问一个问题,就能得到越来越复杂、越来越长的回答。

We started the company as a way to securely store and share data in the cloud. It was a very simple idea, but we kind of cracked a nut or struck a nerve. For Doug out there, we were able to scale up quickly. We pivoted rapidly in the enterprise, and the idea was enterprises would be moving from on-premises systems to the cloud and they would need a better way to store, share, collaborate, and manage all of this unstructured data — their corporate documents, financial documents, marketing assets, research materials — in the cloud securely. So that was the company. We had been flirting with AI products and experiences really since like 2015. If you remember the first rise of — at least in modern times — the AI winter that happened in like the 2015 to 2018 period, where we think it's going to happen now and then it didn't. That was a period where we were like, okay, these very early AI models were showing us signs that, okay, if you looked at an image and you could classify the image well, that's pretty useful if you're in an enterprise, because now maybe you take all of your image data and label it, or maybe you OCR something and you'd be able to pull out the text in there. That's enormously helpful. The problem was it was insanely expensive and you had to have a model for every single use case — a hyper-trained model for each workflow that you wanted to do. So we kind of shelved it. A few years later, we started paying attention to the GPTs. We had some hackathons where people were like, oh, we could do type-ahead in one of our note-taking products. That was sort of early versions. We did some early work in text detection and classification which helped with security use cases. Then obviously the ChatGPT moment hits, and that was the big head-exploding moment. If for no other reason, they figured out a form factor that opened up everybody's mind to: oh, these could be interactive systems that you just ask a question and get an answer back of increasing complexity and length.

Host

对。

Yep.

Aaron

所以我们很快研究了它,全力投入。我们做了整个公司的转型——一切都完全符合学术上你应该做的。我们划出一个团队,把最好的人放进去。我们每天开会,看更新,然后慢慢地、但坚定地搭建出今天我们的 AI 堆栈,以及 Box 智能体。对我们来说,你可以想象用例非常直接。我们坐拥数千亿个文件。每一个文件都包含企业的关键信息。可能是合同、研究文件、营销素材、贷款文件——所有这些关键信息。问题是他们很少知道里面到底是什么。除非你真的去看文档、搜索并找到它,否则你就是不知道里面有什么。所以现在智能体可以被派出去,回答关于这些数据的问题。它们可以预处理数据,从文档中提取元数据,并转化为结构化数据。你可以用智能体自动化步骤和工作流。所以我们搭建了一个平台,基本上让你能针对所有非结构化数据部署智能体。这一直是核心焦点。

So we looked at that very quickly. We jumped all in. We did the whole company pivot — everything was exactly like academically what you should do. We had a team carved out. We put the best people on the team. We met every day, looked at the updates, and then slowly but surely built out what today is our AI stack and then basically the Box agent. For us, you can imagine the use case is very straightforward. We sit on hundreds of billions of files. Every single one of those files contains critical information for an enterprise. That could be their contracts, their research files, their marketing assets, their loan documents — all of this critical information. The problem is they rarely know what's actually inside of it. Unless you literally look at the document and search and find it, you just don't know what's inside. So now agents can go and basically be farmed out to answer questions about that data. They can pre-process it and extract metadata from those documents and turn it into structured data. You can use agents to automate steps and workflows. So we've built a platform that basically lets you deploy agents against all of that unstructured data. And that's been the core focus.

Host

你们在用智能体创建新内容吗?

Are you using agents to create new content?

Aaron

是的。这体现在几种模式里。一是我们有一个在线协作产品,智能体可以在里面生成任意数量的内容。然后我们和 OpenAI 以及 Anthropic 做了大部分——我觉得可能更令人兴奋的——工作,就是如何做高级文档创建、PowerPoint 创建。我们认定他们的技术目前总是能保持前沿。所以我们有一个智能体去和那些系统交互,生成高质量的 PowerPoint 等等。

We are. There are a couple modalities where that shows up. One is we have again an online collaborative product that an agent can just generate any amount of content in. And then we've done most of the — I think probably more exciting — work with OpenAI and Anthropic on just how do you do advanced document creation, PowerPoint creation. We've decided that their tech is at this point can always be frontier. So we have an agent that goes and interacts with those systems to produce a high-quality PowerPoint, etc.

Host

太棒了。你们的标语是“你的业务存在于内容中。用 AI 释放它。”那么,今天人们释放内容的主要本垒打用例是什么?

Awesome. Your tagline is "Your business lives in content. Unleash it with AI." So what are the hero home run use cases for how people are unleashing it today?

企业产品中AI的未来 Future of AI in enterprise products

Host

那么,如果快进几年,你觉得几年后人们会在你们的产品中用 AI 做什么?

And then if you had to fast forward a few years, what do you think people will be doing with AI in your products in a few years?

Aaron

对。所以对传统企业来说,最容易想到的典型场景就是:你有一百万份合同,为什么不弄清楚里面有什么?或者你有一百万份研究文档,能够提取出所有关键的结构化数据,放进数据库,然后能够查询、分析、围绕它自动化工作流。所以这种事情每次都能大获成功,因为它是一个长期存在的问题,人们一直无法投入人力去解决,因为阅读每一份合同、每一份研究文档实在太贵了。也许如果你有贷款文档流程,你可以做到,但大多数其他数据在那个规模上从未被阅读过。然后我认为我们可能同样甚至更兴奋的是,真正等同于我们在编码智能体或其他复杂智能体上看到的东西,即你有长时间运行的智能体,它们只是在执行你的整个工作流或流程。这会是这样的形式:你去一家银行,你在银行开户,他们基本上自动化了每一个可以自动化的步骤,然后在流程中的某些步骤跳出来交给一个人进行额外审查或额外验证。但现在不再是那种一两个星期的来回,而是像一个小时就完成了。这就是大多数企业工作流的梦想状态:如果我们能更快地让客户入职?如果我们能更快地发现研究中的关键数据?如果我们能更快地对安全事件发出警报?所以要做到这一点,你需要这些后台智能体或工作流,为这些流程预先建立好。

Yeah. So the probably the easiest hero for again more of a traditional enterprise to think about is just as simple as you have a million contracts. Why don't you find out what's inside them? Or you have a million research documents, be able to go and pull out all of the critical structured data, put that into a database, and then be able to query, analyze, automate workflows around that. So that's kind of the thing that just knocks it out of the park every single time because it's been a long-standing problem that people have never been able to apply human labor to because it's just too expensive to read every contract, every research document. You know, maybe you could do it if you had like a loan document process, but most other data just never gets read at that scale. And then I think the stuff that we're probably as much if not more excited by is really the equivalent of what we see with let's say coding agents or other complex agents, which is you have long-running agents that are just executing your entire workflow or process. And this would be in the form of you go to a bank and you're onboarding at a bank and they've basically automated every step that is possible to automate and then it jumps out to a person in the steps in the process for extra review or extra verification. But now instead of that one or two week back and forth, it just happens in like an hour. Like that's the dream state of most of these enterprise workflows: what if we could onboard a client faster? What if we could discover critical data inside of our research much more quickly? What if we can alert to a security event much more quickly? So to do that, you need these background agents or workflows that are pre-established for those processes.

Host

完全同意。这是长时间运行智能体的一年。

Totally. It's the year of the long running agent.

Aaron

是的。是的。

It is. It is.

Host

是的。是的。嗯,我很好奇,你做了个类比,提到云代码。在我看来,在编码领域,使用 AI 不仅被接受,而且被拥抱。

Yeah. Yeah. Um I'm curious like you made the analogy to Claude Code. It seems to me that in the coding domain, using AI is like not only accepted, it's embraced.

Aaron

是的。

Yes.

Host

呃,在内容领域,我认为这是很多内容所在的地方。是的。嗯,至少使用 AI 来生产内容。就像有一种几乎过敏的反应,比如推特上所有的 pangram 之类的东西。有那个,你知道,就像工作垃圾这个概念。我很好奇你怎么看工作垃圾,以及几年后这还会是个问题吗?

Uh in the content domain, which is I think a where a lot of the content in in box sits. Yep. um using AI to produce content at least. It's just like there's this almost this allergic reaction to it like all the panggram stuff on Twitter. There's the um you know it's like this concept of work slop. I'm curious what you think about work slop and like will this still be a thing in a few years?

Aaron

我要把 Box 的企业帽子摘掉,现在可能作为一个消费者来即兴聊聊

I'm going to sort of separate the box corporate hat and just now kind of maybe riff as a consumer of

Host

天哪,我希望有个更好的词,但工作垃圾,你知道,在企业背景下。

gosh I wish there was a better term but work slop as you know inside of an enterprise context.

Aaron

这些天我收到的董事会幻灯片完全是由 AI 写的,这让我受不了。

I get board decks that are entirely written by AI these days and it kills me.

Aaron

所以我认为在可接受性上的区别是这样的。实际上这个话题可能更多是象征意义,而不仅仅是垃圾元素,但实际上 AI 的普及几乎和代码类似,除了你知道我们交往的那些顶级工程师,他们对代码有深厚的品味,判断力惊人,对他们来说这既是艺术也是科学。所以把那个群体放在一边。对世界上大多数人来说,代码是一种实用工具。是的。它只是为了完成某件事。你只是想自动化某件事。你只是想放一个界面,有人按个按钮就进入下一步。所以对世界上大多数人来说,代码的价值创造就是自动化,并把它作为一种实用工具。所以,归根结底,我们可能还会用垃圾这个词一段时间,因为前端设计有品味,系统有品味,你不想代码中有漏洞。所以,这还会存在一段时间。但归根结底,如果你能告诉一个智能体,比如,请生成我的整个后端系统或前端系统,那不仅是可以接受的,而且是更可取的,因为那正是阻碍我们前进的事情。所以,我们需要去做。至少在社会运作的方式、世界运作的方式和我们大脑目前运作的方式下是这样。也许这会改变。你知道,当你从某人那里收到一份演示文稿时,仍然有这种联想,就像我在试图决定我是否可以信任那个人去执行那件事或交付那个结果或理解那个话题。所以当你看到工作垃圾时,你会想,我失去了能力去确切知道,比如思考过程有多少是他们自己的,有多少是 AI 的?我甚至应该在乎多少,因为我自己也在做同样的事情。所以我们有这种奇怪的,这是一种非常奇怪的集体问题,就像我在为我的一些头脑风暴和决策做工作垃圾,但当我从别人那里收到时,我会想,嗯,我应该信任你吗?呃,我不知道,你知道,我的意思是,这可能只是作为一个社会,我们必须在接下来的三到五年里继续努力解决,并最终到达另一边,就像你知道,我讨厌用这些完全破旧的类比,但显然当你看到某人的财务模型时,你不在乎,你会说,是的,那显然是由宏生成的,或者你知道,那不是你亲自计算的所有东西,但你展示给我看,我们在讨论它。那么,为什么不能同样存在于战略幻灯片或其他东西上呢?但我认为现在我们正在经历这种演变,比如这个人的角色是什么?内容是什么的代理?它应该是那个人知道多少的代理吗?它是我们认为他们能执行什么的代理吗?我认为我们只是在这个非常混乱的时期,我们必须弄清楚这一点。

So here's the difference I think on the acceptability. So there's probably like more symbolism to this actually topic than just like the slop element, but like actually like diffusion of AI in general sort of almost ties to this code like other than you know the top engineers that we hang out with that like are like they have you know deep taste in the code like like you know and and the judgment is incredible and and like it is them as much an art as is a science. So take that group aside. For most of the world, code is a utility. Yeah. It is just trying to accomplish something. You're just trying to automate something. You're just trying to put a sort of interface up there that somebody presses a button and moves to the next step. So for most of the world, the value creation of code has been to automate things and to be able to have it as a utility. So, so at the end of the day, like like we're and we'll probably still use the term slop for a while because because there's taste in kind of front-end design and there's taste in sort of systems and you don't want to have vulnerabilities in your code. So, that's going to exist for a while. But, at the end of the day, if you can tell an agent like, please go and generate my entire backend system or my front-end system, like it's just it's not only acceptable, it's it's preferable because it's just like that was the thing that was blocking us from moving forward. So, we need to go do that. At least the way society functions and the way the world works and our brains work at the moment. Maybe this changes. You know, when you when you get a presentation from somebody, there's still this association which is like I'm trying to decide if I can trust that person to go and execute on that thing or deliver that result or understand that topic. And so if when you see work slop you're like I my I'm losing my ability to sort of know for a fact that like like how much of the of the thought process was them versus how much was the AI? How much should I even care about that because I myself am doing the same thing. So like we have this weird like it's a it's this very weird sort of like collective issue that we have which is like which is like I'm doing work slop for some of my you know brainstorms and decisions but when I get it from somebody else I'm like hm should I trust you? uh and uh and I don't know I you know I mean it just might be a thing that as a society we have to kind of keep cranking through over the next kind of three to five years and and end up at the other at the other side like like you know I hate to use like these like totally busted analogies but you know obviously you don't care when you see somebody's financial model you're like yeah that was generated clearly by like a macro uh or or you know that was like not you did not personally go and compute all of that but but you're showing it to me and we're talking about it So, why can't the same exist for a strategy deck or or whatnot? But I think right now we're going through this evolution of like what is the person's role? What is the content a proxy for? Is it supposed to be a proxy for how much that person knows? Is it a proxy for what we think that they can go and execute on? I think I think we're just in this very messy period where we have to kind of figure that out.

Host

完全同意。嗯,他读了斯坦·达肯·米勒在《华尔街日报》上的文章吗?我读了,呃,我读了关于它的讨论。

Totally. Um, did he read the Stan Ducken Miller Wall Street Journal? I read the uh I read the discussion about it.

Aaron

是的。那个,呃,对它的反应,但实际上我没有读。但它是不是很草率?

Yeah. The the uh the reaction to it, but I actually I didn't I didn't read it. But was it like very sloppy?

Host

我不觉得它是爱。我喜欢它。所以对我来说,这只是一个很好的反例,我有这种过敏反应。

I didn't think it was love. I loved it. And so to me, it was just a nice counter example of I have this like allergic reaction.

AI检测与计算器类比 AI detection and the calculator analogy

Host

里面有多少个感叹号?

How many exclamation marks were in there?

Aaron

我觉得一个都没有,但它在 Pangram 里显示为 100% AI 生成。

I don't think there were any, but it does show up as 100% AI in Pangram.

Host

好。那破折号呢?

Okay. And dashes?

Aaron

破折号我觉得没问题。你不能那样做。对。但对我来说这是个很好的反例,因为通常我读到明显是 AI 写的东西,就会有一种过敏反应,而读 Stan 那篇时我却没有。我不确定这里面有多少只是因为,你知道,那是 Stan,所以我信任 Stan,而不是——

There, I think they're okay. You can't do that. Yeah. But it was a nice counterexample to me, because normally I read something that's clearly written by AI and I just have this allergic reaction, whereas with the Stan piece I didn't. And I'm not sure how much of that was just, you know, it's Stan, therefore I trust Stan, versus—

Host

对。

Yeah.

Aaron

不,但从心理上说这确实有点怪,因为我会读那些 X 上的文章,现在读它们要花两倍的功夫。我一边读内容,一边还在盘算:这是人写的,还是我就是在读一段 Claude 的提示词?然后我的大脑又开始处理:这该让我对这个人或这条帖子评价更高还是更低?我觉得因为这件事,我们要迎来一些很奇怪的时期了。我可不想当大学老师。我肯定会直接辞职,因为你根本不知道学生到底做了什么。

No, but it is psychologically kind of weird, because I'll read these X articles and I'm now doing twice the amount of work to read these. I'm reading it one, for the substance, and I'm also reading it two, for the calculation of did the person write it or am I just literally reading a Claude prompt. And then my mental processing is now like, should that upweight or lower my judgment of the person or the post? And I think we're in for some weird times because of this. I would hate to be a college professor. I would totally quit, because you're just like, I don't know anymore what you did.

Host

不过计算器似乎是最贴切的类比。

It does seem like the calculator is the closest analogy though.

Aaron

对,只不过计算器在范围上更有限,而且你仍然得自己把很多东西拼起来。这些类比有些正在失效,因为某种程度上这东西一次至少在做 10 件事。不过,是的。

Yeah, except it's just like that was more finite in terms, and you still had to piece together so many more things. And some of these analogies are breaking down, because it's like, well, at some point this thing is doing at least 10 tasks at once. But yeah.

Host

我们来聊聊 harness 吧。Box 是怎么——

Let's talk about harnesses. How does the box—

Aaron

过渡到 harness 很自然。好。

Great transition to harnesses. Okay.

Host

说到计算器,我们来聊聊 harness。

Speaking of calculators, let's talk about harnesses.

为Box构建智能体框架 Building an agentic harness for Box

Aaron

对,这又回到 NeoLab 那种应用层。我们对自己的文件系统、权限结构、搜索引擎有很多了解,当然,我们非常希望所有实验室都能基于我们对这些的理解来训练,因为那只会让外部智能体尽可能好。我们总是跟实验室说,嘿,关于系统怎么运作,你们想要多少数据我们就给多少。但撇开这个不谈,我们对人们在 Box 里做什么、怎么搜索 Box、面对 10 个文件时怎么决定挑哪一个、他们内部判断哪份文档最相关的算法或启发式,有很深的理解。所以这些我们都清楚,我们基本上构建了一个智能体式 harness,试图理解关于我们系统的这一整套领域知识。它显然能访问我们的搜索系统和文件系统。它有一堆机制,可以只提取文档的文本、提取文档的片段、即时对文档做嵌入。所以它有一组可用的工具。实际上它就是一个用来对大型数据集提问的 harness。在我的 Box 账户里,我甚至不知道最新数字,但大概有数千万个文件,因为这是 20 年来积累的一切。但我现在可以用 Box 智能体对所有这些数据提任何问题,它会四处搜索,一次做多次搜索,重新排序,然后很快提取出最相关的信息,有时还会读完整份文档,把所有步骤都走一遍。

Yeah, this is back to the kind of NeoLab applied layer. There's a bunch of things that we know about our file system, permission structures, our search engine, that certainly—and by all means, we would love all the labs to train on our understanding of this, because it would only make external agents as good as possible. We always talk to labs like, hey, we'll give you as much data as you want about how the system works. But that aside, we have a lot of depth of understanding of what do people do in Box, how do they search Box, how do they decide when they look through 10 files which is the one to go pick, what is their internal calculus or heuristic on figuring out the most relevant document to look at. So we know all of that, and we've basically built an agentic harness that attempts to understand that set of domain understanding about our system. It obviously has access to our search system, our file system. It has a bunch of mechanisms for just pulling out the text of a document, pulling out chunks from the document, doing embeddings on the document on the fly. So there's a set of tools that it can use. And effectively it's a harness for asking questions of a large data set. So in my Box account I have—I don't even know the latest number—but on the order of tens of millions of files, just because it's everything that has ever accumulated over 20 years. But I can now ask any question of all that data set using the Box agent, and it goes around, does multiple searches in one, it reranks, it then very quickly pulls out the most relevant information, and then it'll in some cases read the full document, does all the steps.

Aaron

然后我们把它跟另一种做法对比:如果我们直接把 API 给 Claude 或给 OpenAI 会怎样。结果我们在准确率和延迟上看到明显更好的表现,因为还是那句话,我们很清楚怎么针对自己的工作流去调优。所以这就是我们构建出来的 harness。

And then we compare that against, well, what if we just gave Claude our API or gave OpenAI our API, and we see meaningfully better results on accuracy and latency, because again, we know exactly how to tune it for our workflows. So that's effectively the harness that we built out.

哪些评估最重要 Which evals matter most

Host

那哪些 eval 对你最重要?

And then what evals matter the most to you?

Aaron

我有几个纯粹好玩的个人用例,会自己跟踪。但我们有,我不知道,几百个不同的测试,对每一个模型都跑。实际上我们目前有两个 eval。一个是我们发布的叫 complex work eval,是一组特定领域的工作,涵盖生命科学、金融服务、公共部门、科技等。它就是你想象中的那种以文档为中心的 eval。给定这五份文档和这组问题,你的答案会是什么?我们用智能体拿每个模型去测。然后我们还有一个 hold-back eval——其实第一个也是 hold-back——但第二个就是我们自己的 Box 实例,以及 Box 员工怎么使用他们的数据,然后我们再拿每个模型在那上面测一遍。所以我们能大致跟踪所有的渐进进展。模型提升半个点我们都能看到。然后我们根据不同的成本和准确率阈值推出默认模型。同时我们也让客户从我们的模型花园里任选任何模型。

So I have a couple of just funny personal ones that I keep track of, like my own use cases. But we have, I don't know, hundreds of different tests that we do on every single model. We actually have two evals at the moment. One is we put out a thing called the complex work eval, which is a set of domain-specific work in life sciences, financial services, public sector, tech, etc. And it's exactly what you'd think of as a document-centric eval. So given these five documents and this set of problems, what would your answers be? And we test every single model against those with our agent. And then we have a hold-back eval, which—actually the first one is hold-back also—but the second one is just like our Box instance and how Box employees use their data, and then we eval every model again on that. So we're able to roughly keep track of all the incremental progress. We see when things move by half a point in terms of model improvement. And then we roll out default models based on different cost and accuracy thresholds. And then we let customers also choose any model they want from effectively our model garden.

Box用例上的模型竞赛 The model race on Box's use cases

Host

你目前怎么看这场竞赛,就模型在你的用例上的表现而言,各家的马都处在什么位置?

What's your current view of the race and where all the horses are in terms of model performance on your use case?

Aaron

它们或多或少跟代码能力高度相关,只有一个例外,就是在我们某些用例里,Gemini 的表现好得不成比例,远超它在编程上的表现。这可能只是工具使用更好,考虑到 Gemini 生态以及他们要为之构建的东西。它也能解决一大类通用知识工作的用例。但我认为总体上还是跟代码相关。所以 Fable 5.1 显然是最先进的,也是我们见过最好的模型。关于其他模型当然有各种传闻,所以我们就看这场竞赛在这个方向上怎么继续。但基本上,当你看 GDPval、Mercor 的 Apex eval,这些东西大体上都会跟随编程模型。所以我觉得在 Grok、Muse、Fable 这一档,以及 GPT-5.6,不管他们接下来要造什么,大家就是并驾齐驱。现在这就是一场全面的竞赛。

They more or less closely correlate with code, with one exception, which is actually in some of our use cases Gemini is disproportionately better than what you would see from coding. And it might be just better tool use, given the Gemini ecosystem and what they need to build for. It solves a strong set of general knowledge work use cases as well. But I think mostly it correlates to code. So Fable 5.1 was clearly state-of-the-art and the best model that we've seen. There's obviously rumors about other models, and so we'll see how the race continues on this front. But basically by and large, when you look at GDPval, Mercor has their Apex eval, these things will all generally follow the coding models. And so I think we're just neck and neck on like Grok, Muse, the Fable class, and GPT-5.6, whatever they're building next. These are just—it's a total race right now.

模型选择与默认设置 Model choice and defaults

Host

你的客户通常会表达他们想用哪个模型的偏好吗,还是就直接用你们的默认模型?

And do your customers typically express a preference on which model they want to use, or do they just use your default?

Aaron

从用量上看,他们用的是我们的默认模型,因为它简单、效果极好,而且是专门调过的——我们的智能体有几种呈现方式。作为终端用户,你最常体验到的就是搜索、向你的数据提问。但从 token 用量上看,大部分流量走的是我们的工作流智能体或数据抽取。客户会在那里做评估,基本上就是说,好,我要确保在这个成本档位下,数据抽取能达到 98% 的准确率。这时候我们会派一个全职人员进去,帮你理解你的数据环境,对五个不同模型做测试,然后基本上就交给评估结果来决定了。

So by volume they use our default, because it's just easy and it works extremely well and it's tuned for—there are a few ways our agent manifests. The way you'd most commonly experience it as an end user is you'd just be searching and asking questions of your data. But by volume, the volume of tokens tends to go through more of our workflow agents or data extraction. That's where you have customers actually doing evals, basically saying, okay, I want to really make sure that at this cost profile I can get 98% accuracy on data extraction. And that's a place where we'll have an FTE that goes in and helps you understand your data environment, tests against five different models, and then you're just basically at the mercy of the eval.

Host

嗯。在你的客户群里,开放权重模型的采用情况怎么样?

Yeah. What are you seeing in terms of the adoption of open-weight models in your customer base?

Aaron

大概比人们以为的要高,比企业实际想要的要低,比五年后的水平要低得多得多得多得多。所以大概就是这几种说法的混合。

So probably higher than people think, lower than what enterprises actually want, and much much much much much lower than what it'll be in five years. So some mix of that would be the message.

Host

驱动这个决定的主要是成本吗?

And is it primarily cost that's driving that decision?

Aaron

我大概得把 30% 以上归因于那种新鲜感,就是——

I have to probably attribute 30-plus percent to just the sexiness of like—

Host

我想试试 GLM。

I want to try GLM.

Aaron

对,对。我觉得有这部分。我听过——我听过财富 500 强公司的 CIO 说,我们在这儿玩玩开源。我看着就想,我很确定,就你想要的这个成本档位,Gemini 或者 Muse 完全够用,甚至大概 5.6 Luna 或者 Terror 之类的,就是刚搞了疯狂打折的那个,大概也完全够用。但你就是想能说,好吧,我稍微对冲一下,挺酷的,我们还处在那个阶段。随着时间推移,我觉得可以合理推断,你会看到成本出现明显分化,因为你能把那些只有在把推理成本压到极致才划算的工作负载剥离出去,这种情况下开放权重就有经济优势。现在还有这个挑战,就是有时候它 token 效率更低,有时候随机地——我听过一些故事——随机地它会在链条中间突然说起中文,你就想,好吧,这对银行来说会很奇怪。所以这些事我们大概得想办法解决。但长期来看,我觉得必然会是,你会把那些工作负载剥离出去。关于这点我完全认同的一篇比较有意思的文章,是 Decagon 的 Jesse 写的。你大概读过那篇,讲这个悖论:我们会看到成本继续指数级增长。但接下来会发生的是,每一个成熟的使用场景,你都可以剥离到开源。一旦那个使用场景稳定下来,把它转向开放权重模型就开始说得通了,前提是以下两件事之一成立。第一,它确实更便宜;第二,做一些后训练能让你多出 X% 的性能。所以我觉得你会处在一个现实里,我们会——这对媒体来说大概比硅谷的人更让人困惑——因为你会想,等等,Anthropic、OpenAI 等的营收高得离谱,但开放权重不知怎么也在指数级增长,你会想,开放路由怎么会这么快?

Yeah. Yeah. I think there's that. I've heard—I've heard CIOs of Fortune 500 companies say, we're playing with open source here. And I look at that and I'm like, well, I know for a fact that Gemini or Muse would have been just fine at that particular cost profile you're trying to do, or probably even like 5.6 Luna or Terror or whatever, whichever one had the crazy discounting they just did. It probably would have been totally fine. But you want to be able to be like, okay, I'm a little hedged, it's cool, we're at that phase still. Over time I think it stands to reason that you'll see meaningful different costs, because you'll be able to peel off workloads that just only make sense at grinding down to the cost of inference, in which case open weights will have the economic advantage. Right now there's this challenge of sometimes it's more token-inefficient, sometimes randomly—I've heard stories—randomly it'll just speak Chinese mid-chain, so you're like, okay, well that'll be weird for a bank. So we need to probably work on some of those things. But long term I think it has to be the case that you're going to peel off those workloads. One of the more interesting posts on this that I totally subscribe to is Jesse at Decagon. You probably read that post about this paradox of like, we're going to see costs continue to go exponential. But what's going to happen is each use case that matures, you can peel off to open source. And once you have stability in that use case, it starts to make sense to veer it toward an open-weights model, assuming one of two things is true. One, that it's actually literally cheaper, or two, having some post-training gets you X% more performance. And so I think you will just be in a reality where we will—and this is going to be very confusing probably for the press more than people in the valley—because you'll be like, wait a second, the revenue of Anthropic, OpenAI, etc. are off the charts, but somehow open weights is also growing exponentially, and you're like, how is open routing so fast?

Host

是啊。而且就像这块饼增长得太快了。

Yeah. And it's like the pie is growing so fast.

Aaron

但实际发生的是,这里有一种有趣的双重性。甚至不只是水涨船高。而是,不不,我们要么用 Fable 或者 5.6 来做编排,然后把这些长尾任务外包给更便宜的模型。要么反过来,你有一个编排智能体,默认做便宜的事,但偶尔看到某个实在太难的东西,就把它弹给那些更重的模型。所以你可能两边各花 50% 的钱,但开放权重模型上的 token 量是 10 倍。所以大家某种程度上都在赢,但他们为什么赢之间是有相互作用的。

But what's happening is actually there's an interesting duality. It's not even just like rising tide lifts all boats. It's like, no no, we either use Fable or 5.6 for orchestration and then we farm out all these long-tail tasks to a cheaper model. Or the opposite is true, like you have some kind of orchestration agent that by default does the cheaper stuff but occasionally sees something that is just way too hard and then it pops it out to one of these heavier models. And so you might have blended 50% spend on each but 10 times the amount of tokens on the open-weights model. And so everybody's kind of winning, but there's an interplay between why they're winning.

研究员对律师:企业AI访问控制 Memory and personalization

Host

我很好奇你怎么看记忆、定制化或个性化,以及它会走向哪里。因为今天主流的架构似乎还是基于 RAG 的系统——你在 RAG 上玩出花样,但它仍然是上下文查找,权重本身并没有根本改变。看起来——我是说,如果我听我在实验室的朋友们说,持续学习,就是模型权重应该随着它了解你而适应这个想法。我们播客请过 Engram。不知道你认不认识 Dan。

I'm curious about how you think about memory and customization or personalization, and where that's going to go. Because it seems like today the dominant architecture is still kind of a RAG-based system—like you get fancy on the RAG, but it's still context lookup, where the weights themselves aren't fundamentally changing. It seems like—I mean, if I listen to my friends at the labs, continual learning, this idea that the model's weights should adapt as it gets to know you. We had Engram on the podcast. I don't know if you know Dan.

Aaron

我刚被介绍认识他。我听了那期播客。所以我很想成为房间里第四个人。

I just got introduced to him. I listen to the podcast. So I would have loved to have been the fourth person in the room.

Host

太棒了。对,他们——我觉得他们在和客户合作,基本上是把一些上下文烘焙进权重本身。你觉得这会往哪个方向走?

Amazing. Yeah, they—I think they're working with customers to basically bake in some of the context into the weights themselves. What direction do you think this is going to go?

Aaron

你知道,你正好在我跟 Dan 通话之前逮到我。所以我真希望我能先跟他聊过,那样我的回答会优雅得多。我对这个方法极其着迷。我没有任何理由不希望它成功、不希望它存在。我们在 Box 所处的世界里,看到权限、访问控制和数据方面的高度复杂性,这往往是这类方法的一个症结。我会把 Engram 放在一边,因为我确信他们已经想透了。所以我会更泛泛地、从哲学层面来谈。

You know, you're catching me at a time right before I'm actually doing my call with Dan. So I wish I could have talked to him first and then I'll have a way more eloquent answer. I'm extremely fascinated by the approach. I have no reason for not wanting it to work and exist. We live in a world at Box where we see this high degree of complexity on permissions and access controls and data that tends to be sort of the rub on a lot of these types of approaches. And I'm going to put Engram aside because I'm sure they've already thought this through. So I'm going to talk more generically, philosophically.

上下文对权重:决策点 The Researcher vs. the Lawyer: Access Control in Enterprise AI

Aaron

我觉得有时候你会跟一个研究员聊天,他想象世界是按照他的方式运作的,就像:我是研究员,我能访问一切,所以如果我有一个只在我的世界里训练的模型,那会很棒。然后你说,让我给你介绍一位律师,这位律师只有一个小小的访问点,只能看到他们正在处理的五个项目,因为隔壁有人正在做一个与另一家公司竞争的项目,他们不能有任何重叠,无论是他们看到的还是知道的。两堵墙之间不能有任何一份文件传递,必须是那种硬性屏障。所以,当然,你现在仍然可以只为那一个用户训练一个模型,但如果每天他们被添加或移除某些东西,而这些会为他们的理解增加重要背景,那会怎样?而且,我认为持续学习方面可能会有突破,能解决所有这些问题。但这就是为什么以前根本做不到——5 年前你不可能做到,因为成本高得离谱,而且你根本无法理解那些访问控制应该如何运作。但显然,随着成本曲线下降,开放权重变得更便宜、更小、更快、更好,我觉得这变得非常有趣。

I think sometimes you will talk to a researcher that imagines the world working the way they work, which is like, I'm a researcher, I have access to everything, and so if I had a model that was trained just on my world, this would be amazing. And then you're like, let me introduce you to a lawyer, and the lawyer has this tiny little access point of just the five projects they're working on, because somebody one door over is working on a competitive project to another company in the space that they can't have any sort of overlap with, with what they see or what they know. And there can't be a single document that passes between those two walls, and they have to be these kind of hard barriers. So sure, you could still now train a model just for that one user, but what happens if every single day they get added or removed from something that adds important context to what they need to understand? And again, I think there's going to be probably breakthroughs in continual learning that sort of all resolve this. But this is why previously it was just like, there was no way you could pull this off 5 years ago because it would be insanely expensive, impossible to wrap your head around how those access controls are supposed to work. But obviously, as the cost curve goes down, as open weights get cheaper, smaller, faster, better, I think this becomes super interesting.

Box Labs:应用AI团队与研究重点 Context vs. Weights: The Decision Point

Aaron

在播客上有一件事我觉得非常吸引人,我真的需要一张 T 形图来梳理,就是:什么放在上下文里、什么放在权重里,决策点是什么?你必须稍微思考一下,把权重烘焙进去能带来巨大性能提升的地方在哪里。可能有一些很妙的计算,比如当数据的变化率不那么远,但权重的收益能极大改变模型准确率时,你就得找到某种准则。我的意思是,如果你能挥动魔杖,几乎就像——我记得 Karpathy 在之前的一些采访中说过——如果你能几乎移除模型中所有记忆的信息,只让它封装特定的推理能力,比如我们在红杉做事的理念,然后你拥有所有实际内容和查找系统,那几乎就是如果你能挥动魔杖,系统会变成的样子。所以这一点非常有趣,我认为问题将是:在那个层面上,企业有多大不同,还是说实际上是它们字面上的知识产权让它们与众不同?世界上有多少种不同的执行风格,还是说不是,只是对那个特定法律案件的深度知识,以及我如何把它应用到我正在做的另一个项目上——价值就坐落在那里。但再说一次,如果你能等到我和 Dan 的 Zoom 通话,那时我就真的知道答案了。但我是粉丝,因为无论如何都会有——我已经跳到了个人层面,但无论如何,在公司层面,可能有办法采用这种方法。我非常喜欢 Trajectory 或 Applied Compute 在做的事情,或者 Prime Intellect,因为我认为毫无疑问,如果你是礼来,你想要一个关于如何进行药物发现的模型,那可能确实需要比现成产品走得更远、更深或更具体。而且药物发现工作流没有太多政教分离的问题。他们可能希望尽可能多的人能获取尽可能多的信息。所以我认为这将是领域特定的,你会根据不同的垂直领域或用例类型以及过程中防火墙需要设置在哪里,得到不同的结果。

One thing on the podcast that I found very fascinating and I just need like a T-chart honestly is just like, what is the decision point of what goes in context and what goes in the weights? You have to be a little bit thoughtful about where is the massive performance gain that you get by baking in the weights. And there's probably some incredible calculation of like when the rate of change of the data is not so far but the upside of the weights dramatically changes the accuracy of the model, like you'd have to kind of land on some sort of rubric like that. I mean, if you could wave a magic wand, it almost seems like—I think Karpathy has said this in some prior interviews—like if you could almost remove all the memorized information from the models and just have it encapsulate the specific reasoning capabilities, like the ethos of how we do things for example at Sequoia, and then you have all the actual content and a lookup system, that almost feels like if you could wave a magic wand that's what the system would look like. And so that one's super interesting and I think the question will be like how much are enterprises different at that level versus it's actually their literal IP that is what makes them different. Like how many different types of styles of execution are there in the world versus no, it's just like the depth of knowledge about that particular legal case and how do I apply it to this other project I'm working on—that's where so much of the value sits. But again, if you can just like wait till my Zoom call with Dan and then I'll really know the answer. But I'm a fan because no matter what there's going to be—I've jumped right into like the individual, you know, but like no matter what, at a firm level, there's probably ways to take this approach. Like I'm a big fan of what Trajectory or Applied Compute are doing, or Prime Intellect, because I think there's no question that if you're Eli Lilly, you want a model for how you do drug discovery and that probably does need to go farther or deeper or more specific than what you're getting off the shelf. And there's not a lot of church and state problems for drug discovery workflows. They probably want as much of that information available to as many people as possible. So I think it's going to be domain specific and you're just going to have different outcomes based on which vertical or type of use case and where the firewalls need to be in that process.

智能体世界中的记录系统 Box Labs: Applied AI Team and Research Focus

Host

有道理。好的。所以你一开始暗示有一个 Box Labs。Box Labs 在做什么类型的工作?

Makes sense. Okay. So you hinted at the beginning that there's a Box Labs. What type of work is Box Labs doing?

Aaron

所以它相当于 Box Labs——我不知道我们是否已经用了大写 L。但基本上它是我们的应用 AI 团队。

So it's the equivalent of Box Labs—I don't know if we've used capital L yet. But basically it's our applied AI team.

Host

你们团队目前最感兴趣的研究领域是什么?

What research areas are most interesting to your team right now?

Aaron

是的。所以最大的领域——最研究性的,或者你知道,在工程师的连续谱上,有一些集群更偏向研究。在那个集群中,我们花时间的事情仍然是在应用层,但很多是关于如何让智能体在给定问题下再提高 10 个百分点的准确率。所以,如何用给定的数据集构建问题集的地图,以最好地执行那个任务。所以我们花很多时间在那类工作上。例如,我们有一个团队在研究如何在智能体层面、在 harness 层面有效地进行某种自动研究,在给定客户数据的情况下,对回答问题或问题集的准确率进行爬山优化。所以如果你是一家银行,你有一堆贷款文件进来,这些是 100 页的文件,你用现成模型得到 70% 的准确率还是 97%,基本上显然在能否真正自动化那个过程上有天壤之别。所以你必须以某种方式从基础模型爬山到 97%,有很多工作投入到系统中来基本上实现这一点。

Yeah. So the biggest areas that—the most research kind of, or you know, of the continuum of engineers there's some cluster that is more on the research bent. And of that cluster, the things that we spend time on still again is at the applied layer, but it's a lot around how do you take agents and make them another 10 points of accuracy improvement given x problem. So what is—how do you build a map of the problem set with a given set of data to best execute on that task. So we spend a lot of time on that style of work. We have a team for instance working on how do you do effectively at the agent level, at the harness level, some form of auto research on hill climbing on accuracy of answering questions or sets of problems on given a set of client data. So if you're a bank, you have a bunch of loan documents coming in and these are like 100-page documents, whether you're getting 70% accuracy with an off-the-shelf model or 97% is basically obviously a world of difference in can you actually go and automate that process. So somehow you have to hill climb from the base model to the 97% and there's a lot of work going into the system to basically pull that off.

顿悟时刻 Systems of Record in an Agent World

Host

也许拉远一点,你认为——这可以是 Box 的问题,也可以不是 Box 特定的问题——在有智能体的世界里,记录系统(systems of record)的角色是什么?我肯定你看到了 Twitter 上的一些讨论,比如,现在每家软件公司都在试图向我推销他们自己的智能体。我不想要他们另一个智能体。我想要他们的记录系统与我的智能体良好协作。所以你怎么看?

Maybe zooming out, what do you think of as—and this can be a Box question or a non-Box specific question—the role of systems of record in a world with agents? And I'm sure you saw some of the Twitter discourse on like, every software company is trying to sell me their own agent right now. I don't want another agent from them. I want their system of record to work well with my agent. And so how do you think about that?

Aaron

标签云力量。

Hashtag cloud force.

Host

顺便说一句,名字很上口,非常上口。

Catchy name, by the way, very catchy.

Aaron

我的意思是,真的,就像那种事情,头三分钟你会想,天哪,这看起来很有趣。然后四分钟后,尤其是当你看到股票时,你会想,啊,绝妙的一步。这太棒了。我们也要这么做。

I mean, literally, it's like one of those things where the first three minutes you're like, man, that seems funny. And then four minutes later, especially when you see the stock, you're like, Ah, brilliant move. Like this is great. We're doing this.

你必须做的两件事 The Aha Moment

Host

然后有人把这件事说得最到位——他们说,当听到 Matthew McConaughey 亲口说出来时,这事就真的敲定了。

And then when somebody actually said this the best — they were like, when they heard Matthew McConaughey say it out loud, that really sealed the deal.

Aaron

那就是那个顿悟时刻。

That was the aha moment.

Host

那就是那个顿悟时刻。我当时就想,天哪,他居然能卖软件。这真的太不可思议了。他的嗓音太适合推销记录系统和智能体了。

That was the aha moment. I was like, man, he can sell software. It's actually incredible. His voice is so good for selling systems of record and agents.

示例:治理智能体 Two Things You Must Do

Aaron

如果你属于我们这类当代人——你搭建了一个 SaaS 平台,你有一套数据和流程,你的客户在其中运作——那么实际上你只需要做两件事,我认为任何试图只做其中一件而忽略另一件的人都会输。你必须构建一个在你的产品上强到离谱的智能体。这个智能体必须能被证明在使用你的系统时,比现成的智能体强 10 到 20 个百分点。不是因为对方被削弱了,而是因为你把评估做到极致,并且针对你特定的工作流做了如此精细的调优,以至于你能改进你的系统。你必须拥有这个。而且,因为你了解自己的领域——除非你完全在打瞌睡——你很可能有一些别人没想到要围绕其构建产品的用例,因为你每天都和客户交流,你看到他们遇到什么问题,然后你会想,哦,我们完全可以让我们的智能体替你做这件事。我至少有过十几次——我每年大概会以各种身份和几百个客户交流。我至少有二十几次遇到客户提出一个用例,对我来说是个突破时刻,我会想,那简直太疯狂了。

If you're in our sort of contemporary group of people who built a SaaS platform and you have some set of data and workflow that your customers operate in, there are effectively two things you just have to do, and I think anybody attempting to do one over the other is just going to lose. You have to build an agent that is insanely great at your product. That agent has to provably be 10 or 20 points better than an off-the-shelf agent at using your system. Not because it's hobbled the other side — it's just that you are so eval-maxed and so tuned to your particular workflow that you can improve your system. You have to have that. And you probably — because you understand your domain, unless you're totally asleep at the wheel — you probably have use cases that no one has thought to build products around, because you talk to customers every day and you see what they run into, and you're like, oh, we could just have our agent go do that for you. I've had at least a dozen — I probably talk to a couple hundred customers a year in various capacities. I've had at least two dozen times where a customer has a use case that is a breakthrough moment for me, like, that would be actually totally insane.

无头模式与API Example: Governance Agent

Host

比如什么例子?

Like what's an example?

Aaron

不幸的是,既然你让我现场举例,我不知道我的例子能不能配得上我刚才那种兴奋程度,因为它会……

Unfortunately, since you put me on the spot, I don't know if my example will pay off the level of excitement I just had, because it'll be...

Host

让你现场举例。

Put you on the spot.

Aaron

不,不。只是我觉得我想到的那个东西,对播客来说可能会是个“噗”的冷场效果。但基本上有个客户有个想法:他们想要一个后台智能体,试着判断文档何时符合他们的治理政策,比如某些东西是否需要进入某种归档,或者进入某种法律保留之类的。

No, no. It's just like I think the thing I was thinking of is just like I think it's going to be a womp womp for the podcast. But there was basically a customer who had this idea: they wanted an agent in the background trying to figure out when documents met their governance policies, and like does something need to go into some kind of archive or something into some kind of legal hold or whatnot.

Host

哦,酷。

Oh, cool.

Aaron

你看,正是——你看,那正是我担心的那种语气。是的。对。不,你根本装不出来。

And see, exactly — see, that's exactly the voice that was exactly the voice I was worried about. Yes. Yeah. No, you couldn't even pull it off.

Host

好吧。但在我们的世界里,这很棒,因为你会想,哦……

Okay. So but in our world, this is awesome because you're like, oh...

Aaron

对。

Yeah.

Host

不,因为你想,每家公司都有治理负责人。

Like, no, because think about it. Every company has a head of governance.

Aaron

我是真心说的,Cole。

I meant Cole sincerely.

Host

好吧。我知道。我相信你。我是说,听着,你做企业级,所以我觉得这至少有一半是认真的。想象你是一家企业。你有一个合规负责人和一个治理负责人。对吧?他们只能监督整个企业。他们从来没法同时出现在所有地方。现在想象一下,如果他们能坐在员工旁边,基本上能说,哦,你正要去做一些违反我们治理政策的事。所以这个想法就是,哦,如果有一个持续运行的智能体,自动地说,不行,那会违反你的治理政策。而不是让用户自己去预测或理解这些东西。所以总之,这类事情就是:如果你在产品里有一个智能体,你就能比市场上其他人更早、更好地识别出来,或者做一些在平台外可能做不到的事情。

Okay. I know. I believe you. I mean, listen, you do enterprise, so I think that it was at least half serious. Imagine you're an enterprise. You have a head of compliance and a head of governance. Okay? They can only be overseeing the whole enterprise. They've never been able to be everywhere at once. Now imagine if they could sit next to the employee and basically be able to be like, oh, you're about to go do something that breaks our governance policy. So the idea was like, oh, what if there was just an ongoing agent that automatically was just like, nah, that's going to break your governance policy. Instead of the user having to try and predict or understand this stuff. So anyway, those are the kind of things where if you have an agent within your product, you're going to be able to identify sooner and better than the rest of the market, and or just do things that maybe would be impossible to do off platform.

收费站商业模式 Headless and APIs

Aaron

另一方面,显然你必须走向无头化。你真的必须确保你的 API 暴露给 Claude 和 ChatGPT 以及所有不同的平台。你必须确保要么有直接进入确定性 API 的方式,让那些智能体可以通过 MCP 或其他方式使用你的 API 进行调用,要么至少让你的智能体变成无头的,并暴露在那些系统中。然后,这之所以可能还算是个有点难度的争论,唯一的原因是,作为记录系统,你必须确保找到一种方式,在商业上说得通,并且在另一边有价值、有意思。我认为很多不在这些公司里的人搞错的原因,就是低估了那些完全是增量收益的新用例的数量。它们对这些记录系统来说是完全的空白机会。比如在 Salesforce 的例子中,我今天使用 Salesforce 的频率可能比我以前高一个数量级,因为我通过 Claude 或 ChatGPT 用 MCP 接入它,所以我总是在问我们 CRM 系统内数据的问题。

On the other hand, obviously you have to go headless. You literally have to make sure that your APIs are exposed to Claude and ChatGPT and all the different platforms. And you have to make sure that you have either a direct way into deterministic APIs so that those agents can use your APIs to make calls via MCP or whatever, or at least make your agent be headless and be exposed in those systems. And then the only reason maybe this is even remotely a hard debate is you have to just make sure as a system of record that you can find a way where commercially it sort of makes sense on the other side and is sort of valuable and interesting. And the reason why I think a lot of people got that wrong that weren't in these companies was just underestimating the amount of new use cases that are just total upside. They're like complete whitespace opportunities for these systems of record. So like in the Salesforce example, I use Salesforce more today, probably by an order of magnitude than I ever have because I MCP into it via Claude or ChatGPT, and so I just am always asking questions about the data inside of our CRM system.

产品UI与聊天 Toll Booth Business Model

Host

那你觉得这是否意味着记录系统会变得更像收费站生意,以确保它们能抓住这个机会?

And do you think that means the systems of record become more toll booth businesses then, to make sure that they're capturing the opportunity?

Aaron

我不太喜欢这个词。因为没人在收费站有过好体验。

I don't love that term. Because like no one's had a good experience at a toll booth.

Host

我喜欢收费站。

I love toll booth.

Aaron

对,没错。你喜欢它胜过治理智能体。所以我会说,因为它们有深度的目的——组织工作流、管理数据、保护数据、提供护栏——那么真的就是,是的,你必须在另一边有某种以量驱动的商业模式。而且我只是觉得,如果你在为顾客解决真正的问题,它就会赚钱。这听起来很俗,但我跟 LinkedIn 的产品经理说过,如果我能直接 MCP 接入 LinkedIn,我可能愿意多付 10 倍的钱。如果我有办法总是了解,比如,这个 CIO 在做这件事,我需要联系一下之类的,我会掏钱的。

Yeah, exactly. You love it more than governance agents. So I would say that because they have a depth of purpose of organizing the workflow, managing the data, securing the data, providing guardrails, then it's really just — yeah, you have to have some kind of volume-oriented business model on that other side. And I just think there's like — if you're solving real problems for customers, it'll just make money. This is so cheesy, but I've told LinkedIn product managers I'd probably pay 10x more for LinkedIn if I could just MCP into it. If I just had a way of always understanding like, okay, this CIO is doing this thing and I need to reach out or whatever, I'll take my money.

Host

对。

Yeah.

Aaron

所以这些系统实际上基于它们拥有的数据,有着巨大的价值,如果你做得好,客户绝对会找到某种方式来回报你创造的价值。

So these systems actually have a tremendous amount of value based on the data that they have, and customers will absolutely find some way to reward you for that value creation if you're doing a good job.

企业中的智能体 Product UI and Chat

Host

非常有意思。也许相关,我们聊聊产品 UI。大型通用聊天机器人、聊天框智能体。这会不会成为未来人们使用 AI 的主导 UI,尤其是在应用层?

Super interesting. Maybe related, let's talk about product UI. Big generic chatbot, chat box agent. Is that going to be the dominant UI for how people are using AI in the future, especially when it comes to the application layer?

Aaron

这就是为什么我认为应用层有这么大的发展空间,因为那个通用的聊天系统——你问一个问题,得到答案,或者它在后台做一些工作——我认为显然它会是一个常青树,那个 UI 会一直存在。它会极其强大。

This is why I think the applied layer has so much room to run, because probably the universal chat system that you ask a question to, you get an answer back, or it does some work in the background — I think it's going to be obviously that's going to be a mainstay, that UI will always exist. It'll be incredibly powerful.

为何编程扩散快 Agents in the Enterprise

Aaron

横向产品会有,纵向产品也会有,所有人都会有。就像你的产品有一个搜索框,显然它会有。所以对于智能体的一次性请求——去找这个东西、回答这个问题、按需为我产出点什么——这种功能会一直存在。但大多数企业是由这些流程和工作流组成的,它们只是在幕后发生。有时是计算机在跑这些东西,有时是其他人在做,有时本该由人来做,但你永远负担不起让人来做,所以它们干脆就没发生。这和聊天机器人是稍微不同的隐喻——你问一个问题,它回一个答案。那更像是:好吧,我有点想让智能体在后台为我做事——读每一份合同、看每一条日志、对每一个安全事件做分诊——然后我不用聊天,也许聊天只是我用来跟进的方式。我要一个仪表盘,我要一个工作流,我要一个队列,我要一个任务列表。那么挑战就变成了:横向产品是否要承担所有这些组件,把所有这些体验都整合到一个产品里?如果是这样,我想你会开始觉得,天哪,那东西真重,然后我们会开始觉得,哦,这不再是那个简单、轻松、令人愉悦的东西了。所以纵向玩家其实理解流程,能为那个特定工作流呈现出正确的按钮、标签和名称。所以我认为,当智能体在后台做更多工作、做更多异步工作——就像我派出一堆智能体去审查正在发生的事情之类的——这更偏向于那些能理解这些工作流和流程的应用公司。我认为每个领域都会有这些。我们已经知道在法律领域它们会是什么样,有 Harvey、Lor 等等。我们看到它们开始在安全等领域出现。我们看到它们开始在 Cognition 和 Factory 的长时程编程智能体中出现。所以我认为那将是应用 AI 更大的用例之一。最终,五年后,我敢打赌企业里 90% 的 token 都是用户从未启动过的,他们只是看到一个结果。他们只是看到一个任务出现,然后必须去审查它,而它就是那样发生着。

The horizontal products will have it. The vertical products will have it. Everybody will have it. It's just like your product has a search box. Obviously, it does. So that's always going to be here for these sort of one-off asks of an agent: go find this thing, answer this question, or produce something for me on demand. But most of the enterprise is made up of these processes and workflows that are kind of just happening behind the scenes. Sometimes they're happening with computers running these things, or sometimes they're happening with other people doing these things, or sometimes they should be happening with people but you could never afford to have them happen with people, so they just didn't happen. And so that's a slightly different kind of metaphor than a chatbot where you ask a question and it comes back with an answer. That's like, okay, I kind of want agents in the background to do things for me: read every contract, look at every log, triage every security incident, and then instead of me chatting, maybe I'll chat as a means of doing a catch-up. I want a dashboard. I want a workflow. I want a queue. I want a task list. So then the challenge becomes: does the horizontal product take on every one of those components and manifest every one of those experiences in one? In which case I think you'll start to be like, man, that thing is really heavy, and then we'll start to be like, oh, this is no longer this simple, easy, delightful thing anymore. So then the vertical players actually understand the process and can manifest all the right buttons and tabs and the names of the things for that particular workflow. So I think as you have agents that are doing more work in the background, doing more async work—like I farmed out a bunch of agents to review things as they happen or whatnot—that leans more toward the applied companies that can understand those workflows and processes. I think you're going to have these in every field. We already know how they're going to look in legal with Harvey and Lor, etc. We are seeing them start to emerge in areas like security. We've seen them start to emerge in the long-running coding agents with Cognition and Factory. So I think that will be one of the bigger applied AI use cases. And ultimately, in five years from now, I would bet 90% of all tokens in the enterprise are things that a user never kicked off and they just see a result. They just see a task show up and they have to go review it, and it's just happening.

Host

嗯。嗯。有道理。好,我们来聊聊 AI 的扩散。编程智能体——就像砰的一下,2026 年 1 月 1 日发生了,我们见过的最快的东西进入经济的扩散发生了。而 AI 魔法其余部分进入我们其余工作的扩散似乎慢得多。你对此有什么看法?你认为哪些领域会看到更快的扩散,那又会如何发生?

Yeah. Yeah. Makes sense. Okay. Let's talk about AI diffusion. Coding agents—it was like boom, January 1st, 2026 happened and the fastest diffusion of anything into the economy we've ever seen has happened. The diffusion of the rest of the AI magic into the rest of our jobs seems like it's been a lot slower. What are your thoughts on that? And where are the areas where you think we're going to see faster diffusion and how is that going to happen?

知识工作自动化的连续谱 Why Coding Diffuses Fast

Aaron

是的,所以你总是得对比编程和其他一切,才能真正理解其中的差异。在编程里——这又回到关于垃圾内容的效用问题——代码的效用几乎 100% 体现在你能生成的文本量上。显然,大量的价值进入了文本,还有知识、专业能力、会议等等,但最终,文本是产出程序的东西,而程序才是你真正想做的事。如果你能拥有世界上最伟大的程序员,不用睡觉、不用吃饭,他们能凭直觉知道该构建什么,可以整天坐在电脑前,你的价值创造就会和他们坐在电脑前敲代码的小时数 100% 相关。代码行数——理想情况下是好的代码——将与你是否产出了人们想要的软件相关。所以全是文本。模型在这些数据上被超量训练。AI 实验室里的每个人都把编程当作一个竞争性基准,不断试图超越。他们每天都能对它做自己的评估,因为他们自己就是在给模型写代码。而且这是有史以来最技术化的受众:当他们部署一个智能体系统,遇到 bug 或问题,或者某个 MCP 服务器返回连接无效,他们会自己修。他们知道如何分诊问题。他们不会打电话求助。他们只会说,哦对,我没打开那个端口,抱歉,我去修一下。所以这大概有五点。哦,也许还有第六点:这是一个非常非常高薪的垂直领域,只要能获得 10% 或 20% 的生产力提升就自动有价值,更不用说 5 倍生产力提升了。所以,把编程拥有的这五六个有利于自动化的特性拿出来,再和其他所有形式的知识工作对比。你大概会得到一个直方图——我不知道有没有人发表过,也许你可以——关于与编程的相似度,以及随着你扩展出去,哪些领域开始看起来更接近、哪些越来越不像编程。然后你会发现,瞧,法律挺有意思,因为有人坐在电脑前审查法律文件、撰写法律文件、处理大量信息,能创造大量价值。好,所以那个正在爆发。然后你顺着列表往下看。现在拿销售代表来说。好,在相似度上就排得远多了。销售代表的价值创造基本上是说服外部客户从他们那里购买软件、技术或卡特彼勒卡车。这就是销售代表对经济的价值创造。那么,假设我们把世界上最好的自动化带给他们。首先,他们又得弄清楚怎么在技术上把它接起来。他们得确保把所有数据都给它,诸如此类。但无论如何,他们仍然受限于:客户回复他们了吗?客户想见面吗?他们能下周二见还是今天见?客户有预算吗?所有这些其他事情。

Yeah, so you always have to compare and contrast coding versus everything else to really understand the dissimilarity. So in coding—and this is back to the utility point on slop—the utility of code is almost 100% represented by the amount of text that you can generate. Obviously an insane amount of value went into the text, and knowledge and expertise and meetings and everything, but ultimately the text is the thing that produces the program that is actually the thing that you're trying to do. If you could have the world's greatest programmer who never had to sleep, never had to eat, they could intuit what to build and they could just sit on a computer all day long, your value creation would be 100% correlated with how many hours they could sit at that computer and type code. Lines of code—ideally good code—is the thing that will be correlated to whether you produced software that people wanted. So it's all text. The models are hyper-trained on these. Everybody in AI labs treats coding as a competitive benchmark to constantly try and exceed. They get to do their own evals on it every single day because they are the ones coding the models themselves. And it's the most technical audience of all time, where when they deploy an agentic system and they run into a bug or a problem or some MCP server comes back with connection invalid, they fix it. They know how to triage the problem. They don't call it in. They just like, oh yeah, no, I didn't open up that port. Sorry. I'll go fix it. So that's like five things. Oh, and maybe the sixth: it's just a very, very high-paying vertical that is automatically valuable if you could get 10% or 20% productivity gain, let alone 5x productivity gain. So take those five or six things that coding has as beneficial properties to automation, then compare that to every other form of knowledge work. And you'd probably have like a histogram—I don't know if anybody's published this, maybe you can—of the similarity to coding and what are the domains that start to look closer, and less and less like coding as you scale out. And it's like, okay, well lo and behold, legal is kind of interesting because there's a lot of value creation to somebody sitting at a computer reviewing legal documents, writing legal documents, processing large amounts of information. Okay, so that's kind of blowing up. And then you kind of go down the list. Now, let's take something like a sales rep. Okay, so much farther down the list in terms of likeness. The sales rep's value creation is basically convincing an external customer to buy software or technology or a Caterpillar truck from them. That is the value creation to the economy of the sales rep. And so, let's say we brought the world's best automation to them. First of all, again, they'd have to figure out how to technically wire it up. They'd have to make sure they give it all their data, all these kind of things. But no matter what, they're still rate limited and constrained by: did the customer respond to them? Do they want to meet? Can they meet next Tuesday or can they meet today? Does the customer have budget? All of these other things.

编程智能体对知识工作:没有GitHub时刻 The continuum of knowledge work automation

Aaron

这大概就是整个连续谱。一端是被许多外部因素限制住的人,另一端是可以整天坐在电脑前只打文字的人。你自动化事情的能力,就是这条连续谱。所以在真实世界里,我们必须把智能带入这些工作流,方式上要有点贴合他们工作流的形状、贴合他们工作的形状,然后找到办法交付变革管理、交付实施,把数据变成一种格式、放进一个真正能与这些系统配合的环境里。

That's maybe the entire continuum right there. On one end you have somebody rate-limited by so many external factors. On the other end you have somebody who could sit at a computer all day long and just type text. And that is your ability to basically automate things — is that continuum. So for the real world, we have to basically bring intelligence to these workflows in ways that somewhat feel like the shape of their workflow, somewhat feel like the shape of their work, and then find a way to deliver the change management, deliver the implementation, get data into a format and into an environment that actually works with these systems.

访问控制与企业现实 Coding agents vs. knowledge work: no GitHub moment

Aaron

另一个重点是:如果你在 2026 年去找大多数工程师——考虑到最新的现象,也许得减掉两个月——代码就在 GitHub 上。你只要连上 GitHub 就行。记得有那么一段时间,你发布一个编程智能体时,没有注册或登记,就是“把你的 GitHub 给我们”。知识工作里没有这个。知识工作里没有“把你的 GitHub 给我们”。

One of the other big things is: if you go to most engineers in 2026 — maybe minus two months ago, given the latest phenomenon — but the code's in GitHub. You just connect to GitHub. Remember there was this period where, when you launched a coding agent, there was no sign-up or register — it was just, give us your GitHub. That doesn't exist in knowledge work. There's no give us your GitHub for knowledge work.

Host

把你的 Box 给我们。

Give us your Box.

Aaron

嗯,Box 的客户做这一切要容易得多。不幸的是,我们的营收年化流水只有 13 亿。所以这意味着有很多人没用 Box。那他们在用什么?他们的数据在本地部署系统、遗留文件共享、遗留基础设施里,在企业环境里,而这些环境与智能体的沟通并不特别好。所以想想在实施编程智能体与实施其他一切之间的这种区别。

Well, Box customers have a much easier time with all this. Unfortunately, we're only 1.3 billion in revenue run rate. So that means there's a lot of people not using Box. And so what are they using? Their data is in on-premises systems, legacy fileshares, legacy infrastructure, enterprise environments that don't talk to agents particularly well. So just think about that distinction between implementing coding agents versus everything else.

准备两件事:缓慢扩散与应用层 Access controls and enterprise realities

Aaron

还有一件你会觉得无聊到睡着的事:企业里的访问控制完全不同。我其实完全忘了编程那一点。在编程里,你基本上能访问到与你工作相关的大部分东西。在知识工作里,你得说:嘿,Sally,你能帮我打开那个文件共享吗?你能打开那个项目吗,因为我没拿到访问权限。你怎么确保智能体有权限访问那一组东西?所有这些工作都得完成。

Even something again you'll sort of fall asleep about is: access controls in the enterprise are totally different. And I actually totally forgot that point about coding. In coding, you get access to basically most of the stuff ever relevant for your job. In knowledge work, you're like, hey Sally, can you open up that file share for me? Can you open up that project because I didn't get access to it. How do you make sure the agent has access to those set of things? All of that work has to get done.

创始人建议:进入AI对话 Two things to prepare for: slow diffusion and the applied layer

Aaron

所以我认为我们必须准备两件事。第一,硅谷必须准备好,扩散会比他们以为的慢得多——或者比“我们”以为的慢,但我真的觉得是他们,因为我知道要花多久。第二,好消息是,这一切都与应用层相关。

So the thing I think we have to prepare for is two things. One, Silicon Valley has to prepare for diffusion taking a lot longer than they think — or than I'll say we, but I really think they, because I know how long it'll take. And then the second thing is, the good news is this is all correlated to the applied layer.

Host

价值创造。

Value creation.

Aaron

对,价值创造。因为那些有耐心、有完整领域专业知识、有纯粹工作伦理的公司——因为不是所有人都会从闸门里涌出来。你必须出去跑腿、去实地推广。那将是应用层。所以我认为这实际上代表着应用层 AI 的一万亿美元价值:你怎么把技术带给律师、销售代表、生命科学研究员,或者客服团队负责人。这些机会现在就存在。

Yeah, value creation. Because the companies that will just have the patience, the full domain expertise, the sheer work ethic — because it's not like everybody's just from the floodgates. You have to go and pound pavement and get out there. That will be the applied layer. So I think this actually represents a trillion dollars of applied layer AI value: how do you go get the technology to the lawyer, or to the sales rep, or to the life sciences researcher, or to the person that runs the customer support team. That's all opportunity that exists right now.

你在哪里了解AI? Founder advice: getting into the AI conversation

Host

太棒了。最后我想问一些给其他创始人的建议。也许我们先从创始人建议开始,然后是公司建设建议。在创始人方面,看起来你出现在每一张 AI 的股权结构表上,每一家酷炫的新公司——你知道,Engram。你是怎么让自己置身于 AI 对话中心的?

Awesome. I'm going to close by asking some advice for other founders. Maybe let's start with founder advice and then company building advice. On the founder side, it seems like you are in every AI cap table, every cool new company — you know, Engram, you know. How did you kind of get yourself in the middle of the AI conversation?

Aaron

大概有两部分。第一,我对此准备得非常充分,比如做了 20 年非结构化数据。你一下子就能看到智能体在这上面的好处。所以老实说,我们直到现在才进行这场对话,比我期望的要久,因为我们八九年前、十年前就试图进行这场对话了。

There's probably two parts. One, I was just very well primed for it, like working with unstructured data for 20 years. You just can instantly see the benefit of agents on that. So it honestly took longer than I would have wanted that we got to have this conversation, because we tried to have this conversation eight, nine, 10 years ago.

Host

现在它终于发生了。

And now it's finally happening.

Aaron

所以首先,准备得非常充分。我们产品的形态以及人们用我们产品做的事,已经非常适合智能体。所以显然我们得把公司押在这上面,然后你瞧,我们得到了正向反馈循环,客户真的说:是啊,如果我能读每一份文档并回答任何问题,那会非常强大。所以这是第一点,也肯定是我最大的因素,大一个数量级。然后另一点就是,我对这项技术极其着迷。而且它很好玩。

So first of all, just super well primed. Our product sort of shape and what people do with our product already lends itself extremely well to agents. So obviously we had to bet the company on that, and then lo and behold, we had the positive feedback loop of customers actually saying, yeah, that would actually be very powerful if I could go read every document and answer any question. So that's the first and certainly by far my biggest, by a factor of 10. And then the other is just like I'm extremely fascinated by the technology. And it's just fun.

最喜欢的新AI产品 Where do you learn about AI?

Host

你从哪里了解它?

Where do you learn about it?

Aaron

你的播客。Doresh 的播客。Twitter。我是说,可能对脑细胞不幸的是,大概 95% 是 Twitter。我有个习惯,每晚结束时我就刷信息流,就像——我大概看起来像某个悲伤的梗图,一直刷一直刷一直刷,试图三角定位所有信息。

Your podcast. Doresh's podcast. Twitter. I mean, probably unfortunately for brain cells, it's probably 95% Twitter. I have a routine where at the end of each night I just go through the feed, and it's just like — I do look like some sad meme probably, just scrolling and scrolling and scrolling and attempting to triangulate all the information.

Host

太棒了,因为它是 AI 的全球城镇广场。

That's amazing, because it is the global town square for AI.

Aaron

确实是。我想给所有大二或大三学生发紧急警报,就说:在 Twitter 上关注这 20 个账号,并且加入 Twitter,因为这会对你的职业生涯有帮助。仅仅根据你的信息流,你要么领先一年,要么落后一年。而我的信息流非常深入。

It is. And I want to send an emergency alert to everybody who's like a sophomore or junior in college and just be like: follow these 20 accounts on Twitter and also join Twitter, because this will just help your career. You're either a year ahead or a year behind simply based on your feed. And my feed is so wired in.

Host

这是当人们问如何跟上 AI 时我给的第一条建议。就是:关注这一百个账号。

This is the number one advice I give to people when they're asking how to get current on AI. It's like, follow these hundred accounts.

Aaron

就像,但我的意思是,实际上第一步:加入 Twitter。

It's like, but I mean virtually first step: join Twitter.

Host

对。

Yeah.

Aaron

我还会跟 20 岁的人聊天,他们说:是啊,我看到一些文章。我就想,你什么意思?你怎么看到文章?我甚至不知道那是什么意思。你是碰巧有人给你邮件发了篇文章吗?加入 Twitter 吧。你在说什么?所以你必须深入其中。我很享受。很好玩。我什么都玩。

I'll still talk to 20-year-olds that are like, yeah, I see some articles. I'm like, what do you mean? How do you see articles? I don't even know what that means. Do you just get lucky that somebody emailed you an article? Just join Twitter. What are you talking about? So you just have to be wired in. I enjoy it. It's a lot of fun. I play with everything.

个人AI使用与基础设施缺口 Favorite new AI product

Host

你最喜欢的新 AI 产品是什么?请说 Instinct。

What's your favorite new AI product? Please say Instinct.

Aaron

好吧。完全免责声明,我还没用 Instinct 的邀请码,只是因为我积压了大概三个其他个人助理产品。

Okay. Full disclaimer, I have not done the Instinct invite code yet, simply because I have a backlog of like three other personal assistant products.

Host

你得试试 Instinct。

You gotta try Instinct.

Aaron

我知道,我知道。大家都在夸。我很兴奋。我在用。

I know, I know. Everybody's raving. I'm very excited. I'm on it.

Host

而且我不是投资人。所以这完全是真心的。

And I'm not an investor. So this is like totally genuine.

Aaron

哦,你不是?好吧。所以这完全是真心的。好吧。

Oh, you're not? Okay. So this is like totally genuine. Okay.

公司级AI采用与最佳实践 Personal AI Use and Infrastructure Gaps

Aaron

所以我百分之百会——我不知道这期什么时候播出,但我肯定在播出前已经玩过了。我现在正在试用几个预发布版本。我觉得个人助理类的东西非常令人兴奋。至少在我的某些用例中,它仍然显示出浏览器使用的一些局限。我们在这方面还有工作要做。可能整个互联网都需要为他们的产品提供一个 CLI。所以要让这些东西变得完全出色,我们在基础设施层面还有一些基础工作要做。然后我就像你典型的那种被 AI 渗透的知识工作者——每天我都会向五个不同的 AI 系统之一问 20 到 30 个问题,比如做研究、寻找人才、看看竞争对手在做什么、这个市场发生了什么、如何在那里扩张?差不多就是这样。

So I 100% will have—I don't know when this is going to run, but I'm sure I will have played with it by the time it runs. I'm in pre-release in a couple right now that I'm spending some time with. I think the personal assistant stuff is super exciting. At least in some of my use cases, it's still showing some of the limits of browser use as an example. We still have some work to do there. Probably the entire internet needs just a CLI for their product. So we still have some blocking and tackling at the infrastructure level for these things to be totally awesome. And then I kind of look like your average AI-pilled knowledge worker—every day I'm asking one of five different AI systems 20 to 30 questions on just doing research, looking for talent, looking for what a competitor is doing, what's happening in this market, how do you expand there? So kind of that.

Token使用与内部培训 Company-Level AI Adoption and Best Practices

Host

那在公司建设方面呢?你们现在是一家 20 年的公司了,对于那些试图重塑公司、确保 AI 不仅在个人层面、也在公司层面得到最大程度采用的人,你有什么建议?你如何让你的业务对 AI 可读?

What about on the company building side? How do you—you know, 20-year-old company at this point—what's your advice for other people that are trying to reinvent their companies to make sure they have max adoption of AI, not just at the individual level, but also at the company level? How do you make your business legible for AI?

Aaron

是的,同样,其中一些是我们一直以来思考公司信息系统的方式的副产品。毫不夸张地说,如果你有一个关于业务的问题,只要它曾经以非结构化数据的形式被记录过——比如会议记录、项目计划、路线图、演示文稿、财务文件、财务规划会议——它 100% 都在 Box 里。所以我们受益于一个已经为此极度优化的数据架构,因为我们就是这样运营公司的。我们不让任何人用其他东西。所以数据对我们来说非常容易大规模处理。当然,我们还有 Salesforce 和其他核心的规范系统。我们的数据卫生做得相当好,所以在其上运行智能体让我们在运营中更容易做到 AI 优先。也许有一些最佳实践或我们见过的事情:首先,试图找出最高杠杆影响的工作流,并针对它们。我们的 CIO 非常 AI 化。我们有一个小型的 AI 卓越中心。我们雇佣了一些内部 AI 全职员工来帮助这些流程。

Yeah, again, some of this we have as a byproduct of how we've always thought about information systems in the company. So to no exaggeration, if you have a question that you would like to ask about the business that has ever been documented in a form of unstructured data—so a meeting note, a project plan, a road map, a presentation, a financial document, a financial planning session—it's 100% in Box. So we benefit from a data architecture that already is insanely tuned for this because this is how we've run the company. We didn't let anybody use anything else. And so the data is very easy for us to be able to work with at scale. And then of course we have Salesforce and all the other core canonical systems. We've been able to have, I think, pretty good data hygiene, so agents running on top of that makes it a little bit easier to be AI-first in how we operate. Maybe a couple best practices or things that we've seen: first of all, trying to figure out where the highest leverage impact workflows are going to be and trying to target those. Our CIO is very AI-pilled. We have a little bit of a center of excellence on AI. We've hired some internal AI FTEs to help with these processes.

Slack同事智能体 Token Usage and Internal Training

Host

你们有排行榜吗?

Do you have leaderboards?

Aaron

我们不是唯 token 论。我们确实有一个按 token 数量排列的人员名单,但通常是为了去检查:好吧,我们认为这有用吗,或者有没有什么经验可以带回其他部门?所以可能在 Box 排名前三的 AI 用户中,大概有一个是在浪费一半的 token,然后另外两个,无论他们在做什么,我们都需要去为其他人做一个内部培训。所以两周前我们有过这样的事,我说,你能不能把这个特定团队的所有人叫到一个房间里,就展示一下这一个人是怎么用 AI 的?然后六个小时后,他们就在一个房间里,那家伙在做他做的事情的完整演示。向 Mick 致敬。所以这就是我们试图做的事情,就是如何向每个人展示以这种方式工作是什么样子。但尽管我们行动很快——我们在技术栈的某些部分交付的实际面向客户的产品多了两三倍。所以我不在乎写了多少代码,而是我们是否真的交付了客户要求的更多功能?在技术栈的某些部分我们做到了。然后你去和 Anthropic 的朋友聊聊,他们就像——你会说,天哪,我们还有很长的路要走。这个特定的 meta 不断变化。两年前你会说,你只是在 IDE 里用一个插件,你会说,天哪,我们不会那样——我们还没准备好完全 AI 优先。然后你说,好吧,大家都推出 Cursor。然后大家都推出了 Cursor。然后这终于发生了,然后你四处走动,你说,你不只是在 Slack 里工作,只是 @ 机器人来做你的代码。你在干什么?就像我们不断在改变工作流范式是什么。

We don't token max. We do actually have—I guess literally we have a list of people by number of tokens, but it's usually to go and inspect: okay, do we think that's useful or is there a learning there that we should take back to some other function? So probably our top—in the top three all AI users at Box, probably one of them is wasting half the tokens, and then two of them are like, oh whatever they're doing, we need to go do an internal training session for everybody else. So we had this thing like two weeks ago where I was like, can you just get everybody in a room on this particular team and just show them how this one person is using AI? And then six hours later, they were in a room and the guy was doing a full demo of what he was doing. Shout out to Mick. And so that's the kind of stuff that we're trying to do, which is just like how do we show everybody what it looks like to work in this way. But for as fast as we're moving—and we are shipping at some parts of the stack two or three times more actual customer-facing products. So I don't care about how much code, but did we actually deliver more functionality the customers are asking for? In some parts of the stack we're doing that. Then you'll go talk to a friend at Anthropic and they're just like—and you're like, oh my god, we still have ways to go. The particular meta constantly changes. Two years ago you'd be like, you're just using a plugin in your IDE and you're like, man, we're not going to be that—we're not ready to fully be AI-first. And then you're like, okay, everybody roll out Cursor. And then everybody rolls out Cursor. And then that finally happens and then you're like, you're walking around, you're like, you're not only working from Slack, just at-mentioning bots doing your code. Like, what are you doing? It's like we're constantly just changing what the workflow paradigms are on this.

在AI中获胜与创始人经验 Slack Co-worker Agents

Host

你们有没有类似 Slack 同事智能体这样的东西,找不到更好的说法?

Do you guys have like a Slack co-worker agent for lack of a better term?

Aaron

我们有几个是那种形态的。我对 Claude 标签这种形式很兴奋。你还是得把团队结构和数据弄对。但我们有各种方式让人们可以在 Slack 中与智能体协作。我不知道我们是否像 Benny Off 希望的那样被 Slack 渗透,或者像 Anthropic 或 OpenAI 那样,但我们正朝着那个方向前进。

We have a few that have that shape. I'm pretty excited about Claude tag as a form factor. You still have to again get the team construct right and the data right. But we have a variety of ways that people work with agents in Slack. I don't know if it's as Slack-pilled as Benny Off would like us to be, or as Anthropic or OpenAI, but we're heading in that direction.

安全优先业务的约束 Winning in AI and Founder Experience

Host

是的,我认为历史书会记载这个时代的商业艺术。我很想读一读新的《孙子兵法》,结合这里发生的一切,因为我认为我们看到的都是非常非凡的东西。你认为在 AI 时代取胜需要什么,与 AI 之前相比?现在当创始人的感觉与你们创办 Box 时相比如何?

Yeah, I think that history books will be written about the art of business in this time. I would love to read the new The Art of War with everything that is happening here because I think it's pretty extraordinary stuff we're seeing. What do you think it takes to win in AI versus pre-AI? And what does it feel like to be a founder right now versus when you started Box?

Aaron

我既嫉妒又不嫉妒你遇到的年轻创始人,因为你会想,哦,要是能再年轻一次,整个世界都是你的牡蛎,你可以朝任何方向走,你拥有的杠杆。你会见到他们,你会说,天哪,我一周半前看到一个产品,他们做了产品演示,我当时想,这绝对在五年前会是一个 40 人的项目——或者尤其在我们创业时,轻松就是 40 人——而它只是两个人。你就会想,你怎么有这么多能用的标签页,而且每个标签页背后似乎都有非常功能性的东西?这不是假的。所以我非常嫉妒这一点。这太不可思议了,因为你可以从零开始,以这个为设计原则来创办公司。现在我们也会达到那里,因为我们就是会硬着头皮挺过去。

I am both jealous of and also not jealous of the young founders you meet because you're just like, oh to be young again and the whole world is your oyster and you can go in any direction and the leverage you have. You'll meet with them and you're like, oh my god, I saw a product a week and a half ago and they did a demo of the product and I was like, absolutely this would have been a 40-person project five years ago—or especially when we were starting out, easily 40-person—and it was two people. And you're just like, how do you have so many tabs that work and they all seem to have stuff behind the tabs that all seem very functional? Like this is not fake. And so I'm very jealous of that. It's like incredible because you could just start your company from scratch with that as the design principle. Now we will get there because we're just going to muscle through it.

分发作为获胜层 Constraints of a Security-First Business

Aaron

有几件事我们就是做不了。对于取消代码审查以及人们讨论的某些做法,我们感到非常不安,因为如果我们不认真对待数据安全与合规,客户就不可能把他们的数据托付给我们。所以,由于我们在技术栈中的位置和我们的业务性质,我们在生产力上总会打一点折扣。但我真的很羡慕能够轻装上阵。

There's a couple things that we just can't do. We're very uncomfortable with the idea of removing code review and some of these things that get talked about, because our customers can't possibly entrust us with their data security and compliance if we don't take that seriously. So we're always going to have a little bit of a discount on the productivity because of where we are in the stack and what we do as a business. But I'm so jealous of being able to be fresh in that.

Aaron

但另一方面,每一个好点子,立刻就会冒出五个竞争对手。我们以前没有这个问题。我们有过几年好日子,可以专心打磨产品和体验。不会每三天就听到“哦,红杉投了这个,Benchmark 投了那个”,我们当时也没为此抓狂。当然我们也有自己的版本。所以当时我可能也在抓狂,但回过头看,那根本不值得抓狂。现在呢,天哪,每个市场都是真正的竞赛。

And then on the other hand, for every great idea, it's like instantly five competitors. We didn't have that problem. We had a good couple of years where we could just grind on our product and our experience. It wasn't like every three days you were like, "Oh, Sequoia funded this thing and Benchmark funded this thing," and we weren't going nuts with that. Now we had our own version of that. So at the time I probably was going nuts, but in retrospect it was not worthy of going nuts. Now it's like, oh man, this is a real race in every one of these markets.

结束语 Distribution as the Winning Layer

Host

完全同意。

Totally.

Aaron

所以我想我在这点上大概是有共识的:在一个 AI 能更快构建事物的世界里,优势可能转移到那些真正能把产品送到客户手中的人。所以和那些对此非常坚定的创始人聊天很有意思。Scott 或 Matan 就是这样,他们抓住了使命。他们觉得这将是企业级扩散的战场,所以你必须把它带给企业。现在任何搞错使命的人,你就输了。就像游戏结束。抱歉。在应用层, literally 有数万亿到数万亿美元等着被争夺,而那些知道如何组建团队并触达企业的公司将会胜出。这几乎是板上钉钉的。

So I think I'm probably pretty consensus on this, which is in a world where AI builds things so much faster, then probably the shift moves to whoever can actually get it to the customer is in the best position. So it's fun talking to founders that are pretty kind of pled on that. Scott or Matan are like they get the mandate. They're just like this thing is going to be an enterprise diffusion play, and so you have to get it to the enterprise. And anybody who kind of mistakes the mandate right now, you're just going to lose. It's just like it's game over. Sorry. Like there is quite literally trillion to trillions up for grabs at the applied layer, and the companies that know how to build the teams and get to the enterprise will be the ones that win. It's just like obviously guaranteed.

Closing Remarks Closing Remarks

Host

说得好,Aaron。这次对话非常有趣。非常感谢你的参与。

Well said, Aaron. This is a very fun conversation. Thank you so much for joining.

Aaron

谢谢邀请。

Thanks for having me.

互动版:逐字朗读 + 针对本期提问 →