Factory CEO 谈客户痴迷与产出指标及企业 AI 战略

Factory CEO on Customer Obsession vs Output Metrics and Enterprise AI Strategy

马坦·格林伯格 Matan Grinberg · Training Data · 2026-07-21 · 约 52 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Factory 联合创始人兼 CEO Matan 讨论了产出指标相对于客户痴迷等输入指标的重要性,以及 Factory 如何通过模型独立性和模块化在企业 AI 市场中脱颖而出。

Matan, co-founder and CEO of Factory, discusses the importance of output metrics over input metrics like customer obsession, and how Factory differentiates in the enterprise AI market with model independence and modularity.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 16)

全文 · Full transcript(中英对照)

客户至上vs产出指标 Customer obsession vs output metrics

Host

贝佐斯在亚马逊讲的是客户痴迷。但在我们看来,那是一个输入指标。你不想衡量输入指标。你是否客户痴迷并不重要。比如,你可能客户痴迷,但他们却对你申请了限制令,因为他们不喜欢你做的事。我们的工作是打造出足够好的产品,让客户自己对我们痴迷。这就是我们的工作。打个比方,如果你是篮球队教练,你不会在球员上场前说:‘嘿,大家记得要流汗。’那算什么?不,你要说‘得分’。我们需要得分,而在这个过程中你确实可能会流汗。同样,要创造痴迷的客户,你可能需要自己真正痴迷于客户,但产出才是关键。今天我们在演播室请到了 Factory 的 Matan。这是我们第二次邀请 Matan。

Bezos at Amazon, it's customer obsession. But in our mind, that's an input metric. Like, you don't want to measure input metrics. It doesn't matter if you're customer obsessed. Like, you could be customer obsessed and they file a restraining order against you because they don't like what it is that you're doing. Like, our job is to build something so good that our customers themselves become obsessed with us. That is our job. It's like, you know, the analogy is if you're a coach of a basketball team, you don't want to tell your players before they come out there like, 'Hey guys, make sure to sweat.' It's like, what? Like, no. Like, score points. like we need to score points and in doing so yeah you're probably going to sweat and I think similarly to create obsessed customers you probably need to be really obsessed yourself with the customers but the output is what matters. We're here in the studio with Matan from Factory. This is our second time with Matan.

Matan Grinberg

谢谢邀请。

Thanks for having me.

Host

你属于第二次参加 Training Data 的精英小群体。所以,谢谢你。

You're in the small and elite group of second time training data attendees. So, thank you.

Matan Grinberg

哦,是的。

Oh, yeah.

Host

Matan 是 Factory 的联合创始人兼 CEO,该公司制造 droids,即用于软件开发艺术的自主智能体。

Matan is the co-founder and CEO of Factory, which makes droids, which are autonomous agents for the art of software development.

Matan Grinberg

确实如此。

Yes, indeed.

Host

Matan,我们直接进入正题,因为我认为你们在软件开发领域算是一匹黑马。这个市场已经彻底起飞了。有像 Claude Code 和 Cognition 这样的公司已经领先,但你们正在迎头赶上。谈谈竞争格局和 Factory 的独特之处吧。

And Matan, we're going to jump right in because I think you guys are a little bit of a darkhorse candidate in this world of software development. It is a market that has absolutely taken off. There are folks like Claude Code and Cognition and others who who have a lead, but you guys are coming up strong. Talk about the competitive dynamics and what makes Factory special.

工厂历程与模型独立 Factory's journey and model independence

Matan Grinberg

这是一段疯狂的旅程。我们三年前创办了 Factory,确切地说是 2023 年 4 月。当时整个世界,尤其是企业界,连 GitHub Copilot 都还没准备好接受,更不用说完全自主的智能体了。所以头两年我称之为‘沙漠之旅’,因为我们专注于完全自主的智能体,但工程师们没准备好,企业的采购团队也没准备好。现在回想起来,我们确实磨练了技艺,学到了很多如何为企业开发者构建产品,但直到他们准备好接受,确实花了很长时间。现在我们正在更多地崭露头角,而其他一些玩家,比如 Anthropic 或 OpenAI,拥有巨大的分发渠道,正在推出像 Claude Code 或 Codex 这样不可思议的工具。我们在那两年里学到的是,企业真正关心的是他们不希望任何人成为他们的单点故障。他们不希望任何人控制他们的命运。所以真正重要的是模型独立性。每个人都从云服务中吸取了教训,当年 AWS 或 Azure 会说:‘嘿,进来吧,签个三年合同,会很便宜,我们会补贴,会很棒。’然后几年后,续约时他们会把合同价格提高 10 倍。

It's been a wild ride. We started Factory three and a half years ago now. So in April of 2023 when the world and the enterprise in particular was barely ready for GitHub Copilot, let alone fully autonomous agents. And so um I think the first two years it was kind of our journey in the desert is is how I like to refer to it because we were focused on fully autonomous agents but engineers weren't ready procurement teams at the enterprise weren't ready and so I think retrospectively we really like honed our craft and learned a lot about how to build for developers in the enterprise but um you know it it it took a lot of time to actually come around to when they were ready to receive it. And so we're kind of now emerging much more and some of these other players like Anthropic or OpenAI who have a ton of distribution um are going in and you know bringing their incredible tools like Claude Code or Codex. The thing that enterprises are really caring about that we have learned through those two years is they do not want anyone to kind of be their single point of failure. They do not want anyone to kind of control their fate. And so something that really matters is model independence. Everyone learned from cloud where, you know, back in the cloud days, it was like AWS or or Azure being like, 'Hey, you know, come on in, sign this three-year contract. It's going to be so cheap. We're going to subsidize it. It'll be great.' And then a couple years later, when it came time to renewal, they would 10x the the contract.

Host

哈哈,数据引力,现在抓住你了。

Haha, data gravity, we got you now.

Matan Grinberg

是的,抓住你了。你能怎么办?花两年迁移到别家?不可能。现在每个人都有伤疤。所以每个人都知道,Claude Code 很棒,OpenAI 的 Codex 也很棒。但我们不能把命运交到任何一个模型提供商手里。而且你看看模型实验室和云提供商的风险状况。最近云提供商和模型实验室之间有什么戏剧性事件?似乎总是有某种混乱,比如内部斗争,或者与政府或其他实体发生争执。所以,如果你要构建业务中这个非常重要的部分,你希望确保自己对任何变化都有韧性。这就是我们在最初两年里学到的:开发者非常关心模块化。他们希望知道自己可以定制成想要的样子。他们希望知道,如果出现一个更快、更便宜或性能更好的新模型,他们可以热插拔。我认为这就是为什么许多大型企业将他们在 Codex 或 Claude Code 上的势头带到 Factory 的最大原因之一,因为他们从这些出色的模型中获得了性能,但避免了直接来自模型实验室的供应商锁定。

Yeah, we got you. What are you going to do? A two-year migration to go to someone else? Like, no way. Everyone has scars from that now. And so everyone knows look Claude Code is fantastic. Uh codex from openi is fantastic. We cannot put our fate in any one of these model providers hands. Also like you just look at the risk profiles of the model labs versus the cloud providers. What's the last piece of drama that came out of a one of the cloud providers versus like the model labs? It seems like there's kind of always some sort of chaos of you know internal fighting or getting in spats with the government or you know any other entities. And so, uh, if you're going to, you know, build this very important part of your business, you want to make sure that you're robust to any of these changes. And that's something that we've learned over those kind of initial two years is like developers really care about things being modular. They want to know that they can customize it to what they want. They want to know that if there's a new model that comes out that's faster or cheaper or more performant, they can kind of hot swap it in. And that's I think one of the biggest reasons why a lot of the largest enterprises are taking the momentum that they had from a codex or a Claude Code and then are carrying that into factory because they get that performance from these fantastic models but they do it without the vendor lock in that you know the model labs directly from

Host

如果我是企业,我会想:等等,我现在是不是又被 Factory 锁定了?对此你怎么回答?

And if I'm the enterprise I'm going be like wait a minute am I now just getting locked into factory what's the answer to that

Matan Grinberg

问得好,这确实是个好问题,因为你可能会想:好吧,我们只是换了一个锁定点。我们构建的所有模块化设计都意味着,如果有一天你想说:‘嘿,Factory 不再处于前沿了。’无论是你构建的自动化,还是我们帮你创建的技能注册表,我们做的工作都留在你的代码库中,我们创建的任何自动化,其工件也都在你的代码库中。换句话说,我们并没有保留关于你组织的所谓‘部落知识’而不给你。这是我们与客户关系的一部分,我们同样希望确保提供最好的体验。如果我们帮你通过不同模型之间的套利实现成本优化,我们会把这种优化给你,不会把它拿走。我认为这是我们与这些企业建立信任的非常重要的一部分。

So good it's a really good question because that is something that you might think of like okay wait so we're just switching the the lock in point all of the modularity that we build is such that if at some point you wanted to say, 'Hey, you know what? Factory is not staying at the frontier anymore.' Whether it's like the automations that you build or the skills registry that we help you create, the work that we've done stays in your codebase and any of the automations that we've created, the artifacts also live in your codebase. In other words, there aren't really things that we're saying like are tribal knowledge about your org that we're keeping on our side and not giving to you. Um, and that's part of the relationship that we have with customers is like we similarly want to make sure we're providing the best experience possible. If we help you arbitrage between different models to get cost optimization, we're giving you that optimization. We're not taking that away from you. Um, and I think that's a really important part of the trust that we're building with with these enterprises.

先行者与沙漠之旅 Being early and the desert journey

Host

大概几个月前我们聊过,我当时想称赞你两三年前对这个市场有正确的远见,你回答说‘谢谢,但早了两三年就等于错了’。

You and I were talking probably a couple months ago at this point and I was trying to give you credit for having the right vision for this market two three years ago and you responded with something along the lines of thank you but being two three two or three years early is the same as being wrong.

Matan Grinberg

是的。

Yes.

Host

我觉得这个回答在很多方面都很精彩。你能谈谈那两年的‘沙漠之旅’吗?拥有一个后来被证明正确、但一两年内无人欣赏的愿景,是什么感觉?谈谈那段旅程以及它如何塑造了你们公司的 DNA?

Um which I thought was a wonderful response in so many ways. Can you talk about like that those two years in the desert? How did it feel to have this vision that turned out to be right that nobody appreciated for a year or two? Can you just talk about like that journey and what it has done to the DNA of of your company?

Matan Grinberg

是的,当时真的非常困难,因为我之前没有工作过。我从博士退学创办了这家公司,在那两年里,我说服了 20 个我见过的最聪明的人辞掉他们正在做的事情,加入 Factory,加入我们的使命。

Yeah, I mean in the moment it's really really difficult because you know uh I hadn't had a job before. I dropped out of my PhD to start this company and you know over the course of those two years convinced you know 20 of the smartest people that I've ever met to quit what it was that they were doing and you know join factory and join us on this mission.

让利维护信任 Giving back revenue to maintain trust

Matan Grinberg

这些人都有家庭、有孩子,把多年生命投入到这个问题上,一个客户接一个客户地跑。他们还没准备好接受智能体,他们不理解。而且模型当时也不够好。但我认为很大一部分是行为上的。讲个趣事:给开发者发 NPS 调查问卷。如果你给开发者发 NPS 问卷,他们绝对不喜欢你发给他们的任何东西。开发者用脚投票,他们很清楚自己喜欢什么、不喜欢什么。如果你在琢磨他们喜不喜欢,那他们肯定不喜欢。在那段时间,我们学到了很多。我自己以前从没工作过,企业销售对物理学家来说不是本能。但归根结底,这都不重要。早入场没有加分,谁在乎呢?没有安慰奖。要么做成,要么做不成,只有这个重要。对团队来说,那段时间非常艰难。有些时候,我们企业销售做得不错,但产品仍然不好。这是一个非常棘手的处境,因为我们一度营收接近 200 万美元,但产品并不好。我们意识到这一点,因为如果你销售很厉害,你可以签下合同,这绝对能做到。但如果你签了合同,开发者却不喜欢你的产品,那就是一颗定时炸弹,因为最终他们会翻脸,情况会变得非常非常糟糕。我们意识到了这一点,于是主动把所有客户的钱都退了。我记得那是最艰难的决定之一,因为不仅有一个 20 人的团队,他们从各大实验室收到离谱的 offer,有巨大的经济激励去别处,还有很多其他公司做得很好,而他们选择了留下,然后我们却要说:'对了,我们好不容易拿到的那点收入,实际上要退回去,因为我们觉得产品没有让他们的开发者满意。'

These are people with families, with kids, dedicating years of their lives to this problem, going customer after customer. They weren't ready for agents. They didn't get it. Also, the models weren't as performant. But I think a lot of it was behavioral. Even just a fun anecdote: giving developers an NPS survey. If you ever give a developer an NPS survey, they do not like whatever it is you're giving it to them. Developers vote with their feet. They are very clear what they like and what they don't like. If you're wondering if they like it, they definitely don't. During that time, there was a lot we were learning. I myself had never had a job before. Enterprise sales is not something that comes obvious to a physicist. But at the end of the day, it doesn't matter. There's no bonus points for being early. Who cares? There's no consolation prize. It's either you do the thing or you don't. That's all that matters. For the team, it was really tough. There were points where we ended up getting good at enterprise sales, but the product still wasn't good. That's a very tricky position to be in because we ended up getting to a point where we were just under two million in revenue and the product was not good. There was a point in time where we realized this because if you're really good at sales, you can sign contracts. You can definitely do that. But if you're doing that and the developers don't like your product, it's a ticking time bomb because eventually they're going to turn and it's going to be really, really bad. We realized this and we proactively gave all of those customers their money back. I remember that was one of the most difficult decisions to make because not only is there a group of 20 people who are getting ridiculous offers from all the labs, they have huge financial incentives to go elsewhere. There are all these other companies that are doing well and they decided to do this, and then we're going to say, 'Oh yeah, by the way, that little bit of revenue we managed to get, we're actually going to give it back because we don't think the product is making their developers happy.'

Host

你们为什么做出这个决定?

Why did you make that decision?

Matan Grinberg

我们向他们推销了一个美好的愿景,说服他们这是合适的团队,我们会为他们提供解决方案。但我们意识到,我们推销的方式和交付的产品都不够好,不符合我们的一条运营原则。我非常喜欢的一条原则是'创造痴迷的客户'。这有点像翻转了贝佐斯在亚马逊的'客户痴迷',但在我们看来,那是一个输入指标。输入指标是你不想衡量的东西。你是否痴迷客户并不重要。你可能痴迷客户,但他们却因为你做的事而申请限制令。我们的工作是做出好到让客户自己对我们痴迷的东西。这就是我们的工作。打个比方,如果你是篮球队教练,你不会在球员上场前说:'嘿,大家记得流汗。'这算什么?不,要得分。我们需要得分。而在这个过程中,你自然会流汗。同样,要创造痴迷的客户,你可能需要自己非常痴迷客户,但产出才是关键。回到这件事上,我们交付的产品并没有创造痴迷的客户。我们想确保——这是一群我见过的最聪明的人。我们正在接近目标。我们积累了很多直觉,内部开始整合。我们在内部看到,我们的做事方式开始变得更加智能体原生,产品也基本满足了需求,但我们有点领先于客户。我们想维持与客户的信任,这样当产品真正成熟时,我们可以回去对他们说:'嘿,伙计们,这才是真家伙。我保证。'为了建立这种信誉,我们不得不说:'听着,即使你们可能愿意继续,我们也要把钱退给你们,并说,3 个月后,我觉得它会准备好。给我们一些时间,我保证会让你们大吃一惊。'

We sold them on a good vision and convinced them that this is the right team to work with and that we were going to deliver the solution for them. But we realized that the way we had sold them on it and the product we were delivering was not up to snuff in a way that I don't think it would hold true to one of our operating principles. One of our operating principles that I really like is 'create obsessed customers'. This kind of flips over Bezos's thing where Bezos at Amazon it's customer obsession, but in our mind that's an input metric. Input metrics are like you don't want to measure input metrics. It doesn't matter if you're customer obsessed. You could be customer obsessed and they file a restraining order against you because they don't like what you're doing. Our job is to build something so good that our customers themselves become obsessed with us. That is our job. The analogy is if you're a coach of a basketball team, you don't want to tell your players before they come out, 'Hey guys, make sure to sweat.' It's like, 'What?' No, score points. We need to score points. And in doing so, you're probably going to sweat. Similarly, to create obsessed customers, you probably need to be really obsessed yourself with the customers, but the output is what matters. Coming back to this, the product we were delivering was not creating obsessed customers. We wanted to make sure — this was a group of the smartest people I've ever met. We were getting there. We were getting a lot of intuition. Things were starting to come together internally. We could see internally we were starting to become a lot more agent native in how we were doing things and the product was kind of scratching that itch, but we were kind of ahead of our customers. We wanted to maintain trust with our customers so that when it does hit, we can come back to them and say, 'Hey guys, this is the real deal. I promise.' To build that credibility, we had to say, 'Hey, look, even though you were maybe happy to continue, we're going to give you this back and say, 3 months from now, I think it'll be ready. Give us some time, and I promise we will knock your socks off.'

Host

当你们进行那次对话时,客户有什么反应?

How did your customers react when you had that conversation?

Matan Grinberg

有些人说:'哦,太好了。听起来不错。'因为我觉得那并不是他们痴迷的东西。有些人有点困惑。但总的来说,尤其是企业客户,他们不习惯这种事。很多时候企业预算,一旦花掉就没了,没人在乎。有些人甚至不知道有没有退款机制。这很难开口。还有那些相信你的投资者——我记得和 Sean 谈过。Sean 显然一直和公司保持紧密联系,所以他非常理解。但这是一件很可怕的事,要说:'嘿,顺便说一句,还记得那些你说收入在增长的最新情况吗?它马上就要降到零了。'这很吓人。这有点像一次信仰之跃,我们在内部早早看到了信号,知道这是我们需要走的方向。我们需要在产品方法上做出 pivot。但我记得那次全员会议,我们告诉了整个团队。天哪,那是我人生中最糟糕的几个月之一。

Some of them were like, 'Oh, great. Sounds good.' Because I think it wasn't something that they were obsessed with. Some of them were a little bit confused. But generally, especially enterprises, they're not used to these things. A lot of times enterprise budget, once it's gone, it's gone. No one really cares. Some of them didn't even know if they had a mechanism by which to take back the money. It's a difficult thing to tell. Also, investors who believe in you — I remember having the conversation with Sean. I think Sean obviously he's stayed really close with the company, so he was very on the same page. But it's kind of a scary thing to be like, 'Hey, by the way, remember all those updates where you're saying revenue is going up? It's about to go down to zero.' It was a scary thing. It was kind of a leap of faith of like we see the signal internally early of this is the direction we need to go. We need to kind of pivot the approach on the product. But I remember that all-hands where we told the whole team. It was like, oh my, that was like one of the worst months of my life.

团队逆境韧性 Team resilience through tough times

Host

你是如何在那段时期让团队保持凝聚力的?

How did you keep the team together through that?

Matan Grinberg

老实说,团队能留下来的唯一原因是我们早期招聘非常严格,只招那些真正痴迷于使命的人——我们的使命是将自主性带入软件工程。我们确保每个人都有清晰的反馈回路,明白命运掌握在自己手中。这不是那种“我去做就行”的事,而是我们每个人都要为成功出一份力。而且,我认为坦然接受这段艰难时期本身也非常有价值。

I think honestly the only reason the team stayed together is we were so ruthless about hiring early on where it was like people that are genuinely really really obsessed with the mission which our mission is to bring autonomy to software engineering and really really caring about that making sure everyone was also like very clear feedback loops as to like this the fate is in our hands. It's not like this is like oh something that I go do. It's like we all have a part to play in making this work. And I think embracing how much it sucked was also something that was very valuable.

Host

就是坦诚面对。

Just being honest about it.

Matan Grinberg

我们非常坦诚地说:是的,这很糟糕。你看那些竞争对手,营收疯涨,这不好。我们处境很糟,刚不得不退回所有收入,真的得振作起来。现在回想,正是那些时刻建立了最深的纽带。如果你和运动员或学者聊,他们也会说,在压力时期——比如考前突击或高强度训练——我们团队里有划船运动员,我们常拿这个举例:纯粹的痛苦运动,只有一个数字量化你的表现,就是 2000 米的时间。但拥抱痛苦才能创造持久的纽带,之后我们就知道跌入谷底是什么感觉,知道失败是什么感觉。我们公司刚成立时估值只有 500 万,而很多竞争对手如今一出生就是独角兽,他们不知道不是独角兽是什么滋味。但我们经历过那些黑暗时刻,没有一个人离开。

Being super honest about like, yeah, this sucks. Like, oh, look, look at those competitors. Their revenue is going up like crazy. Like, this is not good. Like, we are in a very bad position. Like, we just had to give back all of our revenue. Like, we need to really get our act together. And in the moment, I think retrospectively, those are the moments where really the deepest bonds are made. Like if you talk to people who were like athletes or even like academics or whatever whenever you're in the like stressful period whether it's like cramming before finals or in intense like you know we have some rowers on our team and I think that's an example we always go to like sport pure pain sport it's pain it's literally just there is one number that quantifies your performance it's just what is your time on your 2k or you timing that but embracing that is what creates those enduring bonds such that afterwards like we know what it's like to be at rock bottom. We know what it's like to lose. We know what it's like I mean when we first started the company our valuation was 5 million. Like a lot of our competitors a lot of the companies out there these days they don't know what it's like to not be a unicorn. That's like manifestally that is what they are day one. Whereas like we have been there kind of in those dark moments and not a single person left.

Host

嗯。

Yeah.

Matan Grinberg

这让我们变得非常有韧性和强大,现在情况好多了,但未来还会有低谷,而我们的 DNA 里就有这种韧性,我不确定其他公司有没有。

That makes us so resilient and so strong that you know going forward things are going a lot better now. But there are going to be really bad times, but we have that resiliency in our DNA that I'm not sure some of these other companies do.

Host

我很喜欢。那么跟我们讲讲是什么发生了变化?我很好奇你之前说的“模型变好不是最重要的事”,因为至少在我看来,模型变好就是最重要的事。帮我理解一下。

I love that. So talk us to talk to us about what changed and I'm curious your comment from earlier that the models getting better is not the most important thing that happens because at least in my mind the model's getting better is the most important thing that happens. So just help me understand.

Matan Grinberg

是的,有几件事。首先,我们之前构建的交互模式太激进了。你说得对,我们当时瞄准的是完全自主的智能体,但早了两年,所以错了。完全自主的智能体需要开发者彻底改变行为,而我们试图在他们甚至还没用上 C-pilot 这类工具时就一步到位,跨度太大了。2025 年 9 月 26 日是个重要的日子,我们发布了 Droid CLI。Droid CLI 以之前完全自主智能体没有的方式,在开发者所在的位置与他们相遇。而且它的性能是顶尖的,并且模型无关,可以使用市面上所有模型。9 月 26 日也是我们创业两周年,那时世界已经更习惯使用自动补全之类的工具。到 2025 年底,大多数工程师都在用自动补全工具,很多人开始通过聊天界面让智能体批量修改代码,也就是更智能体式的交互。但我们现在回头看,用这些旧模型做智能体式交互,效果仍然不错。所以最大的变化是开发者,尤其是企业中的开发者,对这种新工作方式更加开放。开发者过去 30 年建立了自己的工作流,可能很固执,很多人说“不行不行,我的技艺 AI 工具永远做不了”。所以很大程度上是理解如何使用这些工具,愿意去尝试,以及直觉上知道需要提供哪些护栏才能成功。

Yeah. So a couple things. So one is the interaction pattern that we were building for before was too ambitious. Like to your point, we were right in that what we were building for was fully autonomous agents, but it was two years too early, which makes it wrong. And fully autonomous agents require a complete change in behavior from the developer. And we were trying to do that out of the box before they were even using tools like C-pilot. It was just too much of a leap. It was too much of a step function jump. So, it's an important day. September 26th, 2025 was when we first put out basically the Droid CLI. And the Droid CLI met developers where they were in a manner that previously these fully autonomous agents did not. And also its performance was like completely state-of-the-art and it was model agnostic. So it could use every model that was out there. September 26 was also 2 years after we initially started. So the world had gotten much more used to using things like autocomplete. By late 2025, most engineers were using an autocomplete tool and many were starting to at the time use like a chat interface to ask an agent to go do changes like wholesale. So like the more agentic interaction. However, what we see is that if you go back now and use in this like agentic interaction some of these older models, they're still good. So the biggest thing that changed was developers and in particular in the enterprise like being open-minded to this new way of working in particular you know developers they've established their workflows over the last 30 years they can be stubborn. A lot of them were like no no no like my craft could never be done by an AI tool. So a lot of it was just like understanding how to work with these tools and having the willingness to go in and try and also intuition about what are the guardrails that you need to provide in order for it to succeed.

Host

嗯,所以我认为是这两方面的结合:模型变好了,所以你需要提供的护栏更少;同时开发者降低了防备,说“好吧,让我试试,它可能会做我不喜欢的事”。另外,当 Andre Karpathy 发推文谈论某件事时,每个工程师突然就会想“好吧,也许这是真的”。Andre 开始发推文谈论这些智能体,早期他并不那么开放,后来他更开放了,这确实改变了一些人的想法。这很有趣,但行为改变的一部分就是:你从信任的人那里听到,你开始看到组织内部那些更原生智能体的人也在用。这些因素加在一起带来了变化。

Um, and so I think it was a combination of both of these things. The model's getting better, so you need to do less in the way of providing guardrails, but also developers lowering their guard and being like, 'Okay, you know what? Let me go try and do these things. It's going to go do things I don't like.' And then also there's a certain degree to which when Andre Karpathy tweets about something, then every engineer suddenly is like, 'Okay, you know, maybe this is true.' And Andre started to tweet about these agents. Early on, he wasn't as open to it. And then him being more open to it genuinely just changed some people's minds. Which is funny, but that's some of the things that go into behavior change is like you hear it from people you trust. You start seeing it from people within your organization who are maybe a little bit more agent native. But that's kind of these things together is what changed that.

Host

现在我们都要上 Slack 了。

And now we're all going to be on Slack.

Matan Grinberg

我们可能会挑战 Slack 的极限,我觉得这会是另一件有趣的事。不过,是的。

We might be pushing the limits of Slack, which I think is going to be another interesting thing. But yeah.

Host

好的。那么 2025 年 9 月你发布了 Droid CLI,你说它是前沿性能。这对你来说意味着什么?

Okay. So September 2025 you launched the Droid CLI. You said frontier performance. What does that mean for you?

Matan Grinberg

基准测试的半衰期很短,一个好的基准测试往往在 3 到 6 个月内就会被刷爆。当时我们发布时主推的一个基准测试是 Terminal Bench,后来它成了一个相当不错的基准。在此之前,领先的是 SWEBench,它基于一些开源项目和已解决的问题示例。但问题在于它非常聚焦于 Python、脚本或单个文件修改,而 Terminal Bench 更侧重于终端环境。

There's like the benchmarks which have a very short half-life. Like anytime there's a good benchmark, it gets benchmaxed within like 3 to 6 months. At the time, I think the one that we kind of championed when we launched and kind of it ended up becoming a pretty good benchmark was terminal bench. So, prior to that, the one that was kind of leading was SWEBench, which was kind of took some open-source projects and some examples of issues that were then solved. The problem with that was it was very focused on like Python and like scripting or like individual file changes whereas terminal bench was more one it was in the terminal settings.

构建智能体框架 Building a great agent harness

Matan Grinberg

所以像调度运行之类的事情,不仅仅是修改代码文件,而是通用的软件开发任务,我们最终在这些任务上达到了前沿性能。现在它已经被基准测试优化到了极致,新发布的模型能达到 90% 的准确率,而且我认为从推出一个好的基准测试到它被纳入训练数据,时间窗口非常短。

So it was things like scheduling runs and things that were not just like changing the code file but general software development tasks and that was something that we ended up having really frontier performance on. Now it's like benchmaxed to the extreme to where it's like models that come out now are like 90% on it and I think there's a very short time horizon from putting out a good benchmark to then it being kind of in the training data.

Host

构建一个优秀的智能体框架需要什么?生态系统中似乎有很多恐慌、不确定性和怀疑,比如我的框架比你的好,你需要拥有模型才能有好的框架,或者实际上如果你不拥有模型,你的框架会更好。抛开基准测试优化不谈,你的思维模型是什么,是什么让你保持在最前沿?

What goes into building a great harness? And it seems like there's almost a lot of FUD in the ecosystem of my harness is better than your harness and you know you need to own the model to have a good harness or actually you have a better harness if you don't own the model. Like what's your mental model for benchmark maxing aside, what keeps you at the frontier?

Matan Grinberg

是的。一些重要的通用因素包括缓存的方式。缓存 token 的成本只有原来的十分之一。所以对于一个给定的框架,一个重要的性能指标是你的 token 缓存率。另一个例子是你在压缩或合并时的表现。通常,当处理长会话时,你会超过模型本身的上下文限制。因此框架会进行某种总结、压缩或合并,随便你怎么称呼。你在合并过程中的表现是决定框架好坏的关键因素。他们为此做的测试叫做“大海捞针”,你有一个很长的线程,可能只有一条信息非常重要。你的框架在合并过程中能多频繁地保留那条信息?其他例子包括工具使用,或者它如何利用环境来验证它正在做的工作。这些都是你可以有单独指标的事情,我们有自己的内部基准来衡量开箱即用的智能体与 Factory 的表现。我认为最初每个人都天真地相信,如果你训练模型并构建框架,你会让它们一起变得更好。

Yeah. So, a couple of general things that matter are the way you do caching. So, cache tokens end up being like a tenth as expensive. And so, one big piece of performance for a given harness is what is your rate of token caching. Another example would be how do you perform while in compression or compaction. So typically when you're dealing with a long session, you're going to exceed the context limit of the model itself. And so the harness will do some sort of summarization, compression, compaction, whatever you want to call it. And the way that you perform during that compaction is a big determining factor of how good your harness is. And tests that they do for that are like, they call it needle in the haystack, where you have some long thread and maybe there's one piece of information that's really important. How often will your harness preserve that through compaction? Other examples are like tool use or how does it use the environment to validate whatever work that it's doing. These are things that you can have individual metrics on and that we have our own internal benchmarks to measure how do the out of the box agents do versus how does factory perform. I think one thing that naively everyone believed initially was if you train the model and you build the harness you're going to make them better together.

Matan Grinberg

是的。这让我在 OpenAI 和 Anthropic 的许多朋友感到懊恼,但事实并非如此。如果你构建一个支持不同模型的框架,那个框架会更好。

Yeah. And much to the chagrin of many of my friends at OpenAI and Anthropic, this is not true. If you build a harness that supports different models, that harness will be better.

Host

直觉是什么?我的直觉是模型和框架的协同设计会让你更好。

What's the intuition? My intuition would be model-harness codesign makes you better.

Host

为什么实际上不是这样的直觉是什么?

What's the intuition for why it's actually not?

Matan Grinberg

这非常类似于大约 10 年前的想法:在早期机器学习时代,GPT-3 之前,如果你说“嘿,我想训练我的个人 AI”,你会给它你所有的数据,因为你希望它了解你。但结果证明,答案是在整个互联网上训练它,它会比只在你数据上训练好得多。所以出现了一种类比:数据之于模型,就像模型之于框架。你让框架接触的模型越多,就越能避免框架过度拟合特定模型的细微差别。你可以从不同模型的某些复杂性中学习,然后在你自己的框架中改进不同模型的表现。这就是为什么我们后来停止了这种做法,因为 Terminal Bench 已经被过度优化了。但最初,每当新的 Opus 或 GPT 模型发布时,它在 Droid 的 Terminal Bench 上的表现会比在 Claude Code 或 Codex 中更好。我认为这让实验室有些沮丧,因为理想情况下你希望它们一起更好,因为那意味着你必须使用他们的框架,而不能使用其他框架。但我认为现实是,拥有多模型框架最终会在所有这些方面达到前沿水平。

It's very analogous to the idea maybe like 10 years ago of if you were to be like, hey, I want to train my personal AI back in early ML days before GPT-3. I want to train my personal AI. I'm going to give it all of my data because I want it to know me. Turns out the answer was train it on the whole internet and it'll be so much better for you than if it were just trained on your data. So there's a sort of analog that emerges where it's what data is to a model, models are to a harness where the more models you expose to a harness, you avoid overfitting that harness to the nuances of that model in particular. And there are certain intricacies about different models that you can learn from and then improve different models' performance in your own harness. This was why for example we kind of stopped doing it because Terminal Bench got so benchmaxed. But initially when every new Opus or GPT model would come out it would perform better on Terminal Bench in Droid than it would in Claude Code or Codex. Which is something that I think was somewhat frustrating to labs because ideally you want it so that it's better together because then it means you have to use their harness and you can't use a different one. But I think the reality is having that multimodel harness ends up getting kind of frontier on all those.

Host

有没有一个好的例子或说明?概念上说得通。有没有简单的方法来说明它?

Is there a good example or illustration of that? Conceptually it makes sense. Is there an easy way to illustrate it?

Matan Grinberg

也许一个好的例子是,如果你熟悉 Opus 和 GPT 5.6 目前的不同行为。

Maybe a good example of it is if you're familiar with the different behaviors of Opus and GPT 5.6 right now.

Host

我熟悉。他在线。

I am. He's on.

Matan Grinberg

好的。Opus 有点像那种超级友好的同事,你说“嘿,我想做这 20 个任务”,他们会说“好的,酷。顺便说一句,其中 5 个任务我发现我们不需要做了。别担心。其他的我已经做完了。用这种方式做的。今晚不太合适。我们早上再继续吧。”然后一起去喝啤酒,随便聊聊。而 GPT 5.6 则像“绝对会做每一个任务,没有什么能阻止我。不完成我就不睡觉。”它非常强迫症且一丝不苟。但有时你想要一个能意识到“你给我的那 20 个任务列表,其实有更好的方法”的模型。5.6 更有条理。如果你为每个模型构建框架,那个框架会有不同的优缺点。例如,智能体通常会生成一个待办事项列表。Claude Code 的框架在某些情况下会非常严格地确保它坚持待办事项列表,因为模型本身通常会偏离。而 Codex 不会这样做,因为模型本身对此非常强迫症。但作为用户,你希望无论使用哪个模型都有相同的体验。你希望确保如果切换到不同的模型,你不会突然丢失正在处理的事情。所以在某些情况下,你真的需要健壮的工具使用。你用来做待办事项列表的工具,你希望非常健壮,确保无论如何,作为用户,你都能看到你的待办事项列表。有些情况下它可能根本没有待办事项列表。

Okay. Opus tends to be like that super friendly colleague where you're like hey I want to go do these 20 tasks. And they're like okay cool. Hey by the way five of those tasks I realized we didn't need to do it. Don't worry about it. I got other these done. Did it this way. Like tonight's not a good time. Let's pick it up in the morning. Yeah. Like let's go get a beer afterwards and hang out. Whatever. Meanwhile, GPT 5.6 is like absolutely I will do every single one of those and nothing will stop me. I'm not going to sleep until it's done. It's like very OCD and meticulous. But sometimes you want one where it actually realizes hey that list of 20 that you gave me actually here's a better way of doing it anyway. 5.6 is more methodical. If you build a harness for each of those there are actually different things that that harness will then be good or bad at. So, for example, one thing that typically agents will do is they'll have a to-do list of like if you have a task, it'll go and generate a to-do list. And the Claude Code harness can in some cases be really strict to make sure it would stick to the to-do list because the model itself would typically wander. Meanwhile, Codex wouldn't do that because the model itself was really OCD about that. But if you're a user, you want to have the same experience regardless. Like you want to make sure if you switch to a different model, you're not going to suddenly lose track of whatever things that you are working on. And so there are certain things where maybe in some cases you really want robust tool use. And there are tools that you use to do these to-do lists. You want really robust tool use and you want to make sure that no matter what, if I'm a user, I want to see my to-do list there. Like there were some cases where it would just not have the to-do list.

Token最大化与成本优化 Token maxing and cost rationalization

Host

好例子。好的,我们聊了一种最大化,基准测试最大化。我们来谈谈 Token 最大化,因为感觉世界变化很大。我们从 Token 最大化走到了现在的成本合理化。这对 Factory 意味着什么?

Good example. Okay. So we talked about one type of maxing, benchmark maxing. Let's talk about token maxing because it feels like the world has changed a lot. We've gone from token maxing to now cost rationalization. What does that mean for Factory?

Matan Grinberg

是的。那么,也许我来梳理一下,这样我们就能对齐看法,了解是什么导致了这种 Token 最大化。大致上,有一个第一阶段,也许第零阶段是没人相信 AI。然后第一阶段每个人都相信 AI,董事会问 CEO 你在 AI 方面做了什么?你的 AI 战略是什么?CEO 说我不知道,我们的 AI 战略是什么?CTO,确保每个人都去用 AI。然后第二阶段,CTO 说,好,我们必须确保每个人都用 AI。让我们把它放进绩效评估里。让我们公开谁用的 Token 最多,因为每个人都很固执。没人想用这东西。他们都持怀疑态度。然后我们进入第三阶段,每个人都看到了这些排名。他们看到这是绩效的一部分,然后他们说,“好,我什么事都用 AI。”这就是第三阶段。这就是 Token 最大化,人们用 Opus 做所有事。比如旧金山天气怎么样?Opus 告诉我,我不知道。比如我们合作的一些银行,他们每月花几十万美元,让人们问诸如天气怎么样,或者给我讲讲 Python 这种你本来可以谷歌的琐碎问题,人们却在问 Opus。嗯,现实是,这是因为我们太担心采用率,所以过度纠正了,我们想不惜一切代价提高采用率。我认为这其实是个不错的做法。可能先这样做,然后再限制使用或让使用更负责任,比一开始就限制说“你只能用它做这个”要快,因为当你面对固执的人时,首先你想证明它有效,然后你才能变得更成熟。Factory 的定位是,我认为我们做的最重要的事情之一是我们有 Factory 路由器,它允许你根据正在执行的任务动态路由到不同的模型。所以,如果你问天气怎么样,你可能不需要人类智能的最前沿来回答你。

Yeah. So, maybe I'll lay this out just so we're all on the same page of the way that we see what's led us to this token maxing. So, loosely there was like this phase one where maybe phase zero was like no one believed in AI. Then phase one everyone believes in AI and then boards were like Mr. CEO what are you doing about AI? What's your AI strategy? And Mr. CEO is like I don't know like what's our AI strategy? CTO, like make sure everyone goes and uses AI. And so then phase two is, you know, CTO is like, okay, we got to make sure everyone uses AI. Let's start putting it in performance reviews. Let's make public like rankings of who's using tokens the most because everyone's stubborn. No one wants to use this stuff. They're all skeptical. And then we enter phase three, which is everyone sees these ratings. They see that it's part of their perfs and they're like, "Okay, I'm going to use AI for everything." And that's kind of phase three. It's this token maxing where people are using like Opus for literally everything. Like what's the weather in SF? Opus tell me I don't know. Like there are banks that we are working with where they are spending literally hundreds of thousands of dollars a month on people asking things like literally what is the weather or like tell me about Python like trivial questions that you could Google people are asking Opus. Um, and the reality is this happened because we were so worried about adoption that we overcorrect and we're like adoption by any means necessary. And I think that's actually it's like a decent approach. Like it's probably faster to do that and then curb usage or make usage more responsible than it is to start limited and be like, you know, you can only use it for this thing because when you have people that are stubborn, first you want to just prove that it works and then you can get kind of more mature about it. Where Factory fits in. I think one of the most important things that we do is that we have the Factory router which allows you to dynamically route to different models based on the task that you're doing. So, you know, if you're asking what the weather is, you probably don't need the very frontier of human intelligence to answer that for you.

Host

或者你真的需要。

Or you really do.

Matan Grinberg

我是说,这取决于情况。我不知道。这取决于你想要什么样的答案。嗯,比如给你一个从分子层面解释天气的完整回答。但是嗯,允许这样做,但更重要的是,对于每个企业来说,还没有人处理但 12 个月后就会面临的问题是,不是每个人都需要相同的 Token。比如说,给某个大银行的每个员工设置一个统一的 Token 上限是没有意义的。所以每个 CIO 都需要回答每个增量 Token 应该放在哪里?而现在,如何做到这一点非常不明确。比如现在我们说,那些用“氛围编码”做仪表盘的 PM 和构建关键基础设施的工程师获得相同的 Token 限制,这可能不是最好的做法。或者类似地,你可能在处理 Cobalt 代码库,Opus 不是最好的模型,而是可能针对该代码库微调的某个模型。路由器的要点是,我们可以适应这些不同的约束,比如你可以说,组织这部分的人只是在“氛围编码”,他们可以用 Gemini Flash;这部分的人在做 Cobalt,我们微调了一个很好的模型来处理 Cobalt,当处理这部分代码库时就路由到那个模型。也许另一部分,我们非常关心可靠性。那么,用 OpenAI 生成代码,用 Anthropic 测试,用 Gemini 审查,诸如此类。我们实际上可以用自然语言接收你的路由过程指令。所以,你甚至可以说一些不是纯粹确定性的东西。甚至可以像这样:“嘿,你知道,Pat,我不知道。我不知道他在做什么。给他用 Flash。”我不知道。或者,你知道,我认为我们真的需要避免他们使用开放模型,因为,不管什么原因,我们不喜欢开放模型在这里的表现。嗯,我们会做内部基准测试,以了解哪些模型更擅长哪些任务。

I mean, it depends. I don't know. It depends on what kind of answer you're looking for. Um, you know, giving you like a full like down to the molecular level of what's happening. But um, you know, allowing that but also more importantly for every enterprise something that no one's dealing with yet but 12 months from now is going to be the case is um not everyone needs the same tokens. Having a blanket kind of token cap for every individual in some large bank let's say makes no sense. So every CIO is going to need to answer for every incremental token where do we put it? And right now it is super not obvious how you would do that. Like right now we're saying oh you know the PMs who are like vibe coding dashboards get the same token limits as like the engineers who are building like critical infrastructure that's probably not the best thing to do or similarly you might be dealing with Cobalt code bases where Opus is not the best model to use but instead maybe some fine-tuned model on that codebase in particular. The point of the router is that we can kind of accommodate these different constraints where maybe you say you know what this part of the org they're just vibe coding they can use Gemini Flash this part of the org they're doing Cobalt, we fine-tuned this great model to work on Cobalt. Let's route to that when we're working on that part of the codebase. Maybe this other part, we really care about reliability. So, let's generate the code with OpenAI, test it with Anthropic, review it with like Gemini, things like that. And we can actually take in your routing procedure instructions in natural language. So, you could even say things like it's not purely deterministic. It can even be like, "Hey, you know, Pat, I don't know. Like, I don't know what he's doing. Like, give him Flash." Like, I don't know. Or, you know, I think we really need to avoid having them use open models because, you know, whatever reason, we don't like the way open models perform here. Um, and we'll do internal benchmarking to know which models are better at which of these tasks.

Host

开放模型现在有多接近?哪个最好?

How close are the open models at this point? Which one's the best?

Matan Grinberg

GLM 5.2 非常棒。嗯,它已经到了我们内部对工程师没有 Token 限制的程度,大约一半的 Token 都用于开放模型。

GLM 5.2 is incredible. Um, it's at the point where internally we have no token limits for our engineers and like half of our tokens are open to open models.

Host

哇。

Wow.

Matan Grinberg

是的。因为它们更快、更便宜。性能也一样好。我认为每个人都搞错的一点是,大家都在拿 GLM 5.2 和最新模型如 Opus 4.8 或 GPT 5.6 比较。但实际上它们应该和 Opus 4.7 或 GPT 5.5 比较。嗯,为什么?因为通常开放模型出现得晚,它们落后一代,前沿模型仍然是前沿。问题是开放模型是否变得和前沿减一代一样好,答案是肯定的,我认为这是一个非常非常有趣的结果。这对消费者很好,我说的消费者不是指个人,而是指 API 的消费者,因为如果你是一个从事软件工程等工作的企业,你的工作是在很高层次上解决问题,如果我们能让你用更快、更便宜但性能相同的模型解决问题,那就意味着你可以解决更多问题,这是一件好事。这是一个非常好的世界,没有智能垄断,而是有一个智能花园,你可以根据需要挑选。我们开玩笑说,关于智能分配这件事,如果你想给你女儿找一个代数家教,你很可能能找到比阿尔伯特·爱因斯坦更便宜的人来当这个家教。

Yeah. Because they're just faster and they're cheaper. They're just as performant. And I think the thing that everyone gets wrong is everyone is comparing like GLM 5.2 to the latest model like Opus 4.8 or GPT 5.6. But really they should be compared to Opus 4.7 or GPT 5.5. Um, why? Because generally the open models come later and they're kind of a generation behind and that's kind of the frontier models will be frontier. The question is are the open models getting as good as like frontier minus one and the answer is unequivocally yes which I think is a really really interesting outcome. It's great for consumers and by consumers I don't mean like individuals I mean the consumers of the APIs because if you're a business that is doing in AR like software engineering your job is at a very high level to solve problems and if we can allow you to solve those problems faster and with cheaper models that are just as performant that means you can solve more problems like that is a good thing and it is a very good world where there is not like a monopoly on intelligence but instead kind of a garden of intelligence that you can pick and choose um you know when you'd like. Something that we joke about is like you know on this intelligence allocation thing. Um if you're trying to get a tutor for your daughter in algebra, you can probably find someone cheaper than Albert Einstein to be that tutor.

模型路由趋势与开源采用 Model routing trends and open model adoption

Host

既然你们做模型路由,如果看一张饼图展示你们客户群当前使用的模型构成,几个月前是什么样?现在是什么样?你觉得一年后会是什么样?

Since you guys do the model routing, if you look at the pie chart that shows the complexion of models being used by your customer base today, what did it look like a few months ago? What does it look like today? What do you think it'll look like in a year?

Matan Grinberg

我先说明一下,目前企业在路由流程上还没有太明确的偏好。这会在未来 6 到 12 个月内发生。但现在,他们只是从没有路由器过渡到有路由器,这是第一个变化。然后才会涉及路由的具体性质。今年年初,不到 1% 的 token 流向了开放模型。第一季度变成了个位数百分比。现在已经达到了两位数的 token 占比。当然,token 占比并不总是等于成本占比,因为开放模型的 token 更便宜。但看到这样的增长还是很惊人的。

I will caveat this with saying that right now enterprises haven't gone too opinionated yet into the routing procedures. This is something that will happen over the next 6 to 12 months. But right now, they're just going from no router to router. That's kind of the first change. Then it's going to be like the exact nature of the routing. At the beginning of the year, it was less than 1% of tokens went to open models. In the first quarter, it became a single-digit percent. It has now crossed into being a double-digit percent of tokens. Now percent of tokens is not always the same as percent of cost because the open tokens are cheaper. But it is pretty crazy to see the growth there.

Host

你的预测是什么?

What's your forecast?

Matan Grinberg

我的感觉是,我们会逐渐趋近于绝大多数 token 来自开放模型,因为它提供了更多选择且更便宜。但这并不意味着那是 token 份额,不一定是杠杆份额,因为可能 1% 的 token 极其有价值,是关键决策,其余则更像是执行型 token 或风险较低的 token。我不认为会达到 100%。

My sense is that we will asymptote towards vast majority being open just because it provides you more optionality and it's cheaper. But that doesn't mean they're going to be like that's a token share, not necessarily of leverage share because maybe there are 1% of tokens that are incredibly valuable and are like very key decision-making and then the rest are more like implementation tokens or kind of lower stakes if you will. I don't think there's going to be a world in which it's ever going to be 100%.

Host

嗯。

Y

Matan Grinberg

我认为智能的前沿本质上对每个企业都始终有价值。因为风险会越来越高,思考的复杂性也会越来越重要,但我们会更好地卸载某些任务。这可以粗略地类比为组织的结构方式:通常工程领导是更有资历的工程师,理论上拥有更多智慧,他们每分钟的脑力在理论上杠杆更高。你甚至可以想象一个人类工程师,追踪他们一天中使用了多少脑力:大部分时间可能很低,但有些时刻会很高,比如他们深度集中思考系统设计问题。所有那些低杠杆的时刻我们想自动化掉,而那些高杠杆的时刻——有时我们称之为“尤里卡时刻”或做高杠杆事情的时刻——如果这些不再是瞬间,而是持续数小时,因为你不用处理其他杂事呢?我认为这就是思考智能分配的方式:如果你是一名工程师,写文档是低杠杆的时间利用。你已经成为你领域的专家,却花数小时写文档——我记得 Stripe 因为拥有出色的文档而获得了巨大的优势——但想象一下,如果那些出色的工程师不用写文档,他们还能做多少其他事情。我们应该生活在一个每个人都能拥有 Stripe 级别文档的世界,这对所有人都有益。然后问题是,那些真正聪明的工程师一旦不用做这些,他们会用时间做什么?

I think the frontier of intelligence will inherently always be valuable for every business. Just because the stakes are going to get higher and the kind of intricacy with which you think is going to be more important but we'll be better at offloading certain tasks. And this is like you can loosely think of this already with the way orgs are structured where you know in general engineering leaders are more tenured engineers who in theory have like more wisdom and each kind of minute of their brain power is higher leverage in theory. And even you know, you can also imagine like consider a human engineer and try mapping over the course of their day like how much brain power they're using and like you know it's probably going to be really low for a lot of it but then there going to be some moments where they're like going pretty high like they're deeply concentrating and thinking about some you know systems design problem or whatever. All of those low-leverage moments we want to automate away and like we want to like those very high leverage moments sometimes like you know we're referring to them as like the eureka moments or the moments where they're like doing something that's very high leverage. What if those aren't just moments but what if those are like hours at a time because you don't have to deal with all the other stuff. And I think that's kind of the way to think about intelligence allocation is if you're an engineer and you were writing docs that is such a low leverage use of your time. Like you've become an expert in your craft and you used to spend hours writing docs like I remember it was actually valuable like I remember Stripe had so much alpha for just having incredible docs but imagine all the other stuff those incredible engineers could do if it wasn't writing documentation like we should live in a world where everyone can have docs as good as Stripe and that is like strictly beneficial for everyone and then the question is okay what do those really smart engineers do with their time once they don't have to do that.

Host

也许现在是讨论商业模式的好时机,尤其是随着开放权重模型的兴起,成本差异——我想这对你们的成本结构意味着非常不同的东西,但给客户带来的价值却非常相似。你们如何看待商业模式和定价?

Maybe this is a good time to talk about business model given that especially with the rise of open weight models, the cost differential, I imagine that means very different things for your cost structure but very similar value delivered to customers. How do you think about business model and pricing?

Matan Grinberg

是的,这更多是客户想要和需要的,而不是我们想要或需要的。例如,我认为目前按用量计费显然是正确的方式。我们希望与他们和我们正在做的事情保持一致。我认为按席位计费不合理,至少对我们来说是这样。我的感觉是,最终我们会转向按结果计费。但我不认为企业已经准备好,我们从最初两年吸取了教训。我们不会强行推行。但我猜测在 2030 年代,情况可能会更接近按结果计费。

Yeah, this is more what our customers want and need as opposed to what we want to need. So for example, I think right now usage-based is clearly the way to go. We want to be aligned with what they are doing and what we are doing. I think seat-based doesn't make sense, at least for what we are doing. My sense is that eventually we will change to outcome-based. Now I don't think the enterprise is ready for that and we've learned our lesson from those first two years. We are not going to impose things right. But my suspicion is that in the 2030s things will probably look more like outcome-based.

Host

嗯。

Yeah.

Host

按结果计费对你的市场意味着什么?结果的定义是什么?

What does outcome-based mean for your market? What would be the definition of an outcome?

Matan Grinberg

也许可以这样理解。目前我们按用量计费:你用的 token 越多,付得越多,我们赚得越多。由于我们与模型无关,通过路由器,我们基本上把 token 大炮指向 OpenAI、Anthropic、AWS、GCP 等任何一家。在某种程度上,这就像一个简化版的市场:买方是想要完成任务的工程师,模型提供商则通过基准测试宣称自己的成本和性能,然后我们决定为特定任务选择谁。未来可能有一种世界,token 以某种方式竞价,比如“这是我们对这个任务的成本。无论怎样,我们都会以这个成本完成任务”,但他们的定价是希望获得利润。定价错误就会亏损,定价正确并中标就能获得正利润。判断任务是否成功的方式是通过一些验证循环,因为现在没有人再使用那种“给我写代码,谢谢”的工具了。通常是“给我写代码,这是我知道它做得好坏的标准”。同样,如果你是一个模型实验室,给你一个任务和验证标准,你就能大致说出你愿意为获得这些 token 支付多少,并希望有一定的利润。在那个世界里,这基本上是从按用量计费动态转向按结果计费的方式。我认为这有很多问题,而且非常前瞻。

So maybe here's a way to put it. So right now we charge usage-based: the more tokens you use, the more you pay, the more we get. Now since we are model independent, we kind of with our router we are kind of pointing a token cannon at either OpenAI, Anthropic, AWS, GCP, you know any one of these people. To a certain degree, this is like a really dumbed down version of a marketplace where right now there is a buy side: an engineer who wants a task done, and then you have the model providers who are saying like either in benchmarks right now they're like we perform at this cost and this performance, and then we determine who we go to for that given task. There's a world in which tokens might bid in a certain way, saying like here is our cost for this task. We will get this task done at this cost no matter what, but they're pricing it such that they hope they can make a margin there. They price it wrong, they're at a negative margin. If they price it right and win the bid, then they get the positive margin. And the way you determine if the task was successful is by some validation loops because no one is using these tools anymore where it's just like write me code, great thank you. It's generally write me code and here's how I know it was done well. And similarly, if you are like a model lab and you are given here's a task, here's the validation criteria. You'll be able to say roughly how much you think you would be willing to pay to get those tokens and you know you want to have some margin on that. And then in that world, that's basically a way to dynamically shift from usage-based to outcome-based. I think that there are so many questions with this and this is very much forward-looking.

任务细分与验证标准 Subdividing tasks and validation criteria

Host

嗯,但我觉得关于如何细分任务有很多问题。你知道,分配任务这件事并不显而易见。

Um, but I think there's a lot of questions about how do you subdivide tasks? You know, divvying that up I think is something that's not obvious.

Matan Grinberg

是的。但随着这些工具变得更好,做这类事情实际上会变得容易得多。

Yeah. But as these tools get better, doing things like that actually become way easier.

Host

嗯,这很有意思。

Yeah. That's fascinating.

Matan Grinberg

是的。

Yeah.

Host

是的。如果你能界定一个任务,然后创建一个竞争性市场,那将是一个迷人的未来版本。

Yeah. If you can scope a task and then create a competitive marketplace, that'd be a fascinating version of the future.

Matan Grinberg

是的。而且作为用户,这会激励你在验证标准上非常彻底。嗯,因为你知道,有些故事是这样的:你让一个智能体修复代码,结果它删掉了你的代码。就好像,解决方案就是全部删掉。

Yes. And as a user, it then creates an incentive to be very thorough in your validation criteria. Yeah. Because like, you know, there are stories of, like, you know, you ask an agent to like fix my code and it deletes your code. It's like, you know, the solution is just get rid of it all.

Host

就像《硅谷》那集一样。

Like that Silicon Valley episode that was.

Matan Grinberg

是的。但所以你需要确保你的测试非常彻底,因为从技术上讲,它可能满足你所有的……

Yeah. But like, so you need to make sure your tests are very thorough because technically it could hit all of your...

Host

……验证会失控。

...validation will go rogue.

Matan Grinberg

是的。完全正确。完全正确。嗯。嗯。参考一下。

Yeah. Exactly. Exactly. Yeah. Yeah. Reference.

软件工厂概念 The concept of a software factory

Host

嗯,也许稍微宏观一点。你把公司命名为 Factory。实际上,在 Factory 之前你叫它 Droid。没错。但你在这个概念流行之前就命名为 Factory,现在感觉每个人都想建一个软件工厂。你认为我们今天在软件工厂建设方面处于什么位置?离软件工厂的终极愿景还有多远?

Um, maybe zooming out a little bit. You named the company Factory. Actually, you named it Droid before before Factory. That's right. But you named it Factory before this concept took off and now it feels like everybody wants to build a software factory. Where do you think we are today in terms of the building of software factories and how close are we to the ultimate vision of a software factory?

Matan Grinberg

是的,每个公司都有一个软件工厂,不管他们知不知道。只不过是一个非常低效的工厂。所以这有点像前工业化时代,人们手工缝制东西或做木工之类的,效率很低。现在,如果你去一个超过一万人的组织,问他们决定和发布一个功能的过程,可能有成百上千人参与其中,而且他们很可能甚至画不出来。他们知道这个流程的可能性很低。嗯,这不是因为他们认为这是正确的方式,而是当今构建大型软件的本质。但有了这些系统,很多隐性知识可以被编码。很多通常需要问一个在这里干了 30 年的专家的事情,需要这个批准那个批准,还有一份文档说我们必须做这个检查清单。它非常依赖人类行为和冗余。其中很多可以自动化,并重新聚焦于真正推动业务的事情。我认为向软件工厂的转变,是为了弄清楚决定我们需要构建什么功能的实际输入是什么,这些输入可能来自客户、市场、公司的产品负责人。我们要非常清楚:这些是我们接收的信号和输入。好的,我们有这些信号。那么构建的过程是什么?嗯,真正绘制出你构建软件的装配线非常重要,因为这样你就能闭环,问:这真的为我们的业务带来了成果吗?

Yeah, everyone has a software factory whether they know it or not. It's just a very inefficient one. So it's kind of like it feels like, you know, pre-industrialization where, like, you know, people were manually, like, you know, sewing things together or woodworking or whatever it might be, and these things are very inefficient. Like right now, if you go to an organization that has more than 10,000 people and you were to ask about the process by which they decide and release a feature, there are like hundreds or maybe thousands of people in that process and most likely they couldn't even draw it for you. Like there's very low likelihood that they would know what that process looks like. Um, that is not because they think that is the right way of doing things. That is just kind of the nature of building large software as it is kind of today. But with these systems, so much tribal knowledge can be codified. So much of this stuff that typically would require, oh, we need to ask this guru who's been here for 30 years who has the wisdom. Oh, we then need this approval and that approval. Oh, and I forgot there was some doc that said we always have to do this checklist. Um, and it relies so much on kind of human behavior and like redundancy. So much of that can be automated and refocused on like what actually moves the needle for our business. And I think this move towards software factories is a move towards how do we figure out what are the actual inputs that determine what features we need to build and that might be inputs from the customers, inputs from the market, inputs from, like, you know, product leaders at the company. And let's be very clear: these are the signals, the inputs that we are taking in here. Okay, great. We have those signals. Then what is the process by which we build this? Um, and really like mapping out the assembly lines of how you are building software is really important because then you get to close the loop and say: did this actually deliver outcome for our business?

Token经济学与分配 Tokconomics and token allocation

Matan Grinberg

之前谈到 tokconomics,如果你是那位 CIO,面临每个增量 token 该放在哪里的问题,实际上两年后的问题会变成每个增量美元该放在哪里。所以你必须问自己:是把那笔增量美元投给人力还是 token?如果是 token,投到组织的哪个部分?这些事情只有当你有了这样的反馈循环时才能真正知道,这些反馈循环会给你例子,比如:顺便说一句,我们基于这些数据做了那些决定,但根本没用。我们添加了新功能,但没人关心。它没有带来更多留存,也没有带来更多使用量或业务想要优化的任何指标。做到这一点的唯一方法是需要更多的严谨和流程。几乎感觉十年后,我们会回顾这个软件时代,感觉就像古代企业不做会计一样。

Talking before about the tokconomics, if you're that CIO and you're faced with that question of where do you put every incremental token, really the question two years from now is going to become where do you put every incremental dollar? And so you're going to have to be asked, do you put that incremental dollar towards headcount or towards tokens? And if tokens, to where in the org. And these are things that you can only really know when you have these kind of feedback loops that give you examples of like, hey, by the way, we made those decisions based on this data. And it did not matter at all. We added these new features and no one cared. It didn't create more retention. It didn't create more usage or whatever metrics that business is looking to optimize. And the only way to do this is like you need kind of more rigor and more process. It almost feels like 10 years from now we're going to look back at this previous era of software and it's going to feel like businesses in ancient times where they didn't do accounting.

Host

这就像《广告狂人》时代的营销,对吧?全是创意,你根本不知道什么真正有效。

It's going to be like marketing in the day of Mad Men, right? Where it's all creative and you have no idea what's actually working.

Matan Grinberg

这毫无道理……就像,哦,我们发布那个功能吧。哦,我觉得效果不错。嗯,我们有一些指标。是的。

It makes no... like it's like, oh yeah, let's ship that feature. Oh, I think it went well. Like yeah, we had I got some metrics on that. Yeah.

Host

就像,不,如果你们读了 Jack Dorsey 关于每家公司都像 AGI 的博客文章,还有一点是,如果你的公司是一个 AGI,你想优化权重。你想弄清楚哪些节点在做什么,哪些是承重的,哪些不是,哪些需要更多 token,哪里需要更多节点?为了做到这一点,你不能凭感觉训练模型。我的意思是,好吧,实际上你有点凭感觉,但更重要的是,你不会凭感觉做反向传播。你是在运行那些实际的计算,观察:当我们改变这个节点时,会发生什么?你可能会凭感觉下注如何改变模型,但你的做法相当数学化。与此同时,在公司里,人们凭直觉决定 token 预算。人们凭直觉裁员,就像,哦,两万人。裁员两万人不可能有科学依据。那就像砍掉一块,看看会发生什么。相反,我认为在这些组织中,他们做事的方式可以更数学化:这部分业务非常重要,给它更多 token 会更好;给它更多人力其实不重要,所以给它更多 token。可能还有其他部分,给更多 token 没用,但更多人重要,因为如果我们与客户建立更多、更深的关系,那很重要。但这些事情我们需要定量洞察,你需要一个软件工厂来做到这一点。否则,你只是凭直觉猜测,效果不会那么好。

It's like, no, if you guys read the blog post that Jack Dorsey put out about how every company is like an AGI, there's also this degree to which if your company is an AGI, you want to optimize the weights. You want to figure out what nodes are doing what things, which are loadbearing, which are not, which need more tokens, where do you need more nodes? And in order to do that, like you don't train a model by vibes. I mean, okay, actually you kind of do, but you don't, I guess more importantly, you don't do backprop in a model by vibes. Like you are running those actual calculations and you are seeing: when we change this node, what happens? Now you might be making bets on how to change the model by vibes, but like you are pretty mathematical in what you are doing. Meanwhile, at companies, you know, people are determining token budgets just by shooting from the hip. People are laying people off by shooting from the hip and just being like, oh yeah, like 20,000. There is no way there is science to laying off 20,000 people. That is just like here is a chunk and let's just see what happens. Instead, I think in these organizations, the way they can do things is much more mathematical: this part of the business matters a lot and does better if we give it more tokens; it doesn't actually matter if we give it more humans, so let's give them more tokens. There might be other parts of the business where actually giving them more tokens doesn't matter, but more people matter because if we build more relationships with our customers and deeper relationships with our customers, that matters. But these are things that we're going to need like quantitative insight on and you need a software factory to do that. Otherwise, you're just like shooting from the hip and just guessing, which won't work as well.

Token与人力成本极限 Token vs headcount spending in the limit

Host

在极限情况下,你认为人们会在 token 上花多少钱,与工程人力相比?

In the limit, how much do you think people will spend on tokens versus on engineering headcount?

Matan Grinberg

这取决于业务。

It'll depend on the business.

核心能力与AI应用 Core competency and AI adoption

Matan Grinberg

我认为每家企业都会有一个平衡点,这取决于具体情况。一个简单的例子是销售人员:如果他们很优秀,可能不需要太多 token,因为他们最大的价值在于与客户面对面交流,理解客户的问题。他们可以用 token 来做 AI 复盘、记笔记、跟进,但这些用量很小。给销售团队增加 token 可能不会改变产出,但增加人手却会。工程团队则不同:你希望人们端到端地拥有成果,给他们更多 token 能让他们产出更多。运营、财务、营销则介于两者之间。每家企业都必须问:我们的核心竞争力是什么?我们过去常看到一些公司说要自建软件开发智能体,比如一家物流公司。六个月后他们意识到这不是核心。这是一个机会,让企业加倍投入核心竞争力,并从外部采购不重要的东西。在互联网早期,披萨店需要工程师来建网站,但那只是短暂时期的副产品。现在有公司帮你建网站,你不需要懂技术。大多数披萨店没有工程部门,这是好事。同样,许多企业不得不为非核心角色招聘,但让企业专注于核心竞争力对消费者有利。我们会看到对真正重要的事情进行无情的重新聚焦。

I think every business will have a balance and it just depends. An easy example is salespeople. They probably don't need that many tokens if they're good, because the most alpha they provide is face-to-face with customers, understanding their problems. They can use tokens for an AI debrief, notes, follow-up, but that's minimal. Adding more tokens to the sales team probably won't change output, but adding more humans will. Engineering teams are different: you want people to own outcomes end-to-end, and giving them more tokens lets them produce a lot more. Operations, finance, marketing are in between. Every business has to ask: what is our core competency? We used to see companies saying they'd build their own software development agents, like a logistics company. Six months later they realize it's not core. This is an opportunity to double down on core competency and procure externally what doesn't matter. In the early internet, a pizza shop needed engineers to build a website, but that was a byproduct of a brief moment. Now companies help you build a website without being technical. Most pizza shops don't have an engineering department, which is good. Similarly, many businesses had to hire for non-core roles, but allowing focus on core competencies is good for consumers. We'll see ruthless refocusing on what matters.

Host

关于这一点,每家公司都必须经历这个重塑过程。10 或 20 年前,人们谈论数字化转型。现在是 AI 转型。几年前,你遇到过一些还没准备好应对自主智能体的组织。你看到你的客户在成熟度曲线上攀升并重塑自我时,有哪些变化?他们有没有什么好的技巧或方法来调整自己?

On that, every company has to go through this process of reinvention. 10 or 20 years ago people talked about digital transformation. Now it's AI transformation. A couple years ago you ran into organizations that weren't ready for autonomous agents. What changes have you seen in your customers as they go up this maturity curve and reinvent themselves? Any good tricks or techniques they use to repot themselves?

Matan Grinberg

令人惊讶的是,举办全公司黑客马拉松的公司最终表现很好。这看起来很简单,但留出一天让每个人都用 AI 来构建,确实能定下基调和节奏。

Surprisingly, companies that do company-wide hackathons end up doing well. It seems trivial, but setting aside a day where everyone builds with AI really sets the tone and pace.

Host

那就让我看看。

So just give me a look.

Matan Grinberg

我试图强迫他用编码智能体来构建东西,但效果不太好。

I tried to force him to build stuff with coding agents. It didn't go so well.

Host

我们会努力的。我们之后再做。

We'll work on it. We'll do after this.

Matan Grinberg

你知道,我们付出了很大的努力。

You know, we gave it a great effort.

Matan Grinberg

但就是这样。只是留出时间去做,即使惨败也没关系。另外,那些允许失败的组织——有些人想从第一天起就做得完全正确,但你肯定会犯错。那些积极投入并拥抱它的组织正在成功。我们最大的客户之一是 EY。他们并不处于 AI 的最前沿,但他们说:‘这很重要。我们不想迟到。我们可能会搞砸,但让我们让工程师们去折腾,构建这些东西,看看哪里会出问题,了解他们喜欢什么。’这非常重要。此外,大胆地重塑流程,说没有什么是神圣不可侵犯的,尝试一些东西,如果不行就把神圣的牛放回去。这是一个决定性因素。而且,当这种动力来自内部——技术团队、个人贡献者或领导层——时,效果比来自董事会要好。

But that's it. Just setting aside the time to do it, even if it fails miserably, is fine. Also, orgs that are okay with failing—some want to do it exactly right from day one, but you're going to make mistakes. The orgs leaning into it and embracing it are succeeding. One of our largest customers is EY. They're not at the absolute frontier of AI, but they said, 'This matters. We're not going to be late. We might mess up, but let's go and get our engineers to mess around, build this stuff, see where it breaks, understand what they like.' That matters a lot. Also, being bold in reinventing processes, saying there are no sacred cows, try something out, if it doesn't work put the sacred cow back. That's a determining factor. And when it comes from within—the tech team, ICs, or leadership—it goes better than from the board.

Host

是的。

Yeah.

Matan Grinberg

如果它来自技术团队、个人贡献者或领导层内部,我们就会看到效果更好。

If it comes from within like the tech team or the IC's or the leadership, that's when we see it go better.

Host

你对未来 12 个月你所在领域最重要的变化有什么预测吗?AI 消费正在疯狂增长,每个人都因为收入飙升而兴奋不已。其中很多是同步使用。

Do you have any predictions for the most important changes that are going to happen in your space over the next 12 months? A lot of AI consumption is going up like crazy and everyone's super excited because revenue is going wild. A lot of this is synchronous usage.

Token与原生智能体软件未来 Future of tokens and agent-native software

Host

换句话说,如果明天所有人都病倒了,很多 Claude Code 的使用量就会归零,因为一切都只是“嘿 Claude Code”或“嘿 Codex”或“嘿 Droid”,对吧?

In other words, if everyone woke up sick tomorrow, a lot of Claude Code usage would be zero because it's all just "Hey Claude Code" or "Hey Codex" or "Hey Droid", right?

Matan Grinberg

我认为在 12 到 24 个月内,90% 的 token 将是异步 token。这些是自主运行的 Droid,它们会像这样:“嘿,我从客户那里发现了一些信号。我们去修复它,或者去创建一个初步解决方案。”我认为这才是真正的智能体原生时代的开始,因为现在我们仍然处于副驾驶模式。如果你对一个智能体说“嘿,帮我去做这个”,它更具智能体特性,因为它不会回来问你一大堆问题,但仍然是你发起的。如果你去过特斯拉的工厂——这也是我们名字的灵感来源之一——那里到处都是机械臂在自动运作。不像有人在那里手动安装零件。这种“黑暗工厂”的概念,灯关着,一切自动进行,这就是软件开发的发展方向。名字的由来也与此相关:埃隆总是说工厂是制造机器的机器。我们深以为然。而且,结合他关于“你注定会成为你名字的反面”的说法,在我们的案例中,Factory 变成了 Artisanal(手工的)。这是一个很好的翻转。

I think in 12 to 24 months, 90% of tokens will be asynchronous tokens. These are droids on their own autonomously being like, "Hey, here's some signal that I found from a customer. Let's go fix it or let's go create a first pass solution to this." And I think that is where the real agent-native stuff begins, because right now we're still in co-pilot mode. If you go to an agent and say, "Hey, go do this for me," it is more agentic because it won't come back and ask you a ton, but it's still you kicking it off. If you've ever been to Tesla's factories, which is one of the sources of inspiration for the name, it's just robotic arms everywhere going and doing stuff. It's not like there are people there attaching the widget to the thing. This idea of a dark factory where the lights are off and things are just happening — that is where software development is going. That's where the name came from: Elon was always talking about the factory is the machine that builds the machine. That's something we took to heart. And I guess also that combined with his whole thing about how you're destined to become the opposite of your name. In our case, Factory becomes Artisanal. It's a good flip there.

乐观未来愿景 Optimistic future vision

Host

你对 Factory 和整个世界最乐观的未来版本是什么?

What's your most optimistic version of the future both for Factory and for the world at large?

Matan Grinberg

我认为短期内会有很多动荡,因为很多公司资源分配非常糟糕。存在大量臃肿。即将发生的调整对很多人来说会非常痛苦。我认为每个 AI 首席执行官都应该比现在承担更多责任,并想办法以某种方式解决和改善这种情况,因为这对很多人来说会很痛苦。不过,我乐观地认为我们实际上可以比想象中更快地解决这个问题。我们只需要现在就开始。从长远来看,我认为这是一件好事,因为我完全不相信人们说的“工程师要消失了”这种鬼话。世界上有大量的问题,其中很大一部分可以用软件解决,而目前只有一小部分问题正在用软件解决。所以短期内,对于某个过度分配了工程资源的问题,我们需要重新分配——这是一种冷酷的说法,意味着有些人会失业。但从长远来看,我们需要工程师。工程师是最优秀的系统思考者和问题解决者之一。有太多可以用软件解决的问题还没有被软件解决。这意味着我们将让这些工程师去解决以前没有被解决的问题。这对世界来说绝对是净收益,因为有很多问题我们还没有解决,还有一些问题我们虽然解决了,但软件非常糟糕。这将使人们能够用出色的软件来解决它们。Factory 的愿景是,我们是那个工厂,让他们能够去构建这些出色的软件来解决不同的问题。这些问题从琐碎的事情,比如政府软件——通常不太好,无论是 DMV 还是 IRS 网站——所有这些体验通常都很差。我们不必那样生活。我们可以生活在一个所有软件都非常出色的世界里。但还有像药物研究这样的事情:解决疾病需要的大量工作不仅仅是生物学问题;很多都需要世界上最好的软件工程师。以前这些问题没有分配足够的资金来吸引最好的工程师。但现在,由于正在发生的事情,我认为我们将更紧密地将资源分配给这些最大的问题。让最聪明的人和最好的问题解决者去解决它们。我认为我们作为行业的工作就是尽快完成这种重新安置和重新分配,不是 10 年,而是可能 6 个月或 1 年。

So, I think short-term there's going to be a lot of turbulence because a lot of companies have misallocated resources pretty poorly. There's been a lot of bloat. The correction that's going to happen will be really painful for a lot of people. I think every AI CEO should bear much more responsibility than they currently are for that, and also figure out ways to address and ameliorate it in some way, because this is going to be very painful for a lot of people. Now, I have optimism that we can actually address that faster than we think. We just need to start now. Longer term, and why I think this is a good thing, is that I don't believe at all the BS that people are saying of "oh engineers are going away." There is a huge number of problems in the world; a large subset of those problems can be solved with software, and a small subset of those problems are currently being solved with software. So in the short term, for a given problem that was overallocated engineering resources, we need to reallocate those — that's a cold way of saying some people are going to lose their jobs. But in the longer term, we need engineers. Engineers are some of the best systems thinkers and problem solvers. There are so many problems that can be solved with software that are not being solved with software. That means we are going to take those engineers and have them go solve problems that previously were not being solved. That is such a net good for the world, because there are so many problems we are not solving, and also problems we are solving but with really shitty software. This is going to enable people to solve them with incredible software. The vision for Factory is that we are the factory that allows them to go build this incredible software to solve these different problems. These problems range from trivial things like government software — typically not very good, whether it's DMV or IRS web — all that stuff is generally a poor experience. We don't need to live like that. We can live in a world where all software is really fantastic. But also things like pharmaceutical research: so much that goes into solving diseases is not just a biology problem; a lot of it requires the best software engineers in the world. Previously those problems haven't allocated the right dollars to attract the best engineers. But now because of what's happening, I think we will be much more closely allocated to these biggest problems. Let's get the best minds and the best problem solvers to solve that. I think it's our job as an industry to do that relocation and reallocation as quickly as possible, so it's not 10 years but maybe 6 months or a year.

结束语 Closing remarks

Host

太棒了。Matan,我认为你的愿景随着时间的推移所展现出的清晰度和一致性一直非常鼓舞人心。而且,自从我们上次录制这期 Training Data 节目以来,看到你作为领导者的成长以及 Factory 作为公司的成长,真的非常令人振奋。所以,感谢你再次加入我们,分享你的近况。

Wonderful. Matan, I think the clarity and consistency of your vision over time has just always been very inspiring. And just seeing how much you've grown as a leader and how much Factory has grown as a company even since the last time we did this Training Data episode. It's truly all inspiring. So, thank you for joining us again to share what you're up to.

Matan Grinberg

非常感谢。谢谢。

I appreciate it a lot. Thank you.

Host

谢谢。

Thank you.

互动版:逐字朗读 + 针对本期提问 →