Gabe Pereyra 谈打造 Harvey:面向法律行业的全栈 AI

Gabe Pereyra on Building Harvey: Full-Stack AI for the Legal Industry

加布·佩雷拉 Gabe Pereyra · Mercor · 2026-09-18 · 约 41 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Harvey 联合创始人兼总裁 Gabe Pereyra 讲述这家法律 AI 公司如何从 GPT-4 演示走向为律所提供全栈 AI。

Harvey co-founder and president Gabe Pereyra shares how the legal AI startup went from a GPT-4 demo to full-stack AI for law firms.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 20)

全文 · Full transcript(中英对照)

起源故事与早年生活 Origin Story and Early Life

Host

Gabe,非常高兴能邀请到你。非常感谢你今天来参加。

Gabe, I'm super excited to have you. Thank you so much for joining today.

Gabe

非常感谢你的邀请。

Thanks so much for having me.

Host

我想我们可以从你的成长故事开始,你在湾区长大,和我一样。所以也许可以聊聊这个,你来自哪里,你成长过程中对什么感兴趣,以及所有那些最终引领你创立 Harvey 的事情。

I figured we can kick it off with a little bit of your origin story growing up in the Bay Area, same as me. And so maybe talk a little bit about that, where you're from, what you were excited about growing up, and sort of all the things that led up to Harvey.

Gabe

是的,我在 Certino 长大。我父母都是计算机科学博士。我妈妈在 Apple 工作了 30 年,领导了自动更正产品的团队。所以每当你的 iPhone 把词更正成“ducking”,那都是她的错。我爸爸是斯坦福和其他地方的数学教授。小时候,我不想学计算机科学。我不想做我父母做的事。我想成为职业足球运动员。所以我非常喜欢足球。高中辍学去巴西踢球。受伤后,仍然不想学计算机科学,就学了金融,想去做投资银行,也许创办对冲基金。然后开始做那个,我就想,“这不是我喜欢的。”我父母从来没有真正给我压力,但他们会说,“你应该试试计算机科学。”我试了。我就想,“哦,我喜欢这个。”然后很快发现了 AI 研究,并开始与 Yoshua Bengio 合作,他是早期先驱之一。

Yeah, so grew up in Certino. Both my parents were PhDs in computer science. My mom worked at Apple for 30 years, led the team that built the autocorrect product. So whenever your iPhone corrects to ducking, that's her fault. My dad was a math professor at Stanford and other places. And as a kid, I did not want to do computer science. I didn't want to do what my parents did. I wanted to play soccer professionally. So I was super into soccer. Dropped out of high school to go to Brazil and play. Got injured and then still didn't want to do computer science, did finance, wanted to do investment banking, maybe start a hedge fund. Then started doing that and I was like, "This is not what I like." And my parents kind of never really pressured me, but they were like, "You should try computer science." And I tried. I was like, "Oh, I love this." And quickly found AI research and started working with Yoshua Bengio, who's one of the early pioneers.

Host

你当时在蒙特利尔吗?

Were you in Montreal there?

Gabe

我当时远程工作,所以我在 USC 读本科,但和他联系上了,然后本科毕业后做了 Brain Residency,有机会和 Hinton、Jeff 等一群人合作,我一直知道我做 AI 研究是为了最终创业。我以为我想创办的公司是个性化教育。我离开 DeepMind 试图创办那个,结果时机不对,市场不对。去了 Meta,在他们的语言模型团队工作,我遇到了 Winston。我们成了最好的朋友、室友,我在头脑风暴创业想法。

I was working remotely so I was at USC for undergrad but connected with him and then did the Brain Residency when I graduated undergrad and got to work with Hinton, Jeff, like a bunch of these folks and always knew that I was kind of like doing AI research to eventually start a company. The company that I thought I wanted to start was personalized education. I left DeepMind to try to start that and it ended up being kind of the wrong time, the wrong market. Went to Meta and was working on their large language model team and I had met Winston. We became best friends, roommates and I was brainstorming startup ideas.

Host

这大概是什么年份,或者你加入 DeepMind 然后 Meta 的年份是什么时候?

And roughly what year was this or what was the year that you like joined DeepMind and then Meta subsequently?

Gabe

本科毕业后。我去了 Brain 一年,然后 DeepMind,然后离开去创业几年,然后加入了 Meta。Meta 大概就在 CO 之后。

Out of undergrad. I went to Brain for a year, then DeepMind, then left to do startups for a couple years and then joined Meta. Meta kind of right around right after CO.

Host

好的。然后 Winston 和我在 COVID 期间是室友。是的。我就在 GPT-3 出来的时候头脑风暴创业想法。

Okay. And then Winston and I were roommates during COVID. Yeah. I was brainstorming startup ideas right as GPT-3 was coming out.

Gabe

GPT-3,对于业内的人来说,很明显事情开始奏效了。

GPT-3 and I for the people in the industry it was clear that things were starting to work.

Host

是的。GPT-2、GPT-3,如果你在那些实验室工作,那就像是一个神圣的时刻,你可以扩展这些语言模型,你可以看到它们变得更好,它们解决了我过去 10 年研究但无法取得进展的任务。所以很明显这会有前途。当时,Winston 在一家大型律师事务所做诉讼。我给他看这些模型,他尝试了一些助理任务,他说,“这些可以部分做我做的事,但不完全是。但这里有东西。”我就说,“这些显然会变得更好。”我们联系了 OpenAI 的 Sam 和 Brad。他们最终成为了我们的种子投资者。当我们刚开始 Harvey 时,还不清楚你能做大型律所和公司法律工作,因为 GPT-3 还不行。但然后他们给我们看了 GPT-4 的早期版本,那就是灵光一现的时刻,就像好吧,我们可以去建立这种类型的公司。

Yeah. And kind of GPT-2, GPT-3 like if you're working at those labs it was kind of this like holy moment of just you can just scale these language models and you could just see them getting better and they were solving these tasks that like I had worked on over the past 10 years and you just like couldn't really make progress. So, it was pretty clear that that was going somewhere. And at the time, Winston was at a large law firm doing litigation. I was showing him these models and he tried doing some of his associate tasks and he was like, "These can kind of do what I'm doing, but not really. But there's something here." And I was just like, "These are clearly going to get much better." And we reached out to Sam and Brad at OpenAI. They ended up being our seed investors. And when we first started Harvey, it wasn't clear that you could do big law and kind of corporate legal work because GPT-3 wasn't there. But then they showed us like an early version of GPT-4 and that was just the light bulb moment of like okay we can just go and build this type of company.

Host

我记得对我来说也是一样,因为 ChatGPT 出来时我在上大学,我痴迷于它,每天都在用,然后 GPT-4 是让我豁然开朗的地方,因为如果我没记错,GPT-3.5 在律师考试中得了大约 10%,然后 GPT-4 得了大约 90%。

I remember it was the exact same thing for me because I was in college when ChatGPT came out and I was obsessed with it using it every day and then GPT-4 was where it clicked cuz if I remember correctly I think GPT-3 got 3.5 got about 10% on the bar exam and then GPT-4 got like 90%.

Gabe

是的,没错。你知道,那就像一个巨大的,如此大的阶跃函数,发生得如此之快,它让我们提前看到了未来几年。

Yeah, exactly. You know, it was just like a massive it was such a large step function that was happening so quickly where it just gave us the early visibility into the years to come.

Host

是的。不,那是一个疯狂的跳跃。

Yeah. No, it was a crazy jump.

早期岁月与打造Harvey Early Days and Building Harvey

Host

那些早期日子是什么样的?所以,你和 Winston 在洛杉矶。我想你们开始组建团队,或者你们什么时候搬到旧金山,事情真正开始起飞?

What did those early days look like? So, you and Winston were in LA. I imagine that you sort of started building out the team or when did you guys move to San Francisco and things really started taking off?

Gabe

是的,所以当我们大概在八月看到 GPT-4 的早期版本时,那是在 ChatGPT 出来之前,从那时起就很清楚,产品就是把模型交到用户手中。所以我们构建了一个我称之为非常糟糕的 ChatGPT 版本,就像一个文本输入框,你发送查询给模型,得到一些输出。我们开始向律师事务所展示,因为这个模型不是公开的。我们不得不去和律师事务所签保密协议来展示。我们不知道哪类律师会感兴趣。所以我们向大型律师事务所、内部团队、初创律师展示。我们以为初创律师会对此兴奋,而大型律师事务所不会。结果恰恰相反,这有点反直觉,但因为我猜大型律师事务所有巨大的业务量,所以他们可能更倾向于有重复的工作流程,他们可以提示模型做好。

Yeah, so as soon as we probably saw the early version of GPT-4 in like August, so before ChatGPT had come out and even from there it was clear that the product was you know just put this model in the hands of a user. And so we built what I like to call just a really shitty version of ChatGPT where it was like a text input box and then you just sent the query to the model and you got some output. And we started showing that to law firms because this model wasn't public. We had to go and like NDA law firms to show it to them. And we kind of didn't know what group of lawyers would be interested. And so we showed it to large law firms, in-house teams, startup lawyers. We thought startup lawyers would be excited about this and then large law firms wouldn't. And it was actually the reverse which was a little counterintuitive but because I guess the big law firms have giant volume and so they're probably more in the mindset of they can have a repetitive workflow that they prompt the model to do well.

早期阶段与首位客户 Early days and first customer

Gabe

我们非常幸运地遇到了 David Wakeling,他是 ANO 的创新负责人。他立刻就看到了这个方向。当时有些律师看到这个会惊呼“天哪”,也有些会说“这是幻觉,没用”。但 David 很有眼光,很早就看到了,直接说“我们愿意在这上面下大赌注”。那算是我们的第一个客户。当时有点疯狂——我刚做出产品的第一个版本,我们招了另外两个人,BK 和我弟弟。然后我们就和 ANO 谈起了百万美元级别的合同,要把产品推广到他们整个律所。所以那时候我们才意识到,这里面确实有非常有意思的东西。

And we got super lucky meeting David Wakeling, who is an innovation leader at ANO. He immediately saw this. At the time, some lawyers would see this and be like, oh my god, and some would be like, this is hallucinated, this is not useful. But David, to his credit, saw this very early and was just like, we're willing to take a very big bet on this. That was kind of our first customer. It was a bit crazy. I had just built the first version of the product. We had hired two other people, BK and my younger brother. And we got into this million-dollar contract negotiation with ANO to roll this out to their entire firm. So that was kind of when we knew there was something really interesting.

Host

那在早期,你们有没有关于定制模型的愿景?或者说,当时感觉更多是在利用 OpenAI、对基础模型做提示,还是在微调、朝着拥有你们自己的智能这个方向走?

And in those early days, did you guys have the vision around customizing models? Or how much did it feel like it was leveraging OpenAI and prompting base models versus fine-tuning and moving towards owning your guys' intelligence?

Gabe

是的,我觉得这其实是我早期犯的一个大错。因为我的研究背景,我当时对公司的想法是:我们需要定制这些模型,我们需要搞清楚怎么做后训练之类的事情。我早期在这上面过度倾斜了。如果你回去看我们早期的一些博客文章,比如我们和 PWC 做的工作,我们当时大谈特谈要帮 PWC 构建一个包含他们税务知识的定制模型,诸如此类。我们确实试验过这个。显然当时开源模型远不及今天。由于我们和 OpenAI 的关系,我们实际上和他们的团队合作做过一些后训练工作,还和他们一起发表了一些成果。但很明显,基础设施还没到位,模型也还没到位。然后我觉得发生的事情是,预训练——你从 Scaling(规模扩张)预训练中获得的收益太大了,以至于我们不得不转变策略。而这是正确的做法:建立 GTM 组织、打造产品、建立一家传统企业级 SaaS 公司。而且,仅仅因为你在用这些模型,周围就有大量应用 AI 的工作要做。早期人们没有意识到,这些模型当时并不像现在这样工作。你不能上传文档,行内引用不工作。所以我们解决了所有这些不是模型训练的问题,但它们像是早期的 harness 工程。但在我早期的脑子里,非常想帮这些公司构建拥有他们全部数据的自己的模型。因为当时显而易见的是,这些公司在我们这个领域拥有无法放进公共模型的独特专业知识,对吧?比如 PWC 的客户数据因为太敏感,不能放进任何通用模型。

Yeah, I think this is actually a big mistake I made early on. Coming from my research background, the way I thought about the company at the time was: we need to customize these models, we need to figure out how to do post-training and things like that. And I overrotated that early on. If you actually go back and look at some of our early blog posts, for example, the work we did with PWC, we were very much talking about how we're going to help PWC build a custom model that has their tax knowledge and things like that. And we kind of experimented with this. Obviously at the time, open source was nowhere close to where it is today. Because of our relationship with OpenAI, we actually did some work with their team to do post-training work, and we published some of this with them. But it was clear the infrastructure wasn't there. The models weren't there. And then I think the thing that happened was pre-training—you were just getting so many gains from scaling pre-training that we had to shift the strategy. And this was the right thing to do: build a GTM org, build a product, build kind of like a traditional enterprise SaaS company. And there was still—just because you were working with the models—a ton of applied AI work to do around it. Early on, people don't appreciate that these models did not work the way they do now. You couldn't upload documents, inline citations didn't work. So there were all these problems that we solved that were not model training, but they were like the early days of harness engineering. But I think in my head early on it was very much like, oh, I want to help these companies build their own models that have all their data. Because the thing that was obvious at the time was these companies have unique expertise in our domain that can't go in a public model, right? PWC's client data can't go into any general model because it's so sensitive.

给应用层创始人的建议 Advice for application layer founders

Host

那把这个推到今天,市场格局当然已经发生了翻天覆地的变化,不仅有了能力强得多的前沿模型,还有了非常出色的开源模型,我们可以在其之上构建。对于试图搞清楚如何应对这一局面、战略应该是什么的应用层创始人,你会给出什么建议?

And then playing that forward to today, of course, the market landscape has changed radically with not only much more capable frontier models, but also incredible open source models that we can work on top of. What would your advice be to application layer founders that are trying to figure out how to navigate this and what their strategy should be?

Gabe

是的,我觉得我们最终的做法——尽管我一开始有点错——可能正是正确的策略:在你被迫考虑成本之类的事情之前,你应该使用闭源模型。我觉得我们用闭源模型打造了一个非常粘性、客户想要的产品,做得不错。我认为除非极端情况,如果你做不到这一点,你就会浪费大量时间去构建所有这些基础设施、用开源做所有这些事情,而这很可能不会带来真正的差别。一个很强的直觉是:我做研究的时候,你可以在基准上提升 1% 就发一篇论文,对吧?很多 ImageNet 的进步就来自人们想,如果我这样初始化模型,得到 2% 的提升,那其实很有意义,因为我把这些叠加起来。但对用户来说,他们分辨不出 1% 的差别。甚至现在,在一些模型变化之间,都超级难分辨。所以建议就是:如果你现在在打造应用层公司,就打造一个不可思议的产品,获得增长,使用闭源模型。然后会有一个点,你达到某个规模——就像 Cursor 很早就碰到了——你会说,我没法以这个规模经济地提供智能服务了,所以现在我得开始考虑这件事。而那个时机会非常明显。

Yeah, I think what we ended up doing, even though I was a bit wrong about it at the start, was probably the correct strategy: until you get forced to think about cost and these things, you want to use the closed source models. I think we did a good job of building a product with the closed source models that was very sticky and that these customers wanted. And I think except in extreme cases, if you can't do that, you are going to waste a bunch of time building all this infrastructure and doing all this stuff with open source that is probably not going to make the difference. One strong intuition is: when I was doing research, you can publish a research paper on a 1% improvement on a benchmark, right? And a lot of the ImageNet progress came from people just being like, if I initialize a model in this way and I get this 2% improvement, that's actually meaningful because I stack all these up. But for a user, they can't tell a 1% difference. And even now, between some of these model changes, it's super hard to tell. So the advice would be: if you're building an application layer company now, build an incredible product, get traction, use closed source models. And there will be a point where you hit the scale—like Cursor hit this very early on—where you're just like, I cannot economically serve intelligence at this scale, and so now I need to start thinking about it. And it will be pretty obvious when that is.

后训练与推理的经济学 Economics of post-training and inference

Host

这非常有见地,我之前没听过,但从财务角度也说得通,因为后训练一个模型真的很贵。那是一笔固定成本投资,你每个季度或每年投入,然后能够在你作为一家企业的所有推理支出上摊销。真的要到你的推理账单进入尤其是九位数——或者至少能看到达到那个量级的前景——的时候,经济账才开始真正说得通。这也解释了为什么 Cursor 是最早构建 Composer 的公司之一,而你们现在当然也在构建自己的模型,看起来是下一家做出这种转变并引领这个类别的公司。你觉得 Harvey 正在走的这条路和 Cursor 相比,最大的不同是什么?因为我认为有很多相似之处——这种应用层的转变——但模型在代码中学习的方式和在法律中学习的方式也有一些差异。所以我很好奇你怎么看这一点。

That's very insightful and I haven't heard that before, but it makes sense also from a financial standpoint because it's really expensive to post-train a model. That is a fixed cost investment that you make either every quarter or every year and then are able to amortize over all of the inference spend that you have as a business. It's really once your inference bill gets into especially the nine figures—or at least line of sight to getting there—that it seems like the economics really start to make sense. And that explains why Cursor was one of the first companies to build Composer, and of course you guys are now building your own models and seem like the next company making that shift and leading that category. What do you think are the largest differences from the journey that Harvey is going down versus Cursor? Because I think there's a lot of similarities—this application layer transition—but there's also some differences between the way that models learn in code versus the way that models learn in law. And so I'm curious how you think about that.

Gabe

是的,这是个好问题。我的意思是,我们通常把 Cursor 和那些编程公司看作很多维度上的领先指标。所以我会说相似之处在于:我认为知识工作会走上代码的道路。在我们的工程组织里,我认为硅谷的大多数工程师越来越不写代码了——你把它委托给模型。现在你越来越多地思考的是,我如何构建好的测试基础设施,我如何协调所有这些智能体来做软件工程。我们开始看到这一点了。我认为法律与代码不同的地方在于,在代码里你可以这样蒙混过关:如果它通过了所有这些测试,形状大体正确,我就不需要检查每一行。而法律工作——你写的合同会被另一个人阅读。所以我认为还需要做很多工作才能完全达到那一步。我会说这是其中一点。

Yeah, it's a good question. I mean, I think we generally look to Cursor and the coding companies as a leading indicator in a lot of dimensions. So I would say the similarity is: I think knowledge work is going to go the way of code. In our engineering org, I think most engineers in Silicon Valley increasingly just don't write code—you delegate this to the models. And now you're increasingly thinking about how do I build good testing infrastructure, how do I coordinate all these agents to do software engineering. And we're starting to see that. I think where legal is different than code is you can get away with this in code where it's like, if it passes all these tests and the shape is generally right, I don't need to check every line. Whereas legal work—you are writing a contract that will be read by another human. And so I think there is still work to do to get that all the way there. So I would say that's one.

代码与法律:数据与确定性 Code vs. Law: Data and Determinism

Gabe

我想说,数据层面最大的区别可能在于代码有一个很好的特性:大量代码是公开的。你有所有这些 GitHub 仓库和拉取请求,所以你就有了一个关于所有这些变更的不可思议的数据集。而在法律领域,公开的示例为零,比如“这是一起并购,以及发生的所有步骤和修改”。

I would say probably the biggest difference on the data side is the nice property of code: there's a huge amount of it that's public. You have all these GitHub repos with all these pull requests, so you just have this incredible data set of all these changes. In law, there are zero public examples of, like, here is a merger and here are all the steps and edits that happened.

Host

嗯,除此之外,代码还完美地内置了单元测试,这是法律领域所没有的。

Well, in addition to that, code also has unit tests perfectly baked in in a way that doesn't exist in law.

Gabe

没错。所以它是确定性的。你可以运行它,可以检查它。它比语言更像是有类型和结构化的。

Exactly. So it's like deterministic. You can run it, you can check it. It's much more like typed and structured than language.

Host

而且在很多方面,从商业角度来看这更令人兴奋,因为它允许你在专有数据上建立护城河,而大多数编码能力都只是来自这些公共语料库。

And in a lot of ways, it feels like that's more exciting from a business standpoint because it allows you to build a moat on proprietary data in a way that most coding capabilities are just coming from these public corpuses.

Gabe

是的。我的意思是,我认为在编码和法律或知识工作中,护城河都来自于:我们都会构建好的合成数据,你可以增强所有这些公共数据。但有些东西是,如果我们看看自己的代码库,即使我拿所有开源代码库,它们也并不能真正代表像你和我这样的公司正在构建的代码库,因为代码库有一些特点,比如你在一个企业环境中有成千上万的人在上面工作,而不是像 TensorFlow 那样的东西。所以会有一些奇怪的细微差别,我想你会明白的。然后当你开始将你的代码库与业务的其他部分连接起来,并开始与业务一起运作时,我认为这最终会因公司而异。然后我认为类似地,对于这些律师事务所,这就是我们开始思考如何为他们构建的地方,他们实际上就像是一个工厂,大规模地解决所有客户问题。那么你如何连接所有这些,然后它如何运作?当我们思考我们正在构建的东西时,它就像是每个律师事务所的基础设施,以这种方式运作,并随着时间的推移进行定制或持续学习。

Yeah. And I mean, I would say where I imagine the moat both in coding and in legal or knowledge work comes from is in both we will build good synthetic data and you can kind of augment all this public data. But there is just something where it's like if we look at our codebase, it's like even if I took all the open-source code bases, they aren't really that representative of the code bases that companies like you and I are building because there is something about a codebase where it's like you have thousands of people working on this within the context of a business versus something like, you know, this is like TensorFlow or something like that. And so there's just these like weird nuances that I think you'll get. And then as you start connecting your codebase to the rest of the business and this starts operating alongside the business, I think that will end up being unique per company. And then I think similarly with these law firms, like that's what we're starting to think about how we build for them where they are really this like factory that is solving all these client problems at scale. And so how do you connect all of that and then how does that operate? And it's like when we think of what we're building, it's like the infrastructure for each of these law firms to operate that way and customize or like do continual learning on that over time.

Host

这非常有道理,因为我认为有时前沿实验室想讲述的故事是,他们只会吸走客户的所有使用数据,但实际上我想你的客户非常关心不仅数据的法律保密性,还关心它为他们的业务建立的护城河,这有点像他们作为一家公司存在的权利的理由,所以你能够帮助他们利用这一点,基于他们的主权模型建立竞争优势。

Makes a ton of sense because I think sometimes the frontier labs want to tell the story that they're just going to suck up all the usage data from their customers when realistically I imagine your customers care a lot about not only the legal confidentiality of their data but also the moat that it builds for their business of sort of being the reason that they have a right to exist as a firm and so you're being able to help them leverage that to build a competitive advantage grounded in their sovereign models.

Gabe

是的。对我们来说,我认为有时人们把这看作某种二元对立,比如实验室与应用层。我们肯定认为这就像你有云提供商。他们正在构建每个人都会访问的基础设施,然后你有像 Salesforce 这样的企业 SaaS 公司,它们构建公司运营所需的所有工具。即使公司同时使用这两者,公司本身仍然是有差异的。我想你会在这里看到,有时当我听到人们谈论开源时,他们说它会完全取代实验室,我认为前沿智能仍然有巨大的价值,我看不到未来我们会不使用 fable 5 或 soul 或任何最好的模型,但确实感觉有这样一个长尾,有很多任务正在变得智能饱和,你不需要它。

Yeah. And for us it's like I think sometimes people talk about this as some binary of like it's the labs versus the application layer. And we definitely think about as like the same as like you have the cloud providers. They're building this infrastructure that everyone is going to access and then you have like enterprise SAS companies like Salesforce which build all of the tools company needs to operate. And even though companies use both of these like the companies themselves are still differentiated. And I think you'll see that here where I think sometimes when I hear people talk about open source they say it's going to like fully replace the labs and I think there is still a ton of value of like frontier intelligence and I don't see a place in the next like in the future where we're not using fable 5 or soul whatever is the best model but it does feel there's this like long tale where that there is there is a lot of task getting intelligence saturated that you don't need it

Host

这引起了极大的共鸣,因为感觉如果我们正朝着一个经济中在推理上花费数十万亿美元的范式发展,那么显然我们不能让所有这些都成为前沿智能。我们需要有很大一部分是前沿智能,但也有一大部分,可以说是更大的一部分,是这种定制的企业特定智能。我知道你们现在也在组建一个研究团队,当然那是你的背景,你在这个领域有很多人脉。那会是什么样子,你如何考虑组建这个团队?

That resonates enormously because it feels like if we are moving towards a paradigm where there's tens of trillions of dollars in spend in the economy on inference then clearly we can't have all of that being frontier intelligence. We need to have a giant portion of it being frontier intelligence but also a giant portion of it arguably a larger portion of it being this custom enterprise specific intelligence. I know you guys are also building out a research team now and of course that's your background and you have a lot of connections in the space. What does that look like and how are you thinking about standing up the team?

Gabe

是的。所以我们刚开始关注这个,这是我最兴奋的事情之一,因为大约有四年的时间,我们不得不真正专注于企业 SaaS 构建业务,这与在 Brain 或 DeepMind 做研究是完全不同的经历,但真正塑造了我对做实验室的看法。我认为我们正在慢慢开始,我们开始招聘后训练人才,甚至在此之前,我认为让我们有信心现在正是时候的事情是,我们过去的大问题总是数据。很难构建这些超级代表客户工作的数据集,我们和你们做了一些工作,但仍然有像获得完全真实的并购数据室之类的东西。我认为今年我们看到,随着模型在我们与你们的合作中变得更好,我们现在可以构建这些非常好的合成数据集,由领域专家指导,然后使用人类专家来策划它们,我们终于到了这样一个点:当我们查看这些数据集并向我们的内部律师展示时,他们说这看起来像我们会做的工作类型。所以这是第一部分。然后第二个问题是,开源基础模型是否足够好,如果你在此基础上进行后训练,你会接近前沿性能,因为你可以有很好的数据,但基础模型不够好,无法让它们达到有趣的地方。我们与许多 NEO 实验室合作验证了这一点。我们得到了非常有希望的结果。这对我们来说是一个信号,让我们开始招聘人员,继续与 Neol 合作,在内部扩大规模。

Yeah. So, we kind of just started focusing on this and that is like one of the things I'm most excited about because there was kind of four years of we had to really focus on kind of enterprise SAS building like the business which was just a completely different experience than doing research at like brain or deep mind but really shaped how I think about doing the lab and I think we're starting slow we're kind of starting to hire post-training talent uh even before that I think the thing that gave us a bunch of confidence that now is the right time is like the big problem we had in the past was always data. it was very hard to build these like data sets that were super representative of the work our clients were doing and we've done some work with you guys but there was still like getting these fully realistic like M&A like a data room or something like that and I think what we've seen this year as the models have gotten better in our work with you guys is we can now build these really good synthetic data sets that are like steered by domain experts and then use human experts to curate them and we finally gotten to this point where when we look at these data sets and we show them to our internal lawyers, they're like this looks like the type of work we would do. And so that was the first piece. And then the second question was, are the open source base models good enough that if you post-trained on this, you would get close to frontier performance cuz you could have great data, but then the base models are not good enough to get them somewhere interesting. And we worked with a bunch of the NEO labs to validate this. And we got super promising results. And that for us was kind of the signal of like let's start hiring folks, keep working with the Neol, scale this up internally

Host

比如应用计算和轨迹,我想我们是所有这些公司的第一个客户。所以我们正在经历完全相同的转变。你们也是,因为我们在内部做了更多的训练。但我认为你在数据方面所描述的非常有道理,因为我觉得两年前每个人都会说,“哦,合成数据。”

With like applied compute and trajectory and I think we were the first customer of all of those companies. So we're going through the exact same transition. You guys are as we do a lot more training internally. But I think what you're describing on the data front makes a ton of sense cuz I feel like two years ago everyone would say, "Oh, synthetic data."

合成数据与杰文斯悖论 Synthetic Data and Jevons Paradox

Host

这对我们需要多少人去整理数据集有什么影响?市场最终呈现的情况是,当我们让人变得更高效,让他们能够合成地填充数据室、更高效地创建任务时,就出现了杰文斯悖论的情况:每个任务的成本下降了,但消耗的数据量却呈指数级增长,因为经济中自动化更多事物的需求并不短缺。所以看起来你们正在经历实验室完全相同的那种转变:合成增强数据集的成本下降,进而推高了对这些数据的总消耗。

How does this impact the amount of humans that we need to curate data sets? And what ended up playing out in the market is as we made humans more and more efficient by giving them the ability to synthetically populate the data room and create tasks much more efficiently, there was a case of Jevons paradox where the cost per task went down, but the amount of data that was consumed went up exponentially because there's not a shortage of demand for automating more things in the economy. And so it seems like you guys are going through that exact same transition of the labs of having the cost of your synthetically augmented data sets go down and then that driving up the total consumption of that data.

Gabe

我认为我们的总成本会大幅上升,就像你说的那样。当我们构建一些合成数据室时,比如一个大型数据室有 8000 万 token,也就是 1 万份文档。当我们问你们,嘿,我们需要律师来审查每一份文档。你们说,是的,大概需要 10 万美元的律师时间。然后我们说,好吧,我们需要数千个这样的数据室来训练模型做尽职调查。

I think the total cost for us is going to go up hugely where it's like exactly what you said when we built some of these synthetic data rooms like a large one is 80 million tokens. is 10,000 documents. And when we asked you guys, hey, we need lawyers to review each of these documents. You were like, yeah, it's going to be like a $100,000 of lawyer time. And then it's like, okay, we need thousands of these to train models to do diligence.

Host

然后还有数百个执业领域。有所有这些不同的国家。所以我们将合成数据视为构建某种脚手架的方式,但你需要用来纠正所有这些的人工时间,正如你所说,将会增加。我认为这得到了价值的支持,比如如果你能在此基础上训练这些模型,那就非常有趣。

And then there's hundreds of practice areas. There's all these different countries. And so we see synthetic data as a way to build kind of like this scaffolding, but then the amount of human time you need to correct all of this is to your point, it's going to go up. And I think it's supported by the value of like if you can train these models on this then it's very interesting.

本地化与法律专业知识 Localization and Legal Expertise

Host

你关于国家的说法也让我非常着迷,因为我认为法律比任何其他领域都更需要极端的本地化,不仅是对美国这样的国家,还有州法律、特定地区、语言等等所有这些不同的东西。甚至经常有各种法律实践,顶级律师知道,但并没有为那些特定地区写下来。那么,你如何考虑捕捉这些,因为它塑造了你们构建模型的数据策略?

The thing you said about countries is also so fascinating to me because I think law more so than any other domain requires extreme localization not just to the country like the US but also to the state law to the like specific locality to the language all these different things. There's often times even all sorts of legal practices that are known by the top lawyers but aren't written down for those given localities. And so how do you think about capturing that as it shapes your data strategy for the models you guys build?

Gabe

是的,我的意思是这真的很难。我喜欢用的一个类比是,当我做工程时,我还可以。我在分布式系统和模型训练方面还行。但如果你看所有的工程实践,有后端、前端、编译器等等所有这些不同的专业。就像随着模型在编码方面变得更好,如果你想让他们擅长编译器,你需要去找编译器专家,让他们为你生成训练数据。而他们与做分布式系统的人是不可互换的,例如。

Yeah, I mean it's really hard. Like one analogy I like to use is like when I did engineering, it's like I was good. I was okay at like distributed systems and like model training stuff. But if you look at like all of the engineering practices, there's like backend, front end, and like compiler and like all these different specializations. And it's like as the models get better at coding, if you want to make them good at compilers, you need to go find the experts in compilers and have them generate training data for you. and they're not exchangeable with someone that did distributed systems for example.

Host

完全同意。

Totally.

Gabe

法律也是一样,你有所有这些执业领域,比如做基金设立的律师与做并购或知识产权诉讼的律师不可互换。然后还有你提到的额外情况,比如现在如果你去不同的国家,在英国的做法与美国不同,足以让这些专家也不可互换。我认为我们在与你们合作进行合成数据方面已经走得很远。但当我想到极限时,我实际上认为这需要来自客户,因为如果你想想每年生成的法律训练数据的成本,就像每年法律服务的成本。所以有 1 万亿美元的法律工作完成。

And then legal is the same where you have all these practice areas where like a lawyer that does fund formation is not interchangeable with a lawyer that does M&A or IP litigation. And then you have the additional thing that you mentioned where it's like now if you go to a different country the way you do this in like the UK is different than the US enough that these experts are also not interchangeable. I think we get very far of doing synthetic data working with with you guys. But when I think of the limit, I actually think this needs to come from the customers where if you think of the cost of the amount of legal uh training data that's generated every year, it's like the cost of legal services every year. And so there's a trillion dollars of legal work done.

Host

如果你把所有这些都视为训练数据,就像我们可能会花费数十亿美元与你们一起构建训练数据。但这仍然只是所有数据的一小部分,而且有来自合伙人的独特专业知识,他们永远不会去标注数据。所以我确实认为这就是为什么你需要弄清楚,那些数据不能进入模型。这就是为什么你需要弄清楚如何为每个律所构建基础设施,让他们的专家将知识放入他们自己的模型中。对我来说,这实际上是我认为在极限情况下唯一的方法。

And if you think of all of that as training data, it's like we'll probably spend billions of dollars on building training like data ourselves with you guys. But it's like that's still a fraction of all of that and there's unique expertise from like partners that are never going to go label data. And so I do think that's why you need to figure and that data can't go in the models. And so that's why you need to figure out how to build infrastructure for each firm to take their experts and put it in their own models. It's like that's to me actually the only way that I think you do this in the limit.

复杂分类体系与专有数据 Complex Taxonomies and Proprietary Data

Host

感觉法律拥有我评估过的所有领域中最复杂的分类法,因为子领域太多了,而且一切都有太多没有写下来的知识。在美国国内以及所有地区,一切都有点不同。嗯,感觉这给了你们机会,基于客户文件的元数据,比如你知道他们在数据室有多少文件?他们关心哪些类型的事情?你如何构建这些非常酷的专有数据集,这是一个如此迷人的问题。

It feels like legal has the most complex taxonomies of any domain that I've really evaluated because there's just like so many sub fields and like everything there's like so much knowledge that's not written down. Everything is a little bit different within the US and then all of the geos. Um and it feels like that gives you guys the opportunity to figure out based on the metadata of customer files of like you know how many files do they have in a data room? What are the types of things they care about? how you can build these really cool proprietary data sets which is such a fascinating problem.

企业级验证器与RL任务 Verifiers and RL Tasks for Enterprises

Host

我感兴趣的另一点是,我想象你们合作的许多企业会有他们可能输出的文件,但他们可能没有干净的验证器,你可以用于评估和 RLVR。你如何看待这种转变?因为这是我们在所有合作企业中发现的,就像我们进去,他们说我们有这些巨大的文件系统和所有电子邮件,我们想把它变成定制智能。然后我们说你需要 RL 任务。很难实现这种转变。但我很好奇你们是怎么想的。

The other thing I'm interested in is I imagine that a lot of enterprises that you work with would have files that they might output but they might not have clean verifiers that you can use for eval and RLVR. How do you think about that transition? Because this is something we're finding with all the enterprises we work with where it's like we go in and they're like we have these giant file systems and all of our emails and we want to like turn this into custom intelligence. Then we're like you sort of need RL tasks. It's hard to make that transition. But I'm curious how you guys are thinking about it.

Gabe

是的。我认为这是一个巨大的挑战,就像你提到的,法律没有单元测试,但我们发现有效的是 LM 作为法官,并开始构建这些评分标准。一个非常酷的技巧是我的弟弟 Julio,他做了很多合成数据工作,他想出的方法是我们构建一些合成数据集的方式是:从评分标准开始,定义你在尽职调查中可能看到的所有问题,然后用它来生成数据,使其包含这些问题,这样当模型生成工作产品时,你知道,哦,这份合同有这个问题,你找到了吗?所以我认为这是这个问题的简单版本。我们开始大量思考的是,我们如何进入一家律所并帮助他们构建这些验证器?然后这越来越会成为产品的一部分,就像如果你想到一家律所,他们每年处理 1 万个客户事项,一个非常有价值的产品是:这是一个自动化系统,无论是智能体还是人类产生工作,我们都可以运行所有这些基于你律所运作方式的检查。但这真的很难。就像我不认为我们找到了简单的方法。

Yeah. I think this is one of the huge challenges where like you mentioned legal doesn't have unit tests but what we found works well is kind of LM as a judge and starting to build these rubrics. Like a really cool trick that Julio, who's my younger brother who does a lot of the synthetic data work, came up with is the way we built some of these synthetic data sets was start with the rubric and define all of these issues that you might expect to see in a diligence and then use that to generate the data so it has those issues so that then when the model generates the work product you you know oh this contract had this issue did you find it? And so I think that's the like simpler version of this problem. What we're starting to think about a lot is how do we go into a law firm and help them build these verifiers? And then increasingly that'll be part of the product where it's like if you think of a law firm and they're doing 10,000 client matters a year, it's like a really valuable product is here's an automated system where whether it's an agent or a human produces work, we can run all these checks that are based on, you know, the way your firm operates. But I it's really hard. Like I don't think we found kind of like an easy way to do this.

构建评分标准的挑战 Challenges of Building Rubrics

Host

是的。而且我认为构建一个伟大的评分标准的难点在于,当律师写法律备忘录时,可能有 10 种不同的东西,10 种不同的方式来做,以及 100 种可能出错的方式。

Yeah. Well, and I think the hard part of building a great rubric is when a lawyer is writing a legal memo, there might be 10 different things, 10 different ways they could go about doing it and 100 ways they could go wrong.

构建评分标准以防奖励黑客 Building Rubrics to Prevent Reward Hacking

Gabe

所以,要构建一个涵盖整个问题空间的评分标准,以防止训练中的奖励黑客行为,这难得出奇。

And so building a rubric that is all-encompassing of the problem space to stop reward hacking when training is shockingly challenging.

Host

太难了。

So hard.

Gabe

嗯,但我同意你的看法,我们需要弄清楚有哪些创造性的策略,能更好地从那些全职员工的大脑中提取信息,无论是通过访谈还是某种更高效的方式,将其转化为干净的验证器。

Um, but I agree with your take that we will need to figure out what are these creative strategies for how can we better extract the information from those FTEs' heads whether via interviews or some more efficient way to turn it into clean verifiers.

Host

是的。我的意思是,我一直思考的是,就像在编程中,你有这些拉取请求,高级工程师会给出修改。在法律领域,对于任何工作成果,都有来自合伙人的这些红线批注。

Yeah. I mean the to me the thing that I've always thought about is like the same way in coding you have like these pull requests where you'll have a senior engineer give edits. What you have in legal is for any work product there's all these red lines from like the partner.

Gabe

那些就相当于你的单元测试。

Those are sort of your unit test equivalent.

Host

是的,没错。我认为最终如果你能达到这样的程度:走进一家律所,帮助他们把资深律师所做的所有修改都转化为专门为他们定制的评分标准。对我来说,这将是解决这个问题的一个非常有趣的方式。我猜这大概还不够,但它会给你一个非常好的基线,就像:这是你的合伙人在他们工作期间传递给律所其他人的所有知识。

Yeah, exactly. And I think eventually if you could get to the point where you could go into a firm and help them take all of the edits their senior lawyers have made and turn those into rubrics specifically for them. Like that to me would be like a really interesting way to like solve this problem. My guess is that's probably not enough, but it would give you kind of this really good baseline of like here's all the knowledge your partners have like transferred to the rest of the firm in the history that they've been working here.

Gabe

这确实很有共鸣。而且我认为它给了我们一个信号:很多时候,模型学习的最佳类比就是人类的学习方式。法律助理的学习方式往往不仅仅是阅读 10 份不同的法律备忘录,而是从他们的合伙人那里获得红线批注。

That definitely resonates a lot. Well, and I think it gives us a signal of often times the best analog for how the models will learn is how humans learn. And the way that the legal associates learn is often not just reading, you know, 10 different legal memos, but it's rather getting the red lines from their partner.

Host

没错。

Exactly.

Gabe

嗯,所以如果我们能找出方法,最好地让这些数据对模型来说既高效又有用,那似乎会非常强大。

Um, and so if we can figure out ways to to best make that data efficient and useful for the models, that seems extremely powerful.

Host

我们已经谈了很多关于你们数据策略的内容。我也很好奇你如何看待你们的算力策略,因为构建法律领域的前沿模型当然是一项巨大的投资。我非常感兴趣你们是如何着手的。

We've talked a lot about your data strategy. I'm also curious how you think about your compute strategy because of course it's a giant investment to build the frontier model for law. I'm super interested in how you guys are going about that.

后训练策略与算力 Post-Training Strategy and Compute

Gabe

是的。所以,在我们讨论后训练策略之前,也许重要的一点是,我认为有时人们有一种误解,认为这是二元的事情,就像你要么后训练这个模型,它要么有效要么无效。我认为我们做得很好的一点是将所有这些基础设施模块化。我们已经构建了大量基础设施,并且在算力方面规模很大,仅推理方面,现在就要为来自不同提供商的这些模型提供服务,以及所有的评估服务基础设施。所以我们现在有了所有这些组件,当我们考虑模型训练时,不再是需要构建一个庞大的法律模型来解决所有问题。而更多的是我们有十几个产品表面领域。在每一个中,我们服务多个模型和多个智能体。因此,我们拥有所有这些数据,可以开始思考在哪里可以服务开源模型,在哪里可以进行路由,然后在某些情况下,我们能否后训练一个模型并将其放入系统中。我认为目前我们会将努力分为两个阶段。一是我们可以为降低成本做哪些后训练工作,所以只需看看这些现有系统。我们正在做的一个项目中的一个很好的例子是,我们成本的一大部分来自一个叫做 vault 的产品部分,它有审查表,主要用例是上传一百万封电子邮件或 1 万、10 万份合同,并分析所有这些并提取字段等。如果你想象在 10 万份合同上运行某个大型模型,并且你想从所有合同中提取一百个字段,就像我们有单个查询可能花费 1 万、2 万美元。所以我们正在做的一些后训练工作的一个例子是,我们能否后训练某个模型,使其能够更高效地完成该任务。因此,在产品不同部分,在成本方面有大量工作。然后第二部分是新能力。我们与你们合作的一个数据集是构建这个尽职调查数据集,在那里你有这些数据室,可能有 1 万份合同,比如 8000 万个 token,它们就是无法放入上下文窗口,所以我们正在做一些工作,我们能否通过强化学习让这些模型能够在这些数据室上更有效地运作。所以我会说,模型方面的努力是如何将这两者——成本节约和创造新能力——结合起来,然后越来越多地在整个产品中提供这些服务。

Yeah. So maybe the important thing before we go into the like post-training strategy is I think sometimes people have this misconception that it's like this binary thing like you either postrain this model and it works or it doesn't. And I think what we've done a good job of is modularizing all this infrastructure. And so we've built a ton and we're at like a large scale compute-wise just from inference now of like serving all these models from all these different providers. all of the like evaluation serving infrastructure. And so we have all these pieces where now when we think of model training, it's not we need to build this one massive legal model that's going to solve all our problems. And it's much more we have a dozen product surface areas. In each of those, we're serving multiple models and multiple agents. And so we have all this data where we can start thinking about where can we serve open source models, where can we do routing, and then in certain cases, can we post- train a model and put it into the system. the big effort like I would say we would split the effort into two phases right now. one is what is the post- training work we can do for cost reduction and so just looking at these existing systems. So one one good example in a project we're working on is a one big part of our cost is we have a part of the product called vault which has review tables which the main use case is upload you know a million emails or 10,000 100,000 contracts and analyze all of them and pull out fields and things like that. And if you imagine running, you know, some large model over a 100,000 contracts and you want to pull out a hundred fields from all of them, it's like we have individual queries that can cost $10,000, $20,000. And so one example of some post-training work we're doing is can we post train some model that can do that task much more efficiently. And so there's a bunch of work on the cost side there um in different parts of the product. And then the second piece is new capabilities. So one data set that we worked with you guys on was building this diligence data set and there you have these data rooms that can be 10,000 contracts like 80 million tokens and they just don't fit in the context window and so we're doing some work can we you know RL these models to be able to operate over these data rooms much more effectively and so I would say the model efforts are how do we combine both these like cost savings and creating new capabilities and then increasingly serve these across the product.

Host

你能分享一下你们目前在推理上的花费大概是多少吗?

Are you able to share what's the ballpark of what you guys spend on inference right now?

Gabe

我的意思是,在数亿美元级别。

I mean it's in the hundreds of millions.

Host

哇。是的。是的。所以这是一个巨大的机会,弄清楚如何优化它。这非常有趣。那么训练算力策略呢?因为当然完全理解,主要成本结构是推理算力,然后弄清楚如何使其更高效并提高性能。但我想你们有办法做到这一点。无论是获取裸金属 GPU,与某个超大规模云服务商合作,还是与 NeoCloud 合作,或者与一个 RAL 即服务供应商合作,后者抽象掉了很多配方构建和后训练的工作。你们现在如何应对这个问题,以及如何看待其发展?

Wow. Yeah. Yeah. So it's a giant opportunity to figure out how to optimize that. That's super interesting. And what about the training compute strategy? Because of course totally understand like the main cost structure is inference compute and then figuring out how do you make that more efficient and also improve performance. But I imagine there's ways that you can go about this. Whether it's um getting bare metal GPUs, working with one of the hyperscalers or working with a NeoCloud uh or working with an RAL as a service vendor that sort of abstracts away a lot of the recipe building and postraining. How are you guys navigating that right now and thinking about how that will develop?

训练基础设施与Harvey的未来 Training Infrastructure and Future of Harvey

Gabe

是的,所以我们开始与一些 NeoCloud 合作,并将训练卸载给他们。我们现在开始与 Fireworks 和 Baseten 合作,既为这些模型提供服务,他们也有良好的训练基础设施,这就是 Cursor 如何做 composer 模型的方式,所以他们有非常好的强化学习基础设施。我认为他们帮助解决了这样的问题:

Yeah, so we started with working with a bunch of the Neols and kind of offloaded the training to them. We're now starting to work with like fireworks base 10 both for like serving these models and then they have good training infrastructure and so that's how cursor did like the composer model and so they have like really good RL infrastructure. I think they help solve the problem of like

Host

如果你在一个堆栈上训练,然后在不同的堆栈上服务,数值不对齐,你就会遇到一堆问题。

if you train on one stack and then serve on a different stack and the numeric aren't aligned you run into a bunch of issues.

Gabe

所以我认为我们想尽可能推进这一点,然后我认为在某个时候,希望我们达到需要构建自己集群的规模,但现在我认为有了这些,我们可以走得很远。

So I think we want to push that as far as it goes and then I think at some point hopefully we get to the scale where you know we need to build our own cluster but right now I think with these we can get pretty far.

Host

是的,我们处于类似的位置,我们在 Fireworks 之上使用 Skyro。嗯,感觉这是一个很好的层次,它给了我们足够的控制,我们可以构建自己的所有配方,但很多信息工作被抽象掉了,这样我们可以专注于我们关心的主题能力。我也感兴趣的是,所有这些对 Harvey 的未来意味着什么,因为当然,业务已经从早期能够回答一些法律问题,发展到如今能够完成大量有经济价值的工作。所有这些后训练能力如何融入那个未来?

Yeah, we're in a similar spot where we're using Skyro on top of Fireworks. Um and it feels like a good layer of it gives us enough enough control where we can be building out all of our own recipes but a lot of the info work is sort of abstracted away so that we can focus on the subject matter capabilities that we care about. I'm also interested in what all of this means for the future of Harvey because of course the business has come such a long way from being able to answer some of the legal questions and the the early days to now being able to complete giant volumes of economically valuable work. How does all of this post-training capability fit into that future?

Gabe

是的。

Yeah.

从应用层到全栈AI From Application Layer to Full-Stack AI

Gabe

所以我想说,现在我们公司正在开始的转型,是从一家应用层公司转向我们所说的全栈 AI 公司。我觉得我们刚起步时,Harvey 这种应用层公司多少带点贬义,就好像,哦,你不过是在调用模型 API。而大约一年半前开始出现的情况,首先是智能体,我觉得人们意识到,光是智能体的基础设施层——甚至不是构建模型,而是如果你要构建这些云沙箱、连接各种工具、让这些智能体运行一小时或一天,还需要 ZDR 之类的——你要构建和提供这些基础设施其实相当复杂。所以我们认为存在应用层,也存在智能体基础设施层。于是我们开始构建所有这些,现在也开始构建模型层,我们会继续与实验室和云提供商合作,但越来越多地,能否提供开源模型?我的感觉是,达到一定规模的应用层公司——就像你看到 Cursor 经历了这个转型——我认为这会因为一系列结构性、竞争性、成本原因而发生。对我们来说,是成本原因和一些性能原因的结合。然后我认为越来越多地,随着我们构建这些,它给了我们帮助客户做同样事情的知识。我们越来越多地有客户问我们,如何拥有堆栈的更多部分?如何定制我们的模型?我们应该训练自己的模型吗?在我们经历这个转型的过程中,我们正在弄清楚哪些部分适合客户拥有,哪些部分适合我们拥有,哪些部分我们会外包给推理提供商、云实验室。然后显然,随着每个人构建更多这种基础设施,这也会变化。

So I would say right now the transition we're starting to make as a company is from an application layer company to what we call a full stack AI company. I think when we started, Harvey application layer company was kind of like a derogatory term, where it's just like, oh, you're just calling this model API. And I think the thing that started happening about a year and a half ago, first with agents, was I think people realized, oh, just the infrastructure layer of agents—like, not even building models, but just if you need to build these cloud sandboxes, connect all of these different tools, have these agents that run for an hour or a day, and you need ZDR and all these things—like the infrastructure you need to build and serve that is actually pretty complicated. And so we think of it as there's the application layer, there's this agent infrastructure layer. And so we started building all of that, and then now starting to build this model layer too, where we'll keep working with the labs and the cloud providers, but increasingly, can you serve open-source models? And so my sense is application layer companies that get to a certain size—like you saw Cursor go through this transition—I think this will happen for just a bunch of structural, competitive, cost reasons. And so for us it's a combination of cost reasons, there's some performance reasons. And then I think increasingly what we're seeing is as we build this, it's giving us the knowhow to help our customers do this. And we're increasingly having customers ask us, you know, how do we own more parts of the stack? How do we customize our models? Should we be training our own models? And as we're going through that transition, we're kind of figuring out here's the parts that it makes sense for customers to own, here's the parts that it makes sense for us to own, here's the part we'll offload to inference providers, cloud labs. And then obviously that will change as everyone builds more of this infrastructure.

定制模型与基础模型 Custom Models vs Base Models

Host

你预计在比如 3 到 5 年后,业务中有多大比例会是针对每个具体律所或企业的定制模型,而不是 Harvey 更多使用的基础模型,无论是你们后训练的开源模型还是前沿模型?

What portion of the business do you expect to be these custom models for each specific law firm or enterprise you work with versus more so the base models that Harvey is using, whether it's your post-trained open source models or frontier models, in say 3 to 5 years?

Gabe

这是个好问题。我不确定我有确切的百分比划分。我想说,我现在思考的方式是,我们开始考虑如何与单个律所,或者可能是律所和客户一起,帮助他们训练模型,无论是和我们一起,还是可能和 Neo 实验室一起。我认为那会有点像现在的强化学习即服务,需要大量数据处理并试图让它工作。但我认为三到五年后,这实际上会消失,只是你作为律所使用产品,它随着你的使用而定制化。所以我认为现在我们正处于这个过渡期,我们做后训练的方式、Neo 实验室的方式、你们构建数据集的方式,我认为这会改变,最终会消失。就像,这是我强烈感受到的事情,当我们创办公司时,人们会说,哦,你在用这个模型,你在用那个模型,最终这会消失,你只会说,哦,我相信 Harvey 会为我做这些决定,然后我知道在底层你们为正确的事情使用了正确的模型。

That's a good question. I don't know that I have the exact percentage split. I would say the way I think about it now is we're starting to think about how do we go with an individual law firm, or potentially a law firm and a client, and help them train a model, whether it's with us, potentially with us and a Neo lab. And I think that will look kind of how RL as a service looks right now, where it's a lot of data munching and trying to get this to work. But I think in three to five years this actually goes away and it's just you use the product as a law firm and it gets customized as you use it. And so I think right now we are in this transition period where the way we're doing post training, the way Neo labs, the way you guys are building data sets, I think this will change and it will eventually go away. Like, this is the thing I felt strongly where when we started the company and people would talk about, oh, you're using this model, you're using this model, like eventually this will go away and you'll just say, oh, I trust Harvey to make those decisions for me, and then I know that under the hood you're using the right models for the right things.

Host

我同意你关于持续学习的观点,当模型生成法律备忘录,合伙人在该备忘录上给出红线修改时。当然,它需要从每一个增量交互中学习,作为它使用的奖励信号。因此,确保你拥有那种特定的企业智能对于实现持续学习至关重要。

I agree with your point on continual learning of as the model generates a legal memo and the partner is giving the red lines on that legal memo. Of course, it needs to be learning from each of those incremental interactions as the reward signal that it's using. And so making sure that you have that specific enterprise intelligence is so critical for enabling continual learning.

Gabe

是的。然后你想要的产品体验就是,我做我的工作,有一个神奇的 AI 系统随着我做工作而变得更聪明,并且总是帮助我完成。但我不会想,哦,我应该后训练这个还是应该放在提示中之类的。

Yeah. And then the product experience you want is just I do my job and there's this magical AI system that just gets smarter as I do my job and is just always helping me do it. But I'm not thinking about, oh, should I post train this or should I put in the prompt or something like that.

法律行业的未来 Future of Legal Industry

Host

这对法律行业或律所的未来意味着什么?

What does this mean for the future of the legal industry or law firms?

Gabe

我开始思考和向律所描述的一种方式是,他们的业务基于杠杆模式。我把他们的杠杆模式想象成一个金字塔,合伙人在顶部,然后是律师助理。律所正在得到的,其实每家公司都在得到的,是现在你将拥有这种疯狂的杠杆,对吧?所以金字塔的底部现在大得多,是智能体,这引发了一系列问题,对吧?你如何组织所有这些律师助理和智能体?你如何连接律所中的所有数据和工具,以便你的律师助理能与这些智能体合作?你如何让持续学习发挥作用?然后显然,你的商业模式也取决于你如何提供这些服务和你的杠杆模式。那么这将如何改变?对我来说,我认为律所将是首批端到端改变的行业之一,就因为他们的计费模式。他们如此基于文本,就像编码完全改变了你的工作方式,但通常编码不会改变公司的商业模式,因为它像是构建产品或其他东西的支持功能。但我认为法律,你将看到那种全面转型。我认为所有这些律所的巨大机会是,历史上你受收入限制,你的收入与员工人数线性增长。现在,如果你能协调所有这些智能体,似乎对于那些弄清楚的律所来说有巨大机会,但这将是一个充满挑战的转型。

One way I've started thinking about and describing it to law firms is like their business is based on the leverage model. And I kind of think of their leverage model as you have this pyramid with partners at the top and then associates. And what law firms are getting, but what every company is getting, is now you're going to have this insane amount of leverage, right? So like there's this much bigger base of the pyramid that is now agents, that raises a bunch of questions, right? How do you organize all these associates and agents? How do you connect all of the data and tools in your firm so your associates can work with these agents? How do you get the continual learning to work? And then obviously like your business model is also predicated on how you deliver these services and your leverage model. And so how is that going to change? And so to me it's like I think law firms will be one of the first industries to change end to end just because of their billing model. They're so text-based, like it feels like coding changed the way you work completely, but usually coding doesn't change the business model of the company because it is like a supportive function to building a product or something else. But I think legal you're going to see that full transformation. And I think the big opportunity for all these law firms is like historically you were revenue limited, like your revenue grew linearly with your headcount. Now, if you can coordinate all these agents, it seems like there's a huge opportunity for these law firms who figure it out, but it will be like a challenging transition.

五年后律师会更多还是更少? More or Fewer Lawyers in 5 Years?

Host

那么,5 年后律师会更多还是更少?

So, will there be more lawyers in 5 years or less?

Gabe

是的,这是我们经常争论的事情。我想我仍然持乐观态度,我认为会有更多的软件工程师和更多的律师。红杉资本的 Constantine 有一条很好的推文,基本上是一个律师数量的图表,比如互联网初期和互联网之后,他的观点基本上是,当互联网被创造时,你也可以提出同样的论点,现在律师不必走到图书馆,不必寻找判例法,应该更容易。但与此同时发生的是,世界变得如此复杂,即使你可以更高效地完成固定数量的法律工作,法律工作的总量和复杂性也增长了。如果我思考这些模型会发生什么,就像你将拥有十亿、万亿个智能体。

Yeah, this is something we debate a lot. I think I'm still on the bullish side of like I think there's going to be more software engineers and more lawyers. And Constantine from Sequoia has this really good tweet that is basically like a plot of number of lawyers like when at the start of the internet and after the internet and his point is basically like you could have made the same argument when the internet was creative like now lawyers don't have to walk to the library and they don't have to like look for case law and like it should be easier. But then what happened in parallel is like the world got so much more complex that even though you could do a fixed amount of legal work more efficiently, the total amount of legal work and the complexity of it grew. And if I think about what's going to happen with these models, it's like you're going to have a billion, a trillion agents.

AI与法律工作的未来 AI and the Future of Legal Work

Host

会有这么多公司,事情会多到你需要这些能理解正在发生什么的专家。编程感觉也是一样。我的意思是,在某个时间尺度上也许不同,但在未来 5 到 10 年,我猜我思考的是,什么时候一家《财富》500 强公司会说我不打算用人类来做这件事,比如赌上公司命运的诉讼或这次并购,而我觉得短期内我看不到这种情况发生。

There's going to be all these companies, there's just going to be so much that you're going to need these experts that can understand what's happening. And it feels like the same thing with coding. I mean on some time frame maybe that's different, but in the next 5 to 10 years I guess the thing I think about is like when is a Fortune 500 company going to say I'm not going to use a human to do this, like bet the company litigation or this merger, and I'm like I don't see that happening anytime soon.

Gabe

是的。这是典型的劳动总量谬误,每次我们经历生产力革命,大家都会猜测我们将无事可做,会出现大规模失业。这可以追溯到卢德分子。但每次实际发生的是,我们从不缺事情做。当我们让每个人的生产力提高五倍时,我们作为一个社会就能完成多得多的事情。

Yeah. It's a classic case of the lump of labor fallacy where every time we've had a productivity revolution everyone has speculated that we're going to run out of things to do and there's going to be mass unemployment. It traces back to the Luddites. But then every time what actually played out is that we had no shortage of things to do. And when we made everyone five times more productive, we were just able to accomplish so much more as a society.

Host

我敢打赌的一点是,5 到 10 年后顶级合伙人赚的钱会比现在更多。对我来说,这就像你在编程领域看到的那样。顶尖工程师现在赚的钱比 5 年前多得多。这如何影响你们在多大程度上专注于大型律所与企业客户的战略?

The one I would bet on is the top partners in 5 to 10 years make more money now than they do. And to me, it's like this is what you've seen with programming. It's like the top engineers are making way more money now than they did 5 years ago. How does it shape your guys strategy of how much to focus on big law firms versus enterprises specifically?

Gabe

我认为我们一直以来的战略思路是,对我们公司最好的结果是为律所和企业创造双赢。所以我认为你不能把它们分开。例如,当我们与这些大型企业合作时,他们一半的支出是外部律师费,一半是内部支出,但重叠很多。我们思考的很多是,如何为律所打造一个能让他们更赚钱的产品,因为他们显然面临来自企业的压力,说我们看到了这项惊人的技术,你们在用这个来转型业务吗?所以我们需要帮助律所应对这一点。然后还有很多内部工作,比如合同,不会外包给律所,但最终我们希望构建一个平台,让每个企业都能在 AI 上运行公司的法务部门,并最终扩展到相邻部门,然后当你需要与律所合作时,最终与相邻的专业服务提供商合作时,你都能做到。但我认为你越能以共赢的方式做到这一点就越好。

I think we've always thought of the strategy of like the best outcome for our company is we create a win for law firms and for enterprises. And so I think you can't decouple these. And so for example, when we work with these large enterprises, half their spend is outside counsel, half their spend is internal, but there's so much overlap. I think a lot of what we think about is like how do we build a product for the law firms that can make them more profitable because they're obviously getting pressure from these enterprises to say you know we're seeing this amazing technology you know are you using this to transform your business so we need to help law firms with that and then there is a bunch of work internally that happens like contracting that doesn't get farmed out to law firms but ultimately we would want to build some platform where it's like every enterprise can run their company's legal and eventually adjacent departments on AI and then whenever you need to collaborate with a law firm and eventually adjacent professional service providers you can do that but I think the more you can do this in a way that everyone wins the better.

Host

完全同意。目前你们是按席位收费、按使用量收费,还是某种组合?

Totally. And right now do you guys have a per seat model or usage based or some combination?

Gabe

我们现在是按席位收费。

We're per seat right now.

Host

你觉得这会随时间演变吗,还是不确定?

Do you see that evolving over time or not sure?

Gabe

我的意思是,我认为产品中可能会有部分会发生变化,最终会采用某种混合模式,我认为会有一些有趣的转嫁成本之类的东西,就像你在很多现有法律科技中看到的那样,例如 Relativity,如果你作为律所使用它,你可以把这些成本转嫁给客户。但我认为对我们来说,主要关注的是,我认为客户现在已经明白,AI 变得比大多数人预期的要昂贵得多。我认为我们想关注的是,如何让这些成本透明,如何让它们可预测,并弄清楚对律所和企业来说什么是最好的商业模式。

I mean I think there will be probably parts of the product where it changes and you'll end up doing some like hybrid and I think there'll be you know interesting like pass through costs and things like this like you see this with a lot of existing legal tech where for example relativity if you're using it as a law firm you can pass these costs to your clients but I think for us the main thing we're focused on like I think what customers have understood now is AI is getting much more expensive than I think most people anticipated I think what we want to focus on is you know how do we make these costs transparent how do we make them predictable and just figure out what is the best business model with for law firms and then for enterprises.

Host

企业重视可预测性以及他们如何规划业务,而你们能够在按席位模式中提供这一点,抽象掉很多幕后发生的事情,即你如何确保为他们试图完成的任何任务提供正确的模型。

Enterprises value predictability and how they plan their business and you guys are able to give them that in the per seat model abstracting away a lot of what happens behind the scenes of how you make sure that you get them the right model for whatever task they're trying to complete.

Gabe

没错。

Exactly.

Harvey名字背后的故事 The Story Behind the Name Harvey

Host

我们之前谈到的另一件事是 Harvey 这个名字的由来。我成长过程中最喜欢的剧是《金装律师》。我当然喜欢哈维·斯佩克特。可以说是史上最好的电视剧角色之一。你和温斯顿是怎么想出这个名字的?

The other thing we talked about before this is the history of naming Harvey. My favorite show growing up was Suits. Of course I loved Harvey Specter. Arguably one of the best TV show characters of all time. What was the history of how you and Winston came up with the name?

Gabe

是的,我们当时一起住在洛杉矶,我记得有大约两周时间,我们向 OpenAI 做了推介。他们说我们会资助这个,我们说好的,我们需要注册公司,然后我们想哎呀,我们需要一个名字才能做这件事。我们会去长时间散步,不停地头脑风暴名字。如果你看大多数法律公司,名字里总是有法律相关的东西。我们说,哦,我们不想叫法律 AI 之类的。我们想要一个更通用的名字。我们最终以 Council AI 注册,我们说,好吧,这够好了,我们想不出什么了。然后有一天,我正坐在客厅里,温斯顿从房间出来,他说,Harvey 怎么样?我们俩都说,哦,太完美了。我认为当时真正引起共鸣的是,我们有一种强烈的直觉,人们会开始像与人互动一样与这些模型互动,而不是像软件。如果你有一个像 Harvey 这样的名字,我们看到了这一点,当人们使用产品时,我们叫它 Council AI,他们不会提到产品。一旦我们开始叫它 Harvey,他们就会说 Harvey 给出了一个很棒的答案,我们说哦,那真的很好。

Yeah, so we were living together in LA and I remember we had like a two-week period where we had pitched to OpenAI. They were like we would fund this and we're like okay we need to incorporate the company and we were like oh shoot we need a name to do this and we would like go for these long walks. We just constantly be like spitballing names. Like if you look at most like legal companies, they always have something legal in the name. And we were like, "Oh, we don't want to be like, you know, legal AI or something like this." Like we wanted this more generic name. We ended up incorporating under Council AI and we're just like, "Okay, that's good enough. We can't like think about we can't figure out anything." And then one day like I was just sitting in the living room. Winston came out of his room and he just goes, "What about Harvey?" And we were both just like, "Oh, it's perfect." And it was just like I think the thing that really resonated at the time was we had this strong intuition that people would like start interacting with these models more like people than like software. And if you had a name like Harvey and we saw this where like when people used the product and we called it Council AI they would not reference the product. As soon as we started calling Harvey they were like Harvey had this great answer and we're like oh that's like really good.

Host

我绝对认为这是我听过的最好的创业公司名字之一。我是超级粉丝。非常感谢你加入我们,Gabe。

I definitely think it's one of the best startup names that I've heard. I'm a big fan. Thanks so much for joining us, Gabe.

Gabe

非常感谢你们邀请我。当然。

Thanks so much for having me. Of course.

互动版:逐字朗读 + 针对本期提问 →