预算有限,如何打造研究实验室

Building a Research Lab on a Budget

加布·佩雷拉 Gabe Pereyra · Training Data · 2026-08-11 · 约 29 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Harvey 联合创始人 Gabe 分享应用层公司如何利用前沿生态系统、构建基准数据集并对开源模型进行后训练,从而与前沿实验室竞争。

Gabe from Harvey shares how application-layer companies can compete with frontier labs by leveraging the frontier ecosystem, building benchmarks, and post-training open-source models.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 11)

全文 · Full transcript(中英对照)

引言 Introduction

Host

接下来,我们很荣幸请到 Gabe。Gabe 是 Harvey 的联合创始人兼总裁。你很久以前是 DeepMind 的研究科学家,后来在 Meta,我想是在你向大学室友展示 GPT-3 的能力之前。然后这个二人组就变成了 Harvey。我们非常高兴你能来。我想在座的各位都是应用层公司,在思考如何开始做自己的研究、后训练自己的模型、创建自己的实验室。所以我认为 Harvey 在这里树立了榜样,我们非常高兴你能来谈谈你们是如何建立 Harvey Labs 的。

Next up, we are honored to have Gabe with us. Gabe is co-founder and president of Harvey. You were a research scientist at DeepMind ages ago, and then at Meta I think before you showed your college roommates what GPT-3 could do. And this duo then became Harvey. We're really excited to have you here. I think everybody here in the audience is an application company thinking about how to start doing their own research, post-training their own models, and create their own labs. And so I think Harvey has really set the example here, and we're really delighted to have you give a talk on how you guys built Harvey Labs.

Gabe

太好了。

Awesome.

预算内建研究实验室 Building a Research Lab on a Budget

Gabe

所以这个演讲的另一个标题是「在预算有限的情况下建立研究实验室」。如果你是应用层公司,与前沿实验室竞争是一场不公平的游戏。有富有的团队,有贫穷的团队,然后还有我们这些应用层的。前沿实验室有更多的钱、更多的人才、更多的算力基础设施和数据。那么你怎么竞争?通过利用前沿生态系统。当我们 4 年前创办 Harvey 时,这些公司大多数要么不存在,要么刚刚起步。所以我们要么自己构建一切,要么在大多数情况下专注于不同的事情,比如建立我们的 GTM 组织和出色的产品。但今天,利用前沿生态系统,我认为你可以与前沿实验室竞争,并构建前沿智能。这个演讲将是我们的高层策略手册。我将谈谈我们如何构建基准和训练数据,如何与 Neo Labs 合作进行后训练,以及如何在生产中服务这些模型。

So the alternative title to this talk is building a research lab on a budget. And it's an unfair game competing with the frontier labs if you're an application layer company. There are rich teams, there are poor teams, and then there's us in the application layer. The frontier labs have more money, more talent, more compute infrastructure, and data. So how do you compete? By using the frontier ecosystem. When we started Harvey 4 years ago, most of these companies either didn't exist or were just getting started. And so we either had to build everything ourselves or in most cases focus on something different. Like building our GTM org and a great product. But today using the frontier ecosystem I think you can compete with the frontier labs and build frontier intelligence. This talk is going to be our high-level playbook for doing this. And I'm going to talk about how we build benchmarks and training data, how we work with the Neo Labs to do post training and how we serve these models in production.

构建基准与训练数据 Building Benchmarks and Training Data

Gabe

首先,你要构建一个基准。如果没有好的基准,你就无法训练模型;如果无法训练模型,你就不需要在生产中服务它们。所以,今年我们发布了三个我们构建的数据集。我们首先构建了 Legal Agent Bench,这是一个大型律师事务所中律师助理会执行的任务分类。这些任务涵盖多个业务领域,而且是复杂的任务,比如起草复杂的基金设立文件、进行判例法研究等等。接着我们构建了一个合同数据集。这让我们能够训练智能体像法务部门那样进行谈判。然后我最兴奋的是我们最近发布的大型尽职调查数据集。我认为这是已发布的最大的强化学习环境之一。这里最大的数据室有 8000 万个 token,这让我们能够研究长上下文、非常复杂的任务。我认为这些数据集最有趣的地方在于我们是如何构建它们的。

To start, you want to build a benchmark. If you don't have a good benchmark, you can't train models, and if you can't train models, you don't need to serve them in production. So, this year we released three of these data sets that we built. We started by building Legal Agent Bench, which is a taxonomy of tasks that associates would do at a large law firm. And these cover multiple practice areas, and they're complex tasks like doing drafting complex fund formation documents, doing case law research, things like that. We followed this up by building a contracting data set. This allows us to teach agents to do negotiation like you would in an in-house department. And then the one I'm most excited about that we recently released is a large diligence data set. This is, I think, one of the largest RL environments that's been released. The largest data rooms here are 80 million tokens, and it lets us do research on long context, very complex tasks. And I think the most interesting thing about these data sets is how we built them.

Gabe

我们在 Harvey 训练模型时一直面临的一个挑战是,我们不能在客户的数据上训练。我们与最大的律师事务所和企业合作,他们的法律数据极其敏感,是特权的,所以你不能把它放入通用模型,甚至不能放入我们自己的模型。那么,在这种情况下,你怎么训练模型?今年开始非常有效的方法是使用领域专家来指导合成数据生成。Reka 的 Brendan 有一个很好的类比:就像工程师现在不写代码,而是「vibe code」并指导这些编码模型一样,我们开始做同样的事情。所以,我弟弟实际上是 Harvey 的律师,他非常擅长使用编码模型。他训练了我们其他律师这样做,他们可以生成极其逼真的数据集,我们可以用于训练和评估我们的产品。

One challenge we always had at Harvey for training models is we can't train on our customers' data. We work with the largest law firms and enterprises, and their legal data is incredibly sensitive. It's privileged, and so you can't put it in generic models. We can't even put it in our models. And so, how do you train models given that? And the thing that started working really well this year is using domain experts to guide synthetic data generation. And Brendan from Reka had a good analogy that the same way that engineers now don't write code, they vibe code and guide these coding models. We're starting to do the same thing. And so, my younger brother is actually a lawyer at Harvey, and he's gotten very good at using the coding models. He's trained our other lawyers to do that, and they can generate incredibly realistic data sets that we can use for training and also evaluating our product.

Gabe

一旦你做到了这一点,合成数据还不够好,但这是一个开始的方式。所以,我们与 Mercor 和 Snorkel 这样的公司合作,他们让你扩展这个过程,构建更大的数据集,特别是用于训练。一旦你做到了,你需要将这些数据集转化为高效的强化学习环境。随着数据集变大,评估变得非常昂贵。例如,在我们的尽职调查数据集中,我们有超过 1000 个单元测试,使用 LLM 作为裁判来评估模型输出。如果你使用最大的模型,并且想要进行强化学习 rollout 等,这会变得非常昂贵。所以有很多工作要做,这里有一些我们与 LangChain 合作做的工作,使这些变得非常高效。

And once you've done that, synthetic data isn't good enough, but it's a way to get started. And so, we work with companies like Mercor and Snorkel, who let you scale up this process and build larger sets particularly for training. Once you've done that, you need to turn these data sets into efficient RL environments. It gets really expensive as these data sets get larger and evaluation gets very expensive. For example, in our diligence data set, we have over 1,000 unit tests that are grading model outputs using LLM as a judge. If you use the largest models and you want to do RL rollouts and things like this, it gets very expensive. And so, there's a lot of work, and here's some we did with LangChain, of making these very efficient.

Gabe

然后我们做的最后一件事,我认为当时有点争议,就是开源其中一些数据集。这样做的动机是,除非有很多人在上面训练,否则很难知道你的数据集是否好。当我在 Google Brain 和 DeepMind 做研究时,最好的数据集是开放的,比如 ImageNet、CIFAR、MNIST,每个人都使用它们,你能够发现所有的问题,我们收到大量的 pull request,我们得到建议。然后我认为越来越多的实验室在报告新模型时会在我们的数据集上进行基准测试。然后最重要的是,我们让 Elon 转发了它。

And then the last thing we did that I think was a little controversial at the time, was open-sourcing some of these data sets. And the motivation for this was it's very hard to know your data set is good unless a lot of people train on it. When I used to do research at Google Brain and DeepMind, the best data sets were open, like ImageNet, CIFAR, MNIST, and everyone used them, and you were able to find all of the issues, and we get a ton of pull requests, we get suggestions. And then I think increasingly we're having the labs when they report new models benchmark on our data set. And then most importantly, we had Elon retweet it.

与Neo Labs合作进行后训练 Working with Neo Labs for Post-Training

Gabe

所以,一旦你构建了基准,现在你有了可以训练模型的东西。现在令人兴奋的是开源模型变得越来越有竞争力。过去,做后训练不值得,因为模型从预训练中改进得太快,你做的任何后训练很快就被下一个预训练模型吸收了。但现在有了像 Kimmy 3、GLM 5.2、NeMo-Megatron、Inkling 等模型,可以将这些非常强大的开源基础模型进行后训练,达到前沿智能的水平。也许不是通用的前沿智能,但如果你像我们一样有特定的任务,它们是有竞争力的。

So, once you've built the benchmark, now you have something to train models against. And the thing that is exciting now is open-source models are getting competitive. In the past, it wasn't worth doing post training because the models were improving so quickly from pre-training that any post training you did quickly got absorbed by the next pre-trained model. But now with models like Kimmy 3, GLM 5.2, NeMo-Megatron, Inkling, and others, it's possible to take these very strong open-source base models and post train them to levels of frontier intelligence. Maybe not general frontier intelligence, but if you have a specific task like us, they are competitive.

Gabe

所以我们推荐的入门方式是与 NeMo labs 合作。他们拥有大量的专业知识和基础设施,可以帮助你确保训练数据集是好的。他们有配方。通常,如果你与他们合作但无法获得更好的结果,那可能是你的数据集有问题,这是一个很好的启动方式。

And so, the way we recommend getting started is working with the NeMo labs. They have a bunch of expertise and infrastructure already in place to help you make sure that your training data sets are good. They have recipes. And usually, if you work with them and you're not able to get better results, there's probably something you're doing wrong with your data set, and this is a very good way to bootstrap it.

Gabe

所以,我们与这些不同的提供商做了一些有趣的工作。Fireworks,我们训练 GLM 5.1 使用 Fable 或可能是 Opus 4.8 作为顾问模型,得到了一些非常有趣的结果。Baseten,我们在 KB 压缩方面做了一些有趣的工作。N gram,我想他在这里,我们在企业搜索和公司知识方面做有趣的工作。Trajectory,我们与他们合作训练 NeMo-Megatron 模型。Applied Compute,我们在 Vault 产品上做了一些有趣的工作。

And so, some of the interesting work we did with these different providers, Fireworks, we got some very interesting results training GLM 5.1 to use Fable or maybe Opus 4.8 as an advisor model. Baseten, we did some interesting work on KB compaction. N gram, who I think is here, we're doing interesting work on enterprise search and firm knowledge. Trajectory, we worked with them to train NeMo-Megatron models. And Applied Compute, we're doing some interesting work on our Vault product.

Gabe

我想我们收到的一个问题是为什么要与多个 NeMo labs 合作。为什么不只选一个?对我们来说,随着我们扩展研究实验室,我们的研究项目比我们内部或仅与一个 Neo lab 合作的带宽要多。每个 Neo lab 都在下不同的赌注。他们对研究有不同的思考方式。

And I think one question we got is why work with multiple NeMo labs. Why not just pick one? And for us, as we're scaling the research lab, we have more research projects than we have bandwidth to do internally or just with a single Neo lab. And every Neo lab is taking a different bet. They have different ways they think about research.

后训练与服务基础设施 Post-training and serving infrastructure

Gabe

我们有不同的开源模型想要训练。我们合作得越多,学到的就越多。而且现在做后训练比以往任何时候都容易。所以,一方面,与 Neo labs 合作,我们在合作中学到了很多。然后我们越来越多地在内部自己做后训练,使用像 Tinker 这样的 API,以及 Fireworks 和 Baseten 构建的基础设施。现在后训练这些模型然后提供服务,从来没有这么容易过。而且我们正在招聘的后训练人才越来越多,市场上也有更多可用的人才。受 Cursor 的启发,这些努力的目标是为我们构建自己的 Composer one 版本。我们如何把我们在合成数据方面所做的工作、用 Mercor 进行扩展、以及与 Neo labs 的合作,打包成一个可以与闭源模型一起提供的模型。

We have different open source models we want to train. And the more we work with, the more we learn. And it's getting easier than ever to do post-training. So one, working with the Neo labs, we're learning a lot in partnership with them. And then we're doing more and more post-training ourselves internally with APIs like Tinker and the infrastructure Fireworks and Baseten have built. It's never been easier to post-train these models and then serve them. And there's increasingly more post-training talent that we're hiring and is available. And inspired by Cursor, the goal of these efforts is for us to build our version of Composer one. How do we package all of the work we've done with synthetic data, scaling it with Mercor, the work with the Neo labs, into a model we can serve alongside the closed source models.

Gabe

现在,一旦你后训练了一个模型,你需要在生产环境中提供服务。这可不是小事。所以我想先谈谈我们的模型服务基础设施,因为我觉得有时候人们仍然把应用层公司想象成你调用一个单一的模型端点,也许就是一个聊天产品。所以过去四年我们解决的一个大问题是,我们在 60 个国家运营。我们有多个产品表面区域。客户有不同的模型偏好。即使只考虑闭源模型,你如何大规模地提供这些服务?所以这个矩阵让你了解我们在考虑大规模服务这些模型时需要处理的所有事情。

Now, once you've post-trained a model, you need to be able to serve it in production. And this is non-trivial. So I want to start first by talking about our model serving infrastructure because I think sometimes people still think about application layer companies as you're calling a single model endpoint and it's maybe a chat product. And so a big problem we've solved over the past four years is we operate in 60 countries. We have multiple product surface areas. Customers have different model preferences. And even with the closed just the closed source models, how do you serve these at scale? And so this matrix gives you a sense of all of the things we need to handle when we're thinking about serving these models at scale.

Gabe

举个简单的例子,对于每个模型家族,我们需要服务多个模型。我们需要跨提供商设置回退,以达到我们的服务级别协议(SLA)。现在有了像 Fireworks.ai 这样的公司,我们可以把开源模型加入这个组合。为了考虑在生产中何时提供服务以及如何管理它们,你需要在考虑后训练之前就建立好这个基础设施。所以这就是我们如何思考,当有一个新模型发布时,无论是开源的、闭源的,还是我们后训练的模型,我们如何决定是否将其投入生产。然后一旦投入生产,我们如何决定是否继续保留它。

And as a simple example for each of these model families, we need to serve multiple of these models. We need to have fallbacks across providers to hit our SLAs. And now with companies like Fireworks.ai, we can add open-source models into this mix. In order to think about when you serve models in production and how you modern them in production, you need to have this infrastructure in place thinking of even before you think about post-training. And so this is how we think about when there's a new model released, whether it's open-source, closed-source, or a model we've post-trained, how we make decisions whether to put it in production. And then once it's in production, how we make decisions whether to keep it in production.

Gabe

在生产前,我们有一组通用的评估。在自动化方面,我们有我提到的实验室基准,每当新模型出现时,它能让我们快速了解它是不是前沿模型,有多强,在法律领域哪些方面表现好。在通用情况下,我们进行人工测试,将这个模型与其他模型并排运行以进行比较。然后对于每个产品表面,我们有关键用户旅程和自动化产品测试,因为一个模型可能在通用情况下很好,但可能不适合特定的产品表面。然后我们还有人工产品测试。这些信号加上关于成本、延迟、区域可用性的启发式规则,共同决定我们是否将模型投入生产。

And so pre-production, we have a set of generic evals. So in terms of automated, we have the lab benchmark that I talked about where whenever a new model comes out, this gives us a very quick sense of is it a frontier model, how strong is it, what areas of legal is it good. We have human testing in the generic case where we run side-by-sides of this model with other models to compare them. And then for every product surface, we have critical user journeys and automated product tests because a model could be very good generically, but it could not be a good fit for a specific product surface. And then we also have human product testing. And together these signals along with heuristics around cost, latency, region availability is how we decide whether we put a model into production.

Gabe

我认为这里的重要一点是,无论模型是否经过后训练,情况都是如此。所以你可以重用这个基础设施,而且你应该在考虑后训练之前就把它准备好。一旦模型投入生产,也是一样,不管模型是否经过后训练。如果我们推出一个大的变更,我们会进行 AB 测试,我们观察参与度来跟踪这个模型是否表现如预期。然后我们看正常运行时间、token 效率等,然后我们称之为产品反馈,我称之为愤怒的客户邮件。所以有所有这些信号告诉你这是否按预期工作。所以你需要先建立好这个基础设施,然后再考虑服务模型。

And I think the important point here is this is the case for post-trained models or non-post-trained models. And so you can reuse this infrastructure and you should have it in place before you think about post-training. Once a model's in production, same thing, doesn't matter if the model's post-trained or not. We do AB testing if we're rolling out a large change, we look at engagement to track if this model's performing how we expect. Um and then we look at things like uptime, token efficiency, and then we call it product feedback, I call it angry customer emails. And so there's all of these signals that tell you this is working as expected. And so you need that infrastructure in place before you think about serving models.

Gabe

然后,在服务模型之前,你还需要做一件事,我称之为简单的开源切换。第一个是查看你服务模型的所有地方,找出产品中哪些地方可以天真地换成开源模型?例如,我们产品中有生成引用的部分,不需要最大的模型,有机会换成 GLM 5.2 并获得成本或性能优势。所以这是第一件事,这样做可以锻炼与闭源模型一起服务开源模型的能力。

And then the thing you need to do even still before serving models is what I call the simple open-source switches. And so the first one is look at all the places you're serving models and find are there places in my product where I can just naively swap open-source models? And so for example, we have parts of our product that generate citations that don't need the largest models and there's opportunities to swap in GLM 5.2 and get cost or performance benefits. And so that's the first thing and doing this builds the muscle of serving open-source models alongside closed-source models.

Gabe

然后第二个现在已经非常流行的是模型路由。所以有些地方你不能天真地换成开源模型,但你可以做的是在某些查询上路由到开源模型。一旦你做到了这一点,你就准备好了一切,可以开始构建后训练的飞轮。一旦你开始在生产中服务它们并收集反馈,这里要提醒的是,你需要非常小心收集反馈意味着什么。在我们的情况下,这并不意味着在客户数据上训练,但我们确实从用户测试和其他事情中获得反馈信号,这些信号可以指导我们如何构建未来的数据集并改进这些模型。

And then the second which has gotten very popular now is model routing. So there could be places where you can't naively swap an open-source model, but what you can do is on certain queries route to open-source models. And then once you've done this, you have everything in place to start building the post-training flywheel. And once you start serving them in production and collecting feedback, and caveat here, you need to be very careful about what collecting feedback means. In our case, it does not mean training on customer data, but we do get feedback signals from our user testing and other things that can inform how we build future data sets and improve these models.

Gabe

这就是我们建立研究实验室的剧本。我们认为未来每家人工智能公司都需要成为一家人工智能公司,并找出这个剧本的某个版本。我认为尽管如此,大多数人仍然在押注反对这个剧本、应用层公司和前沿生态系统。

And that is our playbook for building a research lab. We think in the future every AI company will need to become an AI company and figure out some version of this playbook. And I think despite that, most people are still betting against this playbook and application layer companies and the frontier ecosystem.

Host

但如果我们用这个预算和这个团队赢了,我们就会改变游戏规则。

But if we win on our budget with this team we'll change the game.

Gabe

这是《点球成金》里的一个场景,希望你们看过这部电影,他们在谈论我们刚刚连续赢了 20 场比赛。比利·比恩说:“没关系。如果我们不赢得冠军,没有人会欣赏我们在这里所做的一切。”而那句话是:“但如果我们用这个预算和这个团队赢了,我们就改变了游戏规则。”我认为现在有了前沿生态系统,你们所有人都有机会做同样的事情。去改变游戏吧。谢谢。

This is a scene from Moneyball, which hopefully you've seen this movie, where they're talking about we just won 20 games in a row. And Billy Beane says, "It doesn't matter. If we don't win the championship, no one's going to appreciate what we've done here." And the quote is, "But if we win on this budget with this team we'll have changed the game." And I think now with the frontier ecosystem all of you have the opportunity to do the same. Go change the game. Thank you.

Host

你想做问答环节吗?

Do you want to do Q&A?

Gabe

你介意吗?

You okay with that?

Host

当然可以。

Yeah, of course.

Gabe

好的。

Yeah.

Host

抱歉,谁在说话?嗯,你谈到了拥有数据非常敏感的法律客户的挑战,还谈到了你兄弟的律师和其他内部专家做他们版本的实时编码以产生最佳工作。那是什么样子?你能给我们更多细节,那实际上是什么样子吗?

I'm sorry, who's speaking? Um you talked about the challenges with having legal customers with very sensitive data sets and talked a little bit about your brother's lawyer and other in-house experts doing their version of live coding to produce the best work. What does that look Can you give us a little more detail about what that actually looks like?

Gabe

是的,好问题。法律工作最大的挑战是,当你在大型律师事务所时,你做的工作类型是,例如,你代表一家公司进行并购。你会得到一个数据室,里面有你要收购的公司的所有合同,还有所有这些关于谈判的电子邮件和会议。这些数据都不是公开的,所以你没有像开源 GitHub 仓库这样的类比。

Yeah, great question. So the biggest challenge with legal work is when you're at a large law firm, the type of work you're doing is, for example, you're representing a company doing an M&A. And you get a data room, which is all the contracts of the company you're trying to acquire, plus there's all these emails and meetings about the negotiation. None of that data is public, and so you don't have this analogy of like open source GitHub repos.

法律AI训练的数据生成 Data Generation for Legal AI Training

Gabe

训练时遇到的最大挑战是,你或许能在公开渠道找到一些最终工作成果,比如公开的采购协议,但你没有任何输入数据。过去,找一群律师说“嘿,做一个假的数据室”非常尴尬,因为这些数据室可能有 1 万份合同,而且它们必须相互匹配。所以 Julio 想出了一个非常巧妙的方法来生成这些数据集。他从评分标准出发,埋下了所有这些问题,说:“这是数据室里所有的问题,以及我在场景中要检查的内容。”然后用它来生成数据室,这样你就可以埋下所有这些问题,比如这些合同没有关联、这份合同缺失,并生成所有数据。然后我们用 Mercor Shenorcal 让所有合同看起来逼真,但现在你有了这个输入数据集,你可以让模型生成一份尽职调查备忘录。这种检查方式会说:“嘿,你发现了所有我们埋下的问题吗?”显然,根据你从事的工作,你需要巧妙地创建它,但对我来说,这感觉像是真正的重大突破,因为现在我们可以开始进行这种训练了。以前,我们面临一个先有鸡还是先有蛋的问题:我们去律师事务所说:“嘿,我们可以用你的客户数据训练这个模型。”他们会说:“好,证明给我们看。”我们说:“给我们看数据。”他们说:“不行。”所以现在你可以在这里证明它,而且现在很多律师事务所都表示很感兴趣:“哦,这真的很有趣。我们能用我们的数据来做吗?”

And the biggest challenge you run into when training is you can maybe find some of the final work product publicly, like a public purchase agreement, but you don't have any of the input data. Historically, it was very awkward to get a bunch of lawyers and say, 'Hey, make a fake data room.' Because these data rooms can be 10,000 contracts, and they all need to fit together. So what Julio figured out is a very clever way to generate these datasets. He started from the rubric, planted all these issues, and said, 'Here's all the problems in the data room and what I'm going to check for in a scenario.' Then he used that to generate the data room, so you can plant all these issues like these contracts don't tie together, this contract's missing, and generate all the data. Then we'll use Mercor Shenorcal to make all the contracts look realistic, but now you have this input dataset, and you can have the model generate a diligence memo. This way of checking it says, 'Hey, did you catch all these issues that we know are in here because we planted them?' Obviously, depending on the work you're doing, you need to be clever about how you create it, but that to me feels like the really big unlock because now we can start doing this training. Before, you had this chicken-and-egg problem where we'd go to law firms and say, 'Hey, we can train this model on your client data.' And they'd be like, 'Okay, prove it.' And we were like, 'Show us the data.' And they're like, 'No.' So now you can prove it here, and now we have a lot of interest from law firms saying, 'Oh, this is really interesting. Can we do it with our data?'

招聘策略 Hiring Strategy

Host

好问题。是的。嗯,我们也在为此招聘。我们也在那里构建一个隐含的 AI。从招聘的角度来看,你很难与来自 Anthropic 和其他实验室的研究人员竞争。你是尝试玩那个游戏,还是尝试招聘领域逻辑、律师和经验?哪些地方有效,哪些地方无效?

Good question. Yeah. Um, so we're also hiring for this. We're also building an implied AI there. From a hiring point of view, it's difficult for you to compete with researchers from Anthropic and the labs. Do you try and play that game, or do you try and hire domain logic and lawyers and experience? And where has that worked and where has that not worked?

Gabe

是的,我认为这是我们刚创办公司时犯下的大错误之一,因为我在这些实验室工作过,认识很多人。我会说:“你知道,来和我们一起工作吧。”然后他们拿到了 1 亿多美元的薪酬包,而我们的规模还不足以让这有意义。所以我想说,我们现在的公司规模刚好,我们不是在和前沿人才竞争,但越来越多的人在攻读博士学位,有些人不想在大实验室工作,所以我认为情况正在改变。然后第二点是,早期你需要这些人才的原因不仅仅是做后训练,还需要构建训练基础设施、服务基础设施,所有这些加在一起。我认为能做到这一点的人才非常独特。比如我的老室友是在 OpenAI 负责后训练的人之一,他是我合作过的最好的研究人员之一。但现在你可以使用像 Fireworks、tinker 之类的 API,所以你不需要自己构建训练和服务基础设施。所以我认为这拓宽了人才库,现在获得一些这样的人才感觉非常可行。

Yeah, I think this was one of the big mistakes I made when we first started the company, because I had worked in these labs and so I knew a lot of the folks. I would say, 'You know, come work with us,' and then they were getting these 100 million plus pay packages, and we weren't at the scale where this made sense. So I'd say we're now just getting to the company size where we're not competing with the frontier talent, but there are increasingly folks that are doing PhDs, folks that don't want to work at large labs, and so I think that is changing. And then I think the second thing is a lot of why you needed that talent early on was it wasn't just doing the post-training; it was you need to build the training infrastructure, the serving infrastructure, and all of those things combined. I think the talent to do that was very unique. Like my old roommate was one of the folks who ran post-training at OpenAI, and he was one of the best researchers I've ever worked with. But now you can use things like Fireworks, tinker-like APIs, and so you don't need to build the training and serving infrastructure. So I think it opens up the pool of talent, and it feels very feasible to get some of this talent now.

创建基准评估标准 Creating Rubrics for Benchmarking

Host

是的。嗯,是的,关于基准创建还有一个问题。当你创建这些评分标准时,你仍然必须以某种方式创建它们,使它们能够区分前沿模型,而且你必须与专家一起做这件事。你如何应对这个挑战?所以专家们必须创建这个评分标准,同时知道或试图弄清楚模型在它们上的表现。

Yeah. Uh, yeah, just another question on the benchmark creation. So when you're creating these rubrics, you still have to create them in such a way that they're separating the frontier models, and you have to do that with the experts themselves. How do you navigate that challenge? So the experts kind of have to create this rubric knowing or trying to figure out how the models are going to perform on them.

Gabe

那么,你说的区分是什么意思?

So, what do you mean by separating the...

Host

嗯,我想是的,你必须创建能够充分区分前沿模型的评分标准。它们必须挑战前沿,对吧?

Uh, I guess so, you have to create the rubrics that can adequately separate the frontier models. They have to challenge the frontier, right?

Gabe

是的。

Yep.

Host

嗯,我想你如何应对这一点,尤其是当你有这些内部法律专家,他们可能不完全知道如何为模型创建这样的评分标准?

Um, I guess how do you navigate that, especially when you're having these in-house legal experts who might not exactly know how to create that rubric designed for a model?

Gabe

是的,这是个好问题。所以我们更多考虑的是如何让它们真实地代表我们客户所做的工作?我认为我们选择法律领域的部分原因是,如果你只是构建一个顶级律所正在处理的真实客户案件,前沿模型仍然无法做到这一点。我认为在过去 4 年里我们建立的是……例如,我哥哥是我们的首批员工之一。所以他和模型一起工作了 4 年。他对模型如何工作、如何生成这些数据集的直觉非常好。他训练了一批律师来做这件事。但我想说,重点仍然更多是如何构建一个真正逼真的数据室,然后为尽职调查制定评分标准。然后当我们运行这些模型时,我们发现,好吧,有差距,有改进的空间。但我认为这取决于领域。

Yeah, that's a good question. So we think of it more as how do we make them realistic representations of the work our customers are doing? And I think part of why we picked the legal domain is if you just build a realistic client matter that these top firms are working on, the frontier model still can't do this. And I think the thing over the past 4 years we've built is... For example, my brother was one of our first hires. So he's been working with the models for 4 years. His intuition of how the models work, how to generate these datasets, is incredibly good. And he's trained a bunch of lawyers to do this. But I would say the focus is still more on how do we build a really realistic data room, and then rubrics for the diligence. And then when we run these models, we find, okay, there are gaps and there's room for improvement. But I think it depends on the domain.

调试流水线 Debugging the Pipeline

Host

嘿,呃,来自 Box 的 Shadeen。呃,抱歉,让我先处理那个,然后我会……太好了。我的意思是,如果你想想……顺便说一句,演讲很棒。嗯,来自 Trolysis 的 Ross。如果你想想端到端的流程,就像数据生成、环境部分,然后是算法部分,比如我们如何训练等等,研究部分,然后是基础设施部分。如果有什么不奏效,根据经验,你看到了什么?是在运营上、金钱上还是人力时间上?通常是在数据层?是在研究层?还是在基础设施层?你如何从那里着手?

Hey, uh, Shadeen from Box. Uh, sorry, let me do that one and I'll... Great. I mean, if you think... Great talk, by the way. Um, Ross from Trolysis. If you think of the end-to-end motion, it's like data generation, the environment piece, then there is the algorithm piece of like what how do we train, etc., the research piece, and then there's the infra piece. If something is not working, anecdotally, what have you seen? Is it operationally or dollar-wise or human-hours-wise, is it usually in the data layer? Is it in the research layer? Or is it in the infra layer? And how do you go from there?

Gabe

在某种程度上,谁在编排或调试整个端到端流程?

Who's orchestrating or debugging this whole end-to-end pipeline in some ways?

Host

是的,这是一个很好的问题。我的意思是,这完全是另一个话题。嗯,我不认为有单一的因素。所以我想说,让我们更容易的是我们找到了产品市场契合度。我们有一个在生产中使用的产品。我们一开始主要使用闭源模型。这让我们建立了很多信心,事情是端到端工作的。所以我们有一个产品,人们在使用它,人们已经使用它很多年了。模型以这种方式工作,我们可以监控它们在生产中是否正常工作。所以这就是服务的关键点,你需要很多这样的东西到位。然后我们已经建立了肌肉记忆,新模型出现时,如何将它们放入产品中,无论是闭源还是开源模型。所以你可以考虑整个周期。然后显然我们做了很多工作,比如工具工程、上下文管理等等。所以现在你可以把后训练看作是更广泛系统中的一个小的输入。

Yeah, this is a great question. I mean, this is a whole 'nother talk. Um, I don't think there's a single thing. So I would say what makes it easier for us is we found product-market fit. We have a product that's being used in production. And we did this largely with the closed-source models to start. And so that let us build a lot of this confidence in the thing is working end-to-end. So we have a product, people are using it, people have been using it for many years. The models work in this way, we can monitor whether they're working in production. So that was kind of the point of serving, you need a lot of this in place. And then we have built the muscle of new models come out, how do we put those into the product, whether they're closed-source or open-source models. So you can kind of think about the full cycle. And then obviously we've done a bunch of work like harness engineering and context management and all of these things. And so now you can think of post-training as one small input into this broader system.

后训练与门控 Post-training and gating

Gabe

所以,我们让团队对模型进行后训练,然后把它接入这个更大的系统,你就把它当作另一个新模型。然后,你就能看到各种出问题的信号。我想说的是,在这些步骤中,有一些简单的门控方式。如果你从基准测试开始,后训练一个模型,但它在这些测试上表现不佳,那它就不太可能走得更远。如果它表现良好,然后你把它放进产品里,但有人使用产品时觉得不对劲,因为,比如,在我们构建的基准测试中,仍然有一些它们绝对捕捉不到的东西。最好的例子是,你可以有一个在我们的基准测试上表现非常好的模型,但它们在某种程度上过拟合了,然后你把它放进一个更通用的助手类产品中,当遇到分布外的情况时,它就会崩溃。但是,是的,你需要把所有这些阶段和门控都设置好,然后哪里出了问题就会变得很明显。不过,问得好。

And so, we have our team post-train a model, and then you feed it into this broader system, you know, you treat it just as another new model. And so, then you have all these signals of what's going wrong. And so, I would say, but across these steps, there's kind of easy ways to gate it. So, if you start with the benchmarks and you post-train a model and it doesn't work well on those, then it's unlikely to go farther. If it works well on that, and then you put it in a product, but someone uses the product and they say things feel weird because, for example, on the benchmarks we've built, there are still things that they definitely don't catch. Like, the best example is you can have a model that does very well on our benchmarks, but they've overfit to some degree, and then you put it in a more generic assistant-like product, and it kind of falls apart when you go out of distribution. But, yeah, you need to have all of the stages and gates in place, and then it kind of becomes obvious where things are breaking. But, good question.

Host

嗯。

Yeah.

Host

哦,我有个问题。嘿,很高兴认识你。我来自 Rocks。在基准测试方面,legal agent,我想你们之前有 Big Law,Big Law bench,然后现在是 legal agent。你们对开源基准测试的理念是什么?因为如果你们开源,当然会吸引其他人尝试这个基准测试,但实验室也可以帮助提升。然后如果你们闭源,那就有一个问题:这是不是一个真正的基准测试?所以你们怎么看?你们的理念是什么?

Oh, I had a question. Hey, great to meet you. I'm from Rocks. On the benchmark side, legal agent, I think you guys had Big Law before, Big Law bench, and then now legal agent. What's your philosophy on open sourcing the benchmark? Because if you open source, of course you're getting traction with other people trying out the benchmark, but the labs can help climb as well. And then if you close source, then there's a question of is this a real benchmark? So how do you, what's your philosophy?

Gabe

我认为这是我们思考的一个很好的张力点。我想说的是,即使在我们开源这个基准测试之前,我们就已经与实验室紧密合作。所以我们与他们分享数据,帮助他们改进模型,这也改进了我们的产品。所以我想说,我们的思考方式是,在你们的行业,尤其是法律领域,有一个巨大的优势,就是帮助我们的客户理解不同模型在不同事情上的表现有多好。所以我认为当我创办 Harvey 时,我有一种直觉,哦,我们只要构建最好的模型,然后客户就会满意,因为它是最好的。现在很明显,每个客户都有不同的偏好,不同的模型擅长不同的事情。然后我认为我没有预料到的第二件事是前沿生态系统会变得多么庞大。所以现在有很多公司联系我们说:“嘿,我们有这个新技术,或者我们做了这个。我们能试试吗?”以前我们没有精力,我们只是在做其他事情。但现在有了这个开源数据集,我们可以说:“嘿,去试试这个,如果有效,那就太好了。这是一个值得投资的有趣东西。”然后我认为我们的战略思考方式是,对我们来说有价值的数据将是帮助律师事务所训练他们自己的系统,以及他们如何处理私有数据。然后我们想帮助每个人用合成数据和部分开源数据来改进这些系统,但显然,由于你提到的原因,需要平衡。

I think that's like a good tension that we think about. I would say even before we open source this benchmark, we already work closely with the labs. And so we share data with them to help them improve their models, which improved our product. And so I would say the way we think about it is there is a huge advantage in your industry, particularly in legal, in terms of helping our customers understand how good different models are at different things. And so I think when I started Harvey, I kind of had the intuition of, oh, we'll just build the best model and then customers will be happy because it's the best. And it's very clear now that every customer has different preferences, different models are good at different things. And then I think the second thing that I didn't anticipate is kind of how big the frontier ecosystem would become. And so now we have so many companies reaching out saying, "Hey, we have this new technique or we did this thing. Can we try it?" And before we just didn't have the bandwidth, we were just like we're working on these other things. But now with this open source data set, we can say, "Hey, go try this and if it works, then great. This is something that's interesting to invest in." And then I think the way we think about it strategically is the valuable data for us is going to be helping law firms train their own systems and how they work on private data. And then we want to help everyone improve these systems with synthetic data and some of the data we open source, but obviously a balance for the reasons that you mentioned.

Host

你提到了寻找实验室。在这方面,你还想寻求帮助的未解决问题有哪些?

You mentioned finding a lab. What are the remaining open questions for that you want help on?

Gabe

是的,这是个好问题。我想说一个大的挑战,最大的挑战之一,是我们能生成这些非常逼真的合成数据集,我们可以用人类来增强它们,但这些数据的分布仍然与我们的生产分布不匹配。所以我想说,我们生成的这些数据集在某种程度上是面向未来的。我的意思是,我们可以生成一个非常逼真的数据室,但目前我们的产品并不纯粹用于尽职调查。所以如果一个模型在这方面表现良好,但有人用它来起草电子邮件,也许它就不会那么好用。所以我认为仍然存在这个差距,因为我们不能查看客户数据,那么如何弥合这个差距呢?所以我认为,对我来说,这是数据集方面最大的问题,而且这个问题很有挑战性。我认为还有很多问题,比如,再次以尽职调查数据为例,最大的数据室有 8000 万 token 的上下文。模型在管理这种上下文方面做得不好,那么如何训练这些模型在这种非常复杂的环境中运行呢?所以我认为这是另一个问题,仍然存在很大的性能差距。我想说的第三点是,我们在如何自己后训练我们的后训练模型方面取得了良好进展,但我认为找到某种形式的持续学习是最终目标。不是让我们构建最好的法律模型,而是让我们帮助每个律师事务所或企业客户根据他们所做的工作类型进行定制。所以,思考如何将其操作化,比如你有一个大型律师事务所,每次他们处理客户事务时,他们的 AI 系统都会变得更好,但同时你也要保护客户数据。我认为这是一个巨大的、既涉及技术、又涉及运营和 AI 的挑战,非常有趣。

Yeah, that's a great question. I would say one big challenge, one of the biggest challenges, is we can generate these very realistic synthetic data sets, we can augment them with humans, but the distribution of that data still doesn't match our production distribution. And so I would say these data sets we're generating are somewhat future facing. And so what I mean by that is we can generate a really realistic data room, but right now our product isn't used purely to do due diligence. And so if a model does well on that, but then someone uses it to draft an email, maybe it doesn't work as well. And so I think there's still that gap because we can't look at customer data, and so how do you bridge that gap? So I think that's to me the biggest question on the data set side, and that one is challenging. There is, I think, a bunch of questions around, for example, again with the diligence data, like the largest data room is 80 million tokens of context. The models don't do a good job of managing this context, and how do you train these models to operate in these very complex environments? So I think that's another one, there's still a big performance gap. And I would say the third is we're making good progress on how do we post-train our post-train models ourselves, but I think figuring out some form of continual learning is the end game here. It's not for us to build the best legal model. It's for us to help every law firm or enterprise customer customize it to the type of work they're doing. And so just thinking about how do you operationalize that, where you have a large law firm and every time they work on a client matter, their AI system gets better, but you're also protecting the client data. I think that's kind of one huge, both technical, operational, and AI challenge that is super interesting.

Host

是的。我们结束了。还是继续?我可以继续。我可以。嗯,也许最后一个问题。

Yeah. We're done. Or more? I can do more. I can. Yeah, maybe last one.

Host

是的。所以,你知道,你谈了很多关于基础准备,对吧,与 Front Row Live 竞争,对吧,包括你的开源模型,包括你的基础设施,有点像 Ruby。从产品角度来看,对吧,Coda 作为云产品,他们是通用产品,对吧?那么从产品角度来看,你们如何与 Coda 的云工作流、方法论、效率竞争?你们是怎么考虑的?

Yeah. So, you know, you talked a lot about the foundation readiness, right, to compete with Front Row Live, right, including your open source model, including your infrastructure kind of like Ruby. From product perspective, right, Coda as cloud product, they're general product, right? So, from product perspective, how do you guys compete with the Coda as cloud workflow, methodology, efficiency? How do you think about that?

Gabe

是的。我认为我们正在考虑的重大转变是,我们最初的产品和像 Coda 的 Coda Work 这样的东西都是非常注重个人用户的产品,所以它们关乎个人生产力。而我们正在构建的产品越来越关乎组织生产力。所以,如果你考虑一家大型律师事务所,他们试图解决的问题不是如何让我的每个律师更高效。他们试图解决的问题是我有 10,000 个客户。我正在为所有客户处理项目。我需要确保所有这些客户项目都进展顺利,而且我还能以盈利的方式完成它们。所以,我们为律师事务所构建的很多基础设施是如何操作那台机器?当你开始考虑一个单独的客户项目时,大部分挑战不是如何起草这一部分。而是我为这个项目工作了 6 个月。我有一个由 20、30 人组成的跨所团队。我需要协调他们所有人。我需要确保我在协调所有外部各方。

Yeah. I think the big shift that we're thinking about is like our original product and then things like Coda's Coda Work are very individual-focused products, and so they're about individual productivity. And increasingly, the product we're building is about organizational productivity. And so, if you think of a large law firm, the problem they're trying to solve is not how do I make my individual lawyers more productive. The problem they're trying to solve is I have 10,000 clients. I'm working on client projects for all of them. I need to make sure that all those client projects go very well, and then I can also do them in a way that's profitable. And so, a lot of the infrastructure we're building for the law firms is how do you operate that machine? And when you start thinking about an individual client project, most of the challenge is not how do I draft this one section. It's I'm working on this project for 6 months. I have a team of 20, 30 people across the firm. I need to coordinate all of them. I need to make sure that I'm coordinating all the outside parties.

项目管理与企业复杂性 Project Management and Enterprise Complexity

Gabe

所以,很多工作开始看起来像项目管理,协调这些人类和智能体。这是在团队层面。然后,如果你从组织层面来看,现在你有上千个这样的项目,你需要开始考虑资源分配,我在计费什么,我在推销什么?而在企业层面,情况更复杂,比如一家财富 500 强公司,他们与上千家律师事务所合作。他们内部有上千人。所有系统如何启动工作并交换这些工作?所以我想简单的答案是,而且我认为历史上企业就是这么做的,就是如何深入垂直领域,以横向产品不会的方式。是的,好问题。

And so, a lot of it is starting to look like project management that is orchestrating these humans and these agents. And that's at the team level. Then, if you think about it at the organization level, now you have a thousand of these projects, and you need to start thinking about resource allocation, what am I billing, what am I pitching? And then with enterprises, it's even more complicated where a Fortune 500 company is working with a thousand law firms. They have a thousand people internally. What are all the systems to start work and trade that work? So I would say the simple answer is, and this is historically what enterprise has done, is how do you go hyper vertical into your domain in a way that the horizontal products won't. Yeah, good question.

Host

谢谢你,Gabe。那是一场标志性的演讲。感谢你参加我们的节目。

Thank you, Gabe. That was an iconic talk. Thank you for joining us.

互动版:逐字朗读 + 针对本期提问 →