预算有限的研究实验室:Harvey 的实战手册

Building a Research Lab on a Budget: Harvey's Playbook

加布·佩雷拉 Gabe Pereyra · 红杉资本 · 2026-08-11 · 约 29 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Harvey 联合创始人分享应用层公司如何利用前沿生态系统、构建基准和微调开源模型,与前沿实验室竞争。

Harvey's co-founder shares how application-layer companies can compete with frontier labs by leveraging the frontier ecosystem, building benchmarks, and post-training open-source models.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 11)

全文 · Full transcript(中英对照)

引言 Introduction

Host

接下来,我们很荣幸请到 Gabe。Gabe 是 Harvey 的联合创始人兼总裁。你很久以前是 DeepMind 的研究科学家,后来在 Meta,我想是在你向大学室友展示 GPT-3 的能力之前。然后这对搭档就创立了 Harvey。我们非常高兴你能来。我想在座的各位都是应用层公司,正在思考如何开始做自己的研究、后训练自己的模型,以及创建自己的实验室。我认为 Harvey 在这方面树立了榜样,我们非常高兴你能来谈谈你们是如何建立 Harvey Labs 的。

Next up, we are honored to have Gabe with us. Gabe is co-founder and president of Harvey. You were a research scientist at DeepMind ages ago, and then at Meta I think before you showed your college roommates what GPT-3 could do. And this duo then became Harvey. We're really excited to have you here. I think everybody here in the audience is an application company thinking about how to start doing their own research, post-training their own models, and create their own labs. And so I think Harvey has really set the example here, and we're really delighted to have you give a talk on how you guys built Harvey Labs.

Gabe

太棒了。

Awesome.

预算内建研究实验室 Building a Research Lab on a Budget

Gabe

所以这个演讲的另一个标题是“在预算有限的情况下建立研究实验室”。如果你是应用层公司,与前沿实验室竞争是一场不公平的游戏。有富有的团队,有贫穷的团队,然后是我们这些应用层的。前沿实验室有更多的资金、更多的人才、更多的算力基础设施和数据。那么你如何竞争?通过利用前沿生态系统。当我们四年前创立 Harvey 时,这些公司中的大多数要么不存在,要么刚刚起步。所以我们要么自己构建一切,要么在大多数情况下专注于不同的事情,比如建立我们的 GTM 组织和优秀的产品。但今天,利用前沿生态系统,我认为你可以与前沿实验室竞争,并构建前沿智能。这个演讲将是我们做这件事的高层策略。我将谈谈我们如何构建基准和训练数据,如何与 Neo Labs 合作进行后训练,以及如何在生产中服务这些模型。

So the alternative title to this talk is building a research lab on a budget. And it's an unfair game competing with the frontier labs if you're an application layer company. There are rich teams, there are poor teams, and then there's us in the application layer. The frontier labs have more money, more talent, more compute infrastructure, and data. So how do you compete? By using the frontier ecosystem. When we started Harvey 4 years ago, most of these companies either didn't exist or were just getting started. And so we either had to build everything ourselves or in most cases focus on something different. Like building our GTM org and a great product. But today using the frontier ecosystem I think you can compete with the frontier labs and build frontier intelligence. This talk is going to be our high-level playbook for doing this. And I'm going to talk about how we build benchmarks and training data, how we work with the Neo Labs to do post training and how we serve these models in production.

构建基准测试 Building Benchmarks

Gabe

首先,你要构建一个基准。如果没有好的基准,你就无法训练模型;如果无法训练模型,你就不需要在生产中服务它们。所以,今年我们发布了三个我们构建的数据集。我们首先构建了 Legal Agent Bench,这是一个大型律师事务所中律师助理会执行的任务分类。这些任务涵盖多个业务领域,而且是复杂的任务,比如起草复杂的基金设立文件、进行判例法研究等等。接着我们构建了一个合同数据集。这让我们能够训练智能体像内部法务部门那样进行谈判。然后我最兴奋的是我们最近发布的大型尽职调查数据集。我认为这是已发布的最大的强化学习环境之一。这里最大的数据室有 8000 万 token,让我们能够研究长上下文、非常复杂的任务。我认为这些数据集最有趣的地方在于我们是如何构建它们的。

To start, you want to build a benchmark. If you don't have a good benchmark, you can't train models, and if you can't train models, you don't need to serve them in production. So, this year we released three of these data sets that we built. We started by building Legal Agent Bench, which is a taxonomy of tasks that associates would do at a large law firm. And these cover multiple practice areas, and they're complex tasks like doing drafting complex fund formation documents, doing case law research, things like that. We followed this up by building a contracting data set. This allows us to teach agents to do negotiation like you would in an in-house department. And then the one I'm most excited about that we recently released is a large diligence data set. This is, I think, one of the largest RL environments that's been released. The largest data rooms here are 80 million tokens, and it lets us do research on long context, very complex tasks. And I think the most interesting thing about these data sets is how we built them.

Gabe

我们在 Harvey 训练模型时一直面临的一个挑战是,我们不能在客户的数据上训练。我们与最大的律师事务所和企业合作,他们的法律数据极其敏感,是特权的,所以你不能把它放入通用模型。我们甚至不能把它放入我们自己的模型。那么,在这种情况下你如何训练模型?今年开始非常有效的方法是使用领域专家来指导合成数据生成。Reka 的 Brendan 有一个很好的类比,就像现在工程师不写代码,而是“vibe code”并指导这些编码模型。我们开始做同样的事情。所以,我的弟弟实际上是 Harvey 的律师,他非常擅长使用编码模型。他训练了我们其他的律师这样做,他们可以生成极其逼真的数据集,我们可以用于训练和评估我们的产品。

One challenge we always had at Harvey for training models is we can't train on our customers' data. We work with the largest law firms and enterprises, and their legal data is incredibly sensitive. It's privileged, and so you can't put it in generic models. We can't even put it in our models. And so, how do you train models given that? And the thing that started working really well this year is using domain experts to guide synthetic data generation. And Brendan from Reka had a good analogy that the same way that engineers now don't write code, they vibe code and guide these coding models. We're starting to do the same thing. And so, my younger brother is actually a lawyer at Harvey, and he's gotten very good at using the coding models. He's trained our other lawyers to do that, and they can generate incredibly realistic data sets that we can use for training and also evaluating our product.

Gabe

一旦你做到了这一点,合成数据还不够好,但这是一个开始的方式。所以,我们与 Mercor 和 Snorkel 这样的公司合作,他们让你扩展这个过程,并构建更大的数据集,特别是用于训练。一旦你做到了,你需要将这些数据集转化为高效的强化学习环境。随着数据集变大,评估变得非常昂贵。例如,在我们的尽职调查数据集中,我们有超过 1000 个单元测试,使用 LLM 作为裁判来评分模型输出。如果你使用最大的模型,并且你想进行强化学习 rollout 等操作,这会变得非常昂贵。所以,有很多工作要做,这里有一些我们与 LangChain 合作做的工作,使这些变得非常高效。

And once you've done that, synthetic data isn't good enough, but it's a way to get started. And so, we work with companies like Mercor and Snorkel, who let you scale up this process and build larger sets particularly for training. Once you've done that, you need to turn these data sets into efficient RL environments. It gets really expensive as these data sets get larger and evaluation gets very expensive. For example, in our diligence data set, we have over 1,000 unit tests that are grading model outputs using LLM as a judge. If you use the largest models and you want to do RL rollouts and things like this, it gets very expensive. And so, there's a lot of work, and here's some we did with LangChain, of making these very efficient.

Gabe

然后我们做的最后一件事,我认为当时有点争议,就是开源其中一些数据集。这样做的动机是,除非有很多人在上面训练,否则很难知道你的数据集是否好。当我以前在 Google Brain 和 DeepMind 做研究时,最好的数据集是开放的,比如 ImageNet、CIFAR、MNIST,每个人都使用它们,你能够发现所有的问题,我们收到了大量的 pull request 和建议。然后我认为越来越多的实验室在报告新模型时会在我们的数据集上进行基准测试。最重要的是,我们让 Elon 转发了它。

And then the last thing we did that I think was a little controversial at the time, was open-sourcing some of these data sets. And the motivation for this was it's very hard to know your data set is good unless a lot of people train on it. When I used to do research at Google Brain and DeepMind, the best data sets were open, like ImageNet, CIFAR, MNIST, and everyone used them, and you were able to find all of the issues, and we get a ton of pull requests, we get suggestions. And then I think increasingly we're having the labs when they report new models benchmark on our data set. And then most importantly, we had Elon retweet it.

用开源模型进行后训练 Post-training with Open-Source Models

Gabe

所以,一旦你构建了基准,你就有东西可以训练模型了。现在令人兴奋的是开源模型变得越来越有竞争力。过去,做后训练不值得,因为模型通过预训练改进得太快,你做的任何后训练很快就被下一个预训练模型吸收了。但现在有了像 Kimmy 3、GLM 5.2、NeMo-Megatron、Inkling 等模型,可以拿这些非常强大的开源基础模型,并将它们后训练到前沿智能的水平。也许不是通用的前沿智能,但如果你有像我们这样的特定任务,它们是有竞争力的。

So, once you've built the benchmark, now you have something to train models against. And the thing that is exciting now is open-source models are getting competitive. In the past, it wasn't worth doing post training because the models were improving so quickly from pre-training that any post training you did quickly got absorbed by the next pre-trained model. But now with models like Kimmy 3, GLM 5.2, NeMo-Megatron, Inkling, and others, it's possible to take these very strong open-source base models and post train them to levels of frontier intelligence. Maybe not general frontier intelligence, but if you have a specific task like us, they are competitive.

Gabe

所以,我们推荐的入门方式是与 NeMo labs 合作。他们拥有大量的专业知识和基础设施,可以帮助你确保训练数据集是好的。他们有配方。通常,如果你与他们合作却无法获得更好的结果,那可能是你的数据集有问题,这是一个很好的启动方式。

And so, the way we recommend getting started is working with the NeMo labs. They have a bunch of expertise and infrastructure already in place to help you make sure that your training data sets are good. They have recipes. And usually, if you work with them and you're not able to get better results, there's probably something you're doing wrong with your data set, and this is a very good way to bootstrap it.

Gabe

所以,我们与这些不同的提供商做了一些有趣的工作,Fireworks,我们训练 GLM 5.1 使用 Fable 或可能是 Opus 4.8 作为顾问模型,得到了一些非常有趣的结果。Baseten,我们在知识库压缩方面做了一些有趣的工作。N gram,我想他在这里,我们正在做企业搜索和公司知识方面的有趣工作。Trajectory,我们与他们合作训练 NeMo-Megatron 模型。Applied Compute,我们正在我们的 Vault 产品上做一些有趣的工作。

And so, some of the interesting work we did with these different providers, Fireworks, we got some very interesting results training GLM 5.1 to use Fable or maybe Opus 4.8 as an advisor model. Baseten, we did some interesting work on KB compaction. N gram, who I think is here, we're doing interesting work on enterprise search and firm knowledge. Trajectory, we worked with them to train NeMo-Megatron models. And Applied Compute, we're doing some interesting work on our Vault product.

Gabe

我想我们收到的一个问题是为什么要与多个 NeMo labs 合作,为什么不只选一个?对我们来说,随着我们扩大研究实验室,我们的研究项目比我们内部或仅与一个 Neo lab 合作的带宽要多。每个 Neo lab 都在下不同的赌注,他们对研究有不同的思考方式。

And I think one question we got is why work with multiple NeMo labs. Why not just pick one? And for us, as we're scaling the research lab, we have more research projects than we have bandwidth to do internally or just with a single Neo lab. And every Neo lab is taking a different bet. They have different ways they think about research.

后训练与模型服务 Post-training and Model Serving

Gabe

我们有不同的开源模型想要训练。我们合作得越多,学到的就越多。而且现在做后训练比以往任何时候都容易。所以,一方面,与 Neo 实验室合作,我们通过合作学到了很多。然后,我们自己内部也越来越多地使用 Tinker 这样的 API 以及 Fireworks 和 Baseten 构建的基础设施来做后训练。现在后训练这些模型并部署它们,比以往任何时候都容易。而且我们正在招聘,市场上也有越来越多的后训练人才。受 Cursor 的启发,这些努力的目标是构建我们自己的 Composer 1 版本。我们如何把我们用合成数据做的所有工作、用 Mercor 扩展的工作、以及和 Neo 实验室的合作,打包成一个模型,与闭源模型一起部署。

We have different open source models we want to train. And the more we work with, the more we learn. And it's getting easier than ever to do post-training. So one, working with the Neo labs, we're learning a lot in partnership with them. And then we're doing more and more post-training ourselves internally with APIs like Tinker and the infrastructure Fireworks and Baseten have built. It's never been easier to post-train these models and then serve them. And there's increasingly more post-training talent that we're hiring and is available. And inspired by Cursor, the goal of these efforts is for us to build our version of Composer one. How do we package all of the work we've done with synthetic data, scaling it with Mercor, the work with the Neo labs, into a model we can serve alongside the closed source models.

Gabe

现在,一旦你后训练了一个模型,你需要能够在生产环境中部署它。这可不是小事。所以我想先谈谈我们的模型部署基础设施,因为我觉得有时候人们仍然把应用层公司想象成你调用一个单一的模型端点,也许就是一个聊天产品。所以,过去四年我们解决的一个大问题是,我们在 60 个国家运营。我们有多个产品表面区域。客户有不同的模型偏好。即使只是闭源模型,你如何大规模部署它们?这个矩阵让你了解我们在考虑大规模部署这些模型时需要处理的所有事情。

Now, once you've post-trained a model, you need to be able to serve it in production. And this is non-trivial. So I want to start first by talking about our model serving infrastructure because I think sometimes people still think about application layer companies as you're calling a single model endpoint and it's maybe a chat product. And so a big problem we've solved over the past four years is we operate in 60 countries. We have multiple product surface areas. Customers have different model preferences. And even with the closed just the closed source models, how do you serve these at scale? And so this matrix gives you a sense of all of the things we need to handle when we're thinking about serving these models at scale.

Gabe

举个简单的例子,对于每个模型家族,我们需要部署多个模型。我们需要跨提供商设置回退,以满足我们的服务级别协议。现在有了像 Fireworks.ai 这样的公司,我们可以把开源模型加入这个组合。为了考虑在生产中部署模型以及如何在生产中管理它们,你需要在考虑后训练之前就建立这个基础设施。所以,这就是我们如何考虑当有新模型发布时,无论是开源、闭源还是我们后训练的模型,我们如何决定是否将其投入生产。然后一旦投入生产,我们如何决定是否保留它。

And as a simple example for each of these model families, we need to serve multiple of these models. We need to have fallbacks across providers to hit our SLAs. And now with companies like Fireworks.ai, we can add open-source models into this mix. In order to think about when you serve models in production and how you modern them in production, you need to have this infrastructure in place thinking of even before you think about post-training. And so this is how we think about when there's a new model released, whether it's open-source, closed-source, or a model we've post-trained, how we make decisions whether to put it in production. And then once it's in production, how we make decisions whether to keep it in production.

Gabe

在生产前,我们有一组通用的评估。在自动化方面,我们有我提到的实验室基准测试,每当新模型出现时,它能让我们非常快速地了解它是不是前沿模型,它有多强,它在法律哪些方面表现好。在通用情况下,我们有人工测试,我们会把这个模型和其他模型并排运行来比较。然后对于每个产品表面,我们有关键用户旅程和自动化产品测试,因为一个模型可能在通用方面很好,但可能不适合特定的产品表面。然后我们还有人工产品测试。这些信号加上关于成本、延迟、区域可用性的启发式方法,共同决定我们是否将模型投入生产。

And so pre-production, we have a set of generic evals. So in terms of automated, we have the lab benchmark that I talked about where whenever a new model comes out, this gives us a very quick sense of is it a frontier model, how strong is it, what areas of legal is it good. We have human testing in the generic case where we run side-by-sides of this model with other models to compare them. And then for every product surface, we have critical user journeys and automated product tests because a model could be very good generically, but it could not be a good fit for a specific product surface. And then we also have human product testing. And together these signals along with heuristics around cost, latency, region availability is how we decide whether we put a model into production.

Gabe

我认为这里重要的一点是,无论模型是否经过后训练,情况都是如此。所以你可以复用这个基础设施,而且你应该在考虑后训练之前就把它建好。一旦模型投入生产,也是一样,不管模型是否经过后训练。如果我们推出大的变更,我们会做 A/B 测试,我们看参与度来跟踪这个模型是否表现如预期。然后我们看正常运行时间、token 效率,然后我们称之为产品反馈,我称之为愤怒的客户邮件。所以有所有这些信号告诉你这是否按预期工作。所以你需要先建立这个基础设施,然后再考虑部署模型。

And I think the important point here is this is the case for post-trained models or non-post-trained models. And so you can reuse this infrastructure and you should have it in place before you think about post-training. Once a model's in production, same thing, doesn't matter if the model's post-trained or not. We do AB testing if we're rolling out a large change, we look at engagement to track if this model's performing how we expect. Um and then we look at things like uptime, token efficiency, and then we call it product feedback, I call it angry customer emails. And so there's all of these signals that tell you this is working as expected. And so you need that infrastructure in place before you think about serving models.

Gabe

然后,在部署模型之前你还需要做的事情,我称之为简单的开源切换。第一个是看看你部署模型的所有地方,找出在我的产品中是否有可以天真地替换成开源模型的地方?例如,我们产品中有生成引用的部分,不需要最大的模型,有机会换成 GLM 5.2,获得成本或性能上的好处。这是第一件事,这样做能锻炼我们与闭源模型一起部署开源模型的能力。

And then the thing you need to do even still before serving models is what I call the simple open-source switches. And so the first one is look at all the places you're serving models and find are there places in my product where I can just naively swap open-source models? And so for example, we have parts of our product that generate citations that don't need the largest models and there's opportunities to swap in GLM 5.2 and get cost or performance benefits. And so that's the first thing and doing this builds the muscle of serving open-source models alongside closed-source models.

Gabe

然后第二个现在非常流行的是模型路由。所以有些地方你不能天真地替换开源模型,但你可以做的是在某些查询上路由到开源模型。一旦你完成了这些,你就准备好开始构建后训练的飞轮了。一旦你开始在生产中部署它们并收集反馈,这里要注意,你需要非常小心收集反馈意味着什么。在我们的案例中,这并不意味着在客户数据上训练,但我们确实从用户测试和其他事情中获得反馈信号,这些信号可以指导我们如何构建未来的数据集并改进这些模型。

And then the second which has gotten very popular now is model routing. So there could be places where you can't naively swap an open-source model, but what you can do is on certain queries route to open-source models. And then once you've done this, you have everything in place to start building the post-training flywheel. And once you start serving them in production and collecting feedback, and caveat here, you need to be very careful about what collecting feedback means. In our case, it does not mean training on customer data, but we do get feedback signals from our user testing and other things that can inform how we build future data sets and improve these models.

Gabe

这就是我们建立研究实验室的剧本。我们认为未来每家人工智能公司都需要成为一家人工智能公司,并找出这个剧本的某个版本。我认为尽管如此,大多数人仍然在押注反对这个剧本、应用层公司和前沿生态系统。

And that is our playbook for building a research lab. We think in the future every AI company will need to become an AI company and figure out some version of this playbook. And I think despite that, most people are still betting against this playbook and application layer companies and the frontier ecosystem.

Gabe

但如果我们用这个预算和这个团队赢了,我们就会改变游戏规则。

But if we win on our budget with this team we'll change the game.

Gabe

这是《点球成金》里的一个场景,希望你们看过这部电影,他们谈到我们刚刚连续赢了 20 场比赛。比利·比恩说:“没关系。如果我们不赢得冠军,没有人会欣赏我们在这里所做的一切。”那句话是:“但如果我们用这个预算和这个团队赢了,我们就改变了游戏规则。”我认为现在有了前沿生态系统,你们所有人都有机会做同样的事情。去改变游戏吧。谢谢。

This is a scene from Moneyball, which hopefully you've seen this movie, where they're talking about we just won 20 games in a row. And Billy Beane says, "It doesn't matter. If we don't win the championship, no one's going to appreciate what we've done here." And the quote is, "But if we win on this budget with this team we'll have changed the game." And I think now with the frontier ecosystem all of you have the opportunity to do the same. Go change the game. Thank you.

Host

你想做问答环节吗?

Do you want to do Q&A?

Gabe

你同意吗?

You okay with that?

Host

是的,当然。

Yeah, of course.

Gabe

好的。

Yeah.

Host

抱歉,谁在说话?嗯,你谈到了拥有数据非常敏感的法律客户的挑战,还谈到了你哥哥的律师和其他内部专家做他们版本的实时编码来产生最好的工作。那是什么样子的?你能给我们更多细节,说明那实际上是什么样子的吗?

I'm sorry, who's speaking? Um you talked about the challenges with having legal customers with very sensitive data sets and talked a little bit about your brother's lawyer and other in-house experts doing their version of live coding to produce the best work. What does that look like? Can you give us a little more detail about what that actually looks like?

Gabe

是的,好问题。法律工作最大的挑战是,当你在大型律师事务所时,你做的这类工作,例如,你代表一家公司进行并购。你会得到一个数据室,里面有你要收购的公司的所有合同,还有所有这些关于谈判的电子邮件和会议。这些数据都不是公开的,所以你没有像开源 GitHub 仓库那样的类比。

Yeah, great question. So the biggest challenge with legal work is when you're at a large law firm, the type of work you're doing is, for example, you're representing a company doing an M&A. And you get a data room, which is all the contracts of the company you're trying to acquire, plus there's all these emails and meetings about the negotiation. None of that data is public, and so you don't have this analogy of like open source GitHub repos.

法律训练数据生成 Data Generation for Legal Training

Gabe

训练中遇到的最大挑战是,你或许能在公开渠道找到一些最终工作成果,比如公开的采购协议,但你没有任何输入数据。过去要让一群律师说“嘿,做个假的数据室”非常尴尬,因为这些数据室可能有 10,000 份合同,它们都需要相互匹配。所以 Julio 想出了一个非常巧妙的方法来生成这些数据集。他从评分标准出发,埋下了所有这些问题,说“这是数据室里所有的问题,以及我要在场景中检查的内容”,然后用它来生成数据室。这样你就可以埋下所有这些问题,比如这些合同对不上、这份合同缺失,然后生成所有数据。接着我们用 Mercor Shenorcal 让所有合同看起来真实,但现在你有了这个输入数据集,就可以让模型生成一份尽职调查备忘录。这种检查方式会说:“嘿,你发现了所有这些问题吗?我们知道这些问题在里面,因为是我们埋下的。”显然,根据你做的具体工作,你需要巧妙地创建它,但对我来说,这感觉像是真正的重大突破,因为现在我们可以开始做这种训练了。以前,我们面临一个先有鸡还是先有蛋的问题:我们去律师事务所说“嘿,我们可以用你的客户数据训练这个模型”,他们会说“好,证明给我们看”,我们说“给我们看数据”,他们说“不行”。所以现在你可以在这里证明它。现在我们收到了很多律师事务所的兴趣,他们说“哦,这真的很有趣,我们能用我们的数据来做吗?”

And the biggest challenge you run into when training is you can maybe find some of the final work product publicly, like a public purchase agreement, but you don't have any of the input data. Historically it was very awkward to get a bunch of lawyers and say, 'Hey, make a fake data room.' Because these data rooms can be 10,000 contracts. They all need to fit together. So what Julio figured out is a very clever way to generate these datasets. He started from the rubric and planted all these issues, saying, 'Here's all the problems in the data room and what I'm going to check for in a scenario,' and then used that to generate the data room. So you can plant all of these issues like these contracts don't tie together, this contract's missing, and generate all of the data. Then we'll use Mercor Shenorcal to make all the contracts look realistic, but now you have this input dataset, and you can have the model generate a diligence memo. This way of checking it says, 'Hey, did you catch all of these issues that we know are in here because we planted them?' Obviously, depending on the work you're doing, you need to be clever about how you create it, but that to me feels like the really big unlock because now we can start doing this training. Before, you had this chicken-and-egg problem where we'd go to law firms and say, 'Hey, we can train you this model on this client data.' And they'd be like, 'Okay, prove it.' And we were like, 'Show us the data.' And they're like, 'No.' So now you can prove it here. And now we have a lot of interest from law firms saying, 'Oh, this is really interesting. Can we do it with our data?'

Host

好问题。是的。

Good question. Yeah.

招聘策略 Hiring Strategy

Gabe

嗯,我们也在为此招聘。

Um, so we're also hiring for this.

Host

我们也在那里构建一个隐含的 AI。从招聘的角度来看,你很难与来自 Anthropic 和其他实验室的研究人员竞争。你是尝试玩那个游戏,还是尝试招聘领域逻辑、律师和经验?哪些地方有效,哪些地方无效?

We're also building an implied AI there. From a hiring point of view, it's difficult for you to compete with researchers from Anthropic and the labs. Do you try and play that game or do you try and hire domain logic and lawyers and experience? And where has that worked and where has that not worked?

Gabe

是的,我认为这是我们刚创办公司时犯下的大错误之一,因为我在这些实验室工作过,认识很多这样的人。我会说“你知道,来和我们一起工作吧”,但他们当时拿到的是上亿美元的薪酬包,而我们的规模还达不到那个水平。所以我想说,我们现在刚刚达到一个公司规模,我们不再与前沿人才竞争,但越来越多的人在做博士研究,有些人不想在大实验室工作,所以我认为这种情况正在改变。第二点是,早期你需要这些人才的原因不仅仅是做后训练,还需要构建训练基础设施、服务基础设施,所有这些结合在一起。我认为能做到这一点的人才非常独特。比如我的老室友是在 OpenAI 负责后训练的人之一,他是我合作过的最好的研究人员之一。但现在你可以使用像 Fireworks、tinker 之类的 API,所以你不需要自己构建训练和服务基础设施。所以我认为这拓宽了人才库,现在获得一些这样的人才感觉非常可行。

Yeah, I think this was one of the big mistakes I made when we first started the company because I had worked in these labs and so I knew a lot of the folks. I would say, you know, come work with us, and then they were getting these, you know, 100 million plus pay packages, and we weren't at the scale where this made sense. So I'd say we're now just getting to the company size where we're not competing with the frontier talent, but there are increasingly folks that are doing PhDs, folks that don't want to work at large labs, and so I think that is changing. The second thing is a lot of why you needed that talent early on was it wasn't just doing the post-training, it was you need to build the training infrastructure, the serving infrastructure, and all of those things combined. I think the talent to do that was very unique. Like my old roommate was one of the folks who ran post-training at OpenAI, and he was one of the best researchers I've ever worked with. But now you can use things like Fireworks, tinker-like APIs, and so you don't need to build the training and serving infrastructure. So I think it opens up the pool of talent, and it feels very feasible to get some of this talent now.

基准创建挑战 Benchmark Creation Challenges

Host

是的,关于基准创建还有一个问题。当你创建这些评分标准时,你仍然需要以某种方式创建它们,使它们能够区分前沿模型,而且你必须与专家一起做这件事。你如何应对这个挑战?所以专家们必须创建这个评分标准,同时知道或试图弄清楚模型在它们上的表现。

Yeah, just another question on the benchmark creation. So when you're creating these rubrics, you still have to create them in such a way that they're separating the frontier models, and you have to do that with the experts themselves. How do you navigate that challenge? So the experts kind of have to create this rubric knowing or trying to figure out how the models are going to perform on them.

Gabe

你说的区分是指什么?

What do you mean by separating the...

Host

我想是的,你必须创建能够充分区分前沿模型的评分标准。它们必须挑战前沿,对吧?

I guess so, you have to create the rubrics that can adequately separate the frontier models. They have to challenge the frontier, right?

Gabe

是的。

Yep.

Host

我想你如何应对这一点,尤其是当你拥有这些内部法律专家,他们可能不完全知道如何为模型设计这样的评分标准。

I guess how do you navigate that especially when you're having these in-house legal experts who might not exactly know how to create that rubric designed for a model.

Gabe

是的,这是个好观点。所以我们更多考虑的是如何让它们真实地代表我们客户所做的工作?我认为我们选择法律领域的部分原因是,如果你只是构建一个顶级律所正在处理的真实客户案件,前沿模型仍然无法做到这一点。我认为在过去 4 年里我们建立的东西是,例如,我哥哥是我们的第一批员工之一。他已经和模型一起工作了 4 年。所以他对模型如何工作、如何生成这些数据集的直觉非常出色。他训练了一批律师来做这件事。但我想说,重点仍然更多是如何构建一个真正真实的数据室,然后为尽职调查制定评分标准。然后当我们运行这些模型时,我们发现差距,还有改进空间。但我认为这取决于领域。

Yeah, that's a good point. So we think of it more as how do we make them realistic representations of the work our customers are doing? I think part of why we picked the legal domain is if you just build a realistic client matter that these top firms are working on, the frontier model still can't do this. And I think the thing over the past 4 years we've built is, for example, my brother was one of our first hires. He's been working with the models for 4 years. So his intuition of how the models work, how to generate these datasets is incredibly good. And he's trained a bunch of lawyers to do this. But I would say the focus is still more of how do we build a really realistic data room, and then rubrics for the diligence. And then when we run these models, we find gaps and there's room for improvement. But I think it depends on the domain.

端到端管道调试 End-to-End Pipeline Debugging

Host

嘿,我是 Box 的 Shadeen。

Hey, Shadeen from Box.

Gabe

抱歉,让我先处理那个,然后我会……

Sorry, let me do that one and I'll...

Host

好的。

Great.

Host

我的意思是,如果你想想……顺便说一句,演讲很棒。嗯,我是 Trolysis 的 Ross。如果你想想端到端的流程,就像数据生成、环境部分,然后是算法部分,比如我们如何训练等等,研究部分,然后是基础设施部分。如果有什么不奏效,根据经验,你看到了什么?是运营上的、金钱上的还是人力时间上的?通常是在数据层吗?是在研究层吗?还是在基础设施层?你如何从那里着手?

I mean, if you think... Great talk, by the way. Um, Ross from Trolysis. If you think of the end-to-end motion, it's like data generation, the environment piece, then there is the algorithm piece of like what how do we train, etc., the research piece, and then there's the infra piece. If something is not working, anecdotally, what have you seen? Is it operationally or dollar-wise or human-hours-wise, is it usually in the data layer? Is it in the research layer? Or is it in the infra layer? And how do you go from there?

Gabe

在某种程度上,谁在编排或调试整个端到端流程?

Who's orchestrating or debugging this whole end-to-end pipeline in some ways?

Host

是的,这是一个很好的问题。我的意思是,这完全是另一个话题。我不认为有单一的因素。所以我想说,让我们更容易的是我们找到了产品市场契合度。我们有一个在生产中使用的产品。我们一开始主要使用闭源模型。这让我们建立了很多信心,相信事情是端到端工作的。所以我们有一个产品,人们在使用它,人们已经使用它很多年了。模型以这种方式工作,我们可以监控它们在生产中是否正常工作。所以这就是服务的关键,你需要很多这样的东西到位。然后我们已经建立了肌肉记忆,当新模型出现时,如何将它们放入产品中,无论是闭源还是开源模型。所以你可以考虑整个周期。然后显然我们做了很多工作,比如 Harness 工程和上下文管理等等。所以现在你可以把后训练看作是更广泛系统中的一个小的输入。

Yeah, this is a great question. I mean, this is a whole 'nother talk. I don't think there's a single thing. So I would say what makes it easier for us is we found product-market fit. We have a product that's being used in production. And we did this largely with the closed-source models to start. And so that let us build a lot of this kind of confidence in the thing is working end-to-end. So we have a product, people are using it, people have been using it for many years. The models work in this way, we can monitor whether they're working in production. So that was kind of the point of serving, you need a lot of this in place. And then we have built the muscle of new models come out, how do we put those into the product, whether they're closed-source or open-source models. So you can kind of think about the full cycle. And then obviously we've done a bunch of work like harness engineering and context management and all of these things. So now you can think of post-training as one small input into this broader system.

后训练与门控 Post-training and gating

Gabe

所以,我们让团队对模型进行后训练,然后把它放进这个更大的系统里,就把它当作另一个新模型。然后你就有各种信号来发现哪里出了问题。但在这些步骤中,有简单的办法来设置关卡。比如,从基准测试开始,后训练一个模型,如果它在这些基准上表现不好,那就不太可能继续推进。如果它表现好,然后放进产品里,但有人用了产品后说感觉不对劲,因为比如在我们建的基准上,仍然有一些东西它们肯定捕捉不到。最好的例子是,你可以有一个模型在我们的基准上表现很好,但它们在某种程度上过拟合了,然后你把它放进一个更通用的助手类产品里,一旦超出分布范围,它就会崩溃。但确实,你需要把所有的阶段和关卡都设置好,然后哪里出问题就变得很明显了。不过,问得好。

And so, we have our team post-train a model, and then you feed it into this broader system, treating it just as another new model. And then you have all these signals of what's going wrong. But across these steps, there are easy ways to gate it. So, if you start with the benchmarks and you post-train a model and it doesn't work well on those, then it's unlikely to go farther. If it works well on that, and then you put it in a product, but someone uses the product and they say things feel weird because, for example, on the benchmarks we've built, there are still things they definitely don't catch. The best example is you can have a model that does very well on our benchmarks, but they've overfit to some degree, and then you put it in a more generic assistant-like product, and it kind of falls apart when you go out of distribution. But yeah, you need to have all the stages and gates in place, and then it becomes obvious where things are breaking. But good question.

Host

嗯。

Yeah.

Host

哦,我有个问题。嘿,很高兴认识你。我来自 Rocks。在基准测试方面,法律智能体,我想你们之前有 Big Law,Big Law bench,然后现在是法律智能体。你们对开源基准测试的理念是什么?因为如果开源,当然会有其他人尝试这个基准,从而获得关注,但实验室也能借此提升。如果闭源,那就有个问题:这是不是真正的基准?所以你们……你们的理念是什么?

Oh, I had a question. Hey, great to meet you. I'm from Rocks. On the benchmark side, legal agent, I think you guys had Big Law before, Big Law bench, and then now legal agent. What's your philosophy on open sourcing the benchmark? Because if you open source, of course you're getting traction with other people trying out the benchmark, but the labs can help climb as well. And then if you close source, then there's a question of is this a real benchmark? So how do you... what's your philosophy?

Gabe

我认为这是我们思考的一个很好的张力点。我想说,甚至在我们开源这个基准之前,我们就已经和实验室紧密合作了。我们和他们分享数据,帮助他们改进模型,这也提升了我们的产品。所以我想说,我们思考的方式是,在你们的行业,尤其是法律领域,有一个巨大的优势,就是帮助我们的客户理解不同模型在不同事情上的表现有多好。所以当我创办 Harvey 时,我有一种直觉,哦,我们只要构建最好的模型,客户就会满意,因为它是最好的。现在很明显,每个客户有不同的偏好,不同的模型擅长不同的事情。然后我没想到的第二点是,前沿生态系统会变得如此庞大。所以现在有很多公司联系我们说:“嘿,我们有这个新技术,或者我们做了这个,我们能试试吗?”以前我们没有精力,我们只是在忙其他事情。但现在有了这个开源数据集,我们可以说:“嘿,去试试这个,如果有效,那太好了。这是值得投资的东西。”然后我认为我们战略上的思考是,对我们来说有价值的数据是帮助律师事务所训练他们自己的系统,以及他们如何在私有数据上工作。然后我们想用合成数据和部分开源数据帮助所有人改进这些系统,但显然要平衡你提到的那些原因。

I think that's a good tension that we think about. I would say even before we open source this benchmark, we already work closely with the labs. And so we share data with them to help them improve their models, which improved our product. And so I would say the way we think about it is there is a huge advantage in your industry, particularly in legal, in terms of helping our customers understand how good different models are at different things. And so I think when I started Harvey, I had the intuition of, oh, we'll just build the best model and then customers will be happy because it's the best. And it's very clear now that every customer has different preferences, different models are good at different things. And then I think the second thing I didn't anticipate is how big the frontier ecosystem would become. And so now we have so many companies reaching out saying, "Hey, we have this new technique or we did this thing. Can we try it?" And before we just didn't have the bandwidth, we were just like, we're working on these other things. But now with this open source data set, we can say, "Hey, go try this and if it works, then great. This is something that's interesting to invest in." And then I think the way we think about it strategically is the valuable data for us is going to be helping law firms train their own systems and how they work on private data. And then we want to help everyone improve these systems with synthetic data and some of the data we open source, but obviously a balance for the reasons that you mentioned.

Host

你提到了找实验室。那你还想在哪方面得到帮助?还有哪些未解决的问题?

You mentioned finding a lab. What are the remaining open questions for that you want help on?

Gabe

是的,这是个好问题。我想说一个大挑战,最大的挑战之一,是我们能生成这些非常逼真的合成数据集,我们可以用人类来增强它们,但这些数据的分布仍然不匹配我们的生产分布。所以我想说,我们生成的这些数据集在某种程度上是面向未来的。我的意思是,我们可以生成一个非常逼真的数据室,但目前我们的产品并不纯粹用于尽职调查。所以如果一个模型在这方面表现好,但有人用它来起草邮件,也许它就不那么好了。所以我认为仍然存在这个差距,因为我们不能查看客户数据,那么如何弥合这个差距呢?所以对我来说,这是数据集方面最大的问题,而且这个问题很有挑战性。还有很多问题,比如,再以尽职调查数据为例,最大的数据室有 8000 万 token 的上下文。模型不能很好地管理这种上下文,那么如何训练这些模型在这种非常复杂的环境中运作呢?所以我认为这是另一个问题,仍然存在很大的性能差距。第三点是,我们在如何自己后训练我们的后训练模型方面取得了良好进展,但我认为找到某种形式的持续学习才是最终目标。不是让我们构建最好的法律模型,而是让我们帮助每个律师事务所或企业客户根据他们所做的工作类型进行定制。所以思考如何将其操作化,比如你有一个大型律师事务所,每次他们处理客户事务时,他们的 AI 系统都会变得更好,但同时你也要保护客户数据。我认为这是一个巨大的、既涉及技术又涉及运营和 AI 的挑战,非常有趣。

Yeah, that's a great question. I would say one big challenge, one of the biggest challenges, is we can generate these very realistic synthetic data sets, we can augment them with humans, but the distribution of that data still doesn't match our production distribution. And so I would say these data sets we're generating are somewhat future facing. And what I mean by that is we can generate a really realistic data room, but right now our product isn't used purely to do due diligence. And so if a model does well on that, but then someone uses it to draft an email, maybe it doesn't work as well. And so I think there's still that gap because we can't look at customer data, and so how do you bridge that gap? So I think that's to me the biggest question on the data set side, and that one is challenging. There are also a bunch of questions around, for example, again with the diligence data, the largest data room is 80 million tokens of context. The models don't do a good job of managing this context, and how do you train these models to operate in these very complex environments? So I think that's another one, there's still a big performance gap. And I would say the third is we're making good progress on how we post-train our post-train models ourselves, but I think figuring out some form of continual learning is the end game here. It's not for us to build the best legal model. It's for us to help every law firm or enterprise customer customize it to the type of work they're doing. And so just thinking about how do you operationalize that, where you have a large law firm and every time they work on a client matter, their AI system gets better, but you're also protecting the client data. I think that's one huge, both technical, operational, and AI challenge that is super interesting.

Host

嗯。我们结束了。还是继续?

Yeah. We're done. Or more?

Gabe

我可以继续。我可以……嗯,也许最后一个。

I can do more. I can... Yeah, maybe last one.

Host

是的。所以,你谈了很多关于基础准备,对吧,与 Front Row Live 竞争,包括你的开源模型,包括你的基础设施,比如 Ruby。从产品角度看,对吧,Coda 作为云产品,他们是通用产品,对吧?那么从产品角度看,你们如何与 Coda 的云工作流、方法论、效率竞争?你们怎么考虑这个?

Yeah. So, you know, you talked a lot about the foundation readiness, right, to compete with Front Row Live, right, including your open source model, including your infrastructure kind of like Ruby. From product perspective, right, Coda as cloud product, they're general product, right? So from product perspective, how do you guys compete with the Coda as cloud workflow, methodology, efficiency? How do you think about that?

Gabe

是的。我认为我们思考的重大转变是,我们最初的产品,以及像 Coda 的 Coda Work 这样的东西,都是非常个人化的产品,它们关乎个人生产力。而我们正在构建的产品越来越关乎组织生产力。所以,如果你考虑一个大型律师事务所,他们试图解决的问题不是如何让我的每个律师更高效。他们试图解决的问题是我有 1 万个客户。我正在为所有客户处理项目。我需要确保所有这些客户项目都进展顺利,而且我还能以盈利的方式完成。所以,我们为律师事务所构建的很多基础设施就是如何操作那台机器?当你开始考虑一个单独的客户项目时,大部分挑战不是如何起草这一部分。而是我参与这个项目 6 个月了。我在全所有一个 20 到 30 人的团队。我需要协调他们所有人。我需要确保我协调了所有外部各方。

Yeah. I think the big shift that we're thinking about is our original product and then things like Coda's Coda Work are very individual-focused products, and so they're about individual productivity. And increasingly, the product we're building is about organizational productivity. And so, if you think of a large law firm, the problem they're trying to solve is not how do I make my individual lawyers more productive. The problem they're trying to solve is I have 10,000 clients. I'm working on client projects for all of them. I need to make sure that all those client projects go very well, and then I can also do them in a way that's profitable. And so, a lot of the infrastructure we're building for the law firms is how do you operate that machine? And when you start thinking about an individual client project, most of the challenge is not how do I draft this one section. It's I'm working on this project for 6 months. I have a team of 20, 30 people across the firm. I need to coordinate all of them. I need to make sure that I'm coordinating all the outside parties.

项目管理与企业复杂性 Project Management and Enterprise Complexity

Gabe

所以,很多工作开始看起来像是项目管理,协调这些人类和这些智能体。我觉得这是在团队层面。然后,如果你从组织层面来看,现在你有上千个这样的项目,你需要开始考虑资源分配,我在计费什么,我在推销什么?而在企业层面,情况就更复杂了,比如一家财富 500 强公司,他们与上千家律所合作,内部有上千名员工。所有启动工作的系统,是不是在交易这些工作?所以,简单的答案就是,我认为历史上企业所做的就是,如何深入垂直领域,以横向产品无法做到的方式。是的,好问题。

And so, a lot of it is starting to look like project management that is orchestrating these humans and these agents. And that's at the team level, I would say. And then, if you think about it at the organization level, now you have a thousand of these projects, and you need to start thinking about resource allocation, what am I billing, what am I pitching? And then with enterprises, it's even more complicated where a Fortune 500, they're working with a thousand law firms. They have a thousand people internally. What are all the systems to start work, is trading that work? And so I would say the simple answer is just, and this is I think historically what enterprise has done, is how do you just go hyper vertical into your domain in a way that the horizontal products won't. Yeah, good question.

Host

谢谢你,Gabe。那是一场标志性的演讲。感谢你参加我们的节目。

Thank you, Gabe. That was an iconic talk. Thank you for joining us.

互动版:逐字朗读 + 针对本期提问 →