强化学习环境:智能体数据的新前沿

RL Environments: The New Frontier in Agentic Data

布伦丹·富迪 Brendan Foody · Training Data · 2026-08-12 · 约 26 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Brendan 探讨了强化学习环境如何变革智能体数据,专家构建世界、应用和任务来训练前沿模型。

Brendan discusses how RL environments are transforming agentic data, with experts building worlds, apps, and tasks to train frontier models.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 17)

全文 · Full transcript(中英对照)

引言与背景 Introduction and Context

Host

嗯,Record,我觉得你们在过去 4 个月左右的时间里,营收运行率从 10 亿美金增长到了 20 亿美金。所以这家公司已经起飞了,而且我觉得你正好处在最前沿,看到各家公司如何思考后训练自己的模型、构建自己的智能。所以,感谢你加入我们的对话。嗯,形式上,我们打算先让 Brendan 讲 15 分钟左右的内容。他会特别讲讲强化学习环境,我觉得这是一个前沿话题,探索起来会很有趣。然后我们会在最后留出 15 分钟左右的时间进行问答。所以,请大家把问题先记在心里。我把时间交给 Brendan。

Um, Record, I think you guys grew from a 1 to a 2 billion dollar revenue run rate in the last 4 months or so. Um, so this company's off to the races and I think you were just so front and center to how companies are thinking about uh post training their own models, uh building their own intelligence. So, thank you for joining us for this conversation. Um, format-wise what we're going to do is we've 15 minutes or so of content from Brendan. He's going to talk about uh RL environments in particular, which I think is a, you know, new frontier topic. It'll be fun to fun to explore. And then we're going to leave 15 minutes or so at the end for Q&A again. So, uh please keep please keep questions back pocket. I will turn it over to you, Brendan.

RL环境概览 RL Environments Overview

Brendan

好。所以,我要讲的是强化学习环境。首先,我觉得有必要先介绍一下数据市场的一些背景,以及这段历史如何与 Record 的起源故事联系在一起。事情真正开始于 2020 年,当时是众包数据用于行为克隆的时代。所以,这主要是监督微调数据、输入和输出,以及基于人类反馈的强化学习(RLHF)数据,标注者会从几个模型回复中选出他们更喜欢的那个。我们能够在微调 GPT-3 上取得所有这些进展,在智能体数据的众包时代朝着 ChatGPT 和 GPT-4 迈进。但是,我们看到市场在变化,尤其是当我们进入 2024 年时,这是一个巨大的转变,从低技能的众包行为克隆数据时代,转向智能体数据时代。我们如何找到世界上技能最高的专家,让他们能够协作团队合作,为下一代模型构建前沿评估和强化学习环境。所有软件工程师、律师、医生、银行家等等,他们能够衡量智能的前沿,并帮助利用这些来提升模型能力。所以,Mercor 成长起来,我们的第一个大项目是深度研究。我想第一个突出的强化学习智能体,与所有前沿实验室一起大幅扩展,成为所有领先实验室以及所有领先应用层公司(从 Harvey、Cera、Cognition 到 Ramp)的主要智能体数据供应商。在过去 12 个月里,特别令人兴奋的是,智能体数据范式中的 RLVR 如何演变,也包括了强化学习环境,这些环境有丰富的应用和世界,教智能体如何使用我们每天在笔记本电脑上使用的所有工具。所以,我会谈谈这个,当然还有这项始于前沿实验室的技术如何现在传播到应用层,以及你们在创建公司时正在构建的所有产品。

Sweet. So, I'll be talking about RL environments. Starting out, I figured it's helpful to give a little bit of the background on the history of the data market and how that history ties into Record's origin story. Where things really started in 2020 in the era of crowdsourcing data for behavior cloning. So, this was mainly supervised fine-tuning data, inputs and outputs, and RLHF data where you would have an annotator select from a couple of model responses which they preferred. And we were able to make all this progress in fine-tuning GPT-3, making progress towards ChatGPT and GPT-4 in the crowdsourcing era of agentic data. But, what we saw changing in the market, especially as we uh headed into 2024, was this giant transition away from the low-skilled crowdsourcing era of behavior cloning data and moving towards the agentic era of data. Of how do we find the highest-skilled experts in the world that can work collaboratively in teams to build frontier evals and RL environments for the next generation of models models. All the software engineers, lawyers, doctors, bankers, et cetera that could measure the frontier of intelligence and help to use that to improve model capabilities. And so, Mercor grew up with our first big project being deep research. I guess the first prominent RL agent scaling up dramatically with all of the frontier labs to become the primary agentic data vendor to all of the leading labs and also all of the leading application layer companies ranging from Harvey, Cera, Cognition to Ramp. And what's been really exciting over the last 12 months especially is how RLVR within the agentic data paradigm has evolved to also include RL environments with these rich apps and worlds that teach agents how to use all of the tools on our laptops that we use every day. So, I'll be talking about that and of course how this technology that started in the frontier labs is now getting disseminated to the application layer and all of the products that all of you are building as you work on your company.

RL环境组成 Components of an RL Environment

Brendan

所以,从高层次来看,强化学习环境包括三个部分。第一部分是世界。这包括所有的消息、幻灯片、文档、表格等等,对应你在真实项目或公司中会拥有的一切。第二部分是应用,即流行应用的高保真克隆,如 Salesforce、ServiceNow、Microsoft 365 等,智能体可以通过 MCP、CLI 或 Kua 与之交互。第三部分是任务,我们有提示和验证器。验证器可以是评分标准或单元测试,可用于评估或训练。前沿实验室要自动化你使用 Claude 在笔记本电脑上能做的一切,其障碍是如何覆盖经济中所有世界、所有应用和所有任务的完整分布。所以,为了实现这一点,进行了大规模扩展。嗯,人类在构建这些环境方面一直非常核心,显然模型也在其中有意义地参与。所以我在这里放了一张图,显示过去 24 个月我们人才网络的专家小时吞吐量。嗯,这是一个相当疯狂的轨迹,仅第二季度就有 250 万小时,用于构建所有这些环境的专家时间增长在加速。

So, high level on what an RL environment is is that it includes three parts. The first part is the worlds. So, this includes all of the messages, slides, docs, sheets, etc. that correspond to everything you would have in a real project or company that you're working on. The second part is the apps which is high fidelity clones of popular applications, Salesforce, ServiceNow, Microsoft 365, etc. that agents can interact with via MCP, CLI, or Kua. And then the third part is the tasks where we have prompts and verifiers. Verifiers could be rubrics or unit tests that can be used either for eval or training. And the barrier for frontier labs to automate everything that you can do on your laptop using Claude is how do they cover the full distribution of all of the worlds, all of the apps, and all of the tasks in the economy. And so, there's been this enormous scale out in order to do that. Um, where humans have been really central to how we build these environments, obviously with models in the loop meaningfully. And so I put a graph here of the amount of expert hours that um, of throughput in from our talent network over the last 24 months. Um, and it's a a pretty crazy trajectory with respect to um, 2.5 million hours um, in Q2 alone with growth sort of accelerating on the amount of expert time uh, used to build out all of these environments.

人类为何关键 Why Humans are Essential

Brendan

原因当然是我提到的,我们需要在经济中的每个类别扩展环境分布。很多人可能知道 GDP val,劳工统计局有 205 个领域,涵盖所有不同的工作,但你必须考虑如何拥有所有对应这些工作的应用、所有不同的场景、所有任务。这是一个巨大的构建。在大多数领域,只有人类能衡量前沿,不是每个领域。有少数例外,比如数学,那里有一个非常干净的模拟环境,所以模型能够从它是否得到正确答案中学习,但在大多数领域,比如制作幻灯片,模型很难可靠地识别自己在哪里犯了错误。这就像要求人类给自己的作业打分。所以,这就是为什么让人类创建评分标准非常有价值,就像教授创建评分标准来评论文,或助教评幻灯片一样。类似于我们很多人学习的方式,很大程度上来自周围人的反馈,而不是纯粹地把东西插入计算器或干净模拟。嗯,然后构建这些验证器很难,因为任何时候你制作幻灯片,你需要理解完整的问题空间,比如哪 10 种不同的幻灯片可能是好的路径?你可能犯的几十种错误是什么?你如何构建一个全面的验证器来捕捉这个可能的完整解决方案区域?

The reason being of course as I mentioned we need to scale out the environment distribution across every category in the economy. Many of you might know GDP val where there's 205 domains in the Bureau of Labor Statistics across all the different jobs, but then you have to think through how do we have all of the apps corresponding to all of those jobs, all the different scenarios, all the tasks. Is this enormous build out. Only humans can measure the frontier in most domains, not every domain. There are rare exceptions like math where you have a really clean simulation environment and so uh, the model's able to learn from whether it got the right answer, but in most domains like building a slide deck uh, the model has an incredibly hard time identifying reliably where it made its own mistake. It's as if you would be asking a human to grade their own homework. And so that's why it's really valuable to have a human create a rubric similar to how a professor would create a rubric to grade an essay or a TA would grade that slide deck. Similar to the way that a lot of us learn it's in large part from the feedback we got from those around us rather than uh, purely plugging things into a calculator or clean simulation. Um, and then building these verifiers is hard cuz anytime you're building the slide deck you need to understand the full problem space of what are the 10 different slide decks that, you know, could be a good path to go down? What are the dozens of mistakes you could possibly make? And how do you build a comprehensive verifier that captures this full solution area of what's possible?

RL环境示例 Example RL Environment

Brendan

所以,我要展示一个示例强化学习环境。抱歉,为了给大家分解一下。嗯,这之所以这么酷的部分原因,我稍后会讲到,是我们与实验室合作开发了很多这项技术。这些当然是我们已经开源并发布给世界的,但现在这一切开始传播到应用层公司,它们正在构建并拥有自己的智能。当它们意识到其 AI 战略的三个核心支柱是算力、算法或研究人员,以及它们构建的数据集。而数据往往是最具差异化的因素。所以,这是我们发布的一个法律环境,我们让来自顶级律所(如 Latham and Watkins)的律师写出他们在大型律师事务所工作中处理过的真实项目场景。然后他们为数据室创建一个完整的提纲,对应所有不同的消息、电子邮件、文件、文件大小。我截断了完整的数据室,因为它非常详尽。嗯,当然,模型在循环中有很多参与,他们如何有效地填充这个。

And so, what I'll walk through is a sample RL environment. Excuse me, to also break this down for all of you. Um and part of the reason that this is so cool, which I'll get to in a moment, is that we developed a lot of this technology in collaboration with the labs. These are, of course, ones that we have open-sourced and published to the world, but now that's all starting to get disseminated to the application layer companies that are building and owning their own intelligence. As they realize that the three core pillars of their AI strategy are their compute, their algorithms or researchers, and the data sets they build. And data's often the most differentiating factor. And so, this is one that we published, um as a legal environment, where we have lawyers from top law firms like Latham and Watkins write out a scenario of a real project that they worked on in their big law job. And then they create a full outline for a data room that corresponds to all of the different uh messages, emails, files, size of files. I cut off the full data room cuz it's it's very extensive. Um and of course, there's a lot of model in the loop with how they effectively populate this.

构建RL环境与数据室 Building RL Environments and Data Rooms

Brendan

就像现在的软件工程师不应该完全自己手写代码,他们可能应该编排智能体来高效完成这件事。然后我们把数据室渲染到应用里,也就是你们在这个场景中看到的 Google Workspace 的克隆版,并设置提示词来针对这些数据展开模型轨迹。所以在这个例子里,它是在评估 Star Tanker Tankers International Limited 相对于 Cooper Jefferies Energy Corporation 在《石油与石油法案》下的最大总责任,同时考虑数据室中这个真实场景的所有上下文。然后就像我提到的,类似教授为论文评分制定评分标准,他们有这些关键的评分标准,对应着准确模型响应的特征。而且确保这些评分标准避免奖励黑客行为,并有效对齐,当你展开 100 条轨迹时,确保所有这些分数都准确是极具技术挑战性的。所以有大量的研究、议程质量控制、数据训练等等,都投入到如何解决这个问题,最终产出这些高质量的验证器和排行榜,给你一个聚合的模型分数,反映不同模型在特定领域上的表现。正如我们所见,过去几个月最大的变化之一是 GLM 52 和 Chimera K3 登上了排行榜。这对你们所有人来说都是巨大的机会,因为这给了我们真正实现前沿智能的基础,以及你们所关注的所有特定应用和垂直领域,这并不遥远。

Similar to how a software engineer now should not be coding by hand entirely themselves, they should probably be orchestrating agents to do this very productively. And then we render that data room into the apps, the clones of Google Workspace you can see in this scenario, and have prompts to roll out model trajectories against this. And so, in this one, it's evaluating the maximum total liability for Star Tanker Tankers International Limited compared to Cooper Jefferies Energy Corporation under the Oil and Petroleum Act, considering all of the context from this real scenario in the data room. And then as I mentioned, similar to how a professor would create a rubric to grade an essay, they have these key rubric criteria that correspond to what are the characteristics of an accurate model response. And making sure that these rubric criteria avoid reward hacking and effectively align with the when you roll out 100 trajectories, making sure all of those scores are accurate is incredibly technically challenging. And so there's an enormous amount of research, agenda quality control, training on the data, etc. that goes into how you solve that problem and then ultimately produce these high-quality verifiers and leaderboards that give you an aggregate model score across how well all of the different models are doing on a particular domain. And as we can see, one of the big changes over the last few months is that GLM 52 and Chimera K3 are on the leaderboard. And so that is a huge opportunity for all of you because that gives us the foundation to actually achieve frontier intelligence and all of the specific applications and verticals that you're focusing on that's not too far away.

后训练示例与泛化 Post-training Example and Generalization

Brendan

为了让大家了解具体情况,我来分享一个在 Apex Agents 上进行后训练的示例,这是我刚才展示的数据集或样本,这个例子中有 1800 个任务。这是 GLM 47 的后训练运行,但我们正在为 Chimera K3 重做很多任务,所以很快会给大家更新结果。你可以看到,仅用 1800 个任务和大约 50 万算力,提升就非常显著。公司法从 4.7% 跃升到 26.6%,但请注意,这仅仅是我们给它的 Apex Agents 数据集,而它实际上在 GDP valve 和没有这些数据室的 Apex V1 上泛化得非常好。甚至在其他一些基准上也看到了名义上的提升。

And so to give a little bit of context on what that looks like, I'll share an example of post-training on Apex Agents, which is the data set that I, or the sample I just showed before, where we have 1,800 tasks in this example. This was post-training run of GLM 47, but we're redoing a lot of them for Chimera K3, so we'll have updated results for you all soon. Where you can see the jumps just on 1,800 tasks with about 500k in compute are pretty dramatic. Corporate law going from 4.7% to 26.6%, but notice that this is just Apex Agents data set we gave it, and it actually generalized incredibly well to GDP valve and Apex V1, which doesn't have these data rooms. Even just seeing nominal gains and some other benchmarks as well.

客户合作与未来机遇 Working with Customers and Future Opportunities

Brendan

所以我们正在做很多这样的工作,与像 Harvey 这样的客户合作,我知道他们稍后会展示相关内容,帮助他们构建对应其特定领域的环境,以便他们能在其中构建前沿智能。我认为 Andrew 谈到了 Cursor 是一个很好的第一个例子,展示了应用层公司如何构建行业领先的模型,为客户创造了巨大价值。我相信在未来 12 个月内,会有几十个这样的例子,公司拥有自己的智能,这是他们构建模式的关键来源。

And so, we're doing a lot of this work of working with customers like Harvey, who I know will present on stuff later to help build out the environments corresponding to their specific domain so that they can build frontier intelligence within that. And I think Andrew talked about how Cursor was a great first example of how an application layer company could build an industry-leading model that built an enormous amount of value for their customers. And I believe that over the next 12 months, there is going to be dozens of examples just like that where companies own their own intelligence and that is the key source of the modes that they're building.

高质量数据集整理方法 Ways to Curate High-Quality Data Sets

Brendan

实际上,Josh 和我前几天也谈过这个。有几个策划高质量数据集的方法。我们最常见的三种,我很乐意讨论并发送链接给大家。第一种是按任务,这是最常见的。人们会说:“我真的很喜欢法律环境的数据形态,我们会为每个任务支付 2000 美元来扩大规模。”举个例子,某些前沿实验室可能每月从我们这里购买 5 万个任务。所以规模往往相当惊人。这些任务通常非常复杂,有些甚至需要人类长达一个月才能完成。有时只需要几个小时。这些会是按任务定制的定价。第二种是现成数据,我们投入了数亿美元构建自己的数据集,卖给多个客户。所有这些新的 neo 实验室通常更倾向于现成数据,因为让 10 个不同的实验室都构建自己的数据集没有意义。一次构建、人人可用的东西价值很大。最后一种我们见得少一些,也不再是重点,就是只提供专家,让客户自己按小时模式组织专家。所以我们做一点这个,比如 Harvey 就是这样起步的,我们雇了一些律师,但随着时间的推移,它通常会转向更多这种规模化的数据产品。

And actually, Josh and I talked about this the other day as well. A couple of examples of ways to curate high-quality data sets. The general three that we see most that I'm happy to talk about and send people links to is first by task is the most common. Where people would say, 'I really like this data shape of environments in law and we will pay $2,000 per task to scale this up.' And as an example, certain frontier labs might buy 50,000 tasks a month from us. And so, it tends to be pretty dramatic scale. And these tasks would generally be very complex. Some would even take humans up to a month to complete that given task. Sometimes it would take just a few hours. And these would be sort of custom per task pricing. Second is off-the-shelf data where we have we've invested hundreds of millions of dollars in building our own data sets that we sell to multiple customers. All these new neo labs are generally airing more towards off-the-shelf data because it doesn't make sense for 10 different labs to all be building their own data sets. There's a lot of value to building something once that can then be applied to everyone. And then the final which we see a little bit of, but is less of our focus anymore, is just providing the experts so that customers are able to organize the experts on their own in just an hourly model. So we do a little bit of that when people like that was how Harvey got started with us hiring some lawyers, but it generally moves towards more of these scaled offerings of data over time.

期待与问答引言 Excitement and Q&A Introduction

Brendan

所以,这就是关于如何构建强化学习环境、它们是什么的一些背景,我对我们拥有的所有这些技术感到非常兴奋,这些技术以前仅限于前沿实验室,现在正走向你们所有人。所以,我很乐意回答任何相关问题。

So, that's a little bit of the background of how to build RL environments, what they are, and I'm really excited about all this technology that we have that has previously been limited to the frontier labs all making its way to all of you. And so, happy to answer any questions about that.

问答:数据定价 Q&A: Pricing Data

Host

很好。请讲。

Sweet. Go ahead.

Host

嘿,我是来自 Astro Guide 的 Ali。我的问题有点开放,但很简单。你们如何给数据定价?比如你们如何评估数据的价值?

Hey, I'm Ali from Astro Guide. My question is it's kind of open-ended, but simple. How do you price data? Like how do you value data?

Brendan

所以,有很多不同的方式。我的意思是,最自然的是我们的客户关心模型改进,对吧?所以,我们的客户有一个既定目标,他们想在某个排行榜上达到前沿。所以,我们能够从这对他们值多少钱、我们应该每任务收多少钱、我们认为多少任务能让他们达到那个目标来倒推。所以,当我们考虑像 Nvidia 这样的公司时,他们可能愿意支付 10 亿美元来拥有一个前沿开源模型。所以,有很多复杂性,比如我们如何为所有实现这一目标的成分定价。我们定价的另一种方式是通过我们的成本结构来制作它们,当然当我们有一个任务需要 10 小时的人工时间,我们付给人类每小时 150 美元,可能有 1500 美元的成本基础,然后问题就变成了我们想在此基础上运行多少利润率,基于该特定任务的差异化和前沿程度。但范围非常广。比如我们的任务从 50 美元到 1 万美元不等。

So, there's so many different ways. I mean, the most natural would be our customers care about model improvement, right? And so, our customers have a given goal of they want to be at the frontier on a given leaderboard. And so, we're able to work backwards from how much is that worth to them and how much should we charge per task, how many tasks do we think would get them to that goal. And so, when we think about a company like Nvidia, they're probably willing to pay a billion dollars to have a frontier open-source model. And so, there's a lot of complexity of like how do we price all the different ingredients that go into making that happen. The other way that we price when we look at it through is also our cost structure to make them where of course when we have a task that takes 10 hours of human time and we're paying the human $150 an hour there might be a $1500 cost basis and so then it becomes a question of what margin do we want to run on top of that based on how differentiated and frontier that specific task is. But it's super wide range. Like we have tasks that range from $50 to $10,000.

问答:数据质量与比较 Q&A: Data Quality and Comparison

Host

你们如何看待数据质量?你提到利用人类专家来标注数据,你如何比较人类标注数据和前沿实验室、前沿联盟的判断数据?从你的角度来看,你们如何比较它们?

How do you think about data quality? You mentioned like you know utilize human expert to label data and how do you compare the human label data and you know the frontier lab you know frontier alliance the judgment data? How do you compare them from your opinion?

Brendan

所以第一个问题是我们如何看待质量?第二个问题是我们如何比较偏好标签的判断与自动评分器?

So first question was how do we think about quality? Second one was sort of how do we compare the judgment of preference labels to the auto graders?

Host

是的。

Yeah.

Brendan

所以核心点,通常当人们说数据质量时,他们指的是两件事。

So the core spot which is generally when people say data quality they're referring to two things.

质量:真实性与验证器准确性 Quality: Realism and Verifier Accuracy

Brendan

首先是真实性,其次是验证器的准确性。关于真实性,就像他们想要自动化经济中所有对应公司法的事物,对吧?所以问题是如何确保这真正反映真实律师环境中我们会看到的分布。这也是专家创建大纲并帮助指导数据整理过程的原因之一。环境、应用、任务的真实性都极其重要,还要细致理解驱动整个目标分布中真实性的分类法。第二部分涉及人们思考质量的另一种方式,即验证器的准确性。因为训练这类模型的方式是,你可能会生成 100 条 Kimi K3 的轨迹,然后用这个评分标准给所有这些轨迹打分。可以想象,模型可能走很多不同的路径。所以你要确保这个评分标准的打分方式与人类对这 100 条轨迹进行排序的方式一致。为此我们采用一个叫轨迹分析的过程,我们生成 10 条我们重点改进的模型的轨迹,然后给所有这些轨迹打分,并结合智能体式质量控制系统和一些人工审查,确保所有分数与目标一致。有时你也可以用人类反馈评估或偏好标签作为自动评分器的评估,这是另一种解决方式。

First is realism and secondly is accuracy of verifiers. On realism, it's just like they want to automate everything in the economy that corresponds to corporate law in this case, right? And so it's like how do we make sure that this actually reflects the real distribution of what we would see in a real lawyer's environment. And that's one of the reasons that experts create outlines and help to guide all the processes of the data curation. Realism of the environment, the apps, the tasks, everything is incredibly important, and also granularly understanding the taxonomy that drives that realism across the entire distribution that you're looking for. The second part of it relates to the other way people think about quality, which is the accuracy of the verifiers. Because the way that you would train one of these models is you might roll out 100 trajectories of Kimi K3 and then use this rubric to score all of those trajectories. And as you can imagine, there are so many different paths that a model can go down. And so you want to make sure that the way this rubric is doing the scoring is the same as if we were to just have human stack rank those 100 trajectories. And so what we do for that is a process called trajectory analysis, where we roll out 10 trajectories of the model that we're focused on improving, and then score all of those, and have some combination of agentic quality control systems and some human review go through to make sure that all of the scores align with the goals. And sometimes you can also use human feedback evals or preference labels as an eval for your auto grader, which is the other way related to that that you're able to solve for it.

Host

请继续。

Go ahead.

Host

你觉得合成数据生成在这一切中扮演多大角色,尤其是创建这些大型数据室?

Um, how much do you think synthetic data generation plays into all of this, especially like creating these large data rooms?

Brendan

有趣的是,我认为人们对“合成数据”的含义有很多误解,因为 RLVR 就是对合成数据的一场押注。基本上就是生成一堆合成模型轨迹,而不是让人类写 SFT,然后给所有这些轨迹打分,让模型从这些合成轨迹中学习。所以我认为这是合成数据的第一种使用方式。第二种方式是,模型在填充环境和创建任务方面发挥巨大作用,就像写法律备忘录的律师肯定应该用 Claude 或 ChatGPT 一样。构建这些数据室的专家肯定应该用 Claude、ChatGPT 或任何模型来帮助他们。模型有很多方式可以提高他们的效率。但人类仍然是这个过程中不可或缺且极具差异化的组成部分,因为几乎从定义上讲,你需要人类来衡量超出模型能力边界的东西。比如,你不能直接告诉模型“设计一个法律环境,然后告诉我哪些法律备忘录是好的,哪些是坏的”。那会非常嘈杂,而且没有清晰的信号。你需要有超出该模型能力边界的东西才能可靠地做到这一点。

So the fascinating thing is I think that there's been a lot of misinterpretation of what people mean when they say synthetic data because RLVR is a bet on synthetic data. It's basically let's roll out a bunch of synthetic model trajectories rather than having the humans write the SFT, and then let's score all of them, and let the models learn from all of these synthetic model trajectories. So I think that's the first way that synthetics get used. The second way is that models play a giant role in the way that we populate environments and create tasks, in the same way that a lawyer that is writing a legal memo should definitely be using Claude or ChatGPT to do that. The experts that are building out these data rooms should definitely be using Claude, ChatGPT, or whatever model to help them do that. And there's a lot of ways that the model can make them more efficient. But the reason that humans are still an essential component of the process that's incredibly differentiated is that you need humans almost definitionally to measure what is beyond the frontier of the model capabilities. Like the models, you can't just tell the model like come up with the legal environment and then tell me which of your legal memos are good and bad. It's super noisy and there's not clear signals associated with that. You need something that has capabilities beyond the frontier of that model to do so reliably.

Host

谢谢你做这个。我的问题是,强化学习环境现在似乎很流行,可能已经流行了大约一年。在那之前我并没有怎么听说过,都是人类专家标注。所以为什么现在都是强化学习环境?这是应用公司最需要考虑的事情吗?强化学习环境之后还有什么?

Thank you for doing this. Um, my question is RL environments seem like they're all the rage now and maybe have been for about a year. I hadn't really been hearing about them prior to that and it was all human expert labeling. And so kind of why, why, why is it all about RL environments now? Is that the most relevant thing for application companies to be thinking about? And is there anything after RL environments?

Brendan

嗯,我先说说为什么它变得流行,也许还有 2025 年我们看到的深度研究范式和环境范式之间的一些差异。然后我会谈谈展望未来,我们看到数据领域正在演变什么。具体来说,我认为深度研究环境之所以是第一个,是因为深度研究有搜索工具的使用。所以搜索是模型工作的环境中的工具。但专家不一定在填充应用。所以那是一种更轻量级的强化学习环境,他们只是创建相应的评分标准。再说一次,我之所以谈这些,是因为这已经是几年前的事了,所以不再是超级机密。然后对于 2025 年的应用趋势,我认为它变得巨大是因为人们意识到让模型有用的主要瓶颈是它们如何开始使用代码库中的所有上下文以及我们笔记本电脑上的所有工具,对吧?所以如果我们希望这出现在用户使用分布中,那么我们需要让它进入模型学习的数据分布。所以未来这三个类别中的多样性将继续大幅扩展,但会有一些变化。嗯,也许说出我们最关注的两个变化。第一个是超长时域。比如现在智能体大多没有被训练去做超过 10 小时的事情,我们需要开始构建可能需要人类 100 小时甚至 1000 小时才能完成的任务。所以那将是一个巨大的转变。我们看到的另一个重大转变是引入虚拟同事,这与那相对应。嗯,当人们思考他们的数据分布时,我最喜欢问的一个问题是,他们工作中执行的任务中有多少百分比需要与他人互动。大多数人会说大约 60% 或 70%。有些人说更多,有些人说少一点。但如果你将其映射到评估中有多少百分比衡量模型与他人互动的能力,那大约是 1%,也许 Tau bench 有一点。所以存在巨大的真实性差距,关于你如何实际衡量智能体在工作中与所有不同的人和智能体进行社交互动的表现。

Um, so I'll start with why it's become the rage and maybe some of the differences also between the deep research paradigm and the sort of environments paradigm as we saw it in 2025. And then I'll talk about looking forward what we see evolving in the data landscape. Specifically, I think that the reason the deep research environments were the first was because deep research had tool use with search. So search was the tool in the environment that the model would work in. But the experts would not necessarily be populating apps. So it was sort of a lighter version of an RL environment where they would just create rubrics corresponding to this. And again, I only talk about this stuff because it's a couple of years old at this point, so it's no longer super confidential. And then for the trend of apps in 2025, I think that really became giant because people realized that the primary bottleneck to making the models useful was how they started to use both all the context in the code base and all of the tools on our laptops, right? And so if we want this in the user distribution of usage, then we need to get it in the data distribution that the models are learning from. And so there's going to continue to be this giant scale up of diversity across all three of these categories on a going forward basis, but there's going to be some changes. Um, to maybe name two of those changes that we're thinking about the most. The first one is ultra long horizon. Like right now agents mostly aren't trained to do things that are over 10 hours, and we need to start building tasks for things that might take a human 100 hours or even 1,000 hours to do. And so that's going to be a giant shift. And then the other large shift that we're seeing is introducing virtual co-workers, which corresponds to that. Um, like one of my favorite questions to ask people when they're thinking about their data distribution is what percentage of tasks that they do in their job require interacting with other people. And most people would say like 60% or 70%. Some people say a lot more, some people say a little bit less. But then if you map that on to what percentage of evals measure how well the models can interact with other people, it's like 1%, maybe Tau bench has a little bit of this. And so there's this giant realism gap associated with how you actually measure how well agents engage in social interaction throughout all of the different people and other agents that they need to work with in their jobs.

Host

你谈到了评分标准生成,这就像是你为任务定制的一种验证方式?据我理解,这受限于专家。你有没有发现能够通过你的模型提供一些启发式方法,甚至后训练模型来扩展它?

You talked about rubric generation, which is like is it like a bespoke access of verify that you put task? And from what I understood, that's like bottleneck by experts. Have you found any success of being able to scale that up with your models giving you like some heuristic or even like post any models for

Brendan

我们发现,如果有一个 AI 副驾驶能够与专家合作创建任务和验证器,你可以让它高效得多。这样专家可以与轨迹对话,准确了解正在发生什么以及哪里出了问题。

We have found that you can make it a lot more efficient if you have an AI copilot that's able to work with the expert in creating the task and the verifier. So the expert can talk to the trajectory and understand exactly what's happening and where it's going wrong.

任务创建需人类参与 Task creation requires humans

Brendan

挑战在于,如果你想改进 FABLE,FABLE 无法可靠地写出它犯错的标准。它可能对一半错一半,而那种噪声从训练角度来看是不可用的。所以,这就是为什么最需要人的环节是任务创建。很多环境我们可以大量使用合成数据。让人来写大纲是有帮助的,因为他们熟悉环境,而且他们基于现实,符合真实分布。但任务本身往往确实需要人。少数例外是在代码领域,或者如果你是在蒸馏,比如如果你有一个比 Kimikaze 3 差的模型,那它肯定能从 Kimikaze 3 创建的任务中学习。所以,如果你这样做,是有一些例外的。

The challenge is just that if you're trying to improve FABLE, FABLE cannot reliably write out the rubric criteria for where it's making mistakes. It might get like half of them right and half of them wrong, and that amount of noise is unworkable from a training standpoint. And so, that's the reason that the process that requires humans the most is the task creation. Like a lot of the environments, we can use a lot of synthetic. It's helpful to have humans write the outlines because they're familiar with the environment and they're grounded in reality, a realistic distribution. But with the task, those tend to really require humans. With rare exceptions, in code or if you're sort of distilling from like if you have a model that's worse than Kimikaze 3, then it can definitely learn from tasks Kimikaze 3 is creating. So, there are some exceptions if you're doing it that way.

网络防御RL环境 Cyber defense RL environments

Host

嘿,我是 Cribl 的 Nikhil。当我们考虑某些可证明领域(如网络防御或事件响应)的强化学习环境时,模型或智能体试图在现有系统中寻找漏洞,你是只用人类来编写该环境或设置它,还是也用它来评分?有没有办法扩大规模?

Hey, this is Nikhil from Cribl. So, when we think about RL environments for certain provable domains like cyber defense or incident response, where the model or the agent is trying to find a flaw in an existing system, do you use humans for just authoring that environment or setting it up, or do you also use that for grading? Is there a way to scale that up?

Brendan

我实际上认为网络领域是你不一定总需要人来验证的领域之一,因为你可以有一个攻击者智能体和一个防御者智能体。我认为你说得对,在网络领域,你可以更多地让人来架构一个现实的环境并设置环境,因为你确实需要很多多样性。但在构建验证器方面,它就不那么需要人了。

I actually think cyber is one where you don't necessarily always need humans for the verifiers, because you can have an attacker and a defender agent. And I think you're right in saying that for cyber you can have humans more so architect what is a realistic environment and sort of set up the environment, because you do need a lot of diversity. But then it's less human-intensive with respect to building verifiers.

Host

谢谢。

Thanks.

基础模型局限性 Limitations of base model

Host

你怎么知道什么时候受限于基础模型?

How can you tell when you're limited by the base model?

Brendan

你指的是什么?

What do you mean by that?

Host

嗯,表面上你对这里所有模型使用相同的数据集,它们在这份列表上的最终表现大致相同。但也许如果你尝试一个更小的模型,这可能是一个好的起点,它的表现会低得多。原因是什么?仅仅是参数数量还是……

Well, ostensibly you're using the same data set for all these models here and they somewhat land around the same final performance on this list here. But maybe if you try a smaller model, which is maybe a good place to start, it would land much lower. What is the cause there? Is it just parameter count or...

Brendan

所以参数数量肯定会在模型表现如何方面起作用,就其可训练性而言。我认为主要要看的是 pass@16 和 pass@1 之间的差距。如果你有一个模型,你展开 16 条轨迹,它全部错误,那么模型基本上不太可能从中学习。也许你再展开 100 条轨迹,它只对一条。而理想情况是,pass@1 失败,但 pass@16 时,当你展开 16 条轨迹,它成功一两次,然后模型就能非常有效地从中学习。所以,这通常是我们用来判断基础模型需要多强才能有效从给定数据集中学习的启发式方法。

So the parameter count will definitely play a role in how effectively the model does insofar as how trainable it is. I think that the main thing to look at is generally the gap between the pass at 16 and the pass at 1. If you have a model where you roll out 16 trajectories and it gets all of them totally wrong, then it's sort of hopeless that the model is going to learn from that for the most part. Maybe you roll out another 100 trajectories and it gets one of them right. Versus if you have the ideal case, which is that you have pass at 1 it fails, but then pass at 16 when you roll out 16 trajectories it gets it right once or twice, and then the model is able to learn very effectively from that. So, that's generally the heuristic we would use for how strong the base model needs to be to effectively learn from a given data set.

数据合作建议 Advice on data partnership

Host

酷。也许最后一个问题,很快。对于与你合作数据的公司,你有什么建议?他们应该自己依赖什么作为他们强化学习后训练的一部分,而不是依赖你?互补的是什么?

Cool. Maybe one final question really quick. What advice do you have for companies as they partner with you on data that they should rely on themselves as part of their RL post-training versus relying on you? What's the complementary?

Brendan

嗯,我认为这就是为什么我们业务的很大一部分是定制数据,我们有专门团队,完全专属服务于关键客户,以确保我们构建出世界上最好的数据集,归他们所有。这让他们能够保持与此相关的竞争优势,同时也能受益于我们构建的所有基础设施。我认为有些公司试图在内部构建所有的人才网络和基础设施,但我认为如果你看看前沿实验室和最佳模型的做法,这很好地表明,与一个拥有所有这些规模经济、平台、人才网络等的合作伙伴合作,能带来巨大的规模经济。所以,总之,感谢邀请我。

Well, I think this is why a giant portion of our business is custom data, where we have teams that are siloed and fully exclusive to critical customers to make sure that we build out the best data sets in the world that they own. And that allows them to maintain their competitive advantage associated with this, while also benefiting from all of the infrastructure that we've built. And I think there are some companies that try to build out all of the talent network and infrastructure in-house, but I think if you look at what the frontier labs do and the best models do, it's a pretty good indication that there are so many economies of scale from working with a partner that has all of these economies of scale, the platform, the talent network, etc. So, anyways, thanks for having me.

互动版:逐字朗读 + 针对本期提问 →