Brendan discusses how RL environments are transforming agentic data, with experts building worlds, apps, and tasks to train frontier models.
要点 · TL;DR
强化学习环境是后训练的新前沿,超越了众包行为克隆数据。 RL environments are the new frontier for post-training, moving beyond crowdsourced behavior cloning data.
在大多数领域,人类对于创建验证器和评分标准至关重要,因为模型无法可靠地自我评分。 Humans are essential for creating verifiers and rubrics in most domains, as models cannot reliably grade their own work.
数据是 AI 战略的关键差异化因素,与算力和算法并列。 Data is a key differentiator for AI strategy, alongside compute and algorithms.
核心观点 · Key points
强化学习环境是后训练的新前沿,超越了众包行为克隆数据。 RL environments are the new frontier for post-training, moving beyond crowdsourced behavior cloning data.
在大多数领域,人类对于创建验证器和评分标准至关重要,因为模型无法可靠地自我评估。 Humans are essential for creating verifiers and rubrics in most domains, as models cannot reliably grade their own work.
数据与算力和算法一样,是AI战略的关键差异化因素。 Data is a key differentiator for AI strategy, alongside compute and algorithms.
定制数据合作使公司能够拥有自己的智能,同时利用基础设施和人才网络。 Custom data partnerships allow companies to own their intelligence while leveraging infrastructure and talent networks.
向智能体数据的转变包括在所有经济领域扩展环境,专家工时快速增长。 The shift to agentic data includes scaling environments across all economic domains, with expert hours growing rapidly.
反共识 · Contrarian takes
合成数据常被误解;强化学习验证与奖励本身就是对合成轨迹的押注,而非人类编写的监督微调。 Synthetic data is often misinterpreted; RLVR itself is a bet on synthetic trajectories, not human-written SFT.
大多数评估未能衡量社交互动,尽管60-70%的工作需要它,造成了现实性差距。 Most evals fail to measure social interaction, despite 60-70% of jobs requiring it, creating a realism gap.
前沿实验室并非唯一受益者;像GLM和Kimi这样的开源模型正在追赶,使应用层公司受益。 Frontier labs are not the only ones benefiting; open-source models like GLM and Kimi are catching up, enabling application-layer companies.
强化学习环境的下一转变是超长时域任务(100-1000小时)和虚拟同事。 The next shift in RL environments is ultra-long-horizon tasks (100-1000 hours) and virtual co-workers.
在网络安全领域,验证器可以通过攻击者-防御者智能体自动化,减少人工评分需求。 For cyber domains, verifiers can be automated with attacker-defender agents, reducing human grading needs.
本期章节 · Chapters(共 17)
引言与背景Introduction and Context
RL环境概览RL Environments Overview
RL环境组成Components of an RL Environment
人类为何关键Why Humans are Essential
RL环境示例Example RL Environment
构建RL环境与数据室Building RL Environments and Data Rooms
后训练示例与泛化Post-training Example and Generalization
客户合作与未来机遇Working with Customers and Future Opportunities
高质量数据集整理方法Ways to Curate High-Quality Data Sets
期待与问答引言Excitement and Q&A Introduction
问答:数据定价Q&A: Pricing Data
问答:数据质量与比较Q&A: Data Quality and Comparison
质量:真实性与验证器准确性Quality: Realism and Verifier Accuracy