后训练配方与 RLVR:对话 Nathan Lambert
Post-Training Recipes and RLVR with Nathan Lambert
内森·兰伯特 Nathan Lambert · Interconnects · 2025-07-31 · 约 79 分钟 · 原视频 ↗
打开互动全文版(中英对照 + 朗读 + 问答)→
本期速览 · Overview
Nathan Lambert 讨论 Tulu、RLVR 以及开源后训练方法的演变。
Nathan Lambert discusses Tulu, RLVR, and the evolution of open-source post-training methods.
要点 · TL;DR
- 后训练将复杂的行业配方压缩为可操作的开源方法。
Post-training compresses complex industry recipes into tractable open-source methods. - RLVR 使用可验证奖励来优化数学、代码和精确指令遵循。
RLVR uses verifiable rewards for math, code, and precise instruction following. - 开源模型需要更好的工具使用训练,尤其是搜索和私有数据。
Open models need better tool-use training, especially for search and private data.
核心观点 · Key points
- 后训练是将复杂的行业配方压缩为可操作的开源方法。
Post-training is compressing complex industry recipes into tractable open-source methods. - RLVR 使用可验证奖励来处理数学、代码和精确指令遵循。
RLVR uses verifiable rewards for math, code, and precise instruction following. - 过度优化发生在 RL、RLHF 和 RLVR 中;模型会利用最简单的信号。
Overoptimization occurs in RL, RLHF, and RLVR; models exploit the easiest signal. - 前沿实验室仍使用人类偏好数据;仅靠 AI 反馈可能不够。
Frontier labs still use human preference data; AI feedback alone may not suffice. - 开源模型需要更好的工具使用训练,尤其是搜索和私有数据。
Open models need better tool-use training, especially for search and private data. - 扩展 RL 需要大规模多领域数据集、难度过滤和长运行时间。
Scaling RL requires large multi-domain datasets, difficulty filtering, and long run times.
反共识 · Contrarian takes
- 混合推理模型可能是暂时的;像 O3 这样的纯推理模型可能占主导。
Hybrid reasoning models may be temporary; pure reasoning models like O3 could dominate. - 推理时缩放图是一种“错觉”;它们看起来比实际更容易控制。
Inference-time scaling plots are a 'scope'; they look easier to control than they are. - 推理的并行计算并非变革性;更好的验证器更重要。
Parallel compute for inference is not transformative; better verifiers matter more. - 模型规范比宪法更有用,用于透明度和开发者利益。
Model specs are more useful than constitutions for transparency and developer benefit. - 开源模型可能在个性化和角色训练上获胜,而不仅仅是基准测试。
Open models may win on personalization and character training, not just benchmarks. - Meta 的恐慌按钮是真实的;人才比 GPU 便宜,所以他们雇佣顶尖研究人员。
Meta's panic button is real; talent is cheaper than GPUs, so they hire top researchers.
本期章节 · Chapters(共 19)
- 引言与 Tulu 概述 Introduction and Tulu Overview
- RLVR 与多跳工具使用 RLVR and Multi-Hop Tool Use
- Chatbot Arena 与人类数据价值 Chatbot Arena and Human Data Value
- RLHF 书与 RLVR 书对比 RLHF Book vs RLVR Book
- 混合模型与纯推理模型 Hybrid vs. Pure Reasoning Models
- 模型中的搜索集成 Search Integration in Models
- 强化学习与工具使用 RL and Tool Use
- 工具与模型不确定性 Tools and Model Uncertainty
- 多工具 RL 与研究方向 Multi-Tool RL and Research Directions
- AI 影响力:工件 vs 论文 Impact in AI: artifacts vs. papers
- AI 中的计划与规则 Plans and Rubrics in AI
- 并行计算与验证器 Parallel compute and verifiers
- RL 中的奖励黑客 Reward Hacking in RL
- 训练时间约束与长推理 Training time constraints and long inference
- 有趣研究方向:角色训练与个性 Interesting research directions: character training and personality
- 模型路由与开放模型 Model routing and open models
- 形态因素与市场规模 Form factors and market size
- Meta 的紧急按钮与人才策略 Meta's panic button and talent strategy
- 打造美国版 DeepSeek 与 AI2 之路 Building the American DeepSeek and AI2's path
阅读全文双语转录 →