Tulu 3:大语言模型开源后训练技术

Tulu 3: Open Source Post-Training for LLMs

内森·兰伯特 Nathan Lambert · The Cognitive Revolution · 2024-11-21 · 约 110 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Nathan Lambert 讨论 Tulu 3,一个使用相同基础模型匹配 Meta 后训练性能的开源项目,涵盖 SFT、RLHF 和 RLVR 技术。

Nathan Lambert discusses Tulu 3, an open-source project that matches Meta's post-training performance using the same base model, covering SFT, RLHF, and RLVR techniques.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 33)

全文 · Full transcript(中英对照)

开场与赞助 Introduction and Sponsor

Host

可能不值得花所有时间在偏好调优上,当你可以只是制作更好的数据和更好的流程时,这就是 Tulu 3 的意义所在。受到 Llama 报告和 Chatbot Arena 转变的启发,它再次变成了曲棍球棒曲线,我们有这些增量分数,而 OpenAI 和 Google 的分数在飙升。理念是:我们如何理解开放团队应该做什么,以及在显著增加后训练复杂性时有哪些挑战需要攻克?如果你能让人类和 LLM 都参与偏好数据生成,哪些任务交给人类,哪些交给 LLM?我认为这解决了很多问题。有些事我们肯定希望人类给出答案,但也有很多机械性任务可以外包给 LLM。还有很多预训练没有真正触及,我认为机会很大,因为关键是如何培养性格。性格是你的模型没有估值的东西。我们的模型,如果和 Claude 相比,性格一致性较差。无论你对选举结果感受如何,我想我们都期待结束那些没完没了的筹款邮件和短信。不幸的是,对我来说,即使在选举结束一周多后,它们也没有停止。更不用说那些随时从四面八方涌来的普通商业垃圾邮件了。事实证明,这些噪音大多来自数据经纪人。这些公司不仅收集你的联系方式,还收集从社会安全号码、财务记录到在线购物习惯的一切,而且他们现在与保险公司合作,这可能会影响你的费率。这就是为什么我很高兴现在使用 Incogni。Incogni 联系五类数据经纪人——营销、招聘、财务信息、风险缓解和人肉搜索网站——并要求他们删除你的信息。然后他们持续监控,防止数据重新收集。用 Incogni 拿回你的个人数据。他们的家庭和朋友计划可以保护最多四名家庭成员。他们提供 30 天退款保证。所以用 Incogni 拿回你的个人数据。使用优惠码 Revolution,在下方链接获得年度计划 60% 折扣。网址是 Incogni.com。

It's probably not worth the effort to spend all your time on preference tuning when you can just be making better data and better pipelines, which is what Tulu 3 is about. Inspired by the transition we're seeing with the Llama report with Chatbot Arena, it's turned into a hockey stick again where we have these incremental scores and OpenAI and Google are skyrocketing their scores. The philosophy is: how do we try to understand what the open groups should be doing and where there are hills to climb when you're increasing the complexity substantially of post-training? If you can have humans and LLMs do preference data, which do you send to humans versus LLMs? That, I think, solves a lot of the problems. There are definitely things that we want humans giving the answer on, but there are a lot of mechanical tasks that we can outsource to LLMs. There's a lot more pre-training that is not really touched, and I think that the opportunity is high because the big thing is how do you develop character? Character is something that you don't have a valuation for in your models. Our models, if you compare them to Claude, will not have as consistent of a character. Regardless of how you felt about the outcome of the election, I think we were all united in looking forward to an end to the constant fundraising emails and text messages. Unfortunately for me, they haven't stopped even now, more than a week after the election. And that's to say nothing of the normal commercial spam coming at me from all directions at all times. It turns out that most of this noise is caused by data brokers. These companies aren't just collecting your contact details; they're gathering everything from your Social Security number and financial records to your online shopping habits, and they're now working with insurance companies, which could potentially impact your rates. That's why I'm excited to now be using Incogni. Incogni contacts five types of data brokers—marketing, recruiting, financial information, risk mitigation, and people search sites—and demands that they remove your information. Then they continue monitoring on your behalf to prevent data recollection. Take your personal data back with Incogni. Protect up to four family members with their family and friends plan. They offer a 30-day money back guarantee if you're not satisfied. So take your personal data back with Incogni. Use code Revolution at the link below and get 60% off an annual plan. That's Incogni.com.

欢迎与嘉宾介绍 Welcome and Guest Introduction

Host

欢迎回到《认知革命》。今天的嘉宾是 Nathan Lambert,他是广受欢迎的 Interconnects 通讯的作者,也是艾伦人工智能研究所的机器学习研究员。今天,该研究所发布了 Tulu 3,这是迄今为止最全面的开源项目之一,旨在传播前沿大语言模型后训练技术的理解和实践。通过系统性地使用相同的 Llama 基础模型来匹配 Meta 的后训练性能,并公开分享所有发现和数据,Nathan 和艾伦研究所的团队揭示了大语言模型开发中历史上最不透明的方面之一。这次对话是你在网上能找到的关于这个话题最详细的讨论之一。今天我们涵盖了后训练技术的全谱系,包括监督微调、多种基于偏好的强化学习,以及一种名为可验证奖励强化学习的新技术,该技术奖励模型准确回答具有客观正确答案的问题。每一步,我们都深入探讨使这些技术生效的实际细节、相关的算力需求、数据生成策略以及每项技术的价值,还有用于在探索广阔训练配方空间时衡量性能的实验设计。我们甚至探索了一些迷人的涌现行为,这些行为呼应了我们最近从 OpenAI 的 o1 中看到的前沿推理能力。Nathan 对这项工作的技术和组织挑战的坦诚讨论,包括他承认某些方面尚未被很好理解的时刻,为开发最先进模型所需的条件提供了非常有用的视角。他们最终仅用 10 到 15 人的团队就成功匹配了 Llama 的性能,这一事实使得他们的方法值得仔细研究。现在,这种开放开发能在未来几代模型中持续多久,在我看来仍是一个悬而未决的问题。即使有亿万富翁国家的支持,人类生成的偏好数据和标注对艾伦研究所来说成本过高,而且如果前沿开发者效仿 OpenAI 不发布 o1 风格推理轨迹的做法,这个项目中使用的合成数据生成技术未来可能可用也可能不可用。话虽如此,Nathan 预计社区最终会找到解决办法。就在发布之前,我们看到中国 AGI 公司 DeepSeek 宣布了一个新的 o1 风格模型 DeepThink,它似乎确实展示了其工作过程,并且据报道将在不久的将来开源。这一发展可能会将问题从这种开放强化学习项目能否继续转变为是否应该继续,并且肯定会让任何认为西方公司可以从此一路领先中国公司到 AGI 的人停下来重新思考他们的假设。无论如何,这一集是我们节目迄今为止产出的最高杠杆内容之一。它充满了具体的实践见解,这些见解以前分散在文献中或仅仅隐藏在闭门之后,我非常感谢 Nathan 如此开放且技术细节丰富的对话。如果你觉得这个节目有价值,我们很感激听众花时间在线分享,在 Apple Podcasts 或 Spotify 上留下评论,或者在 YouTube 上留言。你的反馈也随时欢迎通过我们的网站 cognitive revolution.ai 提供。

Welcome back to the Cognitive Revolution. Today my guest is Nathan Lambert, author of the popular Interconnects newsletter and machine learning researcher at the Allen Institute for AI, which today is releasing Tulu 3, one of the most comprehensive open source efforts to diffuse the understanding and practice of frontier post-training techniques for large language models that we have seen to date. By systematically working to match Meta's post-training performance using the same Llama base model and sharing all of their findings and data publicly, Nathan and the team at the Allen Institute have illuminated what has historically been one of the most opaque aspects of large language model development. This conversation represents one of the most detailed discussions of this topic that you can find anywhere online. Today we cover the full spectrum of post-training techniques, including supervised fine-tuning, multiple flavors of preference-based reinforcement learning, and a new technique called reinforcement learning from verifiable reward, which rewards the model for accurately answering questions with objectively correct ground truth answers. At each step, we dig into the practical details that make these techniques work, the associated compute requirements, data generation strategies, and the value derived from each, as well as the experimental designs that are used to measure performance while exploring the vast space of possible training recipes. We even explore some fascinating emergent behaviors that echo the frontier reasoning capabilities that we've recently seen from OpenAI's o1. Nathan's frank discussion of both the technical and organizational challenges of this work, including in a few moments where he acknowledges aspects that are not yet well understood, provides a super useful window into what it takes to develop state-of-the-art models. The fact that they ultimately succeeded in matching Llama performance with a team of just 10 to 15 people makes their approach one to study closely. Now, how long this level of open development can continue into future generations of models remains, in my mind, an open question. Even with billionaire state backing, human-generated preference data and annotations are cost prohibitive for the Allen Institute, and the synthetic data generation techniques used in this project may or may not be available going forward if frontier developers follow OpenAI's lead and choose not to release their o1-style reasoning traces. That said, Nathan expects that the community will ultimately figure something out. And just before publishing, we've seen Chinese AGI company DeepSeek announced a new o1-style model called DeepThink, which does seem to show its work and is reportedly going to be open sourced in the near future. A development that could shift the question from whether or not such open reinforcement learning projects can continue to whether or not they should, and which should definitely cause anyone who thinks that Western companies can maintain a comfortable lead over Chinese companies from here all the way to AGI to stop and rethink their assumptions. In any case, this episode is one of the highest leverage pieces of content we've produced on this show to date. It's full of concrete practical insights that were previously scattered throughout the literature or simply hidden behind closed doors, and I am really grateful to Nathan for such an open and technically detailed conversation. If you're finding value in the show, we appreciate it when listeners take a moment to share it online, post a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. And your feedback is always welcome via our website cognitive revolution.ai.

介绍与艾伦研究所背景 Introduction and Allen Institute background

Host

或者在你喜欢的社交网络上私信我。希望你喜欢这次与艾伦人工智能研究所 Nathan Lambert 的深度对话,深入探讨大语言模型后训练的前沿。Nathan Lambert,来自艾伦研究所,Tulu 的创造者,欢迎来到认知革命。

or by dming me on your favorite social network with that I hope you enjoy this super deep dive into the frontiers of large language model post training with Nathan Lambert of the Allen Institute for AI. Nathan Lambert from the Allen Institute, creator of Tulu, welcome to the cognitive Revolution.

Nathan

谢谢邀请。Nathan 平方播客终于实现了。真不敢相信花了这么久,但我很高兴这一天终于来了。

Yeah, thanks for having me. The Nathan squared pod has happened. I can't believe it's taking this long, but I'm excited that the day is finally here.

Host

要聊的内容很多。我就快速提问,看看能聊多远。首先,跟我们说说艾伦研究所。你们发布了很多东西。我想了解一下背后的理念、资金、GPU 情况——你们是 GPU 富裕还是 GPU 匮乏?以及你们打算如何改变世界?

So lot of ground to cover. I'm just gonna fire questions at you rapid fire and let's see how far we can get. For starters, tell us about the Allen Institute. You guys have put out a lot of stuff. I would love to just get a little context on kind of the philosophy behind it, the funding, the GPUs. Are you GPU rich or GPU poor? And how are you looking to put a dent in the universe?

Nathan

是的,艾伦研究所,我是新来的。我加入刚一年多。研究所实际上已经成立 10 年了。它由保罗·艾伦资助,这对科技界和了解西雅图艾伦研究所的人来说并不意外,因为名字就在那里。首任 CEO Oren Etzioni 也是华盛顿大学的教授,有很多华大关系。我想他和保罗·艾伦是朋友。研究所由此发展而来,旨在为公共利益做 AI 研究,保持非常开放、非常混合的学术产业风格。这大致是第一个十年的故事。现在 Oren 已经离开,离职过程几年前就开始了,新任 CEO 是 Ali Farhadi,也是华大教授,之前曾在苹果工作。他的团队最著名的工作我认为是计算机视觉领域的 YOLO 系列。研究所正在引入新鲜血液,进行转型,构建开放语言模型。我认为很多机构意识到,要在 AI 领域有可信度,你需要在语言建模领域有一定信誉。现在对于 AI 研究来说,就是要表明“我们也能做到”。这缩小了范围,专注于发布优秀模型和构建模型。我认为 OLMo 是这一转变的开端,这是一个很长的故事,我大部分没有参与。我是在那年晚些时候加入的,模型在 2024 年初发布。2024 年则是加倍投入、扩大项目规模的一年,当你有更多人的时候就能做这些。我想我们会讨论后训练方面的事情,OLMo 中也有很多预训练的内容。但这实际上就是开源 AI 中的学术项目在有更多人参与时的样子。全年来看,我们的很多项目都在变大。我们看到了 Molmo,这是一个视觉模型,主要由视觉团队负责,但也涉及语言预训练团队和其他方面。这些项目越来越大,试图在尽可能开放的同时扩大研究范围。所以我们说什么都没有限制。归根结底,非营利组织就是要讲好故事,而故事才能真正改变政策制定和人们看待世界的方式。所以最终,艾伦研究所需要构建好的模型,以便我们能够传达开源的好处。把非营利组织视为讲故事的工作有点简化且有点愤世嫉俗,但当你的资金主要来自科技亿万富翁的遗产时,你不得不实话实说。但势头是好的,所以在这个领域工作很有趣,看到事情暂时还在继续。

Yeah, so Allen Institute, I'm new blood here. I joined just over a year ago. The Institute is actually 10 years old. It was founded, it's funded by Paul Allen, which is not surprising for people in the tech scene and knowing the Allen Institute in Seattle given the name. The original CEO, Oren Etzioni, also was a professor at UW, a lot of UW ties. I think he was buds with Paul Allen. This kind of grew out of that to just do AI research for the common good, make things very open, very hybrid academic industry. That is kind of the story of the first decade roughly. Now there's like Oren left, the process of leaving started a few years ago, and there's new CEO Ali Farhadi, also Professor at UW, was at Apple before. Most prominent work from his group I think is the YOLO line of work in computer vision. It's kind of bringing in new blood to make this transition to build open language models. I think a lot of institutions realize that in order to be credible in AI, you need to have some amount of credit in the language modeling space. Right now for AI research, it is to be like look we can do this too. It is a narrowing of scope to try to release great models and build models. I think OLMo is the start of this, which was a very long story that I was not a part of most of. I joined late in the year and launched in early 2024. Then 2024 has been the year of doubling down on this and scaling up projects that you can do when you have more people. I think we'll talk about the post-training side of things, and there's plenty on pre-training and stuff in OLMo as well. But it's really like what academic projects in open source AI look like when you can bring more people in. Throughout the year, a lot of our projects have been getting bigger. We've seen like Molmo, the vision model, which is primarily on the vision team but gets involved with the language pre-training team and other things. These projects are getting bigger and trying to increase the scope of research while doing as much as we can in the open. So no constraints on what we can say. At the end of the day, nonprofits are about telling a good story, and stories are what can actually change how policy happens and how people view the world. So at the end of the day, Allen Institute needs to build good models so we can communicate why open source has benefits. It's kind of reductionist and a little cynical to view nonprofits as storytelling works, but when your money mostly comes from a tech billionaire's estate, you got to say some of it as it is. But the momentum is good, so it's fun to be in the space and see things continuing for the time being.

Host

我认为故事在生活整体中,尤其是在 AI 时代,确实非常重要。我觉得我们正在召唤这种疯狂的新外星智能,但在大多数角落,我们缺乏一个积极的愿景,关于我们想用它做什么,想让它对生活产生什么影响。有很多非常笼统的讨论,比如“当我们拥有想要的一切时,那该多好”,但问题是,我们能更具体一点吗?所以我确实理解讲故事在这个领域的重要性。你提到了预训练和后训练。就在我们录制今天这个节目时,你刚刚发布的项目(我们提前几天录制,但会尽量在录制时发布)叫做 Tulu。这确实是对后训练技术的深入探讨。请给我们一些背景,为什么做这个项目,它的目标是什么?然后我真的想深入探讨实际的方法论,你学到了什么,并尝试建立自己对后训练前沿的直觉。

I think stories are honestly a super important part of life in general and the AI moment in particular. It feels to me like we are summoning this crazy new alien intelligence and really lack in most corners a positive vision for what we want to do with it, what impact we want it to have on life. There's a lot of very general talk about won't it be great when we have everything we want, and it's like yeah, could we put a little more detail on that? So I do appreciate the importance of storytelling in this space. You mentioned pre-training and post-training. The project that you've just put out today as we record this, a few days in advance but we'll try to publish it right around the time of recording, is called Tulu. It is really a deep dive into post-training techniques. Give us a little bit of context for why this project, what the goals of it are, and then I really just want to go deep on the actual methodology, what you've learned, and try to develop my own intuitions for the state-of-the-art in post-training.

Nathan

是的,我认为这个故事与我们刚才谈到的 AI2 很契合。AI2 的开放后训练方法已经以 Tulu 这个品牌存在了一年多,差不多一年半了。我记得 Tulu 1 是在 Open Assistant 刚出现的时候,当时的问题是:通过混合 2023 年人们发布的所有这些流行数据集和小模型,我们能做什么?如何系统地混合指令数据集,将它们与开放资源结合起来?Tulu 2 是在 DPO 非常流行的时候,那时我们展示了可以将 DPO 扩展到 700 亿参数,并在开放方法中持续获得偏好微调的好处。然后我们做了 Tulu 2.5,试图回答 PPO 与 DPO 的问题。简而言之,如果你参数调得好,PPO 可能稍微好一点,但可能不值得把所有时间花在偏好调优上,而应该专注于制作更好的数据和更好的流程。这正是 Tulu 3 的内容。它受到我们看到的转变的启发,比如 Llama 报告,Chatbot Arena 再次变成曲棍球棒曲线,分数在增长,而 OpenAI 和 Google 的分数再次飙升。Patron 论文,苹果有一篇论文。这些后训练方法比“一个大 SFT 步骤,然后在一些聊天数据上做 DPO 就完事”要复杂得多。我们的理念是,我们知道这些实验室在这样做,我们如何尝试理解开放团体应该做什么,以及在大幅增加后训练复杂性时有哪些困难需要克服?比如 LL 3 有数百人参与,其中很大一部分在后训练上。我们只有 10 到 20 人。问题是,学者们实际上能做什么来进行现代后训练,策划新数据,使用新算法,将事情序列化,但超越那种仅仅提高 Alpaca 或氛围分数的范式,并真正专注于改进数学、改进指令遵循?代码我们做了一些,但不是最大的重点。我认为我们的代码分数很早就基本饱和了。

Yeah, so I think this kind of story fits well with what we were talking about with AI2. AI2's kind of post open post-training recipes have been under this Tulu brand for over a year, almost a year and a half. I think Tulu 1 was back when Open Assistant was new, like what can we do by mixing all these popular datasets that people are putting out in 2023, all these datasets and small models? How do you systematically mix instruction datasets to combine them with open resources? Tulu 2 was around when DPO was very popular, and that was when we showed that you can scale DPO to 70 billion parameters and continue getting benefits from preference fine-tuning in open recipes. Then we went away and did Tulu 2.5, which was kind of like trying to answer the PPO versus DPO question. The TL;DR was if you tune your parameters right, PPO is probably better a bit, but it's probably not worth the effort to spend all your time on preference tuning when you can just be making better data and better pipelines. Which is what Tulu 3 is about. It's kind of inspired by the transition we're seeing with the Llama report, with Chatbot Arena turning into a hockey stick again, where we have incremental scores and OpenAI and Google are skyrocketing their scores again. Patron paper, Apple had a paper. These post-training recipes are just much more complicated than big SFT step, do DPO on some chat data, and be done. The philosophy is kind of like we know these labs are doing it, how do we try to understand what the open groups should be doing and where there are hills to climb when you're increasing the complexity substantially of post-training? Like LL 3 has hundreds of people on it and a large amount on post-training. We have 10 to 20 on ours. It's like what could academics actually do to do modern post-training and curate new data, use new algorithms, sequence things together, but kind of move beyond this paradigm of just increasing like Alpaca or vibe scores and try to get really specific on improving math, improving instruction following, code is something we did a bit but it wasn't the biggest focus. I think our scores for code were mostly saturated pretty early.

后训练与开放模型格局变化 Shifting post-training and open model landscape

Host

但把后训练这个领域转向一个更大的学术进步领域,老实说,人们发布模型和数据集的频率已经放缓了。DPO 时代对于看到开放的后训练模型来说是非常有趣的。那简直是疯狂——每周都有一个新的 DPO 模型在某个方面达到最先进水平。但那是在去年 12 月、去年 1 月,现在感觉更多的是大玩家主导 AI 叙事,而我们正试图重新设定这一点。

But kind of shifting this post-training into a much bigger domain of academic progress, and honestly, there's been a slowdown in how often people are releasing models and datasets. The DPO era was a very fun one for seeing open post-training models. It was bonkers—every week there was a new DPO model that was state-of-the-art on something. But that was like last December, last January, and it feels much now it's much more big players dominating the AI narrative, and it's trying to reset that.

Nathan

是的,所以我在你分享的材料中注意到的一点是,这个项目的目标是击败 Llama 3.1,而且如果我没理解错的话,你们是从几个不同的基础模型开始的,但其中一个基础模型基本上是 Llama 3 基础版,对吧?所以你基本上是在说:好吧,如果我们采用 Meta 最初使用的相同预训练版本,我们能否在后训练方面与你们匹敌或超越你们,进行后训练的对等比较?他们显然发布了一份大报告,分享了很多他们做了什么,但我相信他们没有分享太多或任何后训练数据。请多告诉我们一些:他们分享或不分享什么?

Yeah, so one of the things that I noticed in the materials that you shared is that the goal of the project was to beat Llama 3.1, and you're starting, if I understand correctly, with a couple different base models, but one of the base models is Llama 3 base basically, right? So you're essentially saying: okay, if we take the same pre-trained version that Meta started with, can we match or exceed you, hopefully with post-training, on a post-training apples-to-apples basis? They obviously put out a big report and shared a lot about what they did, but I believe they don't share much or any of the post-training data. Give us a little bit more: what do they share or not share?

Host

我会说他们分享了方法的概要。他们没有很多超参数,但如果不了解他们的代码库之类的东西,就不太清楚。所以他们分享了一个高层次的概要,比如‘这是我们的方法,这是我们做的各种事情’,但没有具体细节,当然也没有数据。所以我们后训练最大的不同点是我们发布了数据。我想我们为这个项目构建了三到六个新数据集,我们把这些与已有的数据集结合使用,并且我们会一次性发布它们。但你说的确实没错:这里的开发模式是,我们有一套关心的评估,涵盖事实性、知识、数学、代码、指令遵循、安全性。而目标,主要的开发目标,是 Llama 3.1 8B,我们实际上训练了大约一千个 8B 模型,试图让这些分数在 8B 上更高,然后在 70B 上验证。我们也在添加 8B 模型,这是我们的主要工作,也就是确保配方可以迁移。所以出于这个原因,这是一个基本目标:好吧,为了知道你在正确的范围内,你需要击败那些公认相当不错的人。

I would say they share an outline of their method. They don't have a lot of hyperparameters, but it's not really clear without knowing their codebase and things like that. So they share a high-level outline like 'here is our approach and here are the types of things that we do,' but not the specific things that we do, and definitely not the data. So the biggest differentiator for our post-training is that we release data. I think we built like three to six new datasets just for this project, and we're using those with a combination of already existing datasets, and we'll release them all at once. But really what you're saying is right: the development model here was we had a set of evaluations we cared about that tried to cover things like factuality, knowledge, math, code, instruction following, safety. And the goal, the primary development target, was Llama 3.1 8B, where we effectively trained about a thousand 8B models to try to get these scores to be higher at 8B, and then validate them at 70B. And we're also adding 8B models, which is our main thing, which is trying to make sure the recipe translates. So for that reason, it's kind of a basic goal: okay, in order to know you're in the right ballpark, you need to beat these people that are known to be pretty good.

Nathan

我想我谈过——关于 Llama 基础模型和指令模型的有趣之处在于,闭源实验室也可以在其上验证他们的流程。所以闭源实验室可以拿 Llama 3 基础版,然后用他们的后训练,与 Llama 也发布的指令模型进行比较。所以从一些业内人士那里听说,他们也说‘是的,Llama 3 指令就像 3.1 指令,最新的东西相当厉害,这是一个非常好的后训练。’我认为显然 OpenAI 的基础设施在 Llama 基础模型上不会那么好用,但 3.1 模型尤其实际上非常好。我认为很明显,在某些方面他们很仓促,或者试图专注于系统而不是新颖性。Llama 3.1 指令主要是 SFT 和 DPO,我知道他们现在正在扩展到更多东西。而我们规模更小,行动更快。我认为我们实际上会做一些事情,我打赌会与 Llama 4 指令或 Llama 3.5 指令之类的东西非常相似。但这就是主题:我们想击败他们的数字,我们做类似的事情,我们有规模更小的优势,我们可能可以更聪明一点。但那是开发目标。而且确实如此——我可以看看日子,我不记得具体日期——就像如果我们四个月前开始这个项目,哦,几个月后我们超过了 Llama 3 指令,然后大约一个半月后我们超过了 Llama 3.1 指令 8B,然后几周后在 70B 上就像节拍器一样稳定。然后每次都是‘哦,哇,这真的在发生。’所以看到这个很有趣,我们非常兴奋地看到人们在此基础上构建什么。我不喜欢——我们稍后可能会讨论从基础版微调与从指令版微调,这是另一个分支。我认为从基础版微调是非常学术性的,对于后训练理解如何从基础模型做到这一点非常重要,而从指令版微调是一个非常重要的领域,尤其是对于新公司和特定领域的事情。因为如果 Llama 要发布一个非常强大的指令模型,你不需要做整套 2-3 套件来让它在你特定领域表现良好。就像我们正在做其中的一部分,但最终那里需要研究。我认为在录制当天,Nexusflow,一家来自伯克利一些人的初创公司,发布了 Starling 模型的团队,他们发布了另一个指令微调版本,我们看到 Neotron 在 Llama 指令上的微调最近非常火爆。所以有一些信息说我们不做这种指令微调,但这是另一个领域。我只是认为,就基本的长期生态系统建设而言,从基础版开始仍然是需要最多透明度的。

I think I've talked about it—the fun thing about Llama base models and instruct is that the closed labs can also validate their pipelines on it. So the closed labs can take the Llama 3 base and then use their post-training and compare to the instruct model that Llama also released. So hearing from some people in industry, they're also like 'yeah, Llama 3 instruct was like 3.1 instruct, the latest stuff was pretty cracked, it was a pretty good post-train.' I think obviously OpenAI's infrastructure isn't going to work as well on Llama base models, but the 3.1 models particularly actually were very good. I think it's clear that there were some ways that they were rushed or trying to focus on systems rather than novelty. It was mostly SFT and DPO for the Llama 3.1 instruct, and I know they are expanding into more things now. And we're kind of being smaller and moving fast. I think we're actually going to do some things that I bet will look pretty similar to the Llama 4 instruct or a Llama 3.5 instruct type thing. But that's kind of the theme: we want to beat their numbers, we're doing similar things, we have the advantage of being smaller, we can be a little bit more clever maybe. But that was the development target. And it really is like—I could look at the days, I don't have the dates in mind—it's like if we started this project four months ago, it's like oh, a few months in we passed Llama 3 instruct, and then like a month and a half later we passed Llama 3.1 instruct 8B, and then like a couple weeks later at 70B it's like just metronomic. And it's like oh my, every time it's like oh wow, this is actually happening. So it's fun to see, and we're very excited to see what people build on it. And I don't like—we might have a discussion later on which is fine-tuning from base versus fine-tuning from instruct, which is another fork. And I think fine-tuning from base is very academic and very important to post-training to understand how to do this from base models, and fine-tuning from instruct is a very important area especially for new companies and domain-specific things. Because if Llama is going to put out a very strong instruct model, you don't have to do the whole 2-3 suite to make it good at your specific domain. It's like we're doing part of this, but eventually there's going to need to be research there. I think on the day of recording, Nexusflow, which is a startup that came from some people from Berkeley, the group that released the Starling models, they released another fine-tune on instruct, and we saw Neotron fine-tune on Llama instruct go very viral recently. So there is some messaging that we're not doing this fine-tune on instruct, but it is another area. I just think in terms of fundamental long-term ecosystem building, the from-base is still the thing that needs the most transparency.

介绍与赞助商 Introduction and Sponsors

Host

革命你可以免费开始,使用我们的链接还能支持节目,所以今天和我一起在 notion.com/cognitive-revolution 试试 Notion AI 吧。The Cognitive Revolution 由 Shopify 赞助。多年来我一直知道 Shopify 是全球领先的电商平台,但直到最近我和 Quickley 的朋友们启动一个项目时,我才真正意识到 Shopify 有多强大。Quickley 是一个紧迫感营销平台,多年来一直为各大品牌开展创新的限时营销活动。现在我们正在合作构建一个 AI 层,利用生成式 AI 将他们的服务扩展到长尾电商企业。由于 Shopify 拥有最大的市场份额、最强大的 API 和最活跃的应用生态系统,我们专门为 Shopify 平台构建。所以,如果你正在建立电商业务,升级到 Shopify,你不仅能享受他们市场领先的结账系统,还能获得越来越强大的前沿 AI 应用库,比如 Quickley,其中许多在发布时将是 Shopify 独占的。The Cognitive Revolution 的听众可以在 shopify.com/cognitive(全部小写)注册每月 1 美元的试用期。没有人比 Shopify 更擅长销售,所以访问 shopify.com/cognitive 升级你的销售吧。那就是 shopify.com/cognitive。

Revolution you can start for free and using our link supports the show so join me in giving notion AI a shot today at notion.com cognitive Revolution the cognitive Revolution is brought to you by Shopify I've known Shopify as the the world's leading e-commerce platform for years but it was only recently when I started a project with my friends at quickley that I realized just how dominant Shopify really is quickly is an urgency marketing platform that's been running Innovative time-limited marketing activations for major brands for years now we're working together to build an AI layer which will use generative AI to scale their service to longtail eCommerce businesses and since Shopify has the largest market share the most robust apis and the most driving application ecosystem we are building exclusively for the Shopify platform so if you're building an e-commerce business upgrade to Shopify and you'll enjoy not only their Market leading checkout system but also an increasingly robust library of cuttingedge AI apps like quickly many of which will be exclusive to Shopify on launch cognitive Revolution listeners can sign up for a $1 per month trial period at shopify.com cognitive where cognitive is all lowercase nobody does selling better than Shopify so visit shopify.com cognitive to upgrade your selling today that's shopify.com cognitive

后训练基础理解 Baseline Understanding of Post-Training

Host

那么我们来深入探讨一下。我有一个大致的了解,我想大多数听众也会有一个大致的了解,当然,预训练,知道那是什么,网络规模的数据,下一个词预测,好的。然后你有典型的流程,可能有一些例外或注意事项,但通常接下来是指令微调,然后是某种偏好调优,这就是基于人类反馈的强化学习(RLHF)发挥作用的地方。然后有不同的算法,你提到了 PPO 和 DPO,可以用来实际进行偏好调优的最后阶段。这是基础理解。给我们下一层次的理解,你知道,之后最重要的是什么。

So let's get into the real nitty-gritty here. I have a general sense, I think most of our listeners will have a general sense of, you know, certainly pre-training, understand what that is, you know, web scale data, next token prediction, okay cool. Then you've got your typical recipe, there may be some exceptions or caveats to this, but typically next comes instruction tuning, and then comes some sort of preference tuning, and that's where reinforcement learning from human feedback fits in. And then there's different algorithms, and you've mentioned PPO and DPO, that can be used to actually do that final stage of preference tuning. That's the baseline understanding. Give us the next level of understanding, you know, what's most important to understand after that.

后训练中开放与封闭配方 Open vs Closed Recipes in Post-Training

Nathan

是的,我想直接跳到真正微妙的地方,你可能甚至没有在你的问题中提到:那就是开源方案在做什么,闭源公司在做什么。在开源方面,我们能够从非常强大的模型(如 GPT-4)的输出上进行训练,这是我们做过的事情。这存在一些法律不确定性,但我们确实在我们的数据集中加入了这样的措辞:你看,你理解你的训练和这些模型的输出。但这确实在某种程度上给了我们优势。Llama 为了获得像 GPT-4 这样的模型,需要构建 Llama 405B,他们需要对各种基础模型进行指令微调,以及所有这些事情,才能让这个蒸馏模型工作。所以他们也在尝试类似的事情,即很多指令都是由非常强大的模型编写的。所以很多数学和代码指令数据现在是由 Llama 405B(对于 Llama)或 GPT 模型(对于 OpenAI)编写的。在这方面,我们做的事情类似。但在开源方面,你可以走一些捷径,比如我们不需要做整个预训练,我们直接使用最适合 SFT 的模型。我认为那个阶段实际上看起来非常相似。我认为我们对使用的提示没有很好的控制。我认为闭源实验室会对分布进行大量过滤和控制。我们做了很多混合,但主要是基于已知数据集的水平,并观察它们在 downstream 评估上的表现。这很大程度上是一个为特定能力添加数据集到指令微调中,并确保它不会降低其他性能的过程。所以如果你有一个像 MATH 这样的数学评估,你可以改进它,但很多归结为格式问题。所以一些数学数据集在补全中会有不同的格式,这会使你的模型更难从中学习,或者可能影响 GSM 或编码之类的东西。所以这就是你在做的事情。然后在最后,你可能还会尝试包含一些更通用的聊天数据。我们包含了来自 Cohere 的 10 万个多语言样本,因为我们知道 Chatbot Arena 有相当多的多语言内容,尽管多语言不是一个评估套件。还有一些其他更边缘的安全问题,比如模型应该如何表现。这就是 SFT 中发生的事情:其中有一些艺术性,以及压制二阶效应。但我认为这在很大程度上与闭源实验室所做的非常相似。然后在偏好调优方面,情况发生了很大转变:闭源实验室使用人类来获取偏好数据。我们没有那么多钱,所以我们使用 LLM 作为裁判来收集我们的偏好数据。我喜欢引用 John Schulman 对此的说法,这是一种精辟的说法:人类偏好数据是高噪声但低偏差,而 LLM 偏好数据是低噪声但高偏差。我们不知道我们得到了什么偏差,但大体上我们仍然可以看到这样的流程:我们从各种模型生成中收集自己的偏好数据,语言模型对其进行标注。偏好数据是好的。这里最大的变化来自像 Hugging Face 上的 UltraFeedback 这样的数据集。我们基本上用我们在 SFT 阶段训练的模型的补全重新做了他们的流程,这比仅仅使用 Hugging Face Hub 上的随机偏好数据带来了有意义的低百分比改进。所以这需要更高的努力:你必须经历生成 LM 补全、确保一切正常、获得多样性、使用 LLM 作为裁判的过程。但这类事情在我们的经验中一次又一次地表明,即使使用 LLM 作为裁判,这种 on-policy 的偏好想法也更好。所以我确实认为,如果你看 Meta 论文中的系统图,他们展示了这一点:他们拿一个新模型,输入它,得到新的偏好数据,然后训练它。区别在于他们做了多次迭代,这可能取决于他们如何获取数据,或者他们的时间线等等。我确实认为多次迭代是我们一次又一次看到的事情,我们没有深入研究,但只是确认了 on-policy 数据确实适用于偏好调优。这不是一个巨大的范式转变,但这是一个完全不同的方法,人们必须去做。我们已经看到一些类似 on-policy 的新在线 PPO 算法变体,每个人都觉得螺旋上升:哦天哪,我们还有另一个 star 算法,我无法关注。但算法正朝着那个方向发展,即更多地从模型生成和标注。但我们的做法,我认为,更加分段。

Yes, so I guess I'm going to jump to the real nuanced thing which you might not even have in your questions: it's like what are open recipes doing and what are closed companies doing. In the open, we have the benefit of being able to train on outputs from very strong models like GPT-4, which is something we did. There's some amount of legal uncertainty there, but we definitely put the wording in our datasets is like, look, you're understanding your training and outputs from these models. But that definitely gives an advantage in some way. Llama, in order to get a model like GPT-4, needs to build Llama 405B, they need to do instruction tuning over their various base models and all of these things to kind of get this distillation model to work from. So they are trying to do a similar thing, which is a lot of the instructions are written by a very strong model. So a lot of math and code instruction data is now written by, for Llama, Llama 405B, for OpenAI, GPT models. And in that way, we are doing something similar. But in the open, you can kind of take shortcuts, which is like we don't need to do this whole pre-training, we just go right to the model which is best for SFT. I think that stage actually looks very similar. I think we don't have as good of control over which prompts we are using. I think the closed labs will do a lot of filtering and controlling of the distribution. We do a lot of mixing, but it's mostly on taking the levels of known datasets and seeing what their performance is on downstream evals. It's largely a process of adding in a dataset to instruction tuning for a specific capability and making sure it doesn't degrade other performance. So if you have like a math eval which is like MATH, you can improve this, but a lot of it comes down to the formatting. So some math datasets will have different formatting in completion that will make it harder for your model to learn from that, or it could affect something like GSM or coding or stuff like this. So these are the kind of things that you're doing. And then also at the end, you're probably going to try to include some more general chat data. We included 100,000 multilingual samples from Cohere because we know Chatbot Arena has a decent amount of multilingual, even though multilingual wasn't an eval suite. And there are some other safety things which are more borderline, which is just like how should a model behave. And this is kind of what happens in SFT: there's some art to it and squashing second order effects. But I think largely that's really similar to what the closed labs are doing. And then at preference tuning, it takes kind of a big turn, which is the closed labs use humans for their preference data. We do not have the money, so we use LLM as a judge to collect our preference data. I like to quote John Schulman saying on this, which is the pithy way to say it: human preference data is high noise but low bias, and LLM preference data is low noise but high bias. We do not know exactly what bias we are getting, but largely we can still see pipelines of we collect our own preference data from various model generations, language models label it. The preference data is good. The biggest change here is from datasets like UltraFeedback which exist on Hugging Face. We essentially redid their pipeline with completions from the models we had trained at SFT, and that gives a meaningful like low percentage improvement over just using a random preference data on the Hugging Face Hub. So it's higher effort: you have to go through the effort of generating completions of LM, making sure that's all right, getting diversity, doing LLM as a judge. But that type of thing gives us, across our experience time and time again, this on-policy idea of preferences even with LLM as a judge was better. So I do think that that's kind of what, if you look at Meta's system diagram in their paper, they show this: they're like we take a new model, we pass it in, we get new preference data, and we train on it. The difference is that they do multiple iterations, which could be based on how they get data, it could be based on how their timeline is and stuff like this. I do think multiple iterations is something we're seeing again and again, and we didn't look at but just kind of checking the box of like okay on-policy data does work for preference tuning. It's not a huge paradigm shift, but it's a whole different approach than people have to do. We've seen some of this in kind of like on-policy, new online PPO algorithm variants that everyone kind of gets spiralized: it's like oh God we have some other star algorithm, I can't pay attention to. But the algorithms were going in that direction where they're doing more of like generation from the model and labeling. But the way that we did it, I think, is a bit more segmented.

训练流程概览:SFT、DPO、RL Training Pipeline Overview: SFT, DPO, RL

Nathan

这有点像 Llama 的做法。从高层看,这和 Llama 3.1 类似,而且我听说他们没有做任何花哨的可验证强化学习或 o1 之类的东西。我们添加的部分——我想我可以具体说,你可以在论文的致谢部分看到,有一个名字和一堆普通的 Grant 乱码——在我发给你的草稿里没有,但那里有一个名字,他告诉我们直接对可验证输出做强化学习。这就是我们做的第三阶段,基本上是从 MATH、GSM8K 这样的数据集取训练集,实际上 IFEval 也是可验证的,你可以检查约束是否满足。IFEval 总结来说是一个评估,其中提示有约束,比如“回答,确保你的回答有 X 个词,确保你的回答有 X 个段落”,这些都可以用 Python 代码验证。所以对于数学和指令遵循这类任务,我们有一个系统,提示带有约束,然后我们用强化学习,如果约束满足就给予奖励。我们在多个模型上看到这能提升 GSM8K 和数学成绩。IFEval 有点棘手——有很多强化学习细节——但我们基本上在最后阶段这样做,以便如果我们的 DPO 降低了数学分数,我们可以把它拉回来。在我们的流程中,DPO 数据非常贴合我们的模型,实际上大部分数学提升来自那里,最后的强化学习阶段影响很小。但如果你从 Hugging Face 拿一个旧的强化学习模型,应用这种可验证奖励的强化学习,可以在 GSM 上获得约 15% 的提升,而其他评估没有太大退化。所以我们只是触及了表面,但 AI 的风向正在朝这个方向吹。有传言说很多大实验室都在做类似的事情——你肯定觉得他们在代码上也在做。我们还没有加入代码。代码解释器?o1 被描述为一个专注于推理任务的大规模强化学习系统。这就像,哦,他们可能也在做类似的事情。我认为我们有一个模型,我们让它对数学做强化学习跑了很久,它开始出现“让我再检查一下答案”这样的行为,在思维链内部重做思维链。这完全就是 OpenAI 向我们展示的“等等,让我检查一下”。这绝对不是 o1,但我认为该领域还有其他论文出现。比如 VinePPO——那一个非常相似。让我把名字说对,方便大家理解。有人讨论过 Quiet-STaR,那更复杂一些。我认为 TRACE 也是一个,也有点难理解,但动机相同。所以文献正在朝那个方向发展,但我们展示了你可以把它用在偏好调优中;你不必只做一个数学模型。你可以在最后添加这类强化学习,它不会搞垮模型。所以总结一下,我们解锁了三个阶段:SFT、DPO 和强化学习。在每个阶段,我们都从行业实践中获取更多优势。SFT 众所周知,它有效。DPO——我们需要发现一些新技巧、参数和不同设置。我们用了长度归一化设置,但这基本已知,不过信息很有用。但强化学习部分,看,这是真的。我们超过了 Llama 的分数。我们用了强化学习,而且强化学习变得越来越相关,这仍然让我震惊。我觉得 2023 年大家都在问“RLHF 要去哪里?”但现在甚至不只是 RLHF——我们这里甚至没有奖励模型。我们可以稍后讨论技术细节,但没有奖励模型,只有一个强化学习价值函数,而且效果很好。这就是训练总结。如果你没听过这些术语,可能有点密集,但也很有信息量。

This is kind of closer to what Llama is doing. At a high level, this is similar to what Llama 3.1 does, and I've heard that they weren't doing any fancy verifiable RL or any o1 stuff on theirs. What we added to this—I think I can specifically say, you can look at the acknowledgement section of the paper and you'll see one name and a bunch of normal Grant garbled text—it's not in the draft I sent you, but there's one name there, and he told us to just do RL on verifiable outputs. So this is the stage three that we did, which is essentially taking the training set from things like MATH, GSM8K, and it's actually like IFEval is verifiable, as you can count if the constraint is satisfied. IFEval as a summary is an evaluation where there are constraints on prompts, like "respond, make sure your response has X words, make sure your response has X paragraphs," and these are things that are verifiable in Python code. So across these things like math and instruction following, we have a system where you have these prompts with constraints, and then we used RL to just give a reward if the constraint was satisfied. We have seen on multiple models that this can improve GSM8K, you can improve math. IFEval is a little bit trickier—there are a bunch of RL details there—but we essentially do this at the last stage to be able to, if our DPO worsened our math scores, we can bring it back up. And with our pipeline where the DPO data is very tuned to our models, we actually get most of the math improvement there, and the last RL stage is pretty minor. But if you take some old RL model off of Hugging Face, you can apply this RL verifiable rewards thing to it and get like a 15% boost in GSM without a ton of degradation in other evals. So we're really just scratching the surface there, but the winds of AI you can see are going this direction. There are murmurs that a lot of big labs do stuff like this—you definitely think they're doing it for code. We haven't added code. Code interpreter? o1 is described as a large-scale RL system specializing in reasoning tasks. It's like, oh, they're probably doing something like this. I think one of our models, we just left running RL for this math really long, and it started doing this like "let me check my answer again," redoing chain of thought within the chain of thought. It's literally the thing that OpenAI was showing us where it's like "wait, let me check that." And this is definitely not o1, but I think we've seen other papers in the space coming out. There's like VinePPO—that's one that's really similar. Let me get their name right so it's easier for people. People have talked about Quiet-STaR, that's a bit more complicated. I think TRACE is also one, which is also a bit hard to understand but motivated by the same thing. So the literature is going in that direction, but we kind of showed that you can do this in your preference tuning; you don't have to just do a math model. You can just add on these types of RL at the end, and it doesn't blow the model up. So to summarize, as we go, we've unlocked these three stages: SFT, DPO, and RL. And at each of them, we're taking more alpha from what industry is doing. Like SFT is known, it works. DPO—there are some new tricks that we needed to uncover, parameters, different settings. We use this like length-normalized setting, but that's mostly known but good information. But the RL stuff is like, look, this is real. We're beating Llama numbers. We used RL, and it's still mind-blowing to me that RL is becoming more relevant. I think 2023 was like, "oh, where's RLHF going?" But now it's not even just RLHF—we didn't even have a reward model in this. We can go through the technical setup later, but there's no reward model; it's just an RL value function, and it's fine. So that's the summary of the training. It's probably somewhat dense if you haven't heard these terms before, but also informative.

Host

嗯,我喜欢这样。我们稍微展开一下,但这是一个很好的初步概述。所以我可能从几个不同角度评论一下,以便更好地理解整体情况。首先,每个阶段投入了多少算力和价值,又产出了多少?也许算力和数据,比如相对数量以及从每个阶段获得的相对价值。

Yeah, well, I like that. Let's unpack it a little bit, but that's a great initial overview. So I'll maybe just comment from a couple different angles to try to understand the overall landscape better. For starters, like how much compute and how much value goes into and comes out of each of these stages? Maybe compute and data, like relative amounts and relative value that you get from each one.

Nathan

是的,这是一个通用指令模型,我必须说明,因为我认为如果你有不同的领域,算力需求会非常不同。我认为 SFT 是我们讨论很多的。数据量大致在增长;我们没有做很多子采样,因为我们在追求高分数。我们最终的混合数据大约有一百万条提示,大部分是单轮。这个模型在多轮对话上不会像 Llama 那么好;我们不是 Meta AI 的团队,我们没那么需要,也没有相应的评估。但 SFT 阶段大约有一百万条提示。如果你用——我可以给出非常具体的吞吐量数字——如果你用 32 块 H100,可以在大约一天内用 RCDE 训练一个 8B 模型。RCDE 不是特别优化,因为它依赖 Transformer。我认为如果你用非常特定的代码,可以获得大约 40% 的加速,这是我们未来可能做的,只支持 MHA 和 Llama 架构,因为支持的架构越少,速度越快。所以大约是一百万条提示。32 块 H100 大约需要 24 小时的训练周期来训练一百万条提示,是的,对于 SFT。我认为市场价格,大概是每块 H100 每小时两美元?据我所知价格已经下降了。所以每小时 60 美元,算力成本大约 1500 美元。我没想到这个数字,但嗯,我觉得合理。大概在 1000 美元左右。因为我确实认为我们可能用一半的提示就能达到 90% 的性能。看看我们的混合数据,我们有大约 30% 以上的数学。我们在数学上很努力,试图达到这些分数。我们最终做了通用数学和评估,然后对数学的特定子集,比如中级代数,当我们表现不好时,我们看了评估,发现哪个微小子集不好,然后专门为它制作数据。所以你肯定不需要所有这些数学和评估。这是一个很好的经验法则。DPO 和偏好调优,我们还没有强烈证明扩展有帮助,比如增加更多提示并持续改进评估。我认为大约几十万条提示,不算很大。DPO 确实使用更多算力,尤其是在 70B 模型上,我认为是因为你需要参考模型和策略模型。不过,我们运行这些任务,大约需要 6 到 12 小时,如果你用……

Yeah, so this is a general-purpose instruction model, which I have to caveat because I think if you have different domains, you have much different compute. I think SFT is something we talked about a lot. The data size largely kept growing; we didn't do a lot of subsampling results because we were searching for high numbers. Our final mix is about a million prompts, most of them single turn. This model will not be as good as Llama at multi-turn; we're not a Meta AI shop, we don't need that as much, we don't have the evals for it. But it's about a million prompts at SFT. If you're using—I can give really specific throughput numbers—if you're using 32 H100s, you can train an 8B model in about a day on RCDE. RCDE is not super optimized because it relies on Transformers. I think you can get about a 40% speedup if you're using really specific code, which is something we might do in the future to only do like MHA and Llama architectures, because if you support fewer architectures, you make it much faster. So that's like a million prompts. So 32 H100s is like a 24-hour train cycle to train on a million prompts, yes, for SFT. And I think market prices, that's what, two—let's say two bucks per H100 per hour? Prices have come down, as I understand. So $60 an hour, $1,500 kind of compute cost. I wasn't expecting this number, but yeah, yeah, I think it's reasonable. It's ballpark like $1,000. Because I do think we could probably get 90% of the performance with half the prompts. I think looking at our mix, we have like 30% plus math. We were going hard on math trying to get these numbers across. We ended up doing like general math and eval, and then we have subsets for specific subsets of math like Intermediate Algebra when we weren't as good. So we looked at the eval and saw which micro subset was not as good and tried to make data specifically for that. So you definitely don't need all of this math and eval. That's a good rule of thumb. DPO and preference tuning, we have had—we haven't been able to show as strongly that scaling helps, like getting more prompts in and keeps improving evals. I think it'll be like on the order of a couple hundred thousand prompts, it's not as big. DPO does use more compute, especially at 70B, I think because you need the reference and policy model. Still, we run these jobs and they take like six to 12 hours if you're using somewhere...

计算成本对比:SFT vs DPO vs RL Compute cost comparison: SFT vs DPO vs RL

Nathan

大概在 16 到 32 块 GPU 之间,它比 SFT 快得多,只是因为数据集小得多。在 70B 模型上,我们添加了常规优化——我们修改了默认的 HuggingFace SFT DPO 实现的前向传播,使其更高效一些。我们缓存了参考对数概率,这样就不必同时在内存中存储两个 70B 模型,因为如果不做这些优化,它就会变得更像 PPO,在 70B 上需要 128 块 GPU——你可以很快看到 PPO 的计算量是如何膨胀的。而且如果你真的想达到最佳绝对分数,它可能会花更长时间。所以我确实认为 SFT 是计算量最大的,因为我们有最多的 token。DPO,我粗略估计是 SFT 的四分之一左右——大概几百美元。而 RL,尤其是在 70B 上,如果长时间运行,可能几乎和 SFT 差不多,但你可能在类似 DPO 的计算量下就能获得大部分收益。所以 RL 曲线看起来非常像老式的 RL 任务:一开始改进最大,然后趋于平稳并上下波动,可能还会略微上升。所以如果你只做一个 epoch——也就是第一次改进——你会省下很多钱。但我们想,哦,我们要追求最佳数字,那就让它再跑几天看看效果。

Between like 16 or 32 GPUs, it's much faster than SFT just because the dataset is much much smaller. I think at 70B we added the normal—we changed our forward pass from the default HuggingFace SFT DPO implementation to make it a little bit more efficient. We do caching reference log probs so you don't have to co-store both 70 models in memory at the same time, because if you don't do these optimizations, it starts to look a lot more like PPO where you need 128 GPUs to do at 70B, which is—you can quickly see how PPO kind of balloons. And it can take longer if you're trying to really get the best absolute scores. So I do think SFT is by far the biggest compute because we have the most tokens. DPO, I would say ballpark a quarter of what we talked about for SFT—it would be a couple hundred bucks. And RL is probably, especially at 70B, could be almost similar to SFT if we really run it for a long time, but you can get most of the benefits probably in a similar amount of compute as DPO. So the RL curves look remarkably similar to kind of old-school RL tasks where at the beginning they get the most improvement, and then it's kind of level and bouncing around, maybe going up a little bit. So if you do like one epoch—which is this first improvement—you're going to save a lot of the money. But we're like, oh, we're trying to get the best numbers, let's let it run for a few more days and see what we're doing.

赞助商插播:甲骨文云基础设施 Sponsor break: Oracle Cloud Infrastructure

Host

嘿,我们稍后继续采访,先听一段赞助商信息。即使你觉得 AI 有点被过度炒作,但它确实突然无处不在——从自动驾驶汽车到分子医学再到商业效率。如果它还没进入你的行业,那也快了,而且速度很快。但 AI 需要大量的速度和算力。那么,如何在成本失控的情况下竞争呢?是时候升级到下一代云了:Oracle 云基础设施(OCI)。OCI 是一个极速且安全的平台,适用于你的基础设施、数据库、应用开发,以及所有 AI 和机器学习工作负载。OCI 的计算成本降低 50%,网络成本降低 80%,所以你能省下一大笔钱。已有数千家企业升级到 OCI,包括 MGM Resorts、Specialized Bikes 和 Fireworks AI。现在,Oracle 为新美国客户提供将当前云账单减半的优惠,但需满足最低财务承诺。优惠截止日期为 12 月 31 日。所以,看看你的公司是否符合条件,请访问 oracle.com/cognitive。网址是 oracle.com/cognitive。

Hey, we'll continue our interview in a moment after a word from our sponsors. Even if you think it's a bit overhyped, AI is suddenly everywhere—from self-driving cars to molecular medicine to business efficiency. If it's not in your industry yet, it's coming, and fast. But AI needs a lot of speed and computing power. So how do you compete without costs spiraling out of control? Time to upgrade to the next generation of the cloud: Oracle Cloud Infrastructure, or OCI. OCI is a blazing fast and secure platform for your infrastructure, database, application development, plus all your AI and machine learning workloads. OCI costs 50% less for compute and 80% less for networking, so you're saving a pile of money. Thousands of businesses have already upgraded to OCI, including MGM Resorts, Specialized Bikes, and Fireworks AI. Right now, Oracle is offering to cut your current cloud bill in half if you move to OCI for new US customers with minimum financial commitment. Offer ends 12/31/2. So see if your company qualifies for this special offer at oracle.com/cognitive. That's oracle.com/cognitive.

Host

我非常兴奋地宣布,我们的新赞助商 80,000 Hours 现在为 Cognitive Revolution 的听众提供免费的一对一职业咨询。80,000 Hours 旨在为那些希望用职业生涯尽可能做最大好事的人提供最佳建议。我们一生通常工作大约 40 年,每年工作约 2000 小时。这是我们大多数人做出积极贡献的最大机会,值得战略性地对待。这就是 80,000 Hours 可以提供帮助的地方。两年前,我亲自使用了他们的职业咨询服务。当时我刚完成 GPT-4 红队项目,我想尽一切可能推动 AI 未来朝着积极方向发展。但能做什么或应该做什么?并不清楚。在与 80,000 Hours 通话后,我获得了与该领域杰出人士的许多联系,在后续对话中,我建立了信心,认为这个播客是我应该追求的项目之一。今天,我很高兴拥有一个由深思熟虑、高潜力人群组成的听众群体,而 80,000 Hours 希望帮助他们。要申请免费的一对一职业咨询,请按照节目说明中的链接操作。网址是 80000hours.org/cognitive-revolution。那就是 80000hours.org/cognitive-revolution。注册免费的一对一职业咨询,找出如何对 AI 未来产生积极影响,我相信你会很高兴这样做。

I am really excited that our new sponsor, 80,000 Hours, is now offering free one-on-one career advising sessions to Cognitive Revolution listeners. 80,000 Hours aims to be the best source of advice for people who want to do the most good that they possibly can with their careers. We typically work for about 40 years in our lifetime, and we work about 2,000 hours per year. That is the single biggest opportunity that most of us have to make a positive contribution, and it's worth being strategic about it. That's where 80,000 Hours can help. I actually used their career advising service myself two years ago. I had just finished the GPT-4 red teaming project, and I wanted to do anything I could to nudge the AI future in a positive direction. But what could or should I do? That was not clear. After my call with 80,000 Hours, I got a number of connections to outstanding individuals in the space, and over the course of the follow-on conversations, I developed the confidence that this podcast was one of the projects that I should pursue. Today, I'm thrilled to have built an audience of thoughtful, high-potential people that 80,000 Hours wants to help. To request a free one-on-one career advising session, follow the link in the show notes. It's 80000hours.org/cognitive-revolution. That's 80000hours.org/cognitive-revolution. Sign up for a free one-on-one career advising session, figure out how you can make a positive impact on the AI future, and I think you'll be glad that you did.

评估改进与后训练设置 Eval improvement and post-training setup

Host

在这个过程中,评估改进怎么样?我猜如果你有一个基础模型,基本上需要做 few-shot;如果进行指令微调,那么你可以做 zero-shot 或 few-shot。你是怎么设置的?在这方面有很多迭代吗?

How about the eval improvement through that process? I guess if you have a base model, you have to basically few-shot if you go to instruction tuning, then you could just do zero-shot or few-shot. How are you setting it? Is there a lot of iteration on it?

Nathan

粗略来说,我认为我们在 SFT 阶段获得了大约 90%的性能,最后 10%来自 DPO 和 RL 的结合。像 AlpacaEval 这样的指标,你从 DPO 和 RL 中获得的收益比 SFT 更多。但到目前为止,我们的 SFT 混合数据在 RVL 套件上已经超过了 Llama 3.1 的分数,但它非常专注于我们的评估。我认为偏好调整在一定程度上软化了模型,使其更易于对话。评估套件很复杂;我无法立即记住所有细节。最好查阅论文。但我们尝试做的是——预训练评估和后训练评估之间存在区别。我确实同意大多数后训练评估应该使用聊天模板,这样你从模型生成 token。很多评估使用思维链,并且上下文样本数量也有一定分布。我认为有些是 zero-shot,有些是 eight-shot,具体取决于领域。可能最难管理的是推理的答案提取,即:答案是否以评估期望的格式出现?特别是数学。Llama 3.1 使用所谓的 Manura 格式和特定提示,而我们使用所谓的 flex 格式,基本上——如果默认写法是'答案在\boxed 中',你也允许'boxed'和'答案:'以及另一种格式。这主要是为了公平对待 Qwen。Qwen 是一个有趣的例子:他们的 72B 模型在 Llama 3.1 设置下得分约为 6,但在我们的设置下得分为 74。所以,我们某种程度上是在与 Qwen instruct 竞争,因此我们的设置不能仅仅照搬 Llama 的做法。我的意思是,所有其他实验室都在做类似的事情,即针对评估调整训练。而且很难知道他们训练了什么。我们做了大量工作来去污染我们用于开发的所有训练数据集。我们还有一个未见过的评估套件。因此,我们使用几种方法检查我们使用的所有训练数据集(包括一些最终数据集)的精确匹配和重叠。我们还将沿途发布其他数据集的去污染版本。例如,流行的数据集如 Open Instruct、NVIDIA 的 Daring Anteater 数据集在 MATH 上存在污染。比如 HuggingFace 的 Numina Math TR TI(用于数学竞赛的工具集成推理)也有 MATH 污染。所以必须移除这些。这就像他们用模型参加 Kaggle 竞赛,而不是为了公平评估。这确实说明了问题。

I would say that for performance roughly, I think we get like 90% of performance at SFT, and then the last 10% a mix of DPO and RL. And like AlpacaEval vibes, things you get the most from DPO and RL rather than SFT. But at this point, our SFT mix beats the Llama 3.1 numbers on the RVL suite, but it's very focused on our evals. And I think the preference tuning kind of softens it a bit to be a nicer model to talk to. The eval suite is complicated; I don't know all the details off the top of my head. It's good to look at the paper. But we tried to do—there's this whole distinction between pre-training evals and post-training evals. And I do agree that most post-training evals should be using a chat template, so you're generating from the model to generate tokens. A lot of them are chain of thought, and there's kind of a distribution over the number of in-context shots. I think there's some that are zero-shot, there are some that are eight-shot, kind of depending on the domain. And probably the trickiest thing to manage is answer extraction for reasoning, which is: does the answer appear in the format that the eval expects? Particularly for math. Llama 3.1 uses what's called the Manura format and a specific prompt, and we use what is called a flex format, which essentially—if the default writing is like 'the answer is in \boxed', you also allow 'boxed' and 'the answer is:' and one other thing. And this is mostly to be fair to Qwen. So Qwen is a fun example where their 72B model gets a score of like 6 with Llama 3.1 setting, but a score of 74 with our setting. So it's like, okay, we're competing with Qwen instruct in a way, so we have to have a setting that is not just what Llama does. And I mean, all the other labs are doing things like this, which is tailoring their training to evals. And it's hard to know what they trained on. We did a lot of work to decontaminate all of our training datasets on the evals that we're doing for development. We also have an unseen eval suite. So we do a few methods to check for exact match and overlap on all of the training datasets that we use throughout it, which is some of the final datasets. And we're also going to release decontaminated versions of other datasets along the way. So popular names like Open Instruct, NVIDIA's Daring Anteater dataset has some contamination on MATH, for example. Like the HuggingFace Numina Math TR TI, which is tool-integrated reasoning for a math competition, has MATH contamination. So you have to remove this. It's like they were using their model for Kaggle competition, not for fair evaluation. And it really goes to show.

数据污染与评估挑战 Data contamination and evaluation challenges

Nathan

第一,我们正在发布数据;第二,我们在展示数据污染有多容易发生。所以,我们需要更多人发布数据,这样我们才能知道我们比较的模型是否在测试集上训练过,无论是我们的开发集还是未见集。我们无法知道。我们发现了很多污染,但我们不知道其他人在做什么。论文里可能有一段讽刺的话:我们不知道这些模型中哪一个在 FS 上训练过。就像,好吧,哪里是最好的?当你发现污染时,你是在别人开源的数据集中发现的,或者你使用某种诊断方法来检测模型中的污染。我见过一些这样的技术。我们正在研究数据集。我们做的最主要的事情是精确提示匹配。如果某个训练数据集与我们的测试集有超过 2%的精确提示匹配——即测试集中提示的精确匹配百分比——我们就认为它被污染了,并将其移除。另一件事是,你能检测出模型是否在特定测试集上训练过吗?今年早些时候,我的大项目是 Reward Bench,旨在建立一个评估奖励模型的生态系统。现在有很多学术论文在讨论它,但我们发现的一个最奇怪的事情是 Reward Bench 提示存在大量污染,这些提示大多来自其他测试集,由 Llama Instruct 使用 Magpie 方法生成。Magpie 是一种合成数据方法,通过操作聊天模板让模型生成与其训练分布一致的提示。这就是你需要做的奇怪思考:模型是否在某个东西上训练过?你无法证明,但如果你有权重,你可以让模型从其分布中生成提示,然后在测试集上检查它们。这很有趣,因为 Magpie 本意是用于生成新数据的训练架构,但当你意识到这可能正是评估所需时,我告诉作者,他说:‘哦,酷,我的论文有了另一个用例。’但我确实认为整个生态系统会继续增长,因为你需要某种真相来源。我认为评估的价值和成本只会越来越高,你需要能够审计模型以在整个行业中进行评估。所以会有新创业公司、学术研究、去污染监管的混合——整个事情有点像公平游戏。

One, we're releasing the data, but two, we're showing how easy it is to have contamination. So it's like, yes, we need more people to release it so we don't know if any of the models we're comparing against trained on test, either on our development set or our unseen set. We just can't know. We found a lot of contamination out there, but we don't know what other people are doing. I think there was probably a snarky paragraph in the paper which is like, we don't know which of any of these models trained on the FS. It's like, okay, where's the best? When you're finding contamination, you're finding it in datasets that other people have open-sourced, or you're using some sort of diagnostic to detect it in the model. I've seen some techniques like that. We're looking at datasets. The biggest thing we're doing is exact prompt match. So if a certain dataset training dataset has more than like a 2% exact prompt match with one of our test sets—that's the percentage of the test set in the prompts exactly—we consider that contaminated and remove them. The other thing is, can you detect if models trained on certain test sets? Earlier this year, my big project was Reward Bench, which is trying to build an ecosystem for evaluating reward models. Now there are lots of academic papers on it, but one of the weirdest things we found is substantial contamination on Reward Bench prompts, which were taken mostly from other test sets, generated by Llama Instruct with the Magpie method. Magpie is a synthetic data method that manipulates the chat template to get the model to generate prompts in distribution of what it was trained on. This is the type of weird head-scratching you need to do: is a model trained on something? You can't prove it, but if you have the weights, you can get the model to generate prompts from its distribution, and then you can check those on test sets. Which is really funny because Magpie is meant as a training architecture for generating new data, but when you realize that might be what you need to do for eval, I told the author and he was like, 'Oh, cool, another use case for my paper.' But I do think that whole ecosystem is going to continue to grow a lot because you need it to have some source of truth. I think evals are only growing in value and cost, and you need to be able to audit the models to assess this across the industry. So there's going to be some mix of new startups, academic research, regulation on decontamination—the whole thing is kind of a fair game.

后训练直觉:指令调优与偏好调优 Post-training intuition: instruction tuning and preference tuning

Host

好的,让我们回到后训练阶段。我希望你能帮我建立直觉,了解模型和模型权重在后训练过程中是如何演变的。关于预训练过程,我有一个很好的小故事:给定一些文本,模型的任务是预测下一个词。下一个词有真实答案,所以它输出一个分布,问题基本上就是如何调整模型中的所有权重数字,让自己更接近正确答案。我们重复这个过程无数次。显然有批处理和其他复杂情况,但基本单元是:你做一个预测,得到一个分数,然后调整自己以更接近正确预测。在后训练的不同阶段,这变得有点复杂,对吧?据我所知,指令微调大部分是相似的——是一样的,但你添加了新的 token,比如‘user’或其他 EOS token。你以特定格式添加这些东西,所以你必须学习一些新 token,但除此之外完全一样。这有点奇怪,有点反直觉,因为对于某些指令任务——这自然过渡到偏好训练——你认为你真正想要的是告诉它什么好什么不好。但通过指令微调,你基本上还是在说这是你应该预测的准确真实 token,所有的权重调整都只是针对那个确切的字符串。

Okay, let's go back to the post-training stages. I want you to try to help me develop my intuition for how the model and the model weights are evolving through that post-training process. I have a pretty nice little story that I tell people about the pre-training process: given some text, the model's job is to predict what comes next. There is a ground truth as to what comes next, so it outputs a distribution, and the question basically amounts to how you can tweak all the little numbers that are the weights in this model to nudge yourself so that you're a little closer to correct. We just do that a ton of times. Obviously there's batching and other complications, but the atomic unit is you make one prediction, you get a score, and you nudge yourself to be closer to having made the right prediction. Now that gets a little more complicated at various phases of the post-training process, right? Instruction tuning, from what I understand, is mostly similar—it's the same, but you add in new tokens like 'user' or some other EOS token. You add these things in a specific format, so you have to learn a little bit of new tokens, but otherwise it's exactly the same. That's kind of weird, a little counterintuitive, in as much as for some instruction tasks—and this is a natural segue to preference training—you think what you really want is to be telling it what's good and what's not good. But with instruction tuning, you're still basically just saying this is exactly the ground truth tokens you should have predicted, and all the weight adjustments are just specifically toward that exact string.

Nathan

是的,对于简单的事情,我认为你不需要很多 SFT 数据,因为你只是在调整权重,使其更专注于这种格式。你需要做足够多,这样模型就会以这种方式输出。但对于像数学这样的特定任务,我认为你可以获得更大的收益,因为模型没有见过很多这类数据,你实际上是在做预训练时做的事情:教模型进行思维链数学。你需要有这些数据来启动。我认为预训练数据可能包含一些,但这是一个我们会看到演变的平衡:你需要多少基本数据来进行指令遵循,然后你需要多少额外数据来获得特定能力?我认为在某个时候,一个很好的参考——我说过很多次,这个项目对读者有好处——是去看 Llama 报告,看看指令微调中每个领域的实例百分比。在指令微调中,数学、编码、推理的比例远高于偏好微调。所以我认为这些是模型需要更多计算来理解基本能力的领域,在 SFT 中。但在偏好微调中,它变成了对比损失函数,要么通过 DPO,它像是对比成对样本。如果你深入研究 DPO 的数学,它实际上是一个奇怪的负负得正,它降低被拒绝响应的负概率,而不是增加被选择响应的概率。如果你深入挖掘,DPO 数学中确实有奇怪的特性。PPO 几乎更直观:你生成新样本,你有一个价值模型为每个 token 分配属性,数值越高越好,然后它试图增加它认为好的东西的可能性。在这种情况下,它由奖励模型引导;在我们的案例中,它由答案是否正确引导,并且变得更加灵活。我认为 DPO 在一定程度上受限于你提供的生成样本,但 RL 在这方面是灵活的,我认为它可以。

Yeah, it's almost like for simple things, I don't think you need a lot of SFT data because you're just manipulating the weights to be more focused on this type of formatting. You need to do enough so that this is the way the model outputs. But for specific tasks like math, I think you can make a lot bigger gains, which is just like the model has not seen a lot of this, and you're really trying to do the same thing that you do during pre-training: teach the model to do chain-of-thought math. You need to have that somewhere to kind of kickstart things. I think pre-training data probably has some of it, but I think that's a balance that we'll see evolve: what is the basic amount that you need to do some instruction following, and then how much more do you need for specific capabilities? I think at some point, a good thing for that—I've said this many times, this project is good for the reader—is go look at the Llama report and look at the percentage of instances per domain at instruction tuning. At instruction tuning, they have way higher math, coding, reasoning than preference tuning. So I think those are the domains where the model needs just more flops to understand the basic capabilities in SFT. But at preference tuning, it becomes a contrastive loss function, either through DPO, which is like pairs and comparing them. If you dig into the DPO math, it's actually some weird double negative where it's decreasing the probability of the negative of the rejected response rather than increasing the probability of the chosen. There are weird oddities in the actual DPO math if you go really deep. PPO is almost more intuitive: you generate new samples, you have a value model that assigns attribution to each of the tokens where higher number is good, and then it tries to increase the likelihood of things that it sees as being good. In this case, it's guided by a reward model; in our case, it's guided by whether the answer is right, and it becomes much more flexible. I think DPO is somewhat restricted to the generations that you give it, but RL in that way is flexible where I think it just can.

后训练中 RL 与 SFT/DPO 对比 RL vs SFT/DPO in post-training

Nathan

RL 更能改变模型的行为。我认为当我们应用 RL 时,模型会采取不同的推理步骤。我们看到的变化类型与 SFT 和 DPO 不同。我确实认为未来会对这些阶段的实际作用有更多理解。有趣的是,SFT 在从聊天到能力的评估中贡献了大部分性能提升,但当你实际做这些实验并观察损失函数时,感觉通过 RL 或这种直接偏好优化之类的方法,你可以对权重做出更多改变。这一点我们早就听说了。我认为 RLHF 领域的领军人物曾说过,RL 只是一个更灵活的损失函数,你可以做更多的 Scaling。我记得 OpenAI 的 Jason Wei 在一次演讲中就是这样说的:RL 有更大的空间,因为损失函数非常不同,我们还没有充分探索它。像 o1 和我们自己的实验让我对此更加乐观。这非常奇怪;你可以做很多奇怪的事情。它显著改变了模型,但验证分数下降得并不多。无论是个人层面还是学术层面,都需要更多的直觉来理解这些参数变化,因为这是微调领域的根本问题。RL 做的事情与之前的损失函数非常不同。从某种意义上说,令人惊讶的是,你只需做 SFT 和预训练,然后直接应用 RL,而 KL 正则化就足以将其稳定住。你几乎完全改变了损失函数,但一切并没有崩溃。这种事情让我非常同情 Dario 等人的想法,即‘他们只是想学习’。这就像有一些复杂的事情在发生,而我对深度学习的直觉还不够深入,无法完全理解。

Kind of change the behavior of the model more. I think when we apply RL, we see that it takes different reasoning steps. We see different types of changes than you see at SFT and DPO. I do think there is going to be a lot more understanding of what these stages actually do. It's kind of funny because SFT is where most of the performance gains come from on evaluations from chat to capabilities, but it just really feels like when you do these and when you look at the loss functions, there's more that you can change about the weights doing RL or doing this kind of direct preference optimization thing. That's something we've heard for a while. I think leaders in RLHF have said like RL is just a more flexible loss function, there's a lot more scale you can do. I think that was the framing that somebody like Jason Wei at OpenAI gave in a talk: look, RL has a lot more legroom because the loss function is very different and we haven't explored it. Things like o1 and our own experiments make me more optimistic in this. It's just very weird; you can do a lot of weird things. It changes the model notably, the validation scores don't go down that much. It will take a lot more intuition both on the individual level and the academic level to understand these parameter changes because it's fundamental to the fine-tuning domain. RL is doing something very different than the previous loss function. In some ways, it's remarkable that you can just do SFT and pre-training, and then you can just throw RL at it, and the KL regularization is enough to just hold it in place. But you just change loss functions 99% of the way there and it doesn't break everything. That type of thing makes me give a lot of sympathy to the whole Dario whatever mindset of 'they just want to learn.' It's like there's something complicated going on that I don't have deep enough intuitions of deep learning to grapple with.

Host

好的,说得很好。让我快速回顾几个要点。首先,KL 正则化:这基本上是一种锚定到早期分布的方法,对吧?所以你的损失函数有多个项,其中一个是‘按照我们希望的方式变得更好’——我们稍后会回到这一点——另一个是‘不要改变太多’。这个直觉基本正确吗?

Okay, that's really good. Let me just go back to a couple points for quick wrap-ups. First, the KL regularization: that is basically a way of anchoring to the earlier distribution, right? So your loss function has multiple terms, and one of them is like 'subject to getting better in the way that we want you to get better' — we'll come back to that in a second — and also 'don't change too much.' Is that basically the right intuition?

Nathan

就是这样。你本质上比较原始模型和新模型的 log 概率,确保 RL 模型和参考模型之间的概率差异不会太大。我们通常的说法是,你有一个 KL 预算,所以你只能对模型进行有限度的改变。一旦达到预算,如果你在做价值更新,通常不会期望模型继续改变,因为它无法做出实质性变化;它只是在同一邻域内移动。我认为这是一个很好的框架。我希望更多的 DPO——DPO 不同,它通过 beta 参数在训练的 epoch 数内控制 KL 距离,而 RL 则是在线 RL,在 KL 方面更开放一些。但我确实认为,总的来说,在后训练中展示更多关于性能与 KL 消耗的图表是非常好的。我认为历史上这些数字听起来完全是随机的,人们也不关注这个。比如 PPO 的 KL 可能在 10 到 20 的量级。这是你在聊天任务上做完整训练时的情况。在 GSM8K 上,比如非常基础的数学,KL 大约是 1,所以与对所有任务做 DPO 相比,我们做 RL 时对模型的改变非常小。但如果你与最佳-of-N 采样等比较,基于奖励模型的最佳-of-N 采样的 KL 消耗也比完整的在线 PPO RL 低得多。我不太清楚 DPO 的直觉位置,但所有这些偏好方法,都是衡量你在偏好调整中改变模型程度的一种方式:在一组受控提示下,KL 差异是多少?有一个问题是,我所说的所有这些数字都是基于不同的训练提示集,所以几乎我们需要一个标准——比如这些是我们评估的提示,这 100 个提示是我们用来评估不同领域 KL 距离的——才能真正看到模型整体移动了多少。但这就是距离,是人们可以查看的一个好方法。

That's what it is. You essentially look at the log probs of the original model and the new model, and you make sure that the difference in probability is not too big between your RL model and your reference. I think the way we phrase it is like you have a KL budget, so you can only change the model so much. Once you reach your budget, you normally don't expect the model to keep changing if you're doing value updates, because it can't make substantial changes; it's just moving around in the same neighborhood. I think that's a nice framing. I wish more DPO — DPO is different where it's like a controlled KL distance through their beta parameter throughout the number of epochs they're doing, whereas RL is like we're doing online RL, which is a little bit more open-ended on the KL side of things. But I do think in general, in post-training, showing more plots of performance versus the amount of KL that you spend is very nice. I think historically these just sound like total random numbers and people not looking at this. It's like PPO KL will be like 10 to 20 scale type thing. This is when you're doing the full thing on chat for example. On GSM8K, like really basic math, KL is like 1, so we're really not changing the model very much doing this RL for math compared to what you would do if you were doing DPO for everything. But if you also compare to best-of-n sampling or something, best-of-n sampling over a reward model also has way lower KL spend than doing this full PPO online RL thing. And I don't know the intuitions off the top of my head for where DPO falls, but it's like all of these preference things, that's kind of a way to measure how much you're changing the model in preference tuning: what is the KL difference across a controlled set of prompts? There's a problem where all these numbers I'm saying are on a different set of training prompts, so it's almost like we need to have a standard — like these are the prompts that we evaluate, these 100 prompts are what we evaluate our KL distances on across different domains — to really see how much the model is moving in general. But that's the distance, that's a good way that people can look at it.

Host

你锚定的那个是原始基础模型还是指令微调模型?

That one that you're anchoring to is the original base model or is it the instruction-tuned model?

Nathan

指令微调模型。而且无论你做多少次迭代,它都保持不变;参考模型在整个过程中保持不变。有些人做实验,先用 SFT 模型作为参考做一轮 PPO,然后以第一个 PPO 模型或 DPO 模型作为参考开始另一轮。所以人们确实会这样做,只是不那么流行。肯定有一些论文是关于移动、重置 DPO 参考以提供更多学习能力的。我一时想不起名字,但这就是人们摆弄的那种想法。

Instruction-tuned. And that just stays no matter how many iterations you're doing; the reference model stays the same the entire time. There are people who do experiments where you can do one round of PPO with your reference as the SFT model, and then you can start another one with your reference as the first PPO model or DPO. So people do things like that, it's just not as popular. There are definitely some papers there which is like moving, resetting your DPO reference to give you more ability to learn. I don't remember the names off the top of my head, but that is the sort of idea that people fiddle with.

Host

而这些 KL 散度,惩罚项——它像一个平方函数,对吧?所以惩罚随着你偏离的程度而增大,这就是预算概念的来源:你有一个限制,超过这个限制,该项就会开始主导。

And these KL divergences, the penalty gets — it's like a square function, right? So the penalty gets bigger the more you sort of drift, and that's kind of where this budget notion comes in: you have a limit as to how far you can move before that term starts to dominate.

Nathan

基本上是这样。在很多 KL 中——在很多损失函数中,我们使用近似 KL 而不是完整 KL。近似 KL 本质上是——我把它调出来了,我有一本愚蠢的 RLHF 书,我正在写,主要是关于 RLHF 基础的笔记,正则化是这样的。很多人在实现中做的是有区别的:它是一个近似 KL。我认为 John Schulman 有一篇博客文章,大家都引用,其中近似方法是,你只需减去策略模型生成的 log 概率,减去 log 概率,这是一个与完整 KL 不同的函数。我不知道这是否是平方和非平方之间的变化,只是凭记忆思考,没有确切的 KL 方程。

Basically. There's an approximation that you use in a lot of KL — in a lot of our loss functions, we use an approximate rather than a full KL. The approximate is essentially — I have it pulled up, I have the silly RLHF book that I'm working on, which is mostly like notes on fundamentals of RLHF, and the regularization is this. There's a difference between what a lot of people do in these implementations: it's an approximate KL. I think that John Schulman has a blog post on this that everyone references, where the approximate is you just subtract the log probs of a generation for the policy model, you do that minus the log probs, which is a different function than the full KL. I don't know if it's a change between square and not, just kind of thinking about it off the top of my head without having the exact KL equation.

Host

有趣。好的,所以所有这些方法可能都有效,并且可能在某种程度上已经有效,各有优缺点。但是,是的,尖锐底部与平坦底部的区别很有趣。我仍然想得到……

Interesting. Okay, so there's presumably all of these things could work and probably have worked to some degree and have some pros and cons. But yeah, that sharp bottom versus kind of flat bottom is an interesting distinction. I still want to get a...

多 token 评估与奖励模型 Multi-token evaluation and reward model

Host

让我们更深入地探讨一下多词元评估和多词元比较是如何工作的。也许我们可以也按 DPO 来分解。从用户的角度来看,我要么看到两个生成结果并选择更喜欢哪一个,要么看到一个并给它打 1 到 7 分,或者看到两个并给它们分别打 1 到 7 分。这就是从中得到的信号。然后,在许多方案中,我们基于此训练一个奖励模型,这样就不需要总是有人类评估者。现在我们有了一个试图像人类一样打分的奖励模型,但它仍然给出类似“这个更好”或 1 到 7 分的评分。我仍在努力向不那么痴迷的朋友们清晰阐述的关键区别是:当这个分数必须通过多词元生成来传递时,它如何转化为更新?在预训练和指令微调中,你预测一个词元,给那个词元打分,然后更新。在这里,分数适用于整个生成结果。可能会有奇怪的分叉,或者如果你多加一个词元,所有东西都偏移了一个词元。当我无法再逐词元比较时,我如何将基础生成与当前策略生成进行比较?

Let's dive a bit more into how multi-token evaluation and multi-token comparisons work. Maybe we can break this down by DPO as well. From a user perspective, I'm either presented with two generations and asked which one I like more, or I'm presented with one and asked to rate it 1 to 7, or two and asked to rate them 1 to 7. That's the signal derived from that. Then in many schemes, we train a reward model on that so we don't always need a human evaluator. Now we have a reward model that attempts to score as a human would, but it still gives a rating like 'this one is better' or a 1-to-7 score. The critical distinction I'm still struggling to articulate clearly to my less obsessed friends is: how does that score translate to updates when it has to go through multi-token generation? In pre-training and instruction tuning, you predict one token, score that token, and update on it. Here, the score applies to the whole generation. There could be weird forks in the road, or if you add an extra token, everything is off by one token. How do I compare a base generation to the current policy generation when I can't do it token by token anymore?

Nathan

在强化学习中,数学本质上是:你根据整个轨迹给出一个标签,然后价值模型(如果你在做 PPO)会接受一个生成结果并输出每个词元的价值。然后你从奖励模型得到标签,即来自环境的奖励,用于更新价值模型以实现逐词元更新。然后策略对批次中的每个词元进行归因,根据它们认为什么会带来更好的长期生成来进行更新。在 PPO 中,这与 DPO 非常不同。我不太理解 DPO 中逐词元发生了什么,因为损失函数基本上将所有词元分组为“选中”或“拒绝”,并试图增加序列中选中和拒绝的对数概率之间的差距。在 DPO 早期,存在关于长度问题的疑问,因为你是在求和所有对数概率,它们都是负数(因为小于 1 的概率的对数是负数)。求和并增加两个负数之间的差距就是损失函数所做的。我对此没有很清晰的直觉,但你可以查看损失函数并看到它们的不同。与 DPO 相比,RL 中有更多的逐词元归因,而两者都与 SFT 有显著不同。

In RL, the math is essentially: you give a label based on the whole trajectory, and then the value model (if you're doing PPO) takes a generation and outputs a value per token. Then you have the label from the reward model, which is the reward from your environment, and that is used to update the value model to have per-token updates. Then the policy takes attribution for every token in the batch, doing updates based on what they think will lead to a better long-term generation. In PPO, that's very different from DPO. I don't have a per-token understanding of what happens in DPO because the loss function essentially groups all tokens into 'chosen' or 'rejected' and tries to increase the margin between chosen and rejected log probabilities across the sequence. Early in DPO, there were questions about length issues because you're summing log probabilities, which are all negative numbers (since log of a probability less than 1 is negative). Summing them and increasing the margin between two negative numbers is what the loss does. I don't have as clear an intuition there, but you can look at the losses and see how they differ. There's more per-token attribution in RL compared to DPO, and both are substantially different from SFT.

PPO 中 token 级价值真值 Ground truth for token-level value in PPO

Host

在 PPO 中,原始的真实值如何将生成级别的分数转化为词元级别的价值?

In PPO, how does the original ground truth translate a generation-level score to token-level value?

Nathan

这取决于你的训练设置,这一点常常被忽略。在某些设置中,你从头开始初始化价值模型,因此需要几百步来用奖励模型的分数预热价值模型。你可以看到,价值模型的损失必须先收敛,策略才会开始显著变化。在其他设置中,你可以从奖励模型或 SFT 模型初始化价值模型,我们在可验证输出上就是这样做的,以使学习更清晰。我不完全知道这会如何改变价值模型,但一定存在某种对数概率与价值之间的映射,这就是为什么你会用模型进行热启动。价值模型通过贝尔曼更新来学习,从最终奖励反向传播。它需要时间才能到达较早的词元,这就是预热所做的。

That depends on your training setup, which is often glossed over. In some setups, you start the value model from scratch, so it takes a few hundred steps to warm up the value model using scores from the reward model. You can see this when the value model loss has to converge before the policy starts changing notably. In other setups, you can initialize the value model from a reward model or an SFT model, which we do for verifiable outputs to make learning cleaner. I don't know exactly how that changes the value model, but there must be some mapping between log probabilities and value, which is why you would warm-start with the model. The value model learns through Bellman updates, backpropagating from the final reward. It takes time to reach the earlier tokens, which is what the warm-up does.

token 价值直觉 Intuition on token values

Host

对于不同的词元获得什么样的价值,是否有直觉或观察?例如,可能“答案”这个词价值低,而实际答案价值高,或者“答案”价值高,因为它知道答案通常紧随其后。

Is there an intuition or observation about what different tokens get what kind of value? For example, presumably the word 'answer' might have low value, and the actual answer has high value, or 'answer' might have high value because it knows the answer normally follows.

Nathan

我认为可能不容易建立直觉,因为特征空间太大了。例如,在“等待”这个词的情况下,它是一个高价值词元吗?但如果价值太高,模型就会一直说“等待等待等待”而从不回答。所以有一些奇怪的事情。此外,所有词元都依赖于前面的词元,所以价值依赖于之前的一切。这不仅仅是词元本身,而是上下文中的词元。

I think it's probably not easy to build intuition because the feature space is so large. For example, in the case of 'wait', is it a high-value token? But if it's too high value, the model would just say 'wait wait wait' and never answer. So there are weird things. Also, all tokens are conditioned on previous tokens, so the value is conditioned on everything that came before. It's not just the token; it's the token in context.

与推理时计算的联系 Connection to inference-time compute

Host

你看过 entropic 项目吗?我想知道是否有联系。

Have you looked at the entropic project? I wonder if there's a connection.

Nathan

我没有深入研究过,但我认为这是一个很好的方向。推理时算力与强化学习密切相关。从强化学习中获得一个能够归因价值的好模型对于在推理上投入更多算力非常有用。我们会看到这些事情继续加强。我把 entropic 更多地归入推理时算力类别,但有类似的直觉。例如,如果你在强化学习中从提示中采样更多完成结果,你就是在探索更多以找到高价值的完成结果。在推理时算力中,很多归结为你如何从语言模型中采样并鼓励它、重新加权或中断它以找到正确的完成结果。这就是蒙特卡洛树搜索之类的东西出现的地方,比如对推理步骤进行搜索或 QAR 类型的方法。

I haven't studied it in depth, but I think it's a good direction. Inference-time compute is very related to RL. Having a good model from RL that can attribute value is very useful for spending more on inference. We'll see these things continue to fortify. I put entropic more in the inference-time compute category, but there are similar intuitions. For example, if you sample more completions from a prompt in RL, you're exploring more to find a high-value completion. In inference-time compute, a lot comes down to how you sample from the language model and encourage it, reweight, or interrupt it to find the right completion. That's where Monte Carlo tree search type things come in, like search over reasoning steps or QAR-type methods.

DPO 与 PPO 对比 DPO vs PPO comparison

Host

在我们继续之前,还有一个关于 DPO 和 PPO 的问题。PPO 有奖励模型。确认一下我的理解是否正确:PPO 有奖励模型,然后有一个将奖励分数转换的过程……

So one more thing on DPO versus PPO before we move on. PPO has this reward model. Just make sure I'm understanding correctly: PPO has the reward model, there's then this process of translating the reward score...

Nathan

正确的直觉是,奖励模型是为了做 PPO 而训练的,但在强化学习框架中,奖励模型实际上在某种程度上就是环境。强化学习中的环境应该是返回奖励的东西。所以奖励模型是一个非常受限的环境,你的动作,你对环境的输入,是提示,而你的动作是补全。这就像一个完全破碎的强化学习环境,但奖励模型并不是训练更新的一部分;它更像是一个我们碰巧训练的孤立东西。

The right intuition is that the reward model is trained to do PPO, but in the RL framing, the reward model is actually the environment in a way. So the environment in RL is supposed to be what returns the reward. So the reward model is a very constrained environment, and your actions, your inputs to the environment, are prompts, and your actions are completions. So it's like a totally broken RL environment, but the reward model is not part of the training update; it's more of an isolated thing that we just happen to train.

Host

明白了。然后与 DPO 相比,在 PPO 中有一个将分数映射或转换为 token 级价值的过程,然后进入反向传播。而在 DPO 中,没有奖励模型,对吧?只是对比两个输出和一些花哨的数学,似乎没有人对此有很好的直觉。也许原作者有;我和他们聊过很多。我觉得六个月前当我还在流行的时候,我对这个有更多的直觉,但我确实认为它仍然可以单独改变 token,就像它在某种程度上做类似的事情,但它不是通过价值模型来做的;它只通过损失来调节,损失看的是 log 概率的总和。所以它有点去掉了中间步骤。

Gotcha. And then to compare against DPO, in PPO there is the process of mapping or translating a score to token-wise value, which then goes into the backprop. Whereas in DPO, there is no reward model, right? And just contrasting two outputs and some fancy math which nobody seems to have a great intuition for. Maybe the original authors do; I've talked to them a lot. I feel like I had more of an intuition for this like six months ago when I was really in vogue, but I do think it can still change the tokens individually, like it is in that way doing something similar, but it's not doing it through a value model; it's only mediated through the loss, which looks at the sums of log probs. So it's kind of removes that intermediate step.

Nathan

是的,所以那更简单,需要更少的算力,需要更少的 GPU。而且无论出于什么原因,似乎总是稍微逊色于……我的意思是,在我们的设置中,我们试图在最后加入 PPO,但我们还没有让 PPO 再次击败我们的 DPO 设置。所以就像,好吧,我们经历了一个大过程,看到 PPO 更好,我仍然认为如果我们继续努力,我们可能能让 PPO 更好,但实验时间和数据如此重要,以至于我们不会花四倍的算力在实验时间上……这不值得。到了这个时候,你会想,哦,我们的 PPO 没那么好。也许一半是因为我们不太知道如何训练奖励模型,因为我们没有花那么多精力在上面,但另一半是因为它更难。所以我确实觉得这很有趣,这是一个很棒的愚蠢辩论,它永远不会消失。

Yeah, so that's simpler, requires less compute, requires less GPUs. And just for whatever reason, seems to consistently fall slightly short of... I mean, in our setup, we tried to throw PPO in at the end, and we haven't gotten PPO to beat our DPO setup again. So it's like, okay, we went through this big process where we saw PPO was better, and I still think like if we kept hammering at it, we could probably make PPO better, but just experimentation time and data is so much more important that it's just like we're not going to take four times the compute in experimental time to... it's not worth it. And at this point, you're like, oh, our PPO's not as good. Maybe it's half that we don't really know how to train a reward model because we haven't been focusing on it as much, but also half of just like it's harder. So I do think that it's funny, it's such a great silly debate, it just never goes away.

Host

是的,但这也是一个很好的宏观纠正,因为尽管这些东西很容易让人钻牛角尖,但数据质量比你选择哪种算法更重要,而且可能比任何事情都重要。

Yeah, but it's a great, very good zoom-out corrective for you to give me there too, because as much as all this stuff is easy to go down a rabbit hole on, data quality matters more than which algorithm you choose, and probably matters more than anything.

Nathan

也许你可以……是的,就像我们从版本 2.2 到 2.3,2.2 到 2.5(DPO 到 PPO)的改进大约是 1%,而 2.2 到 2.3 的改进大约是 14%,这完全是由于数据整理和流程之类的东西。所以,这就是你的数字。如果你想成为一个算法学者,在大多数情况下,你是在争取 1% 到 2% 的提升,而制作非常具体的数据,你可以获得 10 倍的提升。我不惊讶,这很可能也是行业所做的。这些后训练团队就像,你是生成 Python 代码数据的人,你是全职的 Python 代码指令、偏好和提示以及过滤负责人,你这样做,你在这个非常非常具体的事情上生成疯狂好的数据,这些数据可能不在实际套件中。就像那样,即使是我们也没有这样做,我们有做数据的人,但他们不会花一个月时间在这个非常具体的事情上,而且你有 12 个人在做。没有学术实验室会像那些封闭实验室那样深入。

Maybe you could... yeah, it's like we went from our version 2.2 to 2.3, and it's like 2.2 versus 2.5 which is DPO to PPO is like 1%, and then on our B, like 2.2 to 2.3 is like 14%, which is all just like data curation and process and stuff like this. So it's like, okay, there's your number. Like if you want to go off and be an algorithmic academic, like in most cases you're going to be fighting one to two percent versus making really specific data where you care about where you can get like 10x. I'm not surprised that's probably what industry does too. Like these post-training teams are like, you're the person generating Python code data, you are a full-time Python code instruction and preferences and prompts, um, and filtering lead, and you do it, and you generate crazy good data on this one really really specific thing that may or may not be in the actual about suite. It's like that's like even we aren't doing like we have people on data but it's not like they're spending a month on this one really specific thing and you have 12 people doing that. It's like no academic lab will ever do exactly what the close labs are doing in that level of depth.

Host

是的,我的意思是这非常耗费资源。我猜你说过你在这个项目中训练了一千个 8B 的实例,对吧?我想,如果你看我们的验证排行榜,我们有一千多个模型。所以不是所有的都是 8B,有些是测试,但我认为大致如此。如果你想获得这个水平的结果,你可能需要经历这样的过程。如果你好两倍,那就是 500,但这变化不大。

Yeah, I mean it's super resource intensive. I guess you said you trained a thousand instances of 8B in this whole project, right? I'm like, if you look at our Val leaderboard, we have like a thousand plus models. So not all of them are 8B, some of them are tests, but like I think it's ballpark. For if you want to get this level of results, that's probably about the process that you will need to go through. Like you could be twice as good and then it's 500, but that's not that big of a change.

Nathan

是的。

Yeah.

微调的程序性理解 Procedural understanding of fine-tuning

Host

所以我想了解一下这个项目的整体流程,也许我们可以把它和我做的事情以及我在之前节目中谈到的进行对比,那就是针对特定任务微调模型。我对此有一个完整的讲座和指南,我告诉人们通常从 10 个黄金标准示例开始。质量最重要。就像把你的裤子钉在椅子上,努力处理这 10 个示例,直到它们成为你能做到的最好的。这足够小,在很多情况下你可能可以做 few-shot。也许你会想根据情况微调。然后我通常告诉人们期望三轮迭代:你会进行微调,发现一些弱点,然后回来扩充数据。如果总体上不够好,你需要将数据增加 10 倍,这是我通常告诉人们的。如果是一个你之前没有考虑过的边缘情况,你可能可以修补一下……

So I just want to get a little bit of like a procedural understanding for what this, you know, the overall trajectory of the project is, and maybe we can contrast it against something that I do and that I've like talked about in previous episodes, which is just fine-tuning a model for a specific task. I've got a whole, you know, kind of lecture and how-to on that, and I tell people typically start with 10 gold standard examples. The quality matters most. Like just staple your pants to the chair and work on those 10 examples until they're the absolute best you can make them. That's small enough you could probably do few-shot in a lot of cases. Maybe you'll want to fine-tune depending on whatever. And then I typically tell people expect three rounds of iteration where you're going to do that fine-tuning, you're going to find some weaknesses in it, you're going to then come back and augment the data. If it's like generally not good enough, you need to like 10x the data is usually what I tell people. If it's an edge case that you just hadn't considered before, you can maybe just kind of patch with, you know...

项目管理与团队动态 Project management and team dynamics

Nathan

举五到十个例子说明那种情况以及如何正确处理。你知道,三轮,如果有很多不同的边缘情况需要处理,可能更多,通常应该能达到目标,如果你只做一项任务的话。显然,这里的一个巨大区别是你把一般性的讨论变成了整整 15 分钟的对话,部分原因是我也在反思,这非常有趣。我的意思是,这一直延伸到 CEO 那里,他会坐下来问:我们如何管理那些极其兴奋的学生?AI2 与华盛顿大学有很多合作,那里都是最优秀、最有动力的学生,他们想做点什么,但无法完成整个项目。然后还有像我这样的人,或者和我同龄的人,我算是初级到中高级的研究科学家,有博士学位,有一些经验,但也不是教授。我有更多的可见度,所以最终我实际上变得更资深,因为我有资源分配权,也做过一些事情。但还有那些全职研究科学家,他们有一些项目,可以专注于这个并做很多工作。如何平衡激励?这非常困难。所以我认为,对于这个项目,在学术意义上我会是最后作者,但在工业意义上我是第一作者,因为我必须做各种随机的事情。简直不可思议,就像把所有东西都记在脑子里,同时跟踪一切。有一些行政帮助,我们有会议结构来协助,这就像论文的哲学部分:我们试图做什么?我们是否朝着这个方向前进?因为一开始,我想大概有 4 到 7 个负责人,他们做了大量相当于正常学术环境中第一作者级别的工作,比如制作大量数据、运行大量实验、构建整个验证设置。有所有这些事情,一开始我们坐在那里,问大家想做什么?我们还没讨论过这个,但我们真正想看到的一件事是 Llama 2 和 Llama 3 做了拒绝采样:我们能让拒绝采样工作吗?然后有两个人,他们所做的就是:我们能通过拒绝采样让数字上升吗?这是一个负面结果,我们做了很多事情。有一些有趣的事情,比如如果你使用 8B 模型,生成的结果不如我们开始的指令好。所以如果我们生成新的指令,然后对它们进行 SFT,这似乎会让模型变差。但我们有这两个人在做指令微调。有些人开始比较 DPO 的替代方案:DPO、link、norm DPO 等等,弄清楚这些设置是什么。然后有些人问:开放数据集在哪里?我们能拿到吗?让我们再次在开放数据集上训练。我们在哪些方面落后于 Llama?所以它开始于这个算法阶段,然后你得到方向:好吧,我们在哪些方面不如 Llama?事情就是不行吗?让我们砍掉它们。然后它进入第二阶段:让我们尝试获得非常具体的数据。让我们砍掉一些长尾项目,专注于具体数据。然后你继续混合,继续让你的验证套件更稳定。当我们开始项目时,验证套件并不稳定。就像,我们在某方面被 Llama 碾压,但格式完全奇怪。它很晦涩,因为我们试图在我们的代码库中重建整个 Llama 评估套件,所以在早期随着你收敛到验证,有一些演变,但后来它变得更具剥削性,即我们为特定能力构建新数据。所以人们埋头构建数学数据、构建 IFL 数据。然后可能是一个权衡:有些人在做 SFT 数据和 SFT 混合,有些人在做 DPO 和 DPO 混合。然后很快,这个 RL 的东西是否有效,需要两个人自己开始做这个 RL,但作为项目的一部分。然后随着几周几个月的过去,你试图最终确定 SFT。所以 SFT 最终确定了。数据混合大概是一个月前。然后我们必须做更多的去污染。哦,我们必须重新训练它。哦,我们的 70B 超参数错了。哦,我们必须重新训练它。哦,我们必须尝试模型合并。我们必须再训练几次。关于模型合并的快速说明:通过在同一数据集上运行多个种子来合并 SFT 是一个安全的选择,但通常只是通过在多个种子上运行,其中一个 SFT 种子实际上会是最好的。所以如果你想得到一个好的 SFT 模型,你只需要在几个随机种子上训练三到五次。但很有趣,这只是更多的算力。然后在最后一个月,主要是最后的润色,然后全面进行 DPO,这需要我们生成,做这种 on-policy 的事情,这需要大量的 API 积分和大量的动手操作。即你拿你的 SFT 模型,运行补全,然后那个人负责 on-policy 偏好数据。所以你给他们提示和模型,他们做 VM 来生成,然后他们使用 OpenAI API 做 LLM 作为评判,然后他们说这是你的偏好数据集。然后我们有不同的模型:你有 8B、70B,我们有其他模型。所以这是团队中负责制作偏好数据的人,你必须这样做。然后那些人去找训练人员。训练人员做 DPO,然后传给 RL。所以你有这些人自然,概括来说,人们自然收敛到不同领域。这大概有 10 到 15 个人大部分时间积极参与。在这个规模下,我们不需要严格的管理授权。如果你达到 Llama 的规模,100%需要管理授权和规则,谁在什么时候做什么决定。实际上,我算是那个角色,决定最终的 SFT 混合是什么。但组织上有一个有趣的转变,即你无法再扩大规模,因为如果在一个语言建模过程中有超过 10 到 15 个贡献者,就会变得混乱。但如果你增加管理人员,就会有很大的成本,因为你需要做更多的决策授权之类的事情。所以从我的描述中可以看出,很多过程有点混乱和自由形式,只是依靠。

Give it five or 10 examples of that situation and how to handle it correctly. And you know, three rounds, maybe more if there's a lot of different edge cases you want to handle over time, typically should get you there if you're doing one task. Obviously, a huge difference here is you're doing a general turn into a whole 15-minute discussion because partially I'm also reflecting on it, and it's very interesting. As I mean, this goes all the way up to like the CEO is like, I'll sit down, it's like how do we manage the fact that we have extremely excited students? So a lot of AI2 is collaborators with UW, which is like total best of the best, super motivated students want to do something, but they can't do these whole projects. And then you have like people like me, or slightly like people like me that are my same age, so I'm like a junior mid-senior already research scientist, did a PhD, have some experience, but I'm also like not a professor. I have more visibility, so that ends up being like I end up being effectively more senior because I have distribution and I have done things. But you have also these people that are like research scientists full-time, they have some projects, they can focus on this and do a lot of it. It's like how do you balance incentives? It's very hard. So I would say that for this project, in an academic sense I would be last author, but in an industry sense I'm first author because I just have to do all sorts of random things. It's just like incredible, like just hold everything in your head and try to keep track of everything at once. And there's some admin help, and we have meeting structure to help with this, which is like the philosophy section of the paper: it's like what are we trying to do, and are we on track going in this direction? Because when it starts with, I would say there's like four to seven leads that have done a substantial amount of like first-author level work in a normal academic setting, which is like making a lot of data, running a ton of experiments, building the whole validation setup. There's like all these things which is like at the beginning we kind of sit there, it's like what do people want to do? We haven't talked about this, but like one of the things we really want to see is like Llama 2 and Llama 3 did rejection sampling: it's like can we make rejection sampling work? And it's like two people and all they're doing is like can we make the numbers go up with rejection sampling? And this is like a negative result we have, which is like we did a lot of things. There's some interesting things like if you're using an 8B model, the generations are less good than the instructions we start with. So if we're like generating new instructions then we do SFT on them, it like kind of makes sense that it would make the model worse. But it's like we had these two people doing instruction tuning. Some people started with like let's compare the DPO alternatives: DPO, link, norm DPO, whatever all these things are, figure out what these settings are. And then some people that are like where are the open datasets? Can we get them? Let's start training on open datasets again. Where are we falling short of Llama? So it kind of started as like this algorithmic phase, and then you get the bearing of like okay, where are we not doing well with Llama? Are things just not working? Let's kill them. And then it kind of shifts into the second phase of like let's try to get really specific data. Let's kill some of our long-tail projects and work on specific data. And then you continue to mix, you continue to make your validation suite more stable. When we start the project, the validation suite is not that stable. It's like huh, like we're getting crushed by Llama on this thing, but like the formatting is totally bizarre. It's like esoteric because we tried to recreate the entire Llama evaluation suite in our code base, so there's like some evolution there early in the project as you converge on validation, but then it becomes much more exploitative, which is like we build new data for specific capabilities. So people are kind of heads down like building math data, building IFL data. And then it's probably a trade-off of like some people are working on SFT data and SFT mixing, and some people are working on DPO and DPO mixing. And then like soon this whole like does this RL thing work, which takes like two people just start doing this RL thing kind of on their own, but like part of this project. And then as the weeks and months go by, like you try to finalize SFT. So like SFT is finalized. The data mix was probably like a month ago. Then we have to do more decontamination. Oh like oh we have to train it again. Oh our 70B hyperparameters are wrong. Oh we have to train it again. Oh we have to try model merging. We have to train it a few more times. A quick note on model merging: it's like it's a safe bet to merge SFT by running multiple seeds on the same dataset, but it could often be that just by running on multiple seeds, one of your seeds for SFT is actually going to be the best. So if you want to get a good SFT model, you want to just train it three to five times across a few random seeds. But like pretty funny, it's just like even more compute. And then like for the last month it's been mostly like final touches and then full on DPO, which was like we have to generate, we have to do this on-policy thing which takes a lot of API credits and a lot of just hands-on. Which is like you take your SFT model, you run completions, and then those is like the person that just owns on-policy preference data. So you give them prompts and models, and they do like VM to make generations, and then they use the OpenAI API to do LLM as a judge, and then they're like here's your preference dataset. And then you have we have different models: you have 8B, 70B, we have other models. So it's like this is the person in the team that's just owning like they make preference data, and you have to do this. And then those people that goes to a training person. Training person does DPO, they pass it to RL. So you kind of just have these people that naturally, to zoom out, you have people that naturally converge to different areas. And this is probably like 10 to 15 people that are actively involved most of the time. At this size, we did not need strict delegation in management. If you go to the Llama size, 100% chance you need delegation in management and rules who makes what decision when. Effectively I'm like softly that person, which is like making the call on what SFT mix is final. But there's definitely an interesting transition organizationally, which is like you cannot scale it more because it becomes a mess if you're more than 10 to 15 contributors on one of these language modeling processes. But then there becomes a big cost if you add in managers, because then you have to do a lot more of like delegation of decision making and stuff like this. So in some ways you can tell by how I was describing it, a lot of this is like somewhat chaotic and free form and just relying on.

开发流程与团队动态 Development process and team dynamics

Nathan

人们深陷细节之中,并且非常乐意与需要沟通的各种人交流,其中很多是自主进行的。就像我不能站在瓦伦蒂娜身后说‘让你的 DPO 模型去做强化学习’。他们自己就做了。而且很多过程都很混乱。比如我们把站会描述为混乱,这很有趣,因为我们坐下来,写下更新,然后进行 50 分钟混乱的技术更新,完全取决于人们在做什么。这在很多方面都是信息过载。所以我不知道,我认为这准确地反映了混乱,但其中也有某种节奏,比如你的目标是什么?你是否在正轨上?你如何决定砍掉项目的子领域?比如砍掉拒绝采样、砍掉长上下文、砍掉多轮对话。这些事情就是,随着你得到更好的数字,你知道你需要把模型推出来,所以你只能不断减少熵,减少熵。所以它进入了一个阶段,你在中间探索,然后获得动力,最终收敛到一个最终模型。我们可以比 Llama 快得多地做到这一点,因为推出 Llama 对他们来说在法律、战略等方面是更大的问题。所以他们投入更大的投资周期来推出这些模型。我认为甚至比 OpenAI 更甚。OpenAI 和谷歌等公司每隔几个月发布新的 API 模型,而 Llama 一年只发布一两个,并且他们发布权重,模型就固定了。所以思考这种分布很有趣。而且,关于权重固定并发布出去有很多政策讨论,但这种东西也在影响开发周期,我之前没有详细谈过。

People being in the weeds in the details and very happy to communicate with the various people that they know needing it, and a lot of that is autonomous. It's just like I can't be over Valentina's shoulder being like 'get your DPO model, the heish to do RL on it.' It's like they just do that. And a lot of those processes are messy. Like we describe our standup meetings as chaos, and it is very funny because we just sit down, we write our updates, and then we do 50 minutes of just like chaotic technical updates based on what the heck people are doing. It's like just information overload in a lot of ways. So I don't know, I think that accurately reflects the chaos, but there is like some kind of cadence to it of like what is your goal? Are you on target? How do you like make decisions to kill sub areas of the project? Like kill rejection sampling, kill long context, kill multi-turn. These are just things that it's like as you get better numbers, you know you need to get the model out, so you just have to keep reducing entropy and reducing entropy. So it kind of goes to this phase of like you explore in the middle, and then you get momentum and you collapse onto a final model. And we can do that much faster than Llama can, because getting Llama out is a bigger issue for them legally, strategically, and stuff like that. So they do a lot bigger of an investment cycle to get these models out. I think even more so than like OpenAI. It's like OpenAI and Google and stuff release these new API models every couple months, whereas Llama is like you get like one or two Llamas a year and they drop the weights and it's final. So it's like kind of interesting to think about the distributions. And there's I mean there's a lot of policy discussions on like weights being final and out there, but like that type of stuff is also informing the development cycle, which I haven't talked about at length.

Host

是的,听起来很有趣。有很多后续问题。那么当你做日常实验时,什么样的……这让我想起化学。我本科时是化学研究助理,从事反应开发,有一些共同点,比如我们投入一堆试剂,进行反应,几天后回来测量效果如何。有时我们会感到惊讶。这是一个高维空间,所以你可以多加酸、少加酸,做不同用量的实验,然后改变溶剂,以及所有你可以改变的不同东西。听起来有点类似,我猜想那一千件事实际上是 100 个实验,每个有 10 个变体。

Yeah, that sounds fun. A lot of multiple follow-up questions. So when you are doing the day-to-day experiments, what sorts of... It kind of reminds me of chem. I was a chemistry research assistant as an undergrad in reaction development, and it had some commonalities where it was like you know we throw a bunch of reagents in, we kind of run the reaction, then we come back a couple days later and measure how well it works. And at times we were surprised. And it was a high-dimensional space, so you could kind of more acid, less acid, let's do an experiment with varying amounts, and then vary the solvent, and just all these different things that you could vary. It sounds kind of similar, where I'm imagining that like the thousand things were actually like a hundred experiments of 10 variations each.

Nathan

是的,这是一个张力。就像我们在不同阶段如何保持科学性?有些东西我们可以做消融实验,有些像混合就是‘这是我们的流程,我们这样做了’。这非常不学术。但在强化学习方面,比如‘哦,我们可以比较价值模型的不同初始化,你可以比较不同的正则化,这些特定的超参数。’所以两者都有,你需要能在这种混乱中运作的人,直觉上知道往哪里走,但也需要非常科学的人,比如‘根据 X 明确结果,这行还是不行。’所以它涵盖了整个光谱。我同意,我认为我的背景是微机电系统和其他工程类的东西,后来才转向 AI,所以有那种实验室的混乱性质,就像你试图构建这个东西,或者你试图……在这种情况下,就像你在做反应,有时它成功,有时不成功。所以你真的只是……我想也许更好的问法是你在不同探索维度上发现了多少价值?比如你可以改变指令微调的混合比例,或者改变一些超参数,或者你知道是什么东西?但我认为大部分是数据,你找到有效的超参数,它们不会真正改变。但在那之内,我认为仍然有很高的价值。我认为我们在做一个通用模型,但我仍然认为我们可以在混合中加入更多的评估,在不显著改变规模的情况下提高性能。我认为数量……你之前谈到你的配方,针对特定评估所需的数据量实际上并不高。而且你可以利用这个事实做更多事情。还有很多后训练没有被真正触及,我认为机会很大。主要是要建立一个评估反馈循环。就像我们谈到砍掉这些不同的能力,那是因为我们没有喜欢的评估。我 100%确定我们可以改进它们,但在一个分布式环境中,你有一个真理来源,即你的评估,这样更容易。因为大问题是:如何培养性格?我认为性格是你的模型中没有估值的东西。在我们的模型中,如果你把它们与 Claude 比较,它们不会有那么一致的性格。我特别认为对于这些面向许多用户的模型,性格很重要。我觉得这很有趣,但我不知道如何激励合适的人花四个月时间做‘让它像大厨之吻表情符号一样,非常到位’。CEO 会说‘老兄,你到底在干什么?这里发生了什么?’所以在这方面,这是一个我喜欢的分歧。学术界使用评估,但 Anthropic 的 Amanda,比如 matter day 有很大的最终决定权,比如‘这就是 Claude 应该的样子’。

Yeah, so this is a tension. It's like how do we be scientific at the different stages? So different things we can run ablations on, and some of them like mixing are just like 'here is our process, it's like we did this.' It's like it's very much unacademic. But in like the RL side, it's like 'oh we can compare different initializations for the value model, you can compare different regularization, these specific hyperparameters.' So there's definitely both of that, which is like you need people that can operate in this chaos, which is just like intuitively like where do we go, but you also need people who are very scientific as like 'yes no does this work based on X clear results.' So it kind has the full spectrum there. And I mean I agree, I think being my background is in microelectromechanical systems and other kind of engineering stuff before kind of shifting into AI, so there is that kind of messy like in the lab nature of like you just trying to build this thing or you're like trying... In that case it's like you're doing a reaction, it's like it works or it doesn't in some cases. So it's you're literally just like... I guess maybe better way to ask the question is how much juice are you finding in different dimensions of exploration? Like you could vary the mix of the instruction fine-tuning, or you could vary some hyperparameters, or you know what are the sort of things? But I think like most of it data, it's like you kind of find hyperparameters that work and they don't really change. But within that, I think there's still a very high level of juice. I think we're doing a general model, but I still think we could fit more evals into our mix and improve performance without substantially changing the size. I think that like the amount... you were talking about this with your recipe, like the amount of data that you need to target a specific eval is actually not very high. And like there's a lot more that you can do with that fact. There's a lot more post-training that is not really touched, and I think that the opportunity is high. It's mostly about setting yourself to have an eval feedback cycle. Like we talked about killing these different capabilities, and that's because we didn't have an eval that we liked. It's like I'm 100% sure that we could improve them, but it's just much easier in a distributed environment where you have a source of truth which is your evaluation. Because the big thing is like how do you develop character? I think character is something that you don't have valuation for in your models. In our models, if you compare them to Claude, they will not have as consistent of a character. And I especially think for these models that are going to many users, character is important. And I find it very fun, but it's like I don't know how to motivate the right people to do a four-month like 'let's make it like chef's kiss emoji, really spot on.' CEO is gonna be like 'dude what the heck are you doing, what is happening here?' So in that regard, like that is a split that I like. Academics works with the evals, but Anthropic like Amanda, like matter day has a lot of like final say is like 'this is what Claude is supposed to be.'

合成数据与 LLM 作为裁判 Synthetic data and LLM as judge

Host

是的,这在多个方面都非常迷人。好的,我们来谈谈合成数据和 LLM 作为评判者的情况。我明白显然用人类来做这件事是非常耗费资源的。

Yeah, that's really fascinating in multiple respects. Okay, let's talk about the synthetic data and LLM as judge situation. I get obviously it's super resource intensive to create this stuff, you know, with humans doing it.

Nathan

是的,通常只有 1%左右……我们在 API 上花了多少?现在可能更少。已经足够远了,没关系。嗯,每个人看到我们的 OpenAI 账单时都会退缩,但我猜我们在这个项目上花了超过 5 万美元的 LM 作为评判者积分。如果用人类来做,那将是数百万美元。是的,这仍然很痛,当钱流向 OpenAI 而我们做开源研究时,仍然很痛。我的意思是,那大约是 100 亿个 token,因为你知道……LM 作为评判者很奇怪,因为你大部分都丢弃了。它生成一堆 token,你只取一个作为答案。所以它做思维链之类的东西,然后你取那个 token。所以我不知道,很难准确归因于……

Yeah, it's 1% typically is kind of my... yeah, what do we spend on API? Maybe even less at this point. It's far enough in that's fine. Um, like everyone cringes here when we look at our OpenAI bills, but I would guess it's over 50 grand we spent on LM as a judge credits or like about for this project. Which if you're doing that with humans, it would be millions. Yeah, and it still hurts, it still hurts them when it just goes to OpenAI and we're doing open source research. Yeah, I mean that's like 10,000 million tokens because it's like you know couple... The LM as a judge is weird because you throw most of it away. So it generates a bunch of tokens and you take one which is the answer. So it like does chain of thought or something and then you take the token. So I don't know, it's hard to attribute exactly where they're...

计算资源与补偿 Compute resources and compensation

Nathan

是的,听起来你可能做过粗略估算。这仍然远低于 GPU 的成本,这才是关键。所有和我一起工作的人都在面对一个世界,我们的 GPU——我是说,你有腾讯?我没回答那个问题。实际上大约是几千块 H100 的量级,不全是 H100,还有一些其他算力。如果你的项目目标资源就是这个水平,我们瞄准的是每年 GPU 资产超过 1000 万美元的大影响项目。这不是那种每月 5 万美元 API 账单就觉得‘这个月我们真拼了’的情况。那根本不重要。这就很奇怪了:这些人很多是学生,我希望在这种背景下能多付他们一些。而且我也不是前沿实验室付薪的。但这就是为什么实验室的薪酬如此古怪。当他们在 GPU 上花那么多钱时,他们给关键研究人员的薪酬——到头来重要吗?这太离奇又好笑了,但我想对那些赚几百万的人来说是好事。当然,这不会伤害他们。

Yeah, it sounds like you may have done the back-of-the-envelope math. It's still way less than the GPUs, which is the thing. All the people I work with are dealing with a world where our GPU—I mean, you had Tencent? I didn't answer the question. It's effectively on the order of a few thousand H100s, not all H100s; there are some other compute. And if you live in the world where that is the resource your project is targeting, we are targeting big impact projects with $10 million plus assets per year for GPUs. This is not like a $50,000 API bill where we're really pushing it big this month. That doesn't matter. Which then makes it very odd: a lot of these people are students, and I wish I could pay them more in that context. And I also am not paid by a frontier lab. But all this is why compensation is so wonky in the labs. When they're spending so much on GPUs, their headcount for what they're paying their key researchers—does it matter at the end of the day? It's so bizarre and hilarious, but I guess good for the people making millions of bucks. Sure, it doesn't hurt them.

LLM 作为裁判与合成数据 LLM as a judge and synthetic data

Host

再谈谈合成数据和 LLM 作为评判者,你怎么看?我有一些直觉上的疑虑:我们是不是太快就把‘什么对 AI 好’的决定权交给 AI 了?

Staying on synthetic data and LLM as a judge for a second, how would you describe it? I have some intuitive misgivings: are we too quick to delegate to the AIs the decision on what is good for AI to do?

Nathan

有一个项目我不是作者,但我很喜欢它的框架。框架是:如果你可以让人类和 LLM 都做偏好数据,哪些交给人类,哪些交给 LLM?这解决了很多问题。有些事我们肯定希望人类给出答案,但也有很多机械性任务可以外包给 LLM。这篇论文可能只是众多成果之一,我确信实验室也在做类似的事情。问题在于有很多这样的论文,它们好像在说‘为什么我们不能有学术论文证明人类很重要?’所以这里也一样:我们看 GPO 在人类控制数据集上的结果,对比人类偏好和 LLM 偏好,LLM 在 Bows 上的偏好分数更高。然后就想,我们遗漏了什么?我每次看论文都觉得一切似乎没问题,但我不同意你的最终结论,因为我觉得它不够深入。但我也不知道怎么做实验才能更深入。所以这算是我用另一种方式重复你的感受:作为一个开放的研究领域,我们似乎还没有弄清偏好数据对模型做了什么,以及为什么我们不能完全不用人类。所以我有点觉得‘哦,这很糟’,但我确实认为最终我们会逐步加深理解。这又回到了那个观点:人类高噪声、低偏差;机器低噪声、高偏差。最终我们会理解这些偏差和噪声代表什么,但我们现在还不理解。这没关系。我认为 RLHF 已经更多地偏离了安全或偏好领域,现在讨论得少了。当你谈论规范关系和社会学意义上的偏好调优时,接触这些人类与机器偏差就更加重要。如果只是‘让数学数字变大’,我觉得没有明确的指导也没问题。我想看到的是:你给标注者指令,然后看模型在多大程度上反映了这些指令。如果用不同指令收集偏好数据,最终模型会如何变化?这是一个非常复杂的归因问题,但能看到就好了。目前外行人的结论是,至少最好的前沿模型在它们自己创建的偏好数据上表现优于人类,从而在下游评估中比人类偏好数据得分更高。是的,我认为要注意的是我们没有实验室那样的人类偏好数据管道。我们只能用我们能接触到的管道。但这就是结论。他们最初是用人类做的,投入了大量时间、精力、资源、血汗和泪水,做得足够好,以至于现在很难复制他们为人类所做的工作。谁知道呢,他们的人类努力可能比当前的 AI 努力稍好一些,但真的很难复制他们最初动员人类的质量。我显然是在猜测。我在某种程度上认为 RLHF 已经在学术界和工业界之间分叉了。工业界比学术界更关心聊天机器人竞技场和用户留存,这种差距可能会扩大。这没关系。我认为我们永远无法在聊天机器人竞技场上进行爬山优化,因为我们没有 1 亿用户,这也没关系。我想知道他们在做什么,但这需要时间。

There was a project I'm not an author on, but I really like the framing. The framing is: if you can have humans and LLMs do preference data, which do you send to humans versus LLMs? That solves a lot of problems. There are definitely things we want humans giving the answer on, but there are a lot of mechanical tasks we can outsource to LLMs. This paper is probably one of many things that will come; I'm sure labs are doing stuff like this as well. The problem is there are a lot of papers like this where it's like, 'Why can't we have academic papers that show that humans are important?' So the same thing here: we look at our GPO results on a human-controlled dataset versus human preferences versus LLM preferences, and the LLM preference scores on the Bows are higher. And it's like, what are we missing? I always look at the paper and think everything seems fine, but I disagree with your final conclusion because I feel like it's not going deep enough. But I don't know how to do the experiment to go deeper. So that is me reiterating in a different way the thing you are feeling: it just doesn't seem like, as an open research area, we have gotten to the bottom of what preference data is doing to the models and why we can just not use humans. So I kind of think, 'Oh, this sucks,' but I do think eventually we will continue to chip away at that understanding. This goes back to the idea: humans high noise, low bias; machines low noise, high bias. Eventually we will understand what these biases and noises represent, but we do not yet. And it's fine. I think RLHF has shifted more away from safety or what is a preference area that's not discussed as much now. When you're talking about normative relations and sociological things in preference tuning, it is more important to be in touch with these human versus machine biases. If it's literally 'make math number go up,' I'm like, okay, it's fine to not have as clear guidance there. I think the thing I've wanted to see is: you give it instructions to annotators, and I would like to see how well the model reflects the instructions given. If you collect preference data with different instructions, how does that change the final model? It is a very complex attribution, but it would be nice to see. The layman's takeaway at the moment is at least the best frontier models do outperform humans as measured by the preference data they create, giving you downstream better eval scores than human preference data. Yeah, I think the caveat is we don't have the same human preference pipelines that the labs do. This is with whatever pipelines we have access to. But yeah, that's the conclusion. They originally did this with humans and put a ton of time, energy, resources, blood, sweat, and tears into it, and did a good enough job that now it's really hard to replicate what they did for humans. Who knows, their human effort might have been a little bit better than the current AI effort, but it's just really hard to replicate the quality of the human mobilization they had to do to get there in the first place. I'm speculating, obviously. I kind of describe RLHF in some ways has forked between academic and industry. Industry cares much more about Chatbot Arena and user retention than academia does, and that will probably widen. That's fine. I don't think we will ever have the ability to hill climb on Chatbot Arena because we don't have 100 million users, and it's fine. I would like to know what they're doing, but it'll take time.

基于可验证奖励的强化学习 Reinforcement learning from verifiable rewards

Host

另一件我想回头谈的事,一开始我们岔开了,但这可能是你提到的最重要的观察。如果我没记错,是在最后阶段,你在做基于可验证奖励的强化学习。论文里可能会改名,但我们现在就这么叫。这基本上就是:数学题做对了吗?或者代码能运行吗?这显然是一个大趋势,仅仅因为成本和可扩展性。我想知道你是否同意这个描述:我前几天刚跟人说,我们可能很快就能在代码之类的事情上看到超人表现,因为目标和快速反馈使得人类程序员的能力没有完全的上限。而超人诗歌,首先定义不清,其次信号会无限期地嘈杂和缓慢。

Another thing I wanted to go back to that really caught my ear, and then we went in a different direction at first, but might be the most important observation you mentioned. If I recall correctly, it was in that final stage where you're doing reinforcement learning from verifiable rewards. It might get changed in the paper, but that's the name we're going with right now. This is essentially: did you get the math problem correct, or does your code work? This obviously seems like a huge trend just for cost and scalability reasons. I wonder if you would agree with this characterization: I was just telling somebody the other day that we probably should expect superhuman performance on things like code before too long, because the objective and fast feedback are such that there's not totally a limit at what a human programmer can do. Whereas superhuman poetry, first of all, is ill-defined, and second, it's going to be a noisy and slow signal kind of indefinitely.

Nathan

是的,我完全同意。好了,那么在这个过程中,我认为它是在这个基于可验证奖励的强化学习过程中。

Yeah, I agree completely. Okay, so now in doing that process, I believe it was within this reinforcement learning from verifiable reward process.

RL 训练中的涌现推理行为 Emergent reasoning behavior in RL training

Nathan

你开始观察到一种推理行为,它会折返或重新检查结果。这具体是一种涌现现象。

That you started to observe, like, reasoning where it would double back or check its results again. This is specifically an emergent phenomenon.

Host

是的,是的。这就像在一个数学领域,但明确地说,我们训练它的时间远超实际需要。所以在这个训练阶段,正常的评估已经完全崩溃了。模型在正常任务上表现不佳,但它仍然会在我们训练的这个特定验证集上收敛到好的数学答案,不过它的整个思维链过程有点乱套了。当你看到那个在 OpenAI 的 o1 上爆红的相同关键词时,比如“等等,让我检查一下”,这很有趣。我当时就想,这太搞笑了。但同样,这并不那么令人惊讶。这似乎是强化学习会做的事情。o1 发布时的其他例子,比如“哦,看,它改成了法语然后又改回来”,对我来说,这只是模型在做非常奇怪的事情。这并不奇怪,因为这是强化学习部分导致的,因为我们其他的损失函数都不会鼓励这种奇怪的行为。但这也只是一个个例,所以更多是为了好玩和长期拼凑信息,而不是明确地说那个模型和 o1 在做的事情一样。我确实认为同样的训练方法可能适用:他们有一些验证器,然后做大量的强化学习。他们同时在更多领域上做这个,可能混合了确定性和学习到的验证器,并且他们可能用了一些技巧,做了比我们多得多的强化学习训练。但我认为对此感到兴奋并非不合理,而且我不认为有充分理由说 o1 有什么烟雾弹,让他们看起来比实际更复杂。这是一个突破,我确信他们做了很多真正有趣的新颖事情和技巧才达到这一步。但我们已经看到所有实验室一次又一次地发布相同的东西。五个月内,Claude、Anthropic 和 Google 都会推出 o1 的等价物,随便你怎么说。

Yeah, yeah. It's like in one math domain, but to be clear, this is like we trained it for way longer than it's practical. So at this point of training, normal evaluations have totally tanked. So the model is not as good at normal things. It will still converge to good math answers in this one specific validation that we're training on, but its whole chain of thought process kind of got borked. It's just funny when you see the same keyword that everyone was viral about with OpenAI's o1, where it's like "Wait, let me check that." And I was just like, this is so funny. But in the same ways, it's not that surprising. It seems like an RL thing to do. The other examples when o1 came out is like, "Oh look, it changed to French and then it changed back." That's just, to me, it's just models doing really funky things. It's not surprising that it's the RL part of it, because none of our loss functions otherwise are encouraging such weird behavior. But also it's like an N of one, so it's much more just for fun and for piecing things together over the long term than definitively saying that model is anything like what o1 is doing. I do think that the same type of training approach probably applies: they have some sort of verifier, then they do a lot of RL. They do this on many more domains at once, probably with a mix of deterministic and learned verifiers, and they probably do way more RL training than we are doing, with some tricks. But I do think it's not unreasonable to be excited about that, and I don't think there are good reasons that there's some smoke and mirrors about o1 where they made it seem more complicated than it is. It's a breakthrough, and I'm sure they did a lot of really interesting novel things and hacks to get there. But we've seen all the labs release the same things again and again. There's going to be an o1 equivalent from Claude, from Anthropic, and Google within five months, or whatever you want to say.

无监督微调下的纯粹涌现 Purely emergent without supervised fine-tuning

Host

好的,这真的很有趣。你知道,这是关于他们在做什么的有根据的推测。但我只是想确认我正确理解了这种权重观察。这是一种纯粹的涌现现象,因为你没有给它示例数据来学习。

Okay, that is really interesting. You know, informed speculation as to what they are doing. But I just want to make sure I'm understanding this weight observation correctly. This is a purely emergent phenomenon in the sense that this was not something that you gave example data to learn from.

Nathan

我觉得我们的一些做法不太可能……我们没有训练,我们没有对 o1 的思维链进行监督微调或任何类似操作。我们无法获得 o1 的思维链。我们没有对思维链进行任何中间编辑。我认为聪明的人说过,如果你真的把温度调高并做一些事情,他们看到 Llama 也表现出相同的行为。所以在这方面并没有太大不同。我的意思是,即使是愚蠢的 Reflection 70B 模型,原则上也是类似的想法。所以有很多方法可以诱导这种行为。这是一个非常非常非常开放式的行为,它是在某个验证器上使用强化学习损失得到的。所以这就是为什么我当时想,“啊,这太……我们无意中发现了类似 o1 的行为。”就像 OpenAI 把这个东西部署给了 1 亿用户;它必须是一个相当稳定的配方。并不是说他们有一个破解的检查点做了这个,然后他们说,“我们要把这个部署给所有人。”他们肯定能多次做到这种行为。我对它很感兴趣。

I find it very unlikely that some of our... we didn't train, we didn't do SFT on o1 chain of thought or anything. We can't get o1 chain of thought. We didn't do any intermediate edits to chain of thought or anything like that. People that I think are smart have said that they've seen Llama do the same behavior if you really crank the temperature up and do things. So it's not that different in that regard. I mean, even the stupid Reflection 70B model is in principle a similar idea. So there are a lot of ways to induce this behavior. This was a very, very, very open-ended one that was with the RL loss on some verifier. So that's why I was like, "Ah, this is so... we stumble upon o1-like behavior without even meaning it." It's like OpenAI deployed this thing to 100 million users; it has to be a somewhat stable recipe. It's not like they had some cracked checkpoint that did this and they're like, "We're going to deploy this to everyone." They could definitely do this behavior and get it many times. I'm interested in it.

与 grokking 和相变的联系 Connection to grokking and phase changes

Host

是的,这是一个超级迷人的小细节。当然,所有这些事情都在那里,或者即将出现。但我对“grokking”的结果想了很多。对我来说,把行为理解为一个频谱是有道理的:一端是完全随机的鹦鹉,只是 token 之间的原始相关性;另一端是一个完全被理解的、经过足够时间刻入权重的实际算法。当然,对于任何给定的行为,你无法真正分辨是哪种,这几乎是不可能的。但是当你开始从像“你答对了这道题或答错了这道题”这样简单的信号中看到这些,并且开始看到这些定性的行为变化,尤其是当它发生时,就像 grokking 一样,发生在传统上被认为是极端过训练的情况下。这开始描绘出一幅暗示性的画面:可能某种相变正在发生,它进入了一个开始真正推理的领域。这就像奇怪的强化学习表达行为,生成从根本上转变为做奇怪的事情,而大多数互联网文本只是向前推进。这种强化学习行为更加循环和奇怪。这就是为什么我们看不到 o1 的通用推理轨迹很遗憾。一旦我们有了类似的东西,做这些类比就会更有说服力。如果我们能在相同的提示上运行 o1,并查看更多的例子,我认为更容易说,“是的,这是一种非常相似的行为。”我们只需要然后扩展训练方案使其稳定,在每个领域都这样做,并重复进行。这显然不是微不足道的。

Yeah, that's a super fascinating little tidbit. Of course, all these things are kind of in there, or on the verge of being in there. But I wonder a lot about the grokking results. It seems to make sense to me to understand behaviors as kind of on a spectrum from, on one end, fully stochastic parrot, just raw correlations between tokens, and on the other end, a fully grokked actual algorithm that has been traced into the weights through enough time. Of course, for any given behavior, you can't really tell which is which, and it's maybe borderline impossible. But when you start to see these things from such a simple signal as "you got this problem right or you got this problem wrong," and you start to see these qualitative behavior changes, especially when it does, as grokking did too, that happened in what would traditionally be considered an extreme overtraining regime. It does start to paint a suggestive picture that there's potentially some sort of phase change happening, where it's entering into a regime where it's beginning to actually reason. It's like the weird RL expressive behavior, where generation fundamentally shifts to do weird stuff, whereas most internet text is very just going forward. This RL behavior is much more cyclic and strange. That's why it's sad we can't see the o1 reasoning traces for general use. Once we have something like that, making these analogies will be much more compelling. If we could run o1 on our same prompts and look at way more of them, I think it'd be a lot easier to say, "Yeah, this is a very similar behavior." We just have to then scale the training regime to be stable, do this in every domain, and do it repeatedly. Which obviously is not trivial.

无意中观察到的行为 Observing the behavior unintentionally

Host

但这是你在无意中观察到的。你只是 RL 负责人之一,比如 Heish 和 Costa,在做强化学习和基于人类反馈的强化学习时,只是随便看看生成结果。你需要查看生成结果以确保模型仍在正常工作。这是很平常的事情。然后他们说,“哦,这个……我要发推文。这太傻了。o1 阴谋论社区会喜欢这个的。”但就是这样。并不是说我们在生成结果中搜索“等等”这个词。我们找到了几个。我敢肯定,如果我们去找,还会有更多。

But this is something you observed without any intent to see it. You're just one of the RL leads, like Heish and Costa, just poking around generations when you do RL and RLHF. You need to look at the generations to make sure the model is still working. It's just a normal thing. And they're like, "Oh, this is a... I'll tweet this. This is so silly. The o1 tin hat community will love this." But it's in that respect. It's not like we're searching over the generations for the word "wait" or something. We found a couple of them. I'm sure there are many more if we look for it.

Nathan

是的,没有什么能替代深入数据。这个道理可能再怎么重复也不为过。

Yeah, there's no substitute for digging into the data. Can't repeat that mantra enough, probably.

理想化的最终复制流程 Idealized final process for replication

Host

如果你希望描述你学到的一切,并抛开实验和学习的过程,那么理想化的最终流程是什么样的?是不是很简单:如果我要在另一个基础模型上重做你的事情,我基本上需要一天用 32 块 H100 进行监督微调,另一天进行下一阶段,再一天?我认为你需要做几次混合。你需要根据基础模型以及它在每个阶段或多或少需要帮助的能力,进行几次有依据的混合。

If you wish to describe everything that you learned and take away all of the process of the experimentation and the learning, what's the sort of idealized final process? Is it as simple as: if I was going to just redo your thing on a different base model, do I have basically one day of 32 H100s for supervised fine-tuning, another day for the next phase, and another day? I think you need to do a few mixes. You need to do a few informed mixes based on the base model and what capabilities it needs more or less help with at each stage.

微调策略与饱和极限 Fine-tuning strategies and saturation limits

Nathan

你很可能可以从同一个超集开始,然后做几个实验,调整各种行为的多寡。所以大概每个基础模型做几个周期,确保一切正常。如果你真的想获得最佳性能,你可以直接拿这些现成的数据集来用,在给定的基础模型上大概能获得 80%到 95%的性能。所以取决于你在意多少,这已经很不错了。我认为在 SFT、DPO 和 RL 上情况类似。RL 有趣的地方在于,我们不太清楚每个基础模型的天花板在哪里。所以如果你在一个较差的 SFT 模型上做 RL,我们发现比如在 Llama 8B 上,RL 在 GSM8K 上总是饱和在 87 到 88。这就是基础模型的根本限制。无论我们在 SFT 或 DPO 上做什么,我们都能可靠地将 GSM8K 提升到 85 左右,而不会有太多退化。在不同的 Mo 基础模型上,比如我们拿七月份的 Theo 模型,我们从 60 提升到了 75。我们不知道是什么定义了这些饱和极限。但随着你训练阶段越来越多,很多这种“拿现成数据训练”的做法会变得不同。所以如果我们知道真的能大幅提升 GSM8K,那我们就不需要那些 SFT 训练数据了,也许可以把 SFT 预算用在别的地方。但这种复杂性我们还没有深入探索。所以这就是为什么我在反思,并思考下一步该怎么做。如果你真的知道可以在不同阶段恢复某些能力,而不是在每个阶段都追求最大化,那么思考过程会非常不同,而这需要更多我们尚未进行的实验。

You can probably start from the same super set and then do a few experiments with more or less of various behaviors. So probably like a few cycles per base model to make sure things look right. If you really want to get the best performance, you can take these datasets off the shelf and use them, and you'll probably get 80 to 95% of the performance on a given base model. So depending on how much you care, it's pretty fine. I would say similar at both SFT, DPO, and RL. The interesting thing with RL is that we don't quite know how the ceiling is defined per base model. So if you do RL on a less good SFT model, we have found that on Llama for example, on 8B, RL for GSM will always saturate at 87 to 88 GSM8K. That's just a fundamental limit of the base model. No matter what we do at SFT or DPO, we could then get the GSM8K to something like 85 reliably without too much degradation. On different Mo base models, like we take the Theo model for July, it's like we bumped it from 60 to 75. It's like we don't know what defines these saturation limits. But as you get better at having more training stages, a lot of these kind of just take it off the shelf and train will look different. So it's like if we know we can really boost GSM8K, it's like we don't need that SFT training data and maybe we could use that SFT budget in a different way. But that type of sophistication is something that we haven't explored a lot. So that's why it's in my mind to reflect on this and try to think about what you do next. There's a very different thought process if you really know you can recover certain abilities at different times, where it's not just like argmax at every stage, and that takes a lot more experimentation that we haven't done.

Host

世界上大多数人在做的事情,嗯,我不确定,但根据我对大多数实际行业项目的理解,人们并不是在试图创建通用的聊天机器人,对吧?他们试图创建适合特定场景特定需求的东西,并最大化一个或可能几个不同的任务。你会给那些说“好吧,我如何把这个映射到我可能简单得多的情况”的人什么建议?也许我应该每个任务用一个模型,保持简单。如果我有五个不同的任务,它们有点相关,我应该尝试在一个混合数据微调中全部完成吗?

Most people that are doing stuff in the world are, well I guess I don't know this for sure, but my general understanding of most actual industry projects is that people are not trying to create general purpose chat bots, right? They're trying to create something that fits a specific need in a specific context, and kind of maxing one or possibly a few different tasks. What advice would you give to people who are like, okay, how do I map this on to my probably much simpler situation? Maybe I should do one model per task and just keep it simple. Should I, if I have five different tasks that I want to do and they're kind of related, should I try to do them all in one mixed data fine tune?

Nathan

这取决于你使用的接口。如果你真的总能选择正确的模型,你可以每个任务用一个模型。我认为目前大家对通用接口很感兴趣。其中一些会在 RL 阶段改变。我打算在 NeurIPS 上做一个关于 AI 工程师的演讲,我会尝试构建一个世界观,关于如何为不同任务设计这些 RL 验证器。因为我确实认为,如果你有一个验证器,并且分布与你的任务匹配,那么 RL 就会奏效。看到更多工程背景而非研究背景的人尝试适应并直接使用它,会非常有趣。我们有一个早期项目,叫做 LLM Gym,或者一个开源仓库,你可以在其中添加不同的约束,然后直接做 RL,得到模型。你可以添加这些领域,本质上就是为 RL 和语言模型添加领域。所以我认为这是最新的前沿,那些 SFT 单领域、少领域的做法,你可以试试,会发现没那么有趣。但如果我们能用验证器解锁这么多不同细分领域的 RL,那将是微调特定语言模型叙事的下一个时刻。我想通用的答案就是“LLM 作为评判者”,对吧?我的意思是,在可能的情况下你可以更具体。但我总是想到我的公司 Waymark,我们为小企业做视频创作。部分有真实答案,但老实说,我们在具体、客观可验证的事情上并没有太多麻烦。大部分开箱即用,比如确保你提供正确数量的内容和正确的结构。然后真正的问题是,什么是一个好的视频脚本。如果我把你的经验应用到那上面,我基本上会说,是的,用评判者,做一个偏好集,然后就这样做。

It depends on the interface you use. If you really can always just choose the right model, you can do one per task. I think there's a lot of interest in general interfaces right now. Some of this will change in the RL stage. I think I'm going to give a talk at NeurIPS on AI engineers, and I'm going to try to figure out a worldview on how to come up with these RL verifiers for different tasks. Because I do think if you have a verifier and the distribution matches your task, this RL stuff will just kind of work. It would be really interesting to see more engineer-y and less researchy people try to adapt this and just take it. We have early days of what we call like LLM Gym, or an open source repo where you can add different constraints and then just do RL on it and take the model. You could add these domains, essentially adding domains for RL and language models. So I think that's kind of the newest frontier, where some of these SFT single domain, few domain, you could try it and you can see it's not that interesting. But if we can unlock RL with verifiers for so many different niches, it'll be really the next moment in this kind of fine-tuning specific language model narrative. And I guess the general purpose answer there would be LLM as judge, right? I mean, you can get more concrete where possible. But I always think about my company Waymark, we do video creation for small business. There is partially a ground truth, but honestly we don't really have that much trouble with the concrete, objectively verifiable stuff. Mostly it works out of the box, like making sure you're delivering the right amount of content and the right structure. And then the real question is what is a good script for a video. If I'm applying your lessons learned to that, I basically would just say yeah, ver judge it, make a preference set and go with that.

Nathan

是的,我们有……我不知道我有没有精力进行整个多模态讨论,但进入多模态时肯定有不同的指导原则。我认为偏好(preferences)在图像、音频、视频等领域更强大,因为我们的直觉,尤其是人类偏好,在直觉上比文本更具表现力。所以他们说的很多东西都非常以文本为中心,从那个意义上说,能力非常狭窄。

Yeah, we have... I don't know if I have the energy for the whole multimodal discussion, but there are definitely different guidelines as you go multimodal. I think preferences are more powerful in things like images, audio, video, because our intuitions, especially human preferences, are just intuitively much more expressive than in text. So a lot of things they're saying are very text-centric, and in that way capability in a very narrow sense.

Host

是的,有道理。我的意思是,在图像方面我不会期望太多。我觉得你有时确实能得到一些很好的反馈。可能像不同类型物体的检测器。如果你说,确保如果他们有提示,你可以做一个检测器来确保提示中的某些名词或实体出现在图像中。我打赌你可以用这种方法来做精确的图像指令遵循,如果已经有人做了我也不会惊讶。

Yeah, that makes sense. I mean, I wouldn't expect in images. I feel like you do sometimes get some pretty good feedback. Probably do like detectors of different types of objects. If you like, say, make sure if they have a prompt, you can do a detector to make sure that certain nouns or entities in the prompt are in the image. I bet you could have that as a way of doing precise instruction following for images, and I wouldn't be surprised if it's already done.

Nathan

是的,这很有趣。我觉得 Momo 模型在指代方面很有意思。似乎 Claude 现在也有了指代功能。但这很酷……我不需要深入探讨太多。它在机器人技术中的应用:机器人任务以一种很好的方式弥合了 VM 和规划器之间的差距,VM 可以问诸如“我如何”或“我需要什么工具来做 X”之类的问题,然后它可以直接指向它,而不是说出答案。然后规划器知道去拿被指向的东西之类的。这很有趣,我的一些机器人朋友对此很兴奋。

Yeah, that's interesting. I mean, I find the Momo model quite interesting for its pointing. Seems like Claude has now kind of got a pointing function too. But that is a cool... I don't need to dig into that too much. How it works for robotics: the robotics task bridges the gap between VM and planner in a very nice way, where the VM could be asking a question like 'how do I' or 'what tool do I need to do X' and then it could just point at it rather than saying the answer. And then the planner knows to get things that are pointed at or something. That is interesting, and some of my robotics friends are excited about it.

Host

哦,再问一个深入的问题,然后一个宏观的问题。关于策略,我们几次提到了在线数据(on-policy data)的重要性。我的直觉是,你想从模型当前的状态出发。如果你不这样做,效果就不太好。你能再详细说明一下吗?

Oh, one more in the weeds question, then one kind of zoom out. On policy, we've mentioned a couple times the importance of on-policy data. My intuition for that is just like you want to be working from where the model is now. And if you don't do that, it just doesn't work as well. Can you give a little more color on that?

Nathan

是的,这与 PPO 和 DPO 的争论密切相关。很多 PPO 的支持者说,你是根据模型自己生成的东西来评分和更新模型,而不是从别处得到的补全。这才是关键。似乎如果在这些批次和成对比较中,你关注的 token 更接近对数(log),学习信号会更好一些。

Yeah, it's very related to this kind of PPO vs DPO debate. A lot of the proponents of PPO say that you're scoring and updating the model based on things it is generating itself, rather than completions you got from elsewhere. That's really the thing. It seems like it's a bit better learning signal if within these batches and these pairwise comparisons, the tokens you're looking at are closer to the log.

o1 对开放与封闭模型的影响 Implications of o1 for Open vs Closed Models

Host

所以最后,宏观来看:o1 显然已经问世,但我们看不到推理轨迹。你认为这对专有模型(即闭源)与开源模型的未来意味着什么?你发给我的演示文稿中有一张引人注目的幻灯片,基本上直接指出,如果没有前沿模型来进行所有这些生成和评分,开源组织就无法做到这一点。在开源世界里,根本没有替代品。现在我们连 o1 的轨迹都没有。这是否意味着开源追赶的时代结束了,或者你认为未来会怎样?

So finally, the big zoom out I guess is: o1 is obviously out there now, we don't get to see the reasoning traces. What do you think this implies for the future of proprietary, aka closed, versus open? One of the striking slides in the presentation you sent me was basically very directly saying we can't do this as an open organization if we don't have the frontier models to do all these generations, to do these scoring. There's just not a substitute for that in the open world. Now we don't even have the o1 traces. So does this suggest an end of the era of open source catching up, or what do you think is going to be the future?

Nathan

我认为最终会没事的。这需要时间,而且会有所不同,但兴趣非常浓厚。我的意思是,我试图保持中立,但我确实觉得我很有可能在这方面做一个项目。人们对它非常感兴趣,它既令人兴奋又新颖。在某些方面,我们在时间上的劣势较小,因为我们看到了它,而不是像 GPT-4 那样等待。我们进入 GPT-4 时人们才认真起来;我们进入 o1 时也是如此。但我只是认为,一次又一次,我们看到人们在这个领域非常有动力,而且它可能会以与 o1 不同的方式训练,但人们会想办法引出同样的行为。一旦有了存在性证明,它就会有所帮助。我们仍然有 Llama 405B,这是一个非常强大的开源权重模型,用于推理等。这主要需要迭代构建社区趋同的全新基础设施。微调基础设施不会像——好吧,部分会用于构建 o1,但我怀疑会有一个全新的东西。就像什么是面向 o1 类模型的 Transformers 库?那里有一些不同的东西,我认为这对生态系统来说既令人兴奋又成熟,那就是:看,我们需要以完全不同的方式处理这些系统和训练,并且会有更广泛的资源来完成这项工作。再次,悲观的情况是我们不知道沿途的所有步骤,但我很确定,鉴于我们看到的热情,我们会得到接近它的东西。

I think it'll end up being fine. I think it takes time and it will be different, but there's so much interest. I mean, I'm trying to hedge, but I do feel like it's pretty likely that I end up doing a project in this. There's just so much interest in it, and it's both exciting and new. In some ways, we have less of a disadvantage in time because we're seeing it and we're not waiting like GPT-4. We entered like GPT-4 when people got serious; we're entering at o1. But I just think that time and time again, we see that people are very motivated in this area, and it'll probably be trained differently than o1 was, but people will figure out how to elicit the same behavior. It's like once you have an existence proof, it will help. We still have Llama 405B, which is a very powerful open-weight model for reasoning and stuff like this. It mostly just takes iteration on building entirely new infrastructure that the community converges on. The fine-tuning infrastructure is not going to be like—well, part of it will be used in building o1, but I suspect there's going to be a whole new thing. It's like what is the Transformers library for o1-like models? There's something different there that I think is simultaneously exciting and maturing for the ecosystem, which is like: look, we need to approach these systems and training in an entirely different way, and there will be a bigger spectrum of resources that are done to do this. Again, there's the pessimistic case that we don't know all the steps along the way, but I'm pretty sure we'll get something close to it with the amount of excitement that we see.

Host

你认为那会是什么样子?是不是像精心设计的提示来生成这些合成轨迹?我可以想象你拿一个 405,把四个提示串在一起,说“现在换个角度看,现在换个角度看”,然后把它们拼接成某种引导。

What do you think that looks like? Is it sort of elaborate prompting to generate these synthetic traces? I could imagine you take a 405 and you chain four prompts together and say "now look at it a different way, now look at it a different way" and kind of stitch those into some bootstrap.

Nathan

是的,我觉得我快没精力讲完这整件事了。这是我计划写的一篇文章,就像我复现 o1 的计划。我的意思是,你需要生成一些看起来像它的种子数据。你可能需要让一个语言模型查看思维链轨迹,修改并延续它们,然后你这样做,你需要弄清楚如何用这些数据初始化一个模型,然后在正确的领域对某些验证进行某种强化学习。所以这有点像:获取一些看起来不错的初始数据,得到一个能稍微做到这一点的模型,然后你必须——实际上拥有可以验证的东西,并且更新函数强化这种行为,我认为这是最难的部分。我们会看到生成这些数据和验证器的努力,然后把这些东西组合起来,这可能会导致性能的很大差异。我认为我们已经在网上看到有人像开源 o1 复现,而我觉得——我甚至不需要看,直到 2025 年。我甚至不需要认真对待它们。所有已经出来的都不太靠谱,除非是像 Google 和 Anthropic 这样的。

Yeah, I feel like I'm losing the energy to go through this whole thing. This is a post that I plan on writing, like my plan for reproducing o1. I mean, you need to generate some seed data that looks like it. You probably need to have a language model look at chain-of-thought traces and modify and continue them, and then you kind of do that and you need to figure out how to seed a model with those, and then kind of have the right domain to do some sort of reinforcement learning on it on certain verifications. So it's kind of like: get some initial data that looks good, get a model that can do it a little bit, and then you have to—the feedback loop of actually having things that can verify and having the update function reinforce that behavior is, I think, the hardest one. And we'll start to see efforts on generating this data and generating verifiers, and then this putting it together thing, which is where there could be a lot of variability in performance. I think we already see online there's people like open o1 reproduction, and I'm like—I don't even need to look yet until 2025. I don't even need to take them seriously. All the ones that are already out are not that serious unless it's like Google and Anthropic or something.

Host

好吧,我们也许几个月后再来找你。我需要好好理解一下。我暂时处于一种不同类型的后训练环境中,然后这将是下一个要探索的东西,同时还有一些像更好的智能体分类之类的东西,这很有趣,但这是一个很大的思维调整。

All right, well we'll check back in with you in a few months perhaps. I need to wrap my head around this. I'm in a different type of post-training environment for a while, and then that's kind of the next thing to explore along with some like better taxonomies of agents and stuff like this, which is fun but it's a big mental adjustment.

Nathan

是的,这绝对是一个目标丰富的环境。

Yeah, it's a target-rich environment, that's for sure.

Host

是的,太棒了。在我们结束之前,你还有什么其他的结束语或行动号召要分享吗?

Yeah, cool. Well, this has been fantastic. Any other closing thoughts or calls to action you want to share before we break?

Nathan

嗯,我会在我的博客 interconnects 上发布很多这些东西。我已经有另一篇 o1 博文在排队了。我不知道什么时候会发,但很快就会来,很有趣。如果你在听,我会在 NeurIPS 见到很多人。我会在那里,这很令人兴奋。所以谢谢你邀请我。这太棒了。你最后把我累坏了。哦,天哪,我累坏了。

Um, I'll post a lot of these things on my blog, interconnects. Like I already have another o1 blog post in the queue. I don't know when I will send it, but that's to come and it's fun. And I'll see a lot of people around at NeurIPS if you're listening. I'll be there, which is exciting. So thanks for having me. This is a fire. You exhausted me by the end. Oh man, I'm like I'm cooked.

Host

是的,你一直很努力,我当然把你推到了许多不同的小角落。所以谢谢你和我一起做这件事。我确实认为了解你的最新动态有很多价值。所以非常感谢你的时间和精力。就这样,我要说,Nathan Lambert,感谢你成为 Cognoscenti 革命的一部分。

Yeah, well you've been working hard and I certainly pushed you down a bunch of different little dark corners. So thanks for doing it with me. I do think there's a lot of alpha in catching up with what you've been up to. So really appreciate the time and energy. And with that, I will say Nathan Lambert, thank you for being part of the Cognoscenti revolution.

Nathan

是的,谢谢你邀请我。

Yeah, thanks for having me.

互动版:逐字朗读 + 针对本期提问 →