Democratizing AI: The Case for Small Language Models
打开互动全文版(中英对照 + 朗读 + 问答)→斯坦福大学教授 Yann LeCun 探讨小型语言模型对 AI 民主化的重要性,认为通过更多研究,它们可以媲美大型模型的能力。
Stanford professor Yann LeCun discusses the importance of small language models for democratizing AI, arguing that with more research, they can match large models in capability.
即使是开放式问题,模型的多样性也不如我们预期,甚至当你用更高的温度参数多次提问时,它可能也无法产生足够的变化。所以模型输出存在内部同质性,同时我们也发现了模型间的同质性,也就是说,Llama、GPT 和 DeepSeek R1 的行为都惊人地相似。
Even for open-ended questions, the models are not as diverse as we would have expected, to the point that even when you ask multiple times with higher temperature, it may not be able to vary as much. So there's intra-model homogeneity in the model output, as well as we find inter-model homogeneity, meaning you know, Llama, GPT, and DeepSeek R1, they all have strikingly similar behavior.
好的,各位。欢迎收听新一期的 TwiML AI 播客。我是主持人 Sam Charrington。今天邀请到的是 Yejin Choi。Yejin 是斯坦福大学计算机科学系和以人为本人工智能研究院(HAI)的教授和高级研究员。在开始之前,请务必花点时间点击订阅按钮。Yejin,欢迎再次来到播客,好久不见了。
All right, everyone. Welcome to another episode of the TwiML AI podcast. I am your host Sam Charrington. Today I'm joined by Yejin Choi. Yejin is professor and senior fellow at Stanford University in the Computer Science Department and Institute for Human-Centered AI or HAI. Before we get going, be sure to take a moment to hit the subscribe button wherever you're listening to today's show. Yejin, welcome back to the podcast. It's been a while.
哦,是的。谢谢你再次邀请我。
Oh yeah. Thanks for having me back.
当然,当然。我记得我们上次对话是在 2021 年秋天,这在 AI 领域感觉像是很久以前了。我想直接切入正题,请你介绍一下你从那以后一直在做什么。另外,对于没听过那期节目的听众,或许可以先从你的背景说起。
Absolutely. Absolutely. I think we last spoke in the fall of 2021, which seems like ages ago in AI years. I would love to kind of jump in and have you bring us up to date on what you've been working on since then. And actually, for folks who didn't catch that one, maybe start with a little bit about your background.
上次上你节目的时候,我可能还以研究常识知识和推理而闻名,当时我也做了不少自然语言生成的工作。当然,从那以后发生了很多事。最近,我对推理很感兴趣,尤其是让小型语言模型更好地推理。所以我广泛关注大型语言模型、小型语言模型、大型推理模型、小型推理模型,以及如何让模型更好地对齐多元化的规范和价值观。
At the time when I was on your podcast, I was still maybe best known for working on commonsense knowledge and reasoning, and back then I was also working on natural language generation quite a bit. Of course, since then a lot has happened. So more recently, I've been excited about reasoning, especially making small language models reason better. So I'm broadly interested in large language models, small language models, large reasoning models, small reasoning models, and then how we could make models align better for pluralistic norms and values.
很好。是什么驱使你对小型语言模型感兴趣?看起来大部分动作都在大型语言模型上,而我们正在努力让小型模型达到同样的性能水平。你的兴趣具体是由什么驱动的?
Nice. Nice. What drives your interest in SLMs? Seems like a lot of the action is in large language models, and we're working hard to get the smaller ones up to the same level of performance. What's your particular interest driven by?
是的,我们的使命实际上是让生成式 AI 民主化,这样不仅那些能购买大量 GPU 的公司可以创建、采用和提供 LLM 服务,像我这样的学者和同事们也能做到。例如,我们买不起那么多 GPU,那么有没有一些真正有意义且有趣的事情,即使使用更小的模型也能完成?归根结底,我相信这在根本上是可行的。只是这个世界在探索规模扩展带来的效果上投入了太多。而如果我们投入哪怕一小部分,但稍微多一点,我确实认为我们可以从小型语言模型中解锁更多令人兴奋的能力。我的研究部分也源于寻找更好的方式向机器传授智能的愿望。目前这太以数据为中心了。我们可以在播客后面详细讨论。但它太依赖数据了,这几乎是我们知道的唯一教 AI 学习人类知识和智能的方法。但未来,我不知道我们是否能找到解决方案,但作为一名学者,我觉得我们必须尝试找到一个完全更好的解决方案,它更数据高效,能用更少的数据学到更多。
Yeah, so the mission really is democratizing generative AI so that it's not just companies who can purchase a lot of GPUs that are able to create LLMs and adopt LLMs and serve LLMs, but also people like myself and colleagues who are academics. For example, we cannot buy as many GPUs, and is there something really meaningful and fun that we could do even with a smaller counterpart? And at the end of the day, I believe that fundamentally it should be feasible. It's only that the world has invested so much more into exploring what happens when you scale things up so much. Whereas if we invested even a fraction of that investment, but just a little bit more, I do think that we can unlock a lot more exciting capabilities out of small language models. Part of my research is also driven by the desire to find really better ways of teaching intelligence to machines. Currently it's just so data-centric. We can talk about that in more detail later in this podcast. But it's so data dependent, and that's pretty much the only way we know how to teach AI about human knowledge and intelligence. But in the future, I don't know whether we will find the solution or not, but as an academic, I feel like we have to give it a try to find an entirely better solution to this that is so much more data efficient and able to learn so much more with much less data.
当你思考这个领域和行业如何演变,以及你关于投资去向的评论时,你认为为什么会这样?你觉得投资只是迅速追随了有效的方法,而没有花时间退后一步识别所有优化的机会吗?还是你认为小型模型存在特定的障碍,使得它们本质上更具挑战性?
When you think about how the space and the industry evolved, and your comment about where all the investment has gone, why do you think that is? Do you feel like the investment has just kind of followed quickly what works, without us taking time to step back and identify all the opportunities to optimize? Or do you think that there are particular impediments to smaller models that make them inherently more challenging?
肯定存在雪球效应和羊群效应。你看到其他船往哪里走,然后你也想跟着走,因为那是一个安全的选择,尤其是当融资对 AI 来说比以前更容易的时候。因此,那是提升智能的可靠且经过验证的方法。为什么不呢?事实上,我并不反对这种努力。看到规模能解锁多少智能确实很有趣。我很欣赏有些人疯狂地去探索规模带来的前沿。话虽如此,我确实担心每个人都做同样的事情。我认为尝试不同的想法非常重要。尤其是历史上,无论是计算机还是手机,任何创新一开始都非常大,然后随着时间的推移,人们找到了如何让它们更小但更强大的方法。所以同样的事情也一定会发生在生成式 AI 上。事实上,已经有很多研究工作在让模型更小但更强大。我认为如果我们投入更多的心思和努力,我们可以做得更多、更好。
There's definitely a snowball effect and a herding effect. You see other ships going where, and then you want to follow that because it's a safe choice, especially when raising funding is relatively easier than it used to be for AI. Therefore, that's a guaranteed and proven way of increasing intelligence. So why not? And in fact, I'm not against such effort. It's really interesting to see and watch how much intelligence scale can unlock. I appreciate that some people went crazy and found out the frontier of what happens with scale. Having said that, I do worry about everybody trying the same thing. I think it's very important that we try different ideas. Especially historically, whatever innovation happened with computers or phones, they were always very large at the beginning, and then over time people figured out how to make them smaller yet more powerful. So the same thing will definitely happen with generative AI as well. In fact, already there is a lot of research effort that makes models smaller but more powerful. And I think we can do so much more, so much better, if we put more mind and effort into it.
那么你如何看待解决这个问题的不同攻击方向或方法?你觉得哪些已经被探索过了,而哪些机会到目前为止还没有被有效探索?
And how do you think about the different attack vectors or approaches to tackling this problem? What do you feel like is already being explored, and where do you think there are opportunities that really haven't been explored very effectively thus far?
是的,有多个途径。我认为一开始人们试图通过量化或剪枝神经元等方式将大模型压缩成小模型。所以这从某种意义上说是一种基于优化或更机械的优化方法,将大模型变成小模型。而且它确实需要大模型才能制造小模型。就是这样。这没什么错。有这个选项很好。但这不是唯一的方法。所以短期内,我认为新的架构,比如状态空间模型和传统 Transformer 的混合体,例如 NVIDIA 的 Mamba Hybrid,可能是让小型模型更强大的另一种方式。但还有其他方式,比如让数据更好,尤其是提供更强大的数据。这些数据通常必须在互联网数据的边缘,也就是互联网无法提供的那种数据,以便教模型更好地进行某些推理。所以如果我们有更高质量的数据,小型模型通常学得更快。这是另一种方法。
Yeah, so there are multiple routes. I think at the beginning people were trying to think about compressing larger models into smaller models by either quantizing it or pruning some neurons and stuff like this. So that's more like, in some sense, an optimization-based or a little bit more mechanical optimization-based approach to making larger models into smaller models. And it does require larger models in order to make smaller models. So there's that. Nothing wrong with that. It's nice to have that option. But it's not the only way. So in the short term, I think having new architectures, like a hybrid between state space models and conventional transformers, like Mamba Hybrid from NVIDIA for example, could be an alternative way of making small models more powerful. But there can be other ways, such as making data better, especially providing much more powerful data. And this data usually has to be at the outer skirt of the internet data, meaning the kind of data that the internet couldn't quite provide, in order to teach the model to do certain kinds of reasoning better. So if we have much higher quality data, usually small models learn so much faster. So that's another way.
当你说互联网边缘的数据时,有哪些来源的例子?我觉得大家常说我们已经找到了所有公开可用的数据,并基于这些数据训练了所有大模型。而且经常有人提出,未来将来自解锁新类型的数据。例如,视频是人们常说的一个明显例子。你指的是这类东西,还是你对小模型有效的数据有其他想法?
When you say data kind of on the outer reaches of the internet, what are some examples of these sources? Like I think it's commonly thrown around that we found all the data that is available to the public and we train all of the large models based on this data. And it's often proposed that the future is going to come from unlocking new types of data. For example, video I think is the obvious one that people talk about. Is that the kind of thing you're referring to or do you have other ideas about what will be effective for small models?
大致来说,是的。但让我澄清一下我所说的“更好数据”是什么意思,因为互联网数据并不差,至少在数量上。但当我们看 LLM 流程时,预训练模型永远不够好,尽管预训练的规模很大。你必须在大量通常不同于互联网数据的数据上进行后训练。监督微调和基于人类反馈的强化学习(RLHF)需要那些由人类专门为教学 AI 而整理的数据点。这些不是你从互联网上下载的东西,而是你可能付钱让人写的数据点。如今更常见的做法甚至不是众包数据,而是雇佣专家,比如律师、前国际数学奥林匹克竞赛获奖者。这些是真正的专家,你让他们为你写数据。所以也收集了大量专家数据。即便如此,对 AI 来说还是不够,因为 AI 非常依赖数据。所以最近,人们也做了很多自动合成数据生成。合成数据,如果你用普通方式做,只是让 LLM 为你写一些问题或解决方案,通常不够好,或者可能只是重复同样的东西。所以它需要更多的努力,比如设计提示词,并有一个由不同模型组成的流程来改进提示词或改进解决方案,反复迭代。这不像让 ChatGPT 为你写数据那么简单。但如果我们做得相当好,它可以产生互联网上不存在的新数据点。它可能是真正高质量的数据,在性质上与互联网上的数据不同。一个典型的例子是困难的数学解答。互联网数据确实有很多数学,但不一定有很多困难数学问题的解答。所以你必须想出这些解答,要么请专家为你写解答,要么以某种方式使用 LLM 来生成好的解答,尽管它们还不太有能力做到。但你可以使用例如带有验证器的强化学习,它可以进行大量探索,并查看 AI 生成的哪些解答根据验证器恰好是正确的。然后你收集这些数据,这些数据在 RLHF 中被隐式用作好的数据来增强模型行为。但有时人们随后使用这些数据点进行模仿学习,迭代进行。
Roughly speaking, yes. But let me clarify what I meant by better data because internet data is not so bad, at least in terms of quantity. But when we look at the LLM pipeline, the pre-trained model is never good enough despite the scale of pre-training. You have to do post-training on a fairly large amount of data that usually is different from internet data. Supervised fine-tuning as well as RLHF requires data points that are either curated by humans just for the purpose of teaching AI. These are not things you just download from the internet, but you may pay someone to write those data points. The more common practice these days is not even crowdsourcing your data but rather hiring experts like lawyers, former International Math Olympiad winners. These are real experts, and you have them write data for you. So a lot of expert data is being collected as well. And even that is not enough for AI because AI is so data dependent. So more recently, people also do a lot of automatic synthetic data generation. Synthetic data, if you do it in a vanilla way, just ask LLMs to write some problems or solutions for you, often it's not good enough or it could be just repetition of the same thing. So it requires a lot more effort in the way you design prompts and have a pipeline of different models making the prompt better or making the solution better, revising the solution, taking lots of iterations. It's not as simple as just asking ChatGPT to write data for you. But if we do it quite right, it can lead to new data points that didn't exist on the internet. It could be really high-quality data that is qualitatively different from what was on the internet. A prime example is hard math solutions. Internet data does have a lot of math, but it doesn't necessarily have solutions to a lot of hard math problems. So you have to come up with those solutions either by asking experts to write solutions for you or using LLMs in some ways to generate good solutions even though they're not quite capable of doing it yet. But you could use, for example, reinforcement learning with verifiers that can explore a lot of explorations and see which solutions that AI generated happen to be correct based on the verifier. Then you collect that data, which is implicitly used as good data during RLHF to amplify that model behavior. But sometimes people then use those data points to do imitation learning on top, iteratively.
这听起来像是应用了相当广泛的各种方法。你谈到了合成数据生成、模仿学习、强化学习。这些在历史上都是各自的研究领域,并独立实践,而现在你在谈论将它们整合在一起。这是你认为解决这个问题的关键要素吗?整合很多这些想法。
That sounds like the application of a fairly broad variety of approaches. You talk about synthetic data generation, imitation learning, RL. These are all things that historically have been their own field of research and put into practice independently, and now you're talking about integrating them together. Is that a significant element of how you think this problem gets solved? Integrating a lot of these ideas.
是的。事实上,从某种意义上说,人工推理的艺术相当人为,我们必须协调所有这些复杂的、几乎是系统式的研究,以确保事情在正确的时间、以正确的方式、按正确的顺序完成,然后反复迭代。
Yes. And in fact, in some sense, the art of artificial reasoning is fairly artificial in the way that we have to orchestrate all this complex, almost system-style research in order to make sure that things are done at the right time, in the right way, in the right sequence, and then iterate over and over.
你能详细说明一下模仿学习的应用,以及你认为它在流程中如何发挥作用吗?
Could you elaborate a little bit on the use of imitation learning and how you see that playing into the pipeline?
是的。我一时想不起是哪个公司哪个模型,可能是 Llama 3,其白皮书描述了它们如何进行后训练。它们当然重复了类似的过程:首先是预训练,然后是顺序微调、指令微调,以及在一些考试风格数据上的监督训练,最后是强化学习阶段。在 RL 和 SFT 之间进行迭代并不罕见,这样在完成 RL 后,你发现一些好的行为,然后你想强化它。另一个例子:DeepSeek R1 在强化学习之后做了一定量的模仿学习,它们提供了 DeepSeek R1 的蒸馏版本给较小的模型。有趣的是,它们不在这些小规模模型上直接进行强化学习,而是从已经通过强化学习训练的更强模型中进行蒸馏。它们说,如果你只做强化学习,顺便说一句,有时模型会在解决数学问题的中途开始代码切换,突然在中文和英文之间来回切换,或者一些其他对人类读者可能没有意义的外语。所以强化学习只关心你是否得到了正确的最终解答,它不关心你是如何得到的。所以奇怪的行为可能会出现并被强化。所以如果你不想要那样,如果你希望思维链的方式可以被人类解释和验证,那么你只想要那些导致你喜欢正确解答的思维链。所以你可以通过这个经过强化学习的更强模型过滤掉许多这样的输出示例或解答,只收集那些更好的示例来教较小的模型。总的来说,这导致了一个纯粹基于模仿学习的模型,它非常强大,但通常具有你想要的更好行为。
Yeah. I'm blanking on which company which model, it may have been Llama 3 actually, whose white paper describing how they did the post-training. They were repeating, of course, something like: first pre-training, then sequential fine-tuning, instruction tuning, as well as supervised training on some exam-style data, and then finally a reinforcement learning phase. It's not uncommon to take the iteration between RL and SFT such that after doing RL, you find some good behaviors and then you want to drill on that. Here's another example: DeepSeek R1 does some amount of imitation learning after the reinforcement learning in that they provide distilled versions of their DeepSeek R1 in smaller models. What's interesting is that they don't do straightforward RL at those small scale models, but rather do distillation from the stronger model that's already trained through reinforcement learning. They say that if you just do reinforcement learning, by the way, sometimes this model starts code-switching in the middle of solving math problems, suddenly speaking in Chinese and English and back and forth, or some other foreign languages that may not make sense to human readers. So reinforcement learning only cares about whether you got the final solution right or not. It doesn't care about how you got there. So strange behaviors can be emergent and then reinforced. So if you don't want that, if you want interpretability of the way that chain of thought can be interpreted by humans and verified, then you want only those chain of thought that leads to the correct solution that you like. So you can filter out a lot of these output examples or solutions by this stronger model that went through reinforcement learning and collect only those better examples in order to teach smaller models. In general, this leads to a model based purely on imitation learning that's very powerful but often has better behavior that you wanted.
当我们尝试将其应用于更传统的数据(文本数据)时,我不太清楚这与监督微调或基于人类反馈的强化学习(RLHF)等有何区别。你能解释一下在你描述的应用或流程中,是什么让它成为模仿学习吗?
When we try to apply that to more traditional data, textual data, it's not super clear to me how that is distinguished from just supervised fine-tuning or RLHF or something like that. Can you explain what makes it imitation in the application or pipeline you described?
哦,在这个语境下,模仿学习就是指监督微调。但有时人们会在强化学习之前或之后提到模仿学习,因为那时有一些你想模仿的示例轨迹,但这实际上就是监督微调。
Oh, imitation learning just means supervised fine-tuning in this context. But sometimes people say imitation learning when things were performed either right before or after RL, because then there are some example trajectories that you want to imitate from, but it's really just supervised fine-tuning.
好的,明白了。
Okay. Got it.
我们之前谈到,你可能想做的事情之一是用模型为你生成合成数据。你提到这往往行不通。我想你没明说,但我猜我们是在讨论模式坍缩。这让我想起了《人工蜂群思维》那篇论文,它在刚过去的 NeurIPS 上被评为获奖论文之一,真正讨论了这种模式坍缩的影响。你能谈谈那篇论文以及你觉得有趣的地方吗?
We were talking a little bit about one of the things you might want to do is to use the model to generate synthetic data for you. And you talked about how that tends to not work. I don't think you mentioned it, but I'm assuming we're talking about mode collapse here. And that brought to mind the artificial hive mind paper, which was highlighted as one of the award-winning papers at this past NeurIPS, which really talked about one of the implications of this kind of mode collapse. Can you talk a little bit about that paper and what you found interesting about it?
当然。模式坍缩确实是 LLM 生成中的一个真正问题。我们在论文中发现,即使你问开放式问题,比如“讲一个关于时间的笑话”或“说一些关于时间的智慧之言”或“讲一个关于某事的故事”,你会期望没有唯一正确答案,因此语言模型应该能够生成多样化的解决方案。即使你问“给我一个 0 到 10 之间的随机数”,它也不是随机的。
Sure. So mode collapse is a real concern with LLM generation. What we find in our paper is that even when you ask open-ended questions like 'tell me a joke about time' or 'tell me something wise about time' or 'tell me a story about something', you would expect that there's no one good answer, so language models should be able to generate a diverse set of solutions. Even when you ask 'give me a random number between 0 and 10', it's not random.
想想看,对吧?
Just thinking about that, right?
是的。通常是 7,或者如果你问更大的范围,可能是 13。所以它不是随机的,因为数据本身就不是随机的;数据一开始就有偏斜。所以当你从数据分布中采样时,你会得到一个偏斜的分布。这是问题的一部分。即使在预训练之后,更大的问题是在后训练(如顺序微调)之后,模型的输出概率变得更加偏斜,集中在人们倾向于喜欢的刻板答案上。所以我们发现,即使对于这些问题——当然有些问题你根本不应该改变答案,比如事实性问题——你不会改变。但即使是开放式问题,模型的多样性也不如我们预期的那么高,以至于即使你多次询问并提高温度,它也可能无法产生足够的变化。所以模型输出存在模型内同质性,以及模型间同质性,意思是 Llama、ChatGPT 和 DeepSeek R1 都有相似的行为,惊人地相似。有时它们生成的输出几乎逐字相同,这非常奇怪。这就是我们在 NeurIPS 上展示的论文《人工蜂群思维》的要点。这让我有些担忧,因为越来越多的人使用 LLM 在互联网上发帖。我想知道我们的互联网会变成什么样。互联网曾经是人类智能的产物;它真正 encapsulates 了人们写作和思考的截然不同的方式。它是人类智能的历史产物。现在它正变成 LLM 的产物,混合了一些人类智能。但如果它变得更加同质化,不再反映人类思想的多样性呢?我们将失去一些有价值的东西,这是我的担忧。
Yeah. It's usually like seven, or 13 if you ask a bigger range. So it's not random because data is not random; data is skewed in the first place. So when you sample from the distribution of the data, you get a skewed distribution. That's part one of the problem. Even after pre-training, the bigger problem is after post-training like sequential fine-tuning, the output probability of the model becomes even more skewed, zoning in on the stereotypical answers that people tend to like. So we find that even for those questions—of course there are questions for which you shouldn't vary the answer at all, like factual questions—you don't vary that. But even for open-ended questions, the models are not as diverse as we would have expected, to the point that even when you ask multiple times with higher temperature, it may not be able to vary as much. So there's intra-model homogeneity in the model output, as well as inter-model homogeneity, meaning Llama, ChatGPT, and DeepSeek R1 all have similar behavior, strikingly similar behavior. Sometimes they generate output that's almost verbatim identical, which is very strange. So that's sort of the gist of this paper we presented at NeurIPS, 'Artificial Hive Mind'. It's a bit of a concern to me because more and more people use LLMs to post things on the internet. I wonder what happens to our internet. The internet used to be the artifact of human intelligence; it really encapsulates the vastly different ways people write and think. It's a historical artifact of human intelligence. Now it's really becoming the artifact of LLMs mixed with some amount of human intelligence. But what if it becomes more homogeneous and less reflecting the diverse spectrum of human thoughts? We're going to lose something valuable, that's my concern.
你们是否也研究了第二部分,即这些影响,还是主要专注于展示效果以及模式坍缩在这些开放式场景中是如何运作的?
And did you study that kind of second part, the implications of this as well, or were you primarily focused on demonstrating the effect and how mode collapse works in these open-ended scenarios?
我们的研究止于研究这些模型有多同质化,尤其是在后训练之后。预训练模型在这方面更好。不过这次谈话让我想起一项研究,它考察了 ChatGPT 前后 Reddit 论坛的语言使用情况,发现即使是 Reddit 帖子也不如以前多样化了。
Our study stops at just studying how homogeneous these models are, especially after post-training. Pre-trained models are better in this regard. Although this conversation reminds me of a study that looked at the language use of Reddit forums before and after ChatGPT, and they found that even Reddit posts are not as diverse as before.
现在有很多“深入探讨”这个词,以前没有。
There's a lot of 'delve' now that wasn't happening before.
是的,很可能。
Yeah, probably.
你知道吗,实际上每当我看到任何人写作中出现“delve”这个词,我就会想,嗯,你做了什么?
You know, actually whenever I see the word 'delve' in anybody's writing, I'm like, hm, what did you do?
是啊。有趣的是,使用各种 LLM 的人经常看到这种行为。比如你跨多个 LLM 问一个相当开放的问题,你会得到非常相似的回应。在训练数据的背景下,经常有人问:如果模型在吸自己的尾气,或者在这些合成数据上训练,我们如何改进模型?这意味着我们加速了进一步的模式坍缩。但我觉得这篇论文有趣的是,是的,有这个问题,而且还有对读者、生态系统、以及在这个环境中合成数据被作为文章发布的人类的影响是什么?它是否改变了我们的思维方式?你隶属于 HAI,这似乎是跨学科研究这个问题的好地方。我很想听听更多。如果你听说这方面有什么工作,请告诉我。
Yeah. It's interesting that folks who use a variety of LLMs see this behavior a lot. Like you ask a fairly open-ended question across multiple LLMs and you get very similar responses. In the context of training data, the question is often asked: how do we improve models if models are smoking their own exhaust, or training on this synthetic data? The implication is that we accelerate further mode collapse. But what I found interesting about this paper is that yes, there's that, but also what is the impact on the reader, the ecosystem, the humans in this environment where synthetic data is being posted as articles? Is it changing the way we're thinking? You're affiliated with HAI, which seems like the place to cross-interdisciplinarily study this. I'd be super interested in hearing more about that. Let me know if you hear any work in that regard.
哦,是的。我在斯坦福的一半隶属关系是 HAI,即以人为本的人工智能研究所。因此,我一半的研究都与人工智能对人类的影响有关。我个人认为,我们可以从人工智能中获得很多好处,但也存在担忧。当前情况的棘手之处在于,好处和潜在危害并存,而且在某些方面,取决于我们今后如何开展人工智能研究,未来可能会截然不同,这是我的感觉。
Oh yeah. So half of my Stanford affiliation is with HAI, the Human-Centered AI Institute. Therefore, easily half of my research has to do with AI's impact on humanity. I personally think there's a lot of benefit we could get from AI as well as concerns. The thorny thing about the current situation is that both the benefit and potential harms coexist, and in some ways, depending on how we pursue AI research from here on, the future can be drastically different, is how I feel.
一方面,LLM 可能会影响人类智能,导致我们失去个性和多样性。但另一方面,也可能走向相反的未来。至少对一些人来说,他们或许能在 AI 帮助下变得更加专业和富有创造力。而对另一些人来说,他们可能选择过度依赖 AI,失去自己的思考,只会重复 AI 说的话。所以最佳和最差情景之间的差距可能会扩大。无论如何,我认为意识到潜在危害和最坏情况很重要,这样才能采取行动。例如,我们发现 LLM 在微调后变得更加同质化。作为后续研究,我以前的学生 Taylor Sorenson 提出了频谱调优,这是一种后训练方法,让模型保留多种输出方式,而不是只聚焦于后训练数据中的正确答案。所以当我们意识到问题时,就可以尝试寻找解决方案,要么设计新的后训练算法,要么确保后训练数据本身是多样化的。我认为在这个领域还有很多研究要做,以减轻生成式 AI 的潜在担忧。
You know, on one hand, it could be that LLMs influence human intelligence such that we lose individuality and diversity. But on the other hand, it could also lead to a future in which the opposite is true. At least for some humans, I think they might be able to become even more specialized and creative with the help of AI. For many others, they might choose to be overly dependent on AI and lose their own thinking, just repeating whatever AI says. So the gap between the best-case and worst-case scenarios might actually increase. In any case, I believe it's important to be aware of potential harms and worst-case scenarios in order to do something about them. For example, we found that LLMs become more homogeneous after fine-tuning. As a follow-up, my former student Taylor Sorenson worked on spectrum tuning, a post-training method that teaches the model to retain a spectrum of different ways to generate output, instead of just honing in on correct answers from the post-training data. So when we are aware of the problems, we can try to seek solutions, either by designing new post-training algorithms or ensuring that the post-training data is diverse in the first place. I think there is a lot more future research to be done to mitigate the potential concerns about generative AI.
当你开始描述两种可能结果以及差距扩大时,你提到我们开展研究的方式在很大程度上决定了我们会走向哪种结果。你能详细谈谈这一点,以及你认为研究在决定方向上的作用吗?
When you started to describe this idea of two possible outcomes and that gap widening, you talked about the way we pursue research as being kind of central to which outcome we tend towards. Can you elaborate on that and the role you see research having in determining direction?
是的。我确实认为 LLM 在数学数据上表现很好,但数学问题不一定对人类最有益。这样说有点粗糙,但它概括了我对 AI 与人类未来的看法:我们必须更明确地解决具体问题。例如,如果我们关心民主,就需要设计 AI 让人类更民主,通过民主过程相互理解,并处理不同意见,而不是构建优化注意力和参与度的 AI,那可能会加剧紧张。利润激励不一定与人类应追求的目标一致。所以为了得到我们想要的,我们不能只交给几家科技公司。我相信他们有很多善意的人,但在利润和参与度竞争下,事情可能会朝着对人类不利的方向发展。当我思考 AI 民主化时,我喜欢把它看作人类、为了人类、由人类创造的 AI。AI 为人类服务应该为全人类服务,而不仅仅是为某些国家的某些公司工作的人。如果我们不小心,它可能变成 AI 为 AI 服务,甚至更糟,人类为 AI 服务。这就是为什么非营利部门参与设计 AI 的未来非常重要,而不仅仅是营利部门。
Yeah. I do think that LLMs are doing really well on math data, but math problems will not necessarily be the most beneficial for humanity. That's a crude way of saying it, but it encapsulates what I believe about the future of AI on humanity: we really have to work on specific problems more explicitly. For example, if we care about democracy, we need to design AI that makes humans more democratic, helps them understand each other through democratic processes, and work with different opinions, as opposed to building AI that optimizes for attention and engagement, which could increase tension. Profit incentives are not necessarily aligned with what humanity should aspire to achieve. So to get what we want, we cannot just leave it to a few tech companies. I'm sure they have many well-intentioned people, but with profit and engagement competition, things could unfold in a way not beneficial for humanity. When I think about AI democratization, I like to think of it as AI of humans, for humans, and by humans. AI for humans should be for all humans, not just some working for companies in some countries. If we are not careful, it could become AI for AI, or even worse, humans for AI. That's why it's very important for nonprofit sectors to participate in designing the future of AI, not just the profit sectors.
这呼应了这样一种观点:我们这一代最聪明的人专注于让人们点击广告,而不是推动人类和科学进步。你暗示这在 AI 领域也有类似情况,我们需要主动定义自己的未来。
Part of that echoes the idea of the brightest minds of our generation focused on making people click ads, as opposed to advancing humanity and science. You're suggesting there's an AI-oriented aspect to that as well, and we need to be proactive in defining our future.
是的。需要更多投资来支持那些真正思考 AI 对人类影响的研究,而不仅仅是提高数学问题的基准分数。
Yeah. More investment is needed to support research that really thinks about AI's impact on humanity, not just increasing benchmark scores on math problems.
我们回到之前的话题。我们正在讨论小模型,大致谈到了如何让小模型表现更好。但我不记得我们深入探讨过推理和小模型的具体问题,以及推理在小模型或小推理器上的独特特征。你也在研究这个吗?
Let's come back to that. We were in the middle of our small models conversation, talking broadly about making small models perform better. But I don't recall us getting to the specifics of reasoning and small models, and the unique characteristics of reasoning as they pertain to small models or small reasoners. Is that something you're looking at as well?
是的,我对让小型语言模型成为更好的推理者感到非常兴奋,尤其是因为推理是任何未来模型都必须具备的重要智能能力。这也是一个有趣的挑战,因为互联网数据并不能让 LLM 立即很好地推理。它们可以在一定程度上推理,但需要更多的后训练来注入更好的推理能力。我最近的研究重点是让小模型更好地推理。这需要通过在顺序微调阶段提供更好的数据——高质量且多样化的数据——以及其他算法方法,从看似无望的小语言模型中榨出更好的智能。
Yeah, I'm quite excited about making small language models better reasoners, especially because reasoning is such an important intelligence capability that any model in the future has to be good at. It's also an interesting challenge because internet data doesn't really equip LLMs to reason well right away. They can reason to some degree, but they require a lot more post-training to infuse better reasoning capabilities. My recent research focuses on how to make small models reason better. This requires both feeding in better data through sequential fine-tuning phases—data that is high quality and diverse—as well as other algorithmic approaches that can squeeze out better intelligence from even small language models that seemed kind of hopeless.
在数据整理方面,有没有什么具体的技术出现,还是主要靠人工劳动来整理大型数据集以提高质量?你看到了什么?
Are there any specific techniques coming to the fore with regards to the data curation side of things, or is it largely manual labor with humans curating large datasets to increase quality? What are you seeing there?
我认为合成数据前景广阔。我可以举一个我们最近工作的例子,叫做 Prismatic Synthesis。这是一种合成数据生成算法,像棱镜一样散射光线,使其更加多样化。
I think there's a huge future in synthetic data. I can give you one example from our recent work called Prismatic Synthesis. It's a synthetic data generation algorithm that acts like a prism, scattering light to make it more diversified.
所以我们做的基本上是数学问题合成,更准确地说,是方法、问题和解决方案的合成。我们使用 DeepSeek R1 320 亿参数模型作为教师模型。320 亿参数算是中等规模,比现在的中等规模稍大一点,但远不如完整版的 DeepSeek R1——那个最大的模型有 6710 亿参数,比我们选作教师模型的模型大 20 倍。在这项工作中,我们主要专注于使用这个中等规模的教师模型为困难数学问题制作序列微调数据,目标是挑战另一种方案——使用强大 20 倍的教师模型。总的来说,这是一场非常难打的仗。要击败一个强大 20 倍的教师模型非常困难,因为性能差距很大。那么我们究竟如何缩小差距呢?靠的是算法化的数据过滤方法,确保生成数据的多样性。因为无论你的教师模型多好,正如我们在人工蜂群思维论文中展示的,它们都会重复、同质化。所以你必须花大力气让数据多样化。我们多样化的方法是:使用一个代理模型(一个小规模的代理模型,小到只有 15 亿参数,我们直接用网上下载的 Qwen 1.5B)来计算输出相对于输入的梯度向量。这些输出-输入对就是我们刚刚用 DeepSeek R1 320 亿参数模型合成的数据。然后我们查看教师模型生成的每个数据点的梯度表示,再通过 k-means 聚类(一种老派的聚类机制,但在今天仍然有效)来看它们之间的差异。我们做张量化的 k-means 聚类,看看哪些簇过密、哪些数据点过疏。我们非常激进地过滤掉过密的数据点,扔掉合成数据中的绝大部分,只保留那些独特且互不相同的,然后用这些数据点去提示教师模型进行下一轮。我们就这样迭代:过度生成,然后用梯度向量激进过滤,再过度生成、过滤,直到收集到 100 万个数据点。这中间有大量的过度生成和过滤。然后我们发现,这 100 万个数据点实际上比用更强的教师模型(最好的教师模型)生成的 100 万个数据点还要好。
So what we do is basically math problem synthesis, where actually it's more like method, problem, and solution synthesis. We're doing this using DeepSeek R1 32 billion parameter model as the teacher model. 32B parameters is medium size, a little bit bigger than medium size these days, but it's much worse than DeepSeek R1 the full model, the biggest model which is 671B parameter model, so that's like 20 times bigger than the model we choose to use as teacher model. In this work, we primarily focus on making sequential fine-tuning data for hard math problems using this medium-scale teacher model, and then our goal is to compete against the alternative, which is to use a much stronger teacher that's 20 times larger. Now, in general, that's a really difficult game to play. It's really hard to beat a teacher that's 20 times larger because the performance gap is significant. So how on earth do we close the gap? It's algorithmic ways of filtering the data to ensure the diversity of the generated data. Because no matter how good your teacher is, as we demonstrated in our artificial hive mind paper, they're all repetitive, they're all homogeneous. So you have to put a lot of effort into diversifying them. The way we diversify data is we look at the gradient vector of output given input using a proxy model, a small-scale proxy model. It's so small, it's only 1.5 billion parameter model. We just use Qwen 1.5B, just downloaded from the net, and we use it as a proxy model to compute the gradient of output given input. These output-input pairs are the synthetic data that we just synthesized using the DeepSeek R1 32 billion parameter model. So we look at the gradient representation of each data point that this teacher model generated, and then we look at how they differ from each other by doing k-means clustering. It's an old-fashioned clustering mechanism that still works in this modern day. So we do tensorized k-means clustering to see which clusters are overrepresented and which data points are underrepresented. We filter out overrepresented data points really aggressively. Like we throw out the vast majority of all the data that we just synthesized, and only maintain those that are unique and different from each other, and then use those data points to prompt the teacher model in the next round. So we iterate through this: overgenerate, then filter aggressively using gradient vectors, then overgenerate and filter aggressively, until we gather 1 million data points. So it's a lot of overgeneration and filtering. And then we find that that 1 million data points is actually better than the 1 million data points that you generate from the stronger teacher model, the best teacher model.
你能谈谈在这个例子中你看到的提示和回答是什么样的吗?
Can you talk a little bit about the kinds of prompts and responses that you're seeing in this example?
就是一些困难的数学题。在我们的案例中,全是需要很长解答过程的困难数学题。顺便说一句,当我们自动生成所有解答时,我们完全不知道答案是否正确,对吧?所以我们玩了个小把戏。这个把戏很简单。当我们生成题目时——在我们的案例中,题目也是完全合成的。很多其他合成数据通常依赖互联网上已有的题目,这样你只解决合法的问题。但我们的情况是,我们也生成题目,因为我们真的想探索数学推理领域的多样化范围。但这些为虚构题目生成的解答可能正确也可能不正确。所以我们做的是:让模型多次求解同一个问题,然后检查最终答案是否一致。如果不一致,我们就担心质量可能不好。这是一种非常粗糙的数据过滤方法,但对我们的案例来说效果足够好,而且我认为这是一种无需人工验证就能控制合成数据质量的有效方法。
It's just some hard math problem. In our case, it's all hard math that requires a very long solution. And by the way, when we auto-generate all the solutions, we have no idea whether the answer is correct or not, right? So we play a bit of a trick. The trick is simple. When we generate the problem—in our case, it's fully synthetic, even the problems. A lot of other synthetic data usually relies on problems that exist on the internet, so you only solve legit problems. But in our case, we generate the problems as well because we really wanted to explore a diverse scope of the math reasoning domain. But these solutions generated for fake problems may or may not be correct. So what we do is we ask the model to solve the same problem multiple times and then check whether the final answer is identical to each other. If not, then we worry that the quality might be bad. It's a very crude way of filtering data, but it worked well enough for our case, and I think this is a powerful method for controlling the quality of synthetic data without human validation.
明白了。你提到生成一个包含一千个数据点的数据集,而且我想是分多轮进行的。我想了解的是,提示在每轮之间的变化程度如何?你是从一个提示开始,然后每轮基于某种对数数量的变体来发展变异性,而提示本身没有任何变化吗?
Got it. And so you mentioned generating a dataset of a thousand points and it's happening over multiple rounds, I guess. I'm trying to understand the degree to which the prompt is varied across rounds, or are you starting from a prompt and then developing variance based on some logarithmic number of variants with each round without any variance of the prompt?
好问题。关于提示,我们会展示一些例子,比如“嘿,生成这类数学题”。我们在多轮过度生成和过滤中改变的是提示中展示的例子。这些例子来自之前的迭代,希望模型在上下文中看到更新、更多样的例子时能获得更多灵感。
Excellent question. So for prompting, we show some examples: 'Hey, generate math problems of this kind.' And what we change through these multiple iterations of overgeneration and filtering is the examples that we show in the prompt. So the examples come from the previous iterations, in the hope that the model may feel more inspired when provided with newer, more different examples in the context.
在每个阶段,你们会验证问题的答案吗?
And at each stage, are you validating the answer to the question or not?
我们确实会验证我们决定保留的数据。我们会验证。所以最终合成的 100 万个例子,希望大体上都有正确的答案——但不保证。这里的核心思想是,不同于从一个提示开始然后说“生成一百万个例子”并设置一些最终阶段,而是在每个阶段生成一部分,然后从那里扩散开来。你们发现这对保持数据集的一定多样性很有效。所以这是一个简单的想法和简单的方法,这也是这个方法的一大优势。因为原则是:你真的想制作与互联网上容易生成的数据在性质上不同的数据。你需要进入那些相对未被探索的区域。当然,你也需要确保那里的质量也相当不错。但真正重要的是,我们在数据集生产中试图定量增强多样性。我认为仅此一点就能走得很远。如果你真的覆盖了多样化的领域,那么你的模型可以表现得更好,因为归根结底,当前的 LLM,无论多么惊人,都只取决于它们训练所用的数据。所以你必须展示类似的例子。类似的例子越多,模型在测试时处理各种情况就越好。我喜欢这样表述:无论什么分布外的东西,都让它变成分布内的。确保你把所有分布外的都变成分布内的。这就是生成式 AI 的工作原理,甚至自动驾驶汽车也是如此。确保你覆盖了所有行、所有边缘情况,一切。
We do validate whatever we decide to keep. We do validate. So the final 1 million examples that we've synthesized hopefully have, by and large, actually correct answers—not guaranteed. So the core idea here is, as opposed to starting with some prompt and saying 'generate a million examples' with some end stages, at each stage generate some section and kind of disperse from there. And you've found that to work well for maintaining a degree of diversity in the dataset. So this is a simple idea and a simple method, which is the big benefit of this method. Because the principle here is that you really want to make data that's different, qualitatively different from the internet data that's easy to generate. You really need to go to these relatively less explored regions. And then, of course, you need to make sure that the quality is reasonably good there as well. But really, it's diversity that we are trying to quantitatively enhance in the dataset production. And I think that alone can really go quite far. If you really cover diverse ground, then your model can perform so much better, because at the end of the day, current LLMs, no matter how amazing they are, are only as good as the data they were trained on. So you have to show similar examples. The more similar examples, the better for whatever the model may have to deal with during test time. The way I like to put it is: whatever is out of distribution, just make it in distribution. Make sure that you make all the out-of-distribution in-distribution. This is how generative AI works. This is how even self-driving cars work. Make sure that you cover all the rows, corner cases, everything.
而且这应该进入你的训练数据。这与人类学习驾驶的方式非常不同。我们不需要看到大量这些边缘案例,我们直接处理,这是智能的真正奥秘,我希望有一天我们能得到答案。我们如此数据高效。但目前的通用 AI,在当前框架和范式下,唯一的方法是确保所有分布外的情况变成分布内。无论你怎么做,确保它发生。这就是为什么后训练需要整理大量数据,甚至合成和整理,或者两者结合,即便如此也不够。因此,你大规模使用强化学习,因为强化学习是另一种将分布外变成分布内的方法,通过让模型自己探索所有其他未探索的区域。确保在进入测试阶段之前,所有区域都被探索过。
And that should go into your training data. This is really different from how humans learn to drive. We do not need to see a lot of these corner case examples. We just deal with it, which is the real mystery of intelligence that I wish one day we will have some answers for. We're so data efficient. But right now, the general AI, the only way under this current framework and paradigm is making sure that all the out-of-distribution becomes in-distribution. However you do it, just make sure that's going to happen. And so that's why post-training requires curating a lot of data, or even synthesizing and curating, or taking some combination of the two, and even that's not enough. Therefore, you do reinforcement learning at scale, because reinforcement learning is another way of making out-of-distribution in-distribution by having the model explore all these other unexplored areas. Make sure that it's all explored before you get into the testing phase.
每当我参与这些对话,讨论数据的作用以及使用各种技术改进小模型或大模型的数据时,这让我想起几年前你在斯坦福的同事吴恩达开始倡导以数据为中心的 AI。正如你之前提到的,所有 AI 都是以数据为中心的。那么这到底意味着什么?但通过关注训练模型所用的数据而非迭代算法来改进模型的想法,仍然引起共鸣。
Whenever I'm in these conversations and we're talking about the role of data and the idea of using varying techniques to improve the data for small or large models, it brings me back to a few years ago when your colleague at Stanford, Andrew Ng, started planting this banner around data-centric AI. And as you noted earlier, all AI is data-centric. So what does that really mean? But the idea that we're going to improve models by focusing on the data used to train them, as opposed to iterating on the algorithms, continues to resonate.
是的。从这个意义上说,我们并没有走得很远。
Yep. In that sense, we didn't go very far.
嗯。
Yeah.
但我们只是以更实证有效的方式做了更多。而且还有更多可以做的。从这个角度看,那些看似神奇的生成式 AI 前沿模型有点令人失望。但事实就是如此,我认为即使在当前范式下,我们也能做得更好。但这是我们对话中反复出现的主题:必须有一种根本更好的方法,我们能找到吗?在某种程度上,自然找到了解决方案,那就是人类大脑。人类大脑需要的能量如此之少。我们的大脑显然比一个灯泡用的能量还少。
But we're just doing a lot more of it in a more empirically powerful way. And still more can be done. I think it's a bit disappointing when we look at the seemingly magical generative AI frontier models from this lens. But it is what it is, and I think we can do a lot better even following the current paradigm. But this is a recurring theme of our conversation: there must be a fundamentally better way of doing this, and can we find it? In some ways, nature found a solution, which is the human brain. The human brain requires so little energy. Our brain apparently uses less energy than one light bulb.
到目前为止,你提出的一种方法是关注数据,或者创建遵循多样化和分布式模式但质量仍受约束的合成数据。你也在研究将强化学习作为预训练目标的一部分的方法。能谈谈那项工作吗?
So far, you're proposing that one way to do this is to focus on the data, or to create synthetic data that follows a diverse and distributed pattern while still constrained in quality. You are also looking at ways to incorporate reinforcement learning as part of the pre-training objective. Can you talk a little bit about that work?
是的。这是我们最近发表的一篇新论文,大致思路是:在预训练期间,模型被迫完全被动地学习预测下一个词。但如果我们鼓励模型在预测下一个词之前自己思考呢?如果我们鼓励模型通过生成类似思维链的东西来自己思考,然后预测下一个词呢?在这种情况下,奖励——因为是强化学习,我们需要考虑奖励。奖励可以有不同的设计方式,但我们方法的关键思想是让奖励成为有思考与无思考相比预测下一个词的信息增益。所以现在你必须能够比没有思考时更好地预测下一个词。因此,你必须学会更好地思考,使得你的下一个词预测概率比没有思考时的预测概率更高。这样,我们鼓励模型在回答下一个词之前自己思考。
Yeah. So that's a new paper we recently put out, roughly speaking the idea is that during pre-training, the model is forced to be completely passive in the way it learns to predict which token comes next. But what if we encourage the model to think for itself before predicting the next token? What if we encourage the model to think for itself by generating something like a chain of thought and then predict the next token? In that context, the reward—because it's reinforcement learning, we now need to think about reward. The reward could be designed in different ways, but the key idea of our approach is to make the reward the information gain of predicting the next token with thought compared to without thought. So now you have to be able to predict the next token even better than yourself predicting the next token without a thought. So you have to learn to think better so that your next token prediction probability becomes better than your own prediction probability without a thought. That way, we encourage the model to think for itself before answering the next token.
当你提到信息增益,然后你说希望模型更好地预测下一个词。当我听到信息增益时,我想到的是最大化惊喜,即你想奖励那些……我甚至不知道如何描述。不一定是在更准确的意义上更好,而是在更多样化的意义上更好。你是这样想的吗?能详细说明吗?
When you said information gain and then you went on to say you want the model to predict the next token better. When I hear information gain, I think of maximizing surprise, in the sense of you want to give reward to predictions that are... I don't even know how to describe it. Not necessarily better in the sense of more accurate, but better in the sense of more diverse. Is that the way you're thinking about it? Can you elaborate?
是的,实际上在这种情况下我们做的是多样化的反面。因为我们把这种强化学习方法放在预训练框架下,而预训练完全是关于下一个词预测。所以我们遵循那个总体框架。所以一切都是关于下一个词预测,但我们通过在后训练阶段加入奖励来赋予一些强化学习风格,将奖励定义为有好的中间思考时预测下一个词的信息增益。所以我们看的是给定所有先前词加上你的思考后下一个词的条件概率,与给定所有先前词但没有你的思考时预测下一个词的条件概率进行比较。你比较这两个量,这里的挑战是,只有当你将中间思考与所有先前词拼接后,你的中间思考实际上增加了预测下一个词的条件概率时,你才能获得奖励。所以这不是一个容易获得的奖励。
Yeah, actually we kind of do the opposite of diversification in this context. So in this context, what we are trying to do is because we frame this reinforcement learning approach under the pre-training framework, which is all about next-token prediction. So we go with that overarching framework. So it's all about next-token prediction, but we give some RL flavor by incorporating a reward during the last phase of pre-training by defining reward as information gain of being able to predict the next token with good intermediate thought. So what we look at is the conditional probability of the next token given all the previous tokens concatenated with your own thought, compare that with the conditional probability of predicting the next token given all the previous tokens without your thought. So you compare these two quantities, and the challenge here is that you get reward only if your intermediate thought actually increases the conditional probability of predicting the next token when you concatenate your intermediate thought in addition to all the previous tokens. So it's not an easy reward to get.
所以不是关于词的信息增益,而是关于思考相对于生成词的信息增益。
So not information gain with regard to the token, but information gain with regard to the thought relative to generating the tokens.
是的。是的。
Yeah. Yeah.
你刚刚提到这不是一个容易获得的奖励。你预计基于这样的技术,训练的复杂度会增加多少?
You just mentioned it's not an easy reward to get. To what degree do you expect the complexity of training to increase based on techniques like this?
这使得预训练的计算量比以前高得多。所以在我们的工作中,我们做了很多实验。一个实验设置是控制预训练期间有无强化学习的词元数量。如果使用相同数量的词元,那么我们会使用更多的浮点运算,正如你注意到的。另一个实验设置是控制浮点运算,这样我们在预训练的最后阶段使用更少的词元,但使用相同数量的浮点运算。我们发现,令我非常惊讶的是,当你以这种方式完成预训练时,即使在完全受控的设置下,最终的预训练模型在后训练后表现要好得多。
It makes the computation of pre-training much higher than before. So in our work, we do a lot of experiments. One experimental setting is to control the amount of tokens during pre-training with or without RL. So if you use the same amount of tokens, then we use a lot more FLOPs, as you noticed. Another empirical setting is we control for the FLOPs, so that we use way fewer tokens in the last phase of pre-training, but we use identical number of FLOPs. And what we found, to my big surprise, is that when you finish your pre-training in this way, even in the fully controlled setting, the final pre-trained model does much better after post-training.
所以,不仅这个模型在推理基准测试上表现更好——当然会更好,因为它被激励去思考以预测下一个词——而且,如果你应用相同的预训练和后训练方案,比如顺序微调后跟强化学习,性能提升会保留下来。你的模型现在在推理密集型后训练方案下表现甚至更好。这意味着什么?这有点像人类有一个关键期来习得语言,比如,而且很可能在生命早期学习数学和逻辑思维比晚些时候更好。我们在预训练中也经验性地发现了类似的现象。
So not only does this model do better on reasoning benchmarks at that point in time—of course it will, because it was incentivized to think for predicting next tokens—but also, if you apply the same pre-training and post-training recipe, like sequential fine-tuning followed by RL, the performance gain survives. Your model now performs even better with a reasoning-heavy post-training recipe. So what this entails is a bit analogous to how humans have a critical period for acquiring language, for example, and it's probably a good idea to learn math and logical thinking reasonably early in life rather than much later. Something like that is happening even with pre-training, as we empirically found.
有趣。有趣。你知道,我之前提到我们会在结束前回到这个更广泛的 AI 民主化理念。你提到的其中一个话题是多元对齐。你能详细说明一下这是什么意思,以及你为什么对这个想法感到兴奋吗?
Interesting. Interesting. You know, I mentioned that we would come back to this broader idea of democratizing AI before we close out. And one of the topics there that you mentioned is pluralistic alignment. Can you elaborate on what that means and why you're excited about that idea?
是的,这又回到了我之前说的关于“AI 源于人类、为了人类、由人类创造”的观点。所以“AI 源于人类”是指 AI 的起源确实来自人类。在价值观、知识、各种规范方面,AI 学习的内容应该真正反映整个人类。这就是这个想法。当然,互联网并没有均匀地反映所有人类。因此,产生的 AI 在某些方面也有偏见。问题是,我们能做些什么?这真的很难。
Yeah, so that goes to this earlier statement I made about AI of humans, for humans, and by humans. So AI of humans is about the origin of AI being really humans. In terms of the values, the knowledge, the different kinds of norms that AI learns from should really reflect the entirety of humanity. That's the idea. And of course, the internet does not evenly reflect all humanity. Therefore, the resulting AI is also biased in some ways. The question is, is there anything we can do? It's really hard.
对互联网进行去偏以用作训练数据听起来很难。
Debiasing the internet for use as training data sounds hard.
这几乎不可能,对吧?因为你不能回到过去改变历史,让国王和女王的数量相等。人类已经发生的事情已经发生了,你无法改变。
It's almost impossible, right? Because you cannot go back and change history to make an even number of kings and queens. Whatever happened in humanity already happened, so you're not going to change that.
不过,另一方面,我们有统计去偏技术。所以如果你把训练集仅仅看作一个数据集,我们有办法去偏。但让我想到的问题是,这样做你会失去什么?你会不会失去互联网的一些基本特性,正是这些特性让这些我们还不完全理解其工作原理的 LLM 得以工作?
Although, on the other hand, we have debiasing statistical techniques. So if you think of your training set as just a dataset, we have ways to debias it. The question that jumps out to me is, what do you lose in doing so? And do you lose some fundamental aspect of the internet that made these LLMs, which we don't understand how they work, work?
是的。我认为去偏不是正确的答案,因为去偏也不可能,而且有时你确实想保留自己的偏见。例如,如果你是一个有宗教信仰的人,或者你来自一个有着特定规范的国家,你可能希望遵循这些规范,那么也许我们想尊重这一点。顺便说一句,作为人类,当我们与我们知道持有不同价值观的其他人互动时,我们有办法与这个人周旋,以保持礼貌、保持尊重、求同存异,对吧?所以某种程度上,我认为 AI 意识到多元价值观并能够与之周旋非常重要,而不是在所有地方都完全中立,这不仅可能无法实现,而且如果我们想以尊重的方式真正服务不同的文化规范,这可能也不是理想的解决方案。所以在我的工作中,我们从三个不同的角度思考多元对齐:有所谓的奥弗顿多元主义、分布多元主义和可引导多元主义。这些概念需要解释。也许我们先从奥弗顿开始。
Yeah. I don't think debiasing is the right answer, in the sense that debiasing is also impossible, but also that sometimes you do want to maintain your own bias. For example, if you're a religious person, or if you're from a certain country where you have particular norms that you like to go by, then maybe we want to respect that. By the way, as a human, when we interact with other human beings who we know have different values, we have ways to navigate around this person such that we maintain politeness, maintain respect, agree to disagree, right? So to some degree, I think it's very important that AI is aware of diverse values and then be able to navigate around them, as opposed to just being completely neutral everywhere, which may not only be unattainable but also may not be the desirable solution if we are trying to really serve different cultural norms in a respectful manner. So in my work, we think about pluralistic alignment from three different angles: there's something called overton pluralism, distributional pluralism, and steerable pluralism. These concepts require explanations. Maybe let's start with overton.
在奥弗顿窗口的意义上。
In the sense of the Overton window.
是的。所以就像当你问一个政治上棘手的问题时,比如可能有不同的答案,最好的方式可能是让 LLM 呈现所有合理的意见,比如‘答案是人们有不同的看法。这是一种观点,那是另一种观点’,并能够包含所有观点,而不是选择多数意见,因为那会边缘化其他意见。所以能够涵盖所有这些选项。分布多元主义是指当 AI 用于更像决策过程时,比如 AI 在做求职申请筛选,或者 AI 在回答必须以更分类方式回答的问题。你必须选择一个答案;你不能给出所有答案。那么从分布上看,LLM 的分布应该模仿人类决策的分布。所以在每个时间点,AI 可能做出与任何其他人类决策不同的决策。然而,当我们看整体分布时,不是一直追求多数情况(那会导致分布极度偏斜),而是试图至少比现在更均匀分布。当然,人类有偏见,所以我们的决策分布不一定公平或无偏,对吧?但至少不要比这更糟,这就是想法。最后一个是可引导多元主义,即你能够将模型引导到不同的价值观——你的道德框架或价值框架——以服务于你的日常需求,在合理执行的范围内。这意味着在某些场景中,你可能希望模型在操作方式上或多或少地具有多元性。
Yeah. So it's like when you ask a question that's politically thorny, for example, that could have different answers, the best way might be for the LLM to just present all of them—all of the reasonable opinions—as in, 'Hey, the answer is that people have different opinions. Here's one view, there's another view,' and be able to include all of them, as opposed to picking the majority opinion because that marginalizes the rest. So being able to cover all of these options. Distributional pluralism is when AI is made for more like a decision-making process, where maybe AI is doing job application filtering, or AI is answering questions that have to be answered in a more categorical manner. You have to choose an answer; you cannot give all of the answers. Then distributionally, the distribution of LLMs should mimic the distribution of human decisions. So at each point in time, AI might be making a decision that differs from any other human decision. However, when we look at the overall distribution, instead of going for the majority case all the time, which would be distributionally super skewed, the idea is to try to be at least distributionally even compared to now. Of course, humans have a bias, so it's not like our distribution of decisions is necessarily fair or unbiased, right? But at least let's not get worse than that is the idea. Now the last one, steerable pluralism, is that you are able to steer the model to different values—your moral framework or value framework—to serve your day-to-day need within the scope that's reasonable to execute this. Meaning in some scenarios, you might want to be more or less pluralistic in the way the model operates.
所以模型应该能够引导到任何合理的不同价值体系。当然,问题是什么是合理的?因为也许我们不想让模型可引导以支持那些想犯罪的人。罪犯的价值体系可能应该完全排除,但在合理的、合法的、社会可接受的范围内,能够引导你的模型来服务你的价值体系。你在确定实现这些的方法方面进展如何,除了确定多元方法的三个维度之外?
So the model should be able to steer to any different value system that's reasonable. Of course, the question is what is reasonable? Because maybe we don't want to allow the model to be steerable to support people who want to be criminal. Criminals' value system probably should be completely out, but within the reasonable, like legal, socially acceptable scope, the ability to steer your model to serve your value system. How far along are you in identifying ways to do these things beyond identifying the three dimensions of pluralistic approaches?
所以这类研究既需要数据研究,也需要算法研究。令我高兴的是,许多聪明的学者开始在这两个方面开发解决方案。所以有一些新算法可以做到更多元的对齐,比如分布多元主义。其他人也在研究这个。但有人可能会说,前沿模型并不那么差。与开源模型相比,前沿模型通常在多元对齐方面更好,因为开源模型在数据整理和安全护栏等方面投入较少。所以本着让更小的模型,尤其是开源模型,更强大以实现更广泛可及性的精神。
So this sort of research requires both data research as well as algorithmic research. To my delight, a lot of smart academics started developing solutions for both fronts. And so there are some new algorithms that do more pluralistic alignment, for example for distributional pluralism. And other people are working on this. But one could argue that frontier models are not so bad. Frontier models in general are better at this kind of alignment pluralism compared to open-source models that went through less effort in terms of data curation and safety guardrails and everything. So in the spirit of making smaller, especially open-source models, more powerful for wider accessibility.
太棒了。所以我们在年初、2026 年初重新连线。你对今年有什么想法或预测,或者你期待看到什么?
That's awesome. So, we're connecting, reconnecting at the beginning of the year, beginning of 2026. Any thoughts or predictions on what you expect to see happen this year or maybe what you'd be excited about seeing?
是的。我认为社区在小模型上的努力将进一步升级。去年开源社区已经付出了越来越多的努力,当然现在 Nvidia 也大力投资支持开源工作,所以我们会看到更多这样的进展——这是我能做出的一个明显预测。另一个是 AI 在科学领域的应用。我个人对此非常兴奋,因为 AI 对科学领域的积极影响可能非常巨大。如果我们能做得恰到好处,那么医学和人类生活的各个方面都能从 AI 科学中受益。这也是一个非常艰巨的智力挑战,因为它需要能够触及超越互联网数据所反映的人类知识的知识。问题是,AI 只擅长学习人类能够提供的数据。所以这也是一个巨大的智力挑战,我非常期待进一步探索这个方向。
Yeah. I think the community efforts on small models will escalate even further. Last year there were already increasing efforts from the open source community, and of course now Nvidia is also really heavily invested into supporting open source efforts, so we will see a lot more of that — that's one obvious prediction I can make. Another one is the use of AI for science. I'm quite excited about that personally because the positive impact of AI for scientific domains can be really phenomenal. If we know how to do it quite right, then medicine and different aspects of human life could really benefit from AI for science. And that's also a really hard intellectual challenge because it requires being able to reach knowledge that's really above and beyond human knowledge reflected on internet data. The thing is, AI is only really good at learning the data that humans are able to provide. So this is a big intellectual challenge as well, and I'm very excited about pursuing further into that direction.
好的,Yejin,非常感谢你再次加入我们,更新你的工作进展,特别是深入探讨了你如何处理小语言模型的推理。
Well, Yejin, thanks so much for jumping back on with us and giving us an update as to what you're working on and in particular digging into how you're approaching reasoning for SLMs.
非常感谢再次邀请我。谢谢。
Thanks so much for having me again. Thank you.
再见。
Bye-bye.