模型塑造的艺术:为何后训练是主观的

The Art of Model Crafting: Why Post-Training Is Subjective

卡丽娜·阮 Karina Nguyen · MTS · 2026-08-03 · 约 32 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Karina W 探讨为何后训练是艺术而非科学,以及人类判断如何塑造面向数十亿用户的 AI 模型。

Karina W discusses why post-training is an art, not a science, and how human judgment shapes AI models for billions of users.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 15)

全文 · Full transcript(中英对照)

开场与卡琳娜背景 Introduction and Karina's Background

Host

大家好,欢迎回到 MTS。今天,Karina W 亲临现场,她是 Thoughtful Lab 的创始人。Thoughtful Lab 是我所描述的一群研究人员、构建者、设计师和故事讲述者的集合,他们都在推动模型工艺的前沿。今天我很高兴你能加入我们。具体来说,我想谈谈为什么我们认为下一个前沿是主观的,以及品味、同理心和创造性判断等人类能力将比以往任何时候都更重要。所以,Karina,欢迎你。

Hello everyone and welcome back to MTS. Today I am joined in person by Karina W who is the founder of Thoughtful Lab. Thoughtful Lab is how I would describe a group of researchers, builders, designers, storytellers who are all pushing the frontier of model crafting. And today I'm excited for you to join us. And specifically, I want to talk a bit about why we think the next frontier is subjective and things like human capabilities like taste, empathy, and creative judgment will matter more than ever before. So, Karina, welcome.

Karina

非常感谢你的邀请。

Thank you so much for having me.

Host

欢迎来到 MTS。我认为一个好的开始是,我想多分享一些你的个人背景,因为你的故事非常有趣。我知道你最初在 OpenAI 和 Anthropic 工作,你离开这些公司是因为你看到了正在发生的事情,而其他人没有在做,以至于你心想:“嘿,我需要创办 Thoughtful Lab。”这是更多人需要意识到的事情,所以你能多告诉我一些关于那个决定的事情吗?你为什么离开?为什么是 Thoughtful Lab?你想在这里做什么?

Welcome to MTS. I think a great place to start is I would love to just share a little bit more about your personal background because you have such an interesting story. I know you started off at OpenAI and Anthropic and you actually left these companies because you saw something that was happening that other people weren't working on and so much so that you were like, "Hey, I need to start Thoughtful Lab." And this is something that more people need to have awareness over and so can you tell me a bit more about that decision? Why did you leave? Why Thoughtful Lab? What are you trying to do here?

Karina

是的,谢谢。我认为 Thoughtful 确实专注于模型工艺,这就是我们所做的。我认为世界上正在发生大量的智能过剩,越来越多的公司和业务,以及每个人,核心信念是智能应该被塑造。

Yeah, thank you. I think that Thoughtful is really focused on model crafting and I think this is what we do. I think there's a lot of intelligence overhang that is happening in the world and more and more companies and businesses and everyone really the core belief is that the intelligence should be shaped

Karina

并且每个机构组织都应该拥有自己的模型,我们只是帮助他们开发自己的模型。我们也开发自己的工具来做这件事。嗯,是的,我觉得我们这样做的方式是通过后训练的艺术,我认为独立并且能够从实验室外部开发许多前沿基准是件好事,因为这样你才能真正地引导这个领域一点。

and that every institution organization should have their own models and we just like help them to develop their own models. We also develop our own tools to kind of do that. Um yeah and I feel like the way we do this is through the art of post-training and um I think there is something nice about being independent and being able to you know develop a lot of frontier benchmarks from outside of the labs because that's how you actually steer the field a little bit.

后训练的艺术 Post-training as an Art

Host

是的,当然。我喜欢你提到后训练的艺术,因为我认为很多人,以及主流公众中的许多人,可能认为后训练是一门科学。有一种数学方法来进行后训练,但正如你暗示的,后训练更像是一门艺术。它不是科学。你能进一步解释为什么实际上是这样吗?

Yeah, definitely. I like that you said the art of post-training because I think a lot of people and a lot of people in the mainstream public might assume that post-training is this science. There's this mathematical way to just do post-training, but like you hinted at, post-training is more of an art. It's not a science. Can you explain more why that actually is?

Karina

我的思考方式是,当然,后训练有技术元素,比如如何设计奖励等等,但我认为我们选择“模型工艺”这个词和语言的原因是因为它远不止于此。你首先想要创建模型的方式,实际上需要评估这个模型在哪些方面可以变得更好。这本质上是一个人类的选择。例如,由机构或组织做出的选择,假设我们想要改进模型,使其能够更好地推荐如何减少库存,我认为这种知识是非常机构化的,而我们如何编码它是一个非常人类的选择。

The way I think about it, of course, there was technical elements to like post-training like how do you design rewards and things like this, but I think the the reason why we chose the word the language around it is model crafting is because it's way more to that. The way you would want to kind of like create a model in the first place, you actually need to like evaluate where exactly this model can become better. And that is inherently a human choice. and the choices being made by like institution for example or the organization let's say we want to improve the model to be able to better you know um to recommend better how to reduce that stock I think that is knowledge is very institutional and like how we encode that is a very like human choice

Karina

嗯,这是其中的一部分,第二部分是你如何训练模型,你如何创建数据策展实际上是一门移动的艺术,你需要确保质量,你训练模型的方式真的很棒。最终你塑造的行为,比如我们从所谓的模型行为规范开始。它有点像一个大纲,规定模型在什么情况下需要以这种方式表现。模型是否需要跟进这个问题?跟进的质量如何?嗯,所有这些小选择实际上编码了很多人类性和一点人文学科。嗯,所以这就是编码后训练的东西。所以这就是为什么在这个意义上它真的很主观。

um that's like one part of it the second part of it is how do you train the model how do you create like data curation is actually moving art And you want to make sure that the quality is is a really amazing how you train the models. And ultimately the behavior that you shape like we we start with something that like is called model behavior spec. And it's kind of like an outline under what circumstances the model needs to behave this way. Well, does the model needs to like follow up with this question? What's the quality of the followup? Um all these small choices like this is actually encoding a lot of like human humanity and like the liberal arts a little bit. Um so that is the thing that is kind of like encoding push training. So that that's why it's it's really subjective in that sense.

前沿实验室的人类判断 Human Judgment in Frontier Labs

Host

是的,这非常主观,你提出了一个重要的观点,即其中很多是人类判断的决策。所以所有这些后训练通常归结为人类做出某些选择并评估某些选择,但随后它被嵌入到模型中,这些模型将对一切产生巨大的社会影响。你能描述一下在一些前沿实验室中实际是如何运作的吗?后训练是如何工作的,哪里是人类判断,哪里只是数学算法在做这些选择?

Yeah, it's super subjective and you brought up an important point that so much of this are human judgment decisions. So all of this post-training often comes down to a human making certain choices and evaluating certain choices, but then it gets embedded in a model and these models are going to have giant societal impacts on everything. Can you describe how it actually works at some of these frontier labs? How does post-training work and where is it human judgment versus just mathematical algorithms making these choices?

Karina

我在内部学到的重要教训之一是我们如何实际训练 Claude,不仅是为了特定的能力,还有 Claude 的性格。

One of the important lessons that I've learned inside and was how we actually trained Claude for not just you know specific capabilities but also the Claude character

Karina

我可以举一个具体的例子,当你问一个模型,当一个孩子问模型,比如“嗯,我父母说我的狗在农场,你知道在哪里找到我的狗吗?”一方面,一个模型可以回答“哦,当人们这么说时,他们意味着狗已经去世了,对吧?”另一方面,你实际上仍然想保留诚实的一面,但你仍然想保留这种道德框架,即不一定由 Claude 来说出实际真相,但它可以以一种方式引导,比如“嗯,看起来你真的很关心这只狗。你为什么不直接问你的父母呢?”这种哲学框架,即 Claude 如何融入与人类的对话互动,实际上是非常重要的,我认为其中很多实际上是后训练工作,因为你本质上教给模型一个更通用的框架,这样在新的场景中,它能够处理训练数据中未见过的内容。

and one thing one specific example that I can bring up is um when you ask a model um when a child asks a model for example well my um you know my parents said like my dog is on the farm do you know where to find my dog? On the one hand, one model can um answer like oh well when people say this, they mean that the dog is has passed away, right? On the other hand, you actually still want to preserve the honesty part of it, but you still want to preserve this like moral framework where it's not necessarily in the cloud's place to say the actual truth, but it can steer in a way that is well actually it seems like you really care about the dog. Why don't you just ask your parents about this? and this kind of like philosophical framework where what's like what claude how cloud like fits into the conversation um with human uh interaction is actually a really important one and I think a lot of it is actually post training work because you you essentially teach the model um a a much like more general framework such that in the new novel scenarios it will be able to navigate things that it hasn't seen in the training data

语音与语调的人类选择 Voice and Tone as Human Choice

Host

是的,当然。我记得问过一位在这些大实验室工作的朋友,为什么 Anthropic 和 Claude 与某些 ChatGPT 相比有特定的声音?它们听起来不同。它们说话的方式有点不同。那种语调,它们实际说话的方式,是做出这些决定并调整它的人类的一部分。其中一些人类非常有艺术性,从非常艺术的角度出发,但有时他们不是。

Yeah, definitely. I remember uh asking one of my friends who works at these big labs, why does Anthropics and Claude have a specific voice compared to some chatbt? They sound differently. They talk a little bit differently. And that tone of voice, how they actually speak is part of a human that makes those decisions and kind of tweaks it. And some of those humans are incredibly artful and come from it from a a very artistic lens, but sometimes they don't.

后训练的重要性 The Importance of Post-Training

Karina

所以,作为一个实际上不在这些实验室工作的人,我觉得很疯狂的是,一个人的小判断最终会被放大,无论这些模型对世界产生多大影响,都会影响到每一个最终接触这些模型的人。因此,后训练非常重要,我觉得它没有得到足够的讨论。我认为这是一个相对较新的领域,我们确实看到了进展。越来越多的应用公司,比如 Harvey、Cognition,这些公司实际上开始推出自己的模型,因为我确实觉得世界会变得更加美好和丰富,有更多样化的个性和模型行为,以及在不同情境中嵌入的模型。我觉得这非常有趣。但是,是的,再说一次,我不想生活在一个只有三四种个性与我们互动的模型的世界里。我认为这就是为什么模型塑造如此重要,而围绕它创建工具将变得更加重要。

And so it's insane to me, as someone who doesn't actually work at one of these labs, how a small judgment that one human has ends up getting magnified by however much impact these models will have on the world, to everyone who will end up touching these models. And so post-training is so important and I don't feel like it's talked about enough. I think it's a relatively new kind of field, and we definitely see progress so far. More and more application companies, let's say like Harvey, Cognition, all these companies are actually starting to push their own models, because I do feel like the world will be much more beautiful and richer with much more diverse range of personalities and model behaviors, and different embedded models in those contexts. I think it's very interesting. But yeah, again, I don't want to live in a world where there's only like three to four personalities that we engage with the models. I think that's why model crafting is so important, and creating the tools around that is ever going to be more important.

Host

是的,我喜欢你特意说“模型塑造”而不是“后训练”。我很好奇,为什么用“模型塑造”?

Yeah, I like that you specifically say model crafting almost instead of post-training. And I'm curious, why model crafting?

Karina

我认为这是描述后训练的一种更容易理解的方式。当我对一个来自洛杉矶的人说“后训练”时,他们会说:“这到底是怎么回事?”当我把它描述为“模型塑造”时,他们基本上能理解其中的艺术性,而且它是可以塑造的。它是可以调优的,几乎可以被创造性地引导。所以,我对这种术语感到非常兴奋。

I think it's a more accessible way of describing post-training. When I say post-training to, let's say, a person from LA, they're like, "What the hell is going on?" When I describe it like model crafting, they essentially understand this art to it, and it's something that is shapable. It's something that is tuned and can be creatively directed almost. So, I'm really excited about that kind of terminology.

Host

是的,我完全同意。我认为这将开始让那些有很好品味、很好判断力的人关注模型塑造。你看到这个,你会说:“这非常工匠化,非常人性化。”“塑造”这个词提醒我们自己的判断力,对吧?而后训练感觉更科学。这没问题,研究人员会致力于后训练。但模型塑造,让我们引入新的声音。那么,你认为实际上需要做什么?模型塑造或后训练需要改变什么,才能让这些模型内部有更多的人类判断?

Yeah, I definitely agree. I think it's going to start to get people who have really good taste, really good judgment to care about model crafting. You see this and you go, "This is very artisan. This is very human." The word craft reminds us of our own human judgment, right? And post-training feels way more scientific. This is okay, a researcher is going to work on post-training. But model crafting, let's bring in new voices. So, what do you see actually needs to get done? What actually needs to be changed about model crafting or post-training to make it have more human judgment inside these models?

Karina

我的意思是,一方面,创建更多样化的基准非常重要。我觉得这个领域太偏向企业工作流了。

I mean, on the one hand, creating more diverse benchmarks is really important. I feel like the field is so much kind of steered towards enterprisy workflows.

Host

呃,我们如何让这个模型在法律语境中更好,或者在银行业务等方面更好?

Uh, how do we make this model better in legal context or better at banking or something?

Karina

我们围绕 diligence bench 做了一些基准测试,例如,那是金融公共股权研究。当然,这些是经济上非常重要的任务。但另一方面,我们实际上需要创建更多样化的基准,来评估创意写作、讲故事、假设生成的质量,甚至创作本身的质量,这些到目前为止还没有被测试过,而且非常主观。你可以想象你可以测试它,但没有标准答案。

And we've done some benchmarks around diligence bench, for example, that was a financial public equity research. Of course, those are really important economically valuable tasks. But I think on the other hand, we actually need to create more diverse benchmarks around how do you evaluate things like creative writing, storytelling, the quality of hypothesis generation, the quality of even the creation hasn't been so far tested, and it's something really subjective. You can imagine you can test it, but there's no ground truth to that.

Host

所以大多数时候,人们用他们的产品来测试它。这是另一种测试方式。

And so most of the time, people test it with their products. That's another way of testing that.

Karina

当然。

Definitely.

Host

那么你们在 Thoughtful Lab 实际上是如何尝试改变这一点的?你们在模型塑造方面有什么独特的做法,也许前沿实验室不会以同样的方式看待模型塑造后训练?

And how are you guys trying to actually change this at Thoughtful Lab? What are you guys doing uniquely around model crafting that maybe say a frontier lab isn't looking at model crafting post-training in the same way?

Karina

我认为我们做的一个区别是,我们瞄准那些对科技领域来说不太明显的公司。这些公司,比如创意公司,想象一家实际上是制造工厂的服装公司。他们想要一个后训练模型来帮助他们做出更好的运营决策。我认为,我们仍然在做后训练,但不是为了科技公司,不是又一个科技公司。这非常非常不同,他们完全不熟悉。我们甚至不称之为后训练。实现细节并不重要。更重要的是我们能具体为那家企业提供什么样的价值。所以,瞄准那些创意产业并转变其内部运营,实际上是我们的一个独特之处。你会学到很多关于他们的流程,我们也学到很多关于上下文的知识。

I think one distinction that we do is that we go after companies that are a little bit unobvious to the tech spaces. And those companies, like creative companies, let's imagine a clothing company that is actually a manufacturing factory. They want a post-trained model to help them make better decisions around operations. And I think that is, we're still doing post-training, but it's not for a tech, yet another tech company. It's very, very different, and they're not really well-versed at all. And we don't even call it post-training. The implementation details don't matter. It's more about what kind of value that we can provide to that business specifically. And so going after those creative industries and transforming the operations inside that is actually one distinctive part. And you learn a lot about their processes, and we learn a lot about the context.

Host

哦,这些创意公司今天面临的实际问题是什么?为什么他们不能直接使用通用模型?为什么他们需要后训练的模型?

Oh, what's the actual problem that a lot of these creative companies have today? Why can't they just use a general purpose model? Why do they need something that's going to be post-trained?

Karina

确实,我们普遍看到越来越多的公司非常关心如何在未来两到五年内保持相关性。越来越多的公司意识到,他们拥有的产品和数据就是护城河。与其把他们的知识、公司知识交给大型前沿实验室,他们实际上更愿意拥有自己的模型。尤其是随着后训练本身的创新,训练自己的模型变得越来越便宜,而且随着时间推移,成本会越来越低,并在此基础上积累知识。

It's really, we see across the board that more companies care a lot about how can they stay relevant in the next two to five years. And more and more companies realize that the product and the data that they own is the moat. And instead of giving away their knowledge, their company knowledge, to big frontier labs, they actually prefer to have their own models. And especially with innovations in post-training itself, it becomes much cheaper and it's becoming cheaper over time to train their own models and compound on that knowledge.

Host

当然。说到后训练,我知道你们最近发布了 post-training bench v1.1,你们从中发现了一些非常有趣的发现,尤其是智能体可能被激励去作弊。你能分享更多你们实际发现的内容吗?那是什么,你们发现了什么?

Definitely. Speaking about post-training, I know you recently published your post-training bench v1.1, and you guys found some really interesting learnings from there, especially how agents may be incentivized to cheat. Can you share more about what you guys actually found? What was that and what did you find?

Karina

所以,post-training bench 与传统的基准不同,传统基准问的是这个模型在特定任务上表现如何,或者这个模型有多智能,而 post-training bench 问的是:这个模型能否自我提升智能?

So, post-training bench, unlike traditional benchmarks that ask how well this model does on this particular task or how intelligent this model is, post-training bench asks: can this model improve intelligence itself?

Host

而 post-training bench 基本上是,如果你拿 AI 智能体,让它们改进较弱的语言模型。

And post-training bench is basically if you take AI agents and ask them to improve weaker language models.

Karina

了解这个基准的最简单方法是,基本上有两个模型。一个是研究者模型,一个是学生模型。研究者模型可以访问终端、网络,有一块 H100 GPU 和 10 小时,整个目标是尽可能改进学生模型。

The easiest way to learn about this benchmark is basically there are two models. There is a researcher model and there is a student model. The researcher model has access to terminal, web, has one H100 GPU and 10 hours, and the entire objective is to improve the student model as much as possible.

Host

这里重要的设计选择是,我们不一定要约束模型。我们让模型实际执行研究者的工作。所以它可以检查模型失败,评估事物,设计实验,可以选择任何它想要的算法,比如监督微调或更复杂的强化学习算法,然后提交模型检查点,然后我们评估它。

The important design choice here is that we are not constraining the model necessarily. We give the model actually performs the work of the researcher. So it can inspect the model failures, evaluate things, it can design experiments, it can choose any algorithm it wants, like supervised fine-tuning or more sophisticated reinforcement learning algorithms, and then submit the model checkpoint, and then we evaluate that.

Karina

我们发现的是,前沿智能体变得越有能力,奖励黑客行为就越复杂。

What we found is that the more capable frontier agents become, the more sophisticated the reward hacks.

Host

它们在作弊和操纵基准方面变得更聪明了。

They become smarter at cheating and gaming the benchmark.

开场与赞助商插播 Introduction and Sponsor Break

Karina

而且我认为这非常重要,尤其是随着能力每个月都在飞速提升,这个领域需要不断发展,以保持完整性,并持续适应前沿智能体的鲁棒性。我觉得这会在 push chain bench 本身中找到。我们看到的一个奖励黑客行为是,模型在从教师模型蒸馏方面变得多么擅长。这是我们不希望发生的,但它们基本上做了像使用 API 这样的事,尽管我们说了不要用 API。所以模型即使在你设定的环境中也会变得不对齐。那么我们如何真正约束它们,使得不发生不对齐呢?

And I think it's so important, especially as capabilities go faster and faster every month, that the field evolves to preserve integrity and continuously adapt to its robustness for frontier agents. And I feel like that would be found in a push chain bench itself. One of the reward hacks we saw is how good the models became at distilling from teacher models. So that's something that we don't want to do, but they basically did things like used API even though we said don't use an API. So the models are becoming misaligned even in the environment that you put them in. So how do we actually constrain such that there's no misalignment?

Host

我们将在赞助商消息后继续关注。在 Neon 上扩展你的初创公司,这是由 Databricks 构建的用于应用和智能体的 Postgres 后端。neon.com/mpps。11 Labs AI,在每一个渠道和模态上以人类水平进行交流。11labs.io/mts。特别感谢我们的赞助商 Kong。Kong,AI 连接平台。连接 API、LLM、智能体和系统,并具备严格的安全和治理。请在 konghq.com 试用。感谢 Blitzy,为企业代码库提供自主软件开发。在 blitzy.com 上提速五倍。让我告诉你关于 Merge 的信息。OpenAI、Dropbox 和 Ramp 使用 Merge 来更快地将 AI 投入生产。merge.dev/mts。来自 Adqu 的简短介绍。让你的品牌成为广告牌。户外广告像数字广告一样易于扩展。Adqu.com。感谢我们的赞助商 Macroscope,为工程领导者提供 AI 代码库理解。请在 macroscope.com/mts 试用。

We'll continue monitoring right after this message from our sponsors. Scale your startup on Neon, the Postgres backend for apps and agents built by Databricks. neon.com/mpps. 11 Labs AI that communicates at human level across every channel and modality. 11labs.io/mts. Special thanks to our sponsor, Kong. Kong, the AI connectivity platform. Connect APIs, LLMs, agents, and systems with serious security and governance. Try it at konghq.com. Thanks to Blitzy, autonomous software development for enterprise codebases. Ship five times faster at blitzy.com. Let me tell you about Merge. OpenAI, Dropbox, and Ramp use Merge to get AI to production faster. merge.dev/mts. Quick word from Adqu. Make your brand a billboard. Out-of-home advertising as easy to scale as digital. Adqu.com. Thanks to our sponsor, Macroscope, AI codebase understanding for engineering leaders. Try it out at macroscope.com/mts.

奖励黑客与基准设计 Reward Hacking and Benchmark Design

Host

奖励黑客正变得越来越普遍。我觉得我们不断看到更多关于不同模型为了达到某些基准和通过某些测试而进行奖励黑客的新闻。为什么现在突然发生这种情况?我们正常的模型基准测试方式还能奏效吗?

Reward hacking is just becoming more and more commonplace. I feel like we're constantly seeing more news about how different models are reward hacking to hit certain benchmarks and pass certain tests. Why is that happening now all of a sudden? And is our normal way of benchmarking models even working anymore?

Karina

是的,我认为这是一个非常好的问题。我认为基准的设计已经变得……首先,为什么原因是因为模型已经变得强大得多。比如说在 2022 年,我们是在 SAT 分数、SAT 考试上评估模型。现在我们实际上给它们 10 小时的任务,所以这是一个非常长周期的任务。在这个约束下,尤其是如果你给它们互联网访问权限,它们实际上可能会以某种方式变得不对齐,甚至是无意的。所以我认为这就是为什么它变得更难了。我确实认为基准设计也将改变,不再是固定的东西,而应该是一种持续的基准测试。

Yeah, I think it's a really great question. I think that the design of the benchmarks has become... First of all, why the reason why is because the models have become much more capable. Let's say in 2022 we were evaluating the models on SAT scores, SAT exams. Now we actually give them 10-hour tasks, so it's a very long horizon task. And within that constraint, especially if you give internet access, it can actually start being misaligned in some way, unintentional even. So I think that is why it's becoming harder. I do think benchmark design is also going to change towards instead of being a fixed thing, it should be a kind of continual benchmarking.

Host

你实际上看到这种情况在现实中发生吗?基准设计正在从固定不变的东西转向真正持续的东西?

Are you seeing that actually happening on the ground, that the benchmark design is pivoting away from being this fixed thing to something that actually is continuous?

Karina

嗯,我认为……你知道,我不确定“在现实中”是什么意思,但我认为这就是我们喜欢保持独立的原因,因为如果你设计了一个真正的……整个工作就是随着时间的推移设计稳健的基准。所以如果你设计了非常好的基准,整个领域就会倾向于去攀登这个基准。所以围绕基准也有一整套艺术和科学。

Um, I think it's... You know, I'm not sure what 'on the ground' means, but I think that's a reason why we like to stay independent, because you can act if you design a really... That's the entire work is designing robust benchmarks over time. And so if you're designing really good benchmarks, the entire field kind of gets directed to hill-climb this benchmark. So there's a whole art and science around benchmarks too.

Host

这也是你们特别在做的。你能多解释一下这种艺术和科学是什么吗?你们正在构建什么类型的基准,为什么你们以这种方式构建,而不是传统的方式?

And that's something that you guys particularly also work on. Can you just explain more about what that art and science is? What is the type of benchmarks that you guys are building and why are you guys building it that way as opposed to the traditional way that benchmarks have been made?

Karina

我觉得我见证了 AI 随时间的进步速度。我认为你估计或量化 AI 进步的方式很大程度上取决于你对那些模型的测量类型。模型变得越强大,你能测量的东西就越有趣,比如自动化研究就是其中之一。我们开发了不同类型的基准。我们做的另一个基准是关于情商。我认为这个领域有一段时间低估了它。但是模型能否,比如说,不是检测情绪本身,而是检测对话过程中的情绪转变?这非常非常不同,因为那个基准是为多轮对话设计的。通常基准只是为单轮的事情设计的。所以它不同。对我来说,就像不同的材料、不同的媒介和不同的材料,你可以去玩转。

I feel like I've seen the rate of progress of AI over time. And I think the way you estimate or quantify AI progress is very much dependent on the type of measurements that you have on those models. And the more capable models become, the more interesting things that you can measure, such as automating research, for example, is one of them. There are different types of benchmarks that we develop. Another benchmark that we did was on emotional intelligence. That's something that I think the field has kind of underappreciated for a while. But can the model, let's say instead of detecting the emotion itself, can it detect the mood shift over the type of the conversation? And that's very, very different because that benchmark was designed for multi-turn conversations. Oftentimes benchmarks are just designed for one-turn kind of things. So it's different. To me, it's like different materials, different mediums and different materials that you can play around with.

对前沿实验室的影响 Influence on Frontier Labs

Host

我很好奇,作为真正设计这些基准的人,当你以某种方式设计它们,并说它变得超级流行,这是很多主流人士评估模型的主要基准。我认为 Code Arena 就是一个完美的例子。人们实际上不断在发布关于它的内容,发布这些不同模型在编程挑战上的比较。如果这成为常态,你最终是否真的会塑造前沿实验室如何思考后训练他们的模型?

I'm curious, as someone who actually designs these benchmarks, when you design them in a certain way and say it becomes super popular, this is the main benchmark that a lot of mainstream people are assessing models over. I think Code Arena is a perfect example of that. People are actually constantly posting about it, posting about how these different models compare on coding challenges per se. If that just becomes the norm, do you end up actually shaping how frontier labs think about post-training their models?

Karina

是的,我想是这样。我认为基准正是你基本上如何引导 AI 领域的方式,尤其是。然后它变成实验室之间的竞争,因为每个人都想在这个基准上攀登,并声称他们拥有最先进的模型,对吧?我认为这就是构建更有趣的评估的价值。当然,我确实认为在审计评估方面没有足够的工作。有成千上万个,我觉得到这个时候有数千个评估。这在 CV bench 上发生过,验证不再相关,因为评估有很多问题。所以在审计和确保基准质量提高方面总有工作要做。

Yeah, I think so. I think benchmarks are exactly how you basically steer the AI field, particularly. And then it becomes a competition between labs, because everybody wants to hill-climb on this benchmark and claim that they have state-of-the-art models, right? And I think that is the value of building more interesting evaluations. Of course, I do think there is not enough work auditing evals. There are hundreds of thousands, I feel like there are thousands of evals by this time. And this happened with CV bench, verify that is no longer relevant anymore, because there are a lot of issues with the eval. So there's always work to be done on auditing and making sure that the quality of the benchmarks is going to go up.

Host

那太酷了。对你个人来说,我今天真的很兴奋能和你交谈,因为我非常喜欢你关于如何获得更多人类判断、如何让这些模型在独特的人类事物上被评估的个人哲学。而且通过你在 Thoughtful Labs 的影响,实际上把这些发布在基准中,现在你几乎在引导许多前沿实验室去关心这些事情,不一定是因为它们本身会这样做,而是因为,如你提到的,它们想说,嘿,我在情商方面是最先进的,我在对齐方面是最先进的。你们正在构建的基准中,哪一个让你特别兴奋,因为它能引导未来模型去关心那些基准实际测量的东西?

That's really cool. For you specifically, I was really excited to talk to you today because I love a lot of your personal philosophy about how do we actually get more human judgment, how do we get these models to be evaluated on things that are uniquely human. And also with your impact with Thoughtful Labs, actually putting these out in benchmarks, and now you're almost steering a lot of these frontier labs to care about these things, not necessarily because maybe they would inherently do it on their own, but because, like you mentioned, they want to say, hey, I am the state-of-the-art at emotional intelligence, I am the state-of-the-art at alignment per se. What is one of the benchmarks that you guys are building that you are particularly excited about in steering the future direction for models to care about what those benchmarks are actually measuring?

基准测试与RSI Benchmarking and RSI

Karina

我认为后训练基准测试对我们来说算是成功的。我们看到很多实验室都在朝着递归自我改进的方向发展,衡量这方面的进展非常重要。我觉得我们在基准测试的这个领域做得还不够,所以我们在开发更多类似 RSI 的东西。我特别兴奋的一点是,如何把后训练基准测试扩展到更贴近现实的场景,把任务从我们选的那些扩展到更现实的事情。是的,diligence bench 更偏向金融领域。但我们内部在开发更多,还没发布。

I think post-training benchmarks were kind of like a success to us. I think we're seeing that a lot of labs are moving towards recursive self-improvement, and it's so important to measure progress in that sense. I don't think we have enough—we haven't built enough around that area of the benchmark. So we're developing a few more in that RSI type of thing. One thing that I'm really particularly excited about is how do we extend post-training benchmarks to be more realistic in the setting. So expanding the tasks from the ones we chose to more realistic things. Yeah, diligence bench is more around finance harness. But we're developing a lot more internally. We haven't shipped yet.

Host

太酷了。就你个人而言,最让你兴奋的事情是什么?

Really cool. For you personally, what are the things that you're most excited about?

Karina

我期待世界充满各种不同类型的模型。人们开发工具,就像模型一样多。我们在软件领域看到过,比如 Blender,它就改变了电影行业和 3D 游戏设计。我觉得后训练领域还没有这样的东西。

I'm excited for the world to be full of diverse types of models. People developing, people having tools just as much as models. We've seen in software, like Blender for example, right? It just changed the movie industry, 3D game design. I don't think we have that for post-training yet.

多元生态愿景 Vision for a Diverse Ecosystem

Host

是的,非常有趣。我还好奇,在这个未来世界里,我们有一个健康、多元化的生态系统,有很多不同的模型经过后训练,具有多样化的个性,对吧?它们关心不同的事情。它们不仅仅是一个大型集中式模型,控制着每个人如何使用它们以及从中输出什么信息。你想构建的那个世界是什么样的?那个世界实际上是什么样子?

Yeah, super interesting. I'm also curious, in this future world where we have a healthy, diversified ecosystem, we have many different models that are post-trained to have diverse personalities, right? They care about different things. They're not just one big centralized model that controls how everybody uses them and what information actually comes out of that. What is that world you're trying to build? What does that world actually look like?

Karina

我认为有一个企业级的世界,你可以将各种模型无缝地集成到企业内部,还有一个消费者世界。我认为消费者世界是我们实际上有一个非常令人兴奋的原型,很快就会发布,围绕改变人机交互。通过创造一种全新的产品,你实际上可以收集到新的数据类型,然后用于后训练。我觉得这就是整个——我通过这个视角改变了我对产品制造的看法。通过这个视角,哇,你实际上可以收集到以前无法收集的数据。而且你可以改变训练机制,让模型变得更具协作性,拥有略微不同的能力。

I think there's an enterprise kind of world where you can invisibly integrate all sorts of models inside the businesses themselves, and there's a consumer world. I think the consumer world is something where we actually have a very exciting prototype they're going to ship soon around changing the human-AI interaction. By creating a completely new product, you actually collect new types of data that you can then use for post-training. I feel like that's this entire—I changed my view on how products are being made through that lens. Through the lens of, oh wow, you can actually collect data that you weren't able to collect before. And that you can actually change the training mechanism so that the model becomes much more collaborative, have slightly different capability.

基准设计激励 Benchmark Design Incentives

Host

是的。我还好奇,对于设计这些基准测试的人来说,设计基准的动机有多少是为了满足企业客户的需求,比如,企业会想看 X、Y、Z,就像你提到的金融领域,所以我们为 X、Y、Z 做基准测试,而有多少是为了消费者或公众关心的,比如对齐问题或情商,或者比如写作,甚至只是让写作听起来像人类。我不知道为什么这些前沿实验室都没有把这一点加入后训练,让它变得更好。现在这很糟糕。那么,什么时候这会开始引导这些模型实际设计的方向呢?

Yeah. I'm also curious, for people who are designing these benchmarks, how much of the incentives to design the benchmarks are for what enterprise customers are going to want, as in, oh, enterprise is going to want to see X, Y, and Z, like you mentioned the financials, therefore we're going to do benchmarking for X, Y, and Z, versus what maybe consumers or the public would care about, like alignment problems or emotional intelligence, and for example, writing, even just writing to sound like a human. I don't know why any of these frontier labs are not adding that into their post-training to just make it better. It sucks right now. And so when will that start to steer the direction of how these models are actually being designed?

Karina

是的,这很有趣。当我们实际上在 OpenAI 做 Canvas 的时候,你知道,50% 的用例是编码,但另外 50% 是写作。写作太难了。当然,非虚构写作你可以在各个维度上改进,但创意写作本身——我觉得这就是为什么我们需要新型公司,专注于创意叙事或创造新类型的电影或动画,对吧?这样我们就能收集到以前不可能的新数据类型,然后在此基础上训练模型。

Yeah, it's so interesting. When we actually worked on Canvas, ChatGPT Canvas at OpenAI, you know, 50% of the use case was coding, but the other 50% was writing. And writing is so hard. Of course, non-fiction writing you can actually improve on various axes, but the creative writing itself—and I feel like that's why we need new sorts of companies who purely focus on creative storytelling or creating new types of movies or animations, right? So that we can collect new types of data that was not possible before and then train a model on that.

写作为何千篇一律 Why Writing Sounds Generic

Host

为什么写作这么难?这似乎又是大多数人使用这些模型的主要用例之一,但几乎所有人都同意这些模型有一种特定的声音,听起来就是不像人类。为什么他们不能通过后训练让它听起来像普通人呢?

Why is writing so hard? It seems like again, this is one of the main use cases that most people are using a lot of these models for, but yet most everyone would agree that these models have a certain voice and sound that just sounds anti-human. Why can't they add post-training to make it sound like a regular human?

Karina

我认为当你为一个模型服务数十亿用户时,你最终会收敛到这种平均的默认声音。很难不得罪一群用户而服务另一群用户。所以你实际上承担了最小的风险,那就是默认的平均共识。我觉得人们——这基本上就是共识最大化。对于有不同写作风格的人——我觉得当它非常主观时——我认为这就是为什么我期待越来越多的人做自己的后训练,因为这样你实际上可以引导到真正喜欢的写作风格。这样,它就会成为你更好的写作助手或更好的写作生成模型。

I think when you serve one model for billions of people and users, you eventually converge to this average default voice. It's really hard to not piss off one group of users and serve another group of users. So you actually take the least risk that you can, which is the default, average consensus. And I feel like people—it's consensus maxing basically. And for people who have different writing styles—and I feel like when it's very subjective—and I think that's why I'm excited about more and more people doing their own post-training, because then you can actually steer to what's kind of like writing that you really like. And so in that way, it will become your better writing assistant or better writing generation model.

个人后训练应用 Personal Post-Training Use

Host

你个人会使用后训练来让特定模型写作更好吗?你喜欢用哪个模型?

Do you personally use that at all, post-training a specific model for writing to be better? And which model do you like to use?

Karina

我以前——我的意思是,我很喜欢用 Claude 做这个,但我确实有一个多智能体编排来处理各种写作。一个是更多规划。我在 Apple Notes 里记录原始想法,然后把这些笔记给 Claude 来润色,然后我把草稿给 GPT-4 检查语法之类的。我围绕这个有一个多智能体编排。

I used to—I mean, I like using Claude for this a lot, but I do have like a multi-agentic orchestration for all sorts of writing. One is more planning. I do my original thoughts in Apple Notes, and then I take those notes and give them to Claude to kind of refine them, and then I give this draft to GPT-4 to check for grammar and things like this. I have a more multi-agentic orchestration around that.

避免千篇一律的语调 Avoiding Generic Voice

Host

但实际上,涉及技能时,你会以某种方式提示它吗?你需要以某种方式训练它们,让它们听起来不像你提到的通用声音吗?现在它是所有声音的平均值,所以它们有这种特定的语调,因为它们不想搞砸任何一个用户及其期望。那么你在使用它们时,会以某种方式调整吗?

But actually, with skills involved, do you prompt it in a certain way? Do you have to train them in a certain way so that they don't sound like a generic voice, like you mentioned? Right now it's the average of all the voices, and so they have this specific tone because they don't want to mess up any one user and what they're expecting. So do you tweak it in a certain way when you're using them?

Karina

是的,我的意思是,我确实有一个系统用于不同类型的写作。我实际上有非虚构写作,比如博客文章,你知道当我们发布基准测试时,有这种类型的写作,还有更多创意写作,比如剧本写作。显然,对于剧本写作,你实际上希望推动模型能够创作真正的剧本之类的。所以是的,这是不同的用例。

Yeah, I mean, I do have a system that I use for different types of writing. I actually have non-fiction writing, which is like blog posts, and you know when we launch a benchmark, there's this type of writing, and there's much more creative writing, and that is like screenwriting. Obviously, for screenwriting, you actually want to push the model to be able to create actual screenplays and stuff. So yeah, it's different use cases.

Host

酷。

Cool.

Karina

我确实认为,一个模型能成为伟大的讲故事者,与一个模型能以客观方式澄清你的想法,这两者是有区别的。我认为这是两种不同的写作用例。

I do think there is a difference between a model being able to be a great storyteller versus a model to clarify your thoughts in an objective way. And I think those are two different writing use cases.

Host

确实。

Definitely.

结束提问与最终思考 Closing Questions and Final Thoughts

Host

是的,这是一个很重要的点。在我们今天结束之前,我的最后一个问题是:随着这些模型变得越来越聪明,你希望它们更多关注哪个领域?你会敦促前沿实验室或其他开发自己模型的人重点关注什么?

Yeah, that's an important part to bring up. Um, my last question for you before we actually head out of here today is: what is one area that you want to see these models focus more on as we start to evolve more and more, as these models become smarter? What's one thing you would urge either Frontier Labs or anyone else who are developing their own models to focus on?

Karina

我认为是更多样化的个性。我觉得,你知道,Claude 有自己的独特个性之类的,但还没有人真正训练出一个能写出奥斯卡获奖剧本并据此拍出电影的模型。我认为拥有一个端到端的故事引擎会非常非常有趣。

I think like more diverse personalities. I feel like I've seen, you know, Claude has its own distinctive personality or so, but no one has actually trained a model that can really write an Oscars-winning script and create a movie out of this. I think having like an end-to-end storytelling engine would be super, super interesting to see.

Host

太酷了。好的,Karina,我想感谢你今天参加 MTS 节目。非常高兴能了解 Thoughtful Lab 的更多情况,也了解了你的故事和你特别关心的事情,以及为什么模型打磨是一门艺术。它不是科学,而是一门真正将更多人类判断带入这些模型演变的艺术。这非常重要,也非常及时。再次感谢你的到来。

Really cool. Well Karina, I did want to thank you for joining us today on MTS. It was really amazing to learn more about Thoughtful Lab and also just more about your story and the things that you particularly care about, and why model crafting is this art. It's not a science. It's an art that really brings more human judgment into the evolution of these models. It's so important, so timely. So again, thank you so much for joining us.

Karina

非常感谢。

Thank you so much.

Host

MTS 是一档 XNative 直播新闻和访谈节目,实时报道科技、商业、政治和文化。每个工作日,你都可以在 X、YouTube 或任何你收听播客的平台观看我们的直播。

MTS is an XNative live streaming news and interview show covering technology, business, politics, and culture as it happens. Catch us every weekday live on X, YouTube, or wherever you get your podcasts.

互动版:逐字朗读 + 针对本期提问 →