The Pitfalls of Bad Benchmarks in AI Models
打开互动全文版(中英对照 + 朗读 + 问答)→Serge 首席执行官 Edwin Shan 讨论了优化有缺陷基准(如 El Marina)的后果,用户倾向于选择花哨、冗长的回答而非准确的回答。
Edwin Shan, CEO of Serge, discusses the consequences of optimizing for flawed benchmarks like El Marina, where users prefer flashy, verbose responses over accurate ones.
作为 Serge 的 CEO,Edwin Shan 对基础模型生态系统中发生的一切都了如指掌。据报道,Serge 估值 240 亿美元,与所有顶级实验室密切合作,帮助他们改进模型。Edwin 对许多不同事物都有非常有趣的见解。我是 Jacob Efron,今天在《无监督学习》中,我们讨论了 Edwin 与顶级实验室合作的经验,以及他迄今看到的他们采取的不同方法。我们讨论了模型的现状、他对强化学习、RL 环境、RL 即服务以及整个初创生态系统的看法。我们讨论了拥有模型品味意味着什么,以及真正优秀的研究人员会做什么。我们还谈到了 Edwin 的观点如何从“一个模型统治一切”转变为“一群不同的模型”。与一个每天都能看到未来被创造的人进行了一次精彩的对话。话不多说,有请 Edwin。Edwin,非常高兴你能来参加播客。非常感谢你的到来。今天有很多事情我想和你探讨,但我想从一个地方开始:为了准备这次采访,我听了你参与的一些播客,你说过很多有趣的事情。你一直强调的一个问题是,作为模型构建者,选择错误的基准会带来后果,我注意到你特别提到了 LMSys 和其他一些人们正在优化的东西。所以希望你能详细谈谈这些优化带来的后果。
As a CEO of Serge, Edwin Shan gets a front row seat to everything happening in the foundation model ecosystem. Serge, which is worth a reported $24 billion, works super closely with all the top labs on improving their models. And Edwin had a really interesting perspective on a bunch of different things. I'm Jacob Efron and today on Unsupervised Learning, we talked about Edwin's learnings from working with the top labs and the divergent approaches he's seen them take so far. We talked about where models are today, his view on RL, RL environments, RL as a service, and that whole startup ecosystem. We talked about what it means to have model taste and what really good researchers do. And we also hit on how Edwin's changes views from there being one model to rule them all to there being a constellation of different models. Just a fascinating conversation with someone who gets to see the future being created every day. Without further ado, here's Edwin. Well, Edwin, really excited to have you on the podcast. Thanks so much for coming on. There's a lot of things I'm excited to explore today with you, but one place I figured I'd start: I listened in prep for this to a few podcasts you've been on and a lot of things you said that were pretty interesting. One thing you've talked about pretty consistently is just the consequences of choosing bad benchmarks as model builders, and I think you've specifically called out LMSys and some of the other things that folks are optimizing for. So I'd love if you could just elaborate a bit on the kind of consequences of some of these optimizations.
我认为看待 LMSys 的方式应该是:当你针对 LMSys 进行优化时,你基本上是在优化点击诱饵。你对 LMSys 的心智模型应该是,用户登录 LMSys,输入提示,然后看到你的回答,他们的目标就是投票选出哪个更好。但他们实际上并没有仔细阅读你的回答。他们只是快速浏览两个回答,两秒钟后就选出哪个更好。所以他们并没有仔细阅读,没有注意到“哦,这个回答遵循了我所有的指令,完全准确,研究充分,质量很高”等等。他们只是非常快速地扫过,大概一两秒,然后想:“哪个给我印象最深?”他们优化的是“哪个回答有更多表情符号?哪个吸引了我的注意力?”吸引注意力的东西是很多表情符号、很多格式。所以包含大量 Markdown、大量标题的回答,他们会自然偏好看起来更长的回答。如果更长,那当然看起来更有专业感。他们不会仔细阅读。例如,我最近看了一个数据集,他们在线发布了很多这样的数据,我看到了很多令人难以置信的错误。其中一个错误是:提示是“告诉我 1452 的所有除数”,一个回答是“哦,1452 只有一个除数,那就是 1”。另一个回答完全正确。实际上问题可能是“1452 的小于 6 的所有除数”,显然 1、2、3、4 和 6 都是除数。但如果你看数据,猜猜评分者更喜欢哪个?用户更喜欢那个完全错误的回答。再想想,这是一个相当简单的数学问题。但如果你问一个更高级的提示,你会自己去研究吗?你会去核实事实吗?不会。作为用户,你不会这么做。所以最终结果是,用户基本上是在优化哪个回答吸引了他们的注意力。这几乎就像小报。有趣的是,这似乎出现在许多不同的消费者用例中。我记得大概一年前有一项关于 ChatGPT 医疗回答的研究,ChatGPT 的回答评分远高于医生的回答,但后来他们分析原因,发现仅仅是因为回答更长。
I think the way you should think about LMSys in particular is that when you optimize for LMSys, you're basically optimizing for clickbait. The mental model you should have in your head of LMSys is that users go on to LMSys, they issue a prompt, and then they get to your responses, and then their goal is to basically vote on which one is better. But they actually aren't carefully reading your responses. What you're doing instead is, okay, they just scroll through the two responses. After two seconds, they just pick which one is better. So they're not carefully reading them. They're not seeing, oh yeah, this one followed all of my instructions. This one was completely accurate. This one was well researched. It was high quality, and so on. They're not doing any of that. What they're doing is again, they're reading through very very quickly, like maybe one, two seconds, and they're like, "Okay, which one impressed me the most?" What they're going to optimize for is, "Okay, which response had more emojis? Which one caught my attention?" The things that catch your attention will be a lot of emojis, a lot of formatting. So responses that contain a lot of markdown, a lot of headers, and so on, they will just naturally prefer responses that seem longer. If it's longer, then sure, it has a sheen of expertise to it. They're not reading these things through. For example, I was actually looking through one of the datasets recently. They published a bunch of these online, and I saw all these incredible errors. One of the errors was: I think the prompt was "Tell me all the divisors of 1452," and one of the responses said "Oh, there's only one divisor of 1452, and that's the number one." And the other one got it completely correct. I think the problem was actually "What are all the divisors under six of 1452?" Obviously one, two, three, four, and six are all divisors. So that was the other response. But if you look at the data, guess which one the rater preferred? The user preferred the one that was completely wrong. Again, if you even think about it, this is a fairly simple mathematical question. But if you ask a more advanced prompt, are you going to research it? Are you going to go fact check it yourself? No. That's just not what you're going to do as a user. So what ends up happening is that users again, they're basically optimizing for whichever one caught their attention. It's almost like a tabloid. And it's interesting. It seems like something that pops up in a bunch of different consumer use cases. I remember there was like maybe a study of ChatGPT for medical responses. I forgot maybe a year ago or so, and the responses were rated way higher than physician responses, but I think when they actually unpacked why, it was literally just because the responses were longer.
是的。没错。
Yeah. Yeah. Exactly.
这是我们看到的一个大问题:即使你没有故意在 LMSys 数据上训练,比如你把一堆模型放到网站上,然后有 10 个不同的模型,你进行 A/B 测试,选出表现最好的。用户自然偏好更长的回答,这是认为回答高质量的最简单方式。所以我们看到很多模型,当你经历这个过程时,你自然就会得到比原来冗长两三倍甚至四倍的回答。显然,这可能不是人们唯一优化的东西,而且还有很多东西被忽略了。你还见过哪些其他例子,说明没有优化正确的东西会让模型走上错误的道路?
That's one of the big things that we've seen is that even when you haven't intentionally trained on LMSys data, like maybe you put a bunch of your models on the site and then okay, you have 10 different models that are on the site and you're just going to AB test which one and pick the one that performs the best. Again, users just naturally prefer responses that are longer. It's like the easiest way to think that response is high quality. And so we've seen a lot of models. I just think when you go through that process, you just naturally end up with responses that are two, three, four times as verbose as ones that aren't. And obviously this isn't the only thing that people are optimizing for probably, and leaving things on the table. What are some other examples you've seen of where not optimizing for the right thing really leaves a model on the wrong track?
我认为问题尤其出现在当你同时优化错误的数据、优化错误的目标函数,并且没有正确的测量手段时。例如,今年早些时候,我们开始与一个新团队合作。很多研究人员告诉我们,他们怀疑自己的模型在变差,但没有定量证据,因为他们没有正确的测量手段。所以我们为他们深入分析了模型,发现实际上在过去 6 到 12 个月里,他们的模型确实退步了。
I think the problems especially arise when you are both optimizing on the wrong kind of data, you're optimizing for the wrong objective function, and then in conjunction with that you don't have the right measurements in place. So for example, earlier this year we were actually starting to work with a new team. A lot of the researchers were telling us like yeah they were suspecting that their models were getting worse but they didn't have any quantitative evidence of it because they didn't have the right measurements in place. So we basically dug into their models for them and what we found is that actually over the past 6 to 12 months their models had actually regressed.
所以实际情况是,他们从所谓专家级程序员那里收集的数据——那些标注员——根本没有执行代码。他们没有仔细检查代码是否正确。相反,他们做的事情和 Elmsus 案例非常相似:提供的训练数据充满了华丽的辞藻和夸大的声明,比如“哦,是的,我为你生成了这个超棒的程序,它能做 ABC 这些事情。”而他们本应花时间执行代码——你知道,执行代码通常很难。你需要安装这些库,准备好所有基础设施,确保你真正理解这门语言,等等。所以代码实际上要么完全错误,要么充满了各种微妙的 bug,人们直到后来才会发现。结果 6 到 12 个月后,由于他们没有实际衡量模型是否在改进,他们一直在为这个错误的目标优化。这有点疯狂,因为整个行业都在前进。编程是一个非常热门的领域,其他所有人都在进步,而某些团队却在倒退,因为你没有正确的数据,没有正确的衡量标准。我认为这是一个真正值得担忧的问题:行业中有太多人没有关注他们收到的数据质量,也没有关注他们是否在衡量正确的东西。
And so what happened was the data they were gathering from supposedly expert coders, the raters, they just weren't executing the code. They weren't checking the code carefully to see that it was actually correct. And so what they were doing instead was, again very similar to the Elmsus use case, they were giving training data that was essentially full of flowery language and grandiose claims, like "Oh yeah, here I produced this amazing program for you that does ABC things." And they had taken the time to execute — you know, executing code can often be pretty hard. You have to install these libraries, you have to have all this infrastructure in place, you have to make sure that you actually understand the language, and so on. And so the code was actually just completely wrong or full of all these subtle bugs that people wouldn't notice until later. And so at the end of 6 to 12 months, because they didn't have any actual measurements in place to see whether or not the models were actually improving, they were optimizing for this. And it's kind of crazy because really the whole industry is moving forward. Coding is such a hot area where everybody else is progressing, and then whatever teams are basically making negative progress because you don't have the right data, you don't have the right measurements. I think it's a real concern where enough people in industry just aren't paying attention to the quality of data that they're receiving and whether or not they're measuring the right things.
我的意思是,有趣的是,在模型不断改进的情况下,如果你在数据方面走捷径,事情实际上会变糟。而且,在那个例子中,让我印象深刻的是,问题的一大半在于你甚至六个月都不知道你的模型是否在变好。我认为这是我们在整个行业看到的一个有趣现象,现在情况并不总是那么清晰。我认为有一代模型改进中,模型在某些方面变得更好是显而易见的,但现在感觉有时更难判断了。最好的公司都在做什么,来逐月判断他们的模型是否在变好?历史上,所有公司都非常关注基准测试——那些由研究社区创建的非常学术的基准测试。并非所有实验室都意识到基准测试的问题。几乎很容易针对基准测试进行优化。归根结底,模型非常擅长爬山非常具体、非常狭窄的目标函数。所以最终结果是,模型在基准测试上取得了巨大进步。有时这是虚假的,因为基准测试数据实际上就在训练数据中,而人们没有意识到。或者即使不是虚假的,比如他们确实在基准测试上有所改进,但他们没有意识到的是,因为你狭隘地专注于基准测试,却没有在基准测试之外设置衡量标准,人们不会意识到他们确实在某个狭窄问题上变得非常擅长,但他们的模型在更现实的问题上实际上变得更差了。也许一个简单的例子是想象为 SAT 优化。作为一个高中生,你当然可以花几百个小时优化 SAT。SAT 是一套非常狭窄的问题集。它涉及阅读理解、类比和词汇,但它并不能真正衡量你写得好不好,或者你在现实世界中解决复杂问题的能力,或者你在 SAT 领域之外做其他事情的能力。所以最终结果是,我们实际上会看到公司,比如许多前沿实验室,突然在基准测试上取得令人印象深刻的分数,而我们自己试用模型时却说:“不,你们的模型实际上变差了很多。”甚至在他们告诉我们他们针对基准测试优化之前,这种情况就已经很普遍了:我们时不时会试用一个前沿实验室给我们的模型,然后问:“为什么质量突然下降了?”我们会说:“哦,你们可能收集了大量合成数据来改进基准测试。”他们会问:“你们是这么看的吗?”我们会说:“是的,比如我们在 XYZ 基准测试上突然翻倍了,但我们没有意识到所有其他后果。”所以我认为这是一个真正值得担忧的问题。然后,最好的公司已经意识到,衡量模型性能的唯一方法是进行适当的人工评估。具体做法是,你让高质量的人——非常勤奋的人,既关注回答内容,又非常老练、有品味,以确保“是的,这是一个有趣的个性。这是前沿实验室希望他们的模型模仿的风格。”这基本上是在模仿真实用户的真实世界体验——所有多样性、所有复杂性、所有现实世界的混乱,而没有基准测试的人为概念。同时,它也依赖于来自经验丰富且值得信赖的标注员的质量,而不是那些不关注的匿名人士。进行这些严格的人工评估的过程一直是所有前沿实验室的黄金标准。除了更加关注之外,你注意到一个好的评估者需要具备什么品味?
I mean, it's interesting how things can actually go south given all the model improvements that are happening if you're taking shortcuts on the data side. And then also, I was struck by in that example, a big part of the problem is just not even knowing for six months if your models are getting better. I think this is something interesting that we've seen across the industry now where it's not always clear. I think there was a generation of model improvements where it was blindingly obvious that models were getting better in some ways, and now it feels like it's a little more challenging to tell at times. What are the best companies doing to figure out, on a month-to-month basis, if their models are getting better? So historically, all the companies have paid really hard attention to benchmarks — these very academic benchmarks that have been created by the research community. And not all the labs realized the problems with benchmarks. It's almost like very easy to optimize for benchmarks. Models, at the end of the day, are very good at hill climbing very concrete, very specific, very narrowly defined objective functions. And so what would end up happening is that the models would make a ton of progress on the benchmarks. And sometimes it would be fake because the benchmark data would actually be in training data and people wouldn't realize. Or even if it wasn't fake per se, like they were actually improving on the benchmarks, what they wouldn't realize is that because you've narrowly focused on benchmarks but don't have measurements in place outside of the benchmarks, people wouldn't realize that sure they got really good at some narrowly defined problem, but their models were actually getting worse on more real world problems. Maybe an easy example would be imagine optimizing for the SAT. As a high school student, sure you can spend hundreds of hours optimizing for the SAT. The SAT is a very narrowly defined set of problems. It's reading comprehension and analogies and vocabulary, but it's not really measuring your ability to write well or your ability to perform complex problem solving in the real world or your ability to do all these other things outside of the SAT's domain. And so what would end up happening is we would actually see companies, like a lot of frontier labs, they would suddenly get all these impressive scores on the benchmarks, and we would just play around with the models themselves and say, "No, your models only got a lot worse." Even before they told us that they were optimized for the benchmarks, it became prevalent enough that every now and then we would start playing around with a model that one of the frontier labs would give us and ask, "Why did this suddenly drop in quality?" And we'd say, "Oh, you guys probably gathered a lot of synthetic data to improve on your benchmarks." And they'd ask, "Is that what you guys have seen?" And we'd say, "Yeah, like we suddenly doubled the performance on XYZ benchmark and we didn't realize all the other consequences." So I think that's a real concern. And then basically what the best have realized is that the only way to measure the performance of their models is to run proper human evals. The way it works is you're essentially asking people — high quality people, people who are very diligent, people who are both paying attention to the content of the responses, but then are also really sophisticated and have a lot of taste to make sure, "Yeah, this is a fun personality. This is the kind of style that the frontier labs want their models to emulate." And it's basically mimicking the real user real world experience — all the diversity, all the complexity, all the messiness of the real world without the artificial concept of benchmarks. And then it's also relying on the quality from really experienced and really trusted raters, as opposed to just anonymous people who aren't paying attention. This process of going through these rigorous human evals has been a gold standard for all the frontier labs. Beyond paying a little more attention, what have you noticed about what makes a really good evaluator have taste?
所以从高层次来看,可能包括以下三点。第一当然是专业知识本身。比如,如果你在评判我们的许多评估集,它们非常高级。它们可能衡量模型进行代数拓扑研究的能力,或者衡量模型使用 PyTorch 的能力。所以首先,你需要非常聪明、在评估领域有深厚专业知识的人。这是第一点。
So at a high level, it might be the following three things. One would be certainly expertise in itself. Like if you're judging a lot of our evaluation sets, they are very advanced. So they might be measuring a model's ability to perform algebraic topology research, or they might be measuring a model's ability to use PyTorch. So first of all, you just need really intelligent people with a lot of expertise in the domain that they're evaluating. So that's one piece.
另一个维度是所谓的“品味与精致度”。归根结底,当你让模型写代码时,你不仅关心正确性,还关心代码是否写得优雅、设计是否良好。在创意写作中,它是否是一篇真正优秀的文章,引入了新想法、文笔出色,并且没有那种“AI 味”?这就是精致度。另一个是创造力。我认为人们常常低估了提示词的重要性。他们以为评估模型就是评估回复,但要准确衡量模型,你需要覆盖你想要模型擅长的整个分布的提示词。以创意写作为例,如果你的提示集是一千个故事或文章,而且都用同样的方式表述,而不是模拟真实世界中人们与模型交互的长尾分布,那根本行不通。人们常常低估了创建好提示词和发挥创造力的难度。就像那个现象:如果你让我说出 50 种食物,除非我使劲想或者限定范围,否则真的很难。同样,以这种方式发挥创造力也出奇地难。所以创造力是另一个维度。
Another would be this notion of sophistication and taste. At the end of the day, you want models to be — when you ask them to write code, you don't just care about correctness; you care about whether it's well-written and well-designed. Or in creative writing, was it a really well-written essay that introduced new ideas, had great prose, and doesn't feel like AI slop? So that's sophistication. Another is creativity. I think people often underestimate how important prompts are. They think evaluating models is just about evaluating the response, but to measure models properly, you need prompts that span the entire distribution of what you want the model to be good at. For creative writing, if your prompt set is a thousand stories or essays all phrased the same way, as opposed to the long tail of how people actually interact with models in the real world, that won't cut it. People often underestimate how difficult it is to create good prompts and to be creative. It's like that phenomenon: if you ask me to name 50 foods, it's really hard unless I think hard or constrain the list. Similarly, it's surprisingly hard to be creative in this fashion. So creativity is another piece.
是的。
Yeah.
最后,第四个维度是遵循指令的能力。当客户要求你评估模型时,他们通常有非常具体的标准——风格指南、特定个性,或者更看重 XYZ 标准而非 ABC。这些指令往往非常复杂,所以人们需要非常擅长遵循它们。说到当前模型改进的一个主要主题:强化学习环境和奖励模型现在非常热门。你能谈谈这种转变,以及你是如何支持客户应对的吗?
And finally, the fourth piece is the ability to follow instructions. When clients ask you to evaluate models, they often have very specific criteria — a style guide, a certain personality, or they care about XYZ criteria more than ABC. Often these are very complex instructions, so people need to be very good at following them. Moving to one of the main themes of model improvement these days: RL environments and reward models around them are all the rage. Can you talk about that transition and how you've supported your customers with it?
当然。我认为我们的环境是一种延续——训练范式的下一步。历史上,大量工作投入了 SFT、RLHF,然后是验证器。我们的环境是这一进程的下一步。我相信未来还会有其他东西。我们的环境非常有趣,因为我们已经研究了相当长一段时间——大概一两年。有趣的是,行业其他公司才刚刚开始跟进。例如,我们与 Meta 出色的智能体团队合作已超过一年。这个团队创建了 GAIA 基准测试,并刚刚开源了他们的智能体研究环境平台。他们确实看到了未来的浪潮。我们的环境有有趣的差异。要支持它们,你需要尽可能模拟真实世界的丰富世界。就像创建提示词出奇地难,因为你希望它们有创意、多样且像真实世界一样丰富,而不是合成基准提示词。如果让人们毫无约束地创建提示词,你得不到多样性。同样,创建这些世界时,你需要底层实体——人、企业、工具、互动、消息、Slack 消息、电子邮件、日历事件——都以复杂的方式模仿真实世界。我们很大一部分工作是构建工具、基础设施、质量控制和数据测量来确保这一点。然后我们创建这些世界中的所有工具——MCP 服务器、浏览器、代码执行——让模型能在世界中运行并执行提示词。接着我们创建提示词本身,测试模型表现,设计任务来挑战前沿模型的极限。还有一个重要的测量和内省方面:当模型失败时,我们深入探究原因。所以支持这一切需要大量的基础设施和工具。
Yeah, definitely. The way I think about our environments is that they're a continuation — the next step in training paradigms. Historically, a lot of work went into SFT, then RLHF, then verifiers. Our environments are the next step in that progression. I believe there will be other stuff in the future. Our environments are really interesting because we've been working on them for quite a while — maybe one or two years. It's interesting that the rest of the industry has only just started picking up on them. For example, we've been working with the amazing agents team at Meta for over a year. This is the team that created the GAIA benchmark and just open-sourced their agent research environment platform. They really saw the wave of the future. Our environments are interestingly different. To support them, you need rich worlds that simulate the real world as best as possible. Just like creating prompts is surprisingly hard because you want them creative, diverse, and rich like the real world, not synthetic benchmark prompts. If you let people create prompts without constraints, you don't get diversity. Similarly, when creating these worlds, you need underlying entities — people, businesses, tools, interactions, messages, Slack messages, emails, calendar events — all mimicking the real world in complex ways. A big part of our effort is building the tooling, infrastructure, quality control, and data measurements to ensure this. Then we create all the tools in these worlds — MCP servers, browsers, code execution — so models can run within the world and execute prompts. Then we create the prompts themselves and test how models perform, coming up with tasks that test the limits of frontier models. There's also a big measurement and introspection aspect: when a model fails, we dig in to understand why. So there's a lot of infrastructure and tooling to support this.
你做这项工作已经有一段时间了。你在这个过程中学到了什么?有什么让你惊讶的?或者有没有一开始做错、后来改进的地方?
You've been doing this work for a while. What have you learned along the way? What surprised you? Maybe something you initially got wrong and have gotten better at?
我们试图提出的一个有趣观点是,关注模型的轨迹非常重要,这样才能理解它们成功和失败的原因。
One interesting viewpoint we've tried to propose is that it's very important to pay attention to model trajectories to understand why they are succeeding and why they are failing.
人们常常低估模型为了得到正确答案而进行奖励黑客的程度。
People often underestimate the amount to which models can reward hack themselves to the correct answer.
是的,我觉得有很多非常有趣的例子。比如,我最喜欢看这些例子,因为它们经常以各种疯狂的方式表现。同样,我认为人们实际上低估了模型可能失败的多种方式,以及这说明了模型的底层能力。人们有一种观念:好吧,我只要给一个最终奖励,如果我拥有奖励,一切都会顺利。但实际上,模型往往会以非常奇怪的方式偏离,或者根据底层能力的不同表现出不同类型的智能。如果你没有充分加固这些底层能力,模型可能在短期内表现良好,但基本上你会在未来遇到很多问题。
Yeah, I feel like there's a lot of very funny examples of that. Like, one of my favorite things is looking at some of these examples because they often just perform in all these crazy ways. And similarly, I think people have actually underestimated the myriad of different ways that models can fail and what that says about the model's underlying capabilities. People have this notion that, okay, I'm just going to give a final reward and if I own a reward, everything's going to work out. But what happens in practice is that models tend to just deviate in these very odd ways, or they can show different types of intelligence depending on what the underlying capabilities are. If you don't shore up those underlying capabilities enough, models may seem to perform well in the short term, but you're basically going to face a lot of problems down the road.
这感觉像是你在实验室工作中反复出现的主题,对吧?有些容易爬坡的事情,但如果你没有以深思熟虑的方式配合正确的评估,它实际上并不能服务于你的最终目标。所有这些都联系在一起,因为我们在基准测试中看到过,人们最终会黑掉这些基准测试,所以他们以为自己在进步,因为每个人都在关注的数字在上升,而他们的底层模型并没有变得更智能。
This feels like a theme in a lot of the work you do with the labs, right? There are easy things to hill climb on, but if you're not doing it in a deeply thoughtful way with the right evaluations, it doesn't actually serve the end purpose you're going for. All these things tie together because we've seen it with benchmarks where people will end up hacking these benchmarks, so they think they're making progress because the numbers everybody's paying attention to are going up, when their underlying model isn't becoming materially more intelligent.
是的,我认为所有这些都联系在一起。
Yeah, I think all these things tie together.
显然,随着强化学习环境成为改进模型的主流方式,涌现了一波新的创业活动。我觉得大概有 30 家 YC 公司在尝试构建强化学习环境。有很多这种“强化学习即服务”的公司。我想知道你怎么看待这种活动,以及作为后续,感觉每次实验室改进模型的主要方式发生转变时,有些人就会视之为介入的机会。但显然,前一波的人也有持续性。所以我想知道你对此有何反思。
And obviously, alongside RL environments becoming a more consensus way to improve models, there's been a flurry of net new startup activities. I feel like there's probably 30 YC companies trying to build RL environments. There's a lot of these RL as a service companies. I wonder what you make of that activity, and as a follow-up, it feels like every time there's a transition in the main type of ways labs are improving their models, some people see that as an opportunity to step in. But obviously there's persistence from people that were there in the previous wave. So I'm wondering how you reflect on that.
我认为硅谷让我抓狂的一点是这种“转向文化”,人们不断转向任何看起来最热门的话题或能带来最高估值的东西,而不是构建他们真正相信的东西。这几乎是一个有趣的现象:硅谷很多人谈论华尔街追逐金钱,但硅谷在做什么?追逐同样的东西——估值、风投——转向不是因为你有惊人的好主意,而是因为 YC 告诉你这样做才能实现目标市场契合,并在融资时向风投展示一些收入。所以我真的很不喜欢这种文化。
I think one of the things that always drives me crazy about Silicon Valley is this pivot culture where people are constantly pivoting to whatever seems the latest hot topic or whatever seems to drive the highest valuations, instead of building things they really materially believe in. It's almost like a funny thing where a lot of people in Silicon Valley talk about Wall Street chasing money, but what is Silicon Valley doing? Chasing the same thing—valuations, VCs—pivoting not because you have some amazing great idea, but because that's what YC told you to do to achieve target market fit and get some revenue to show VCs when fundraising. So I really dislike that culture.
嗯,让我印象深刻的是,显然没有人清楚强化学习环境作为改进模型的主流方式会持续多久。也许它会持续一段时间,也许不会。但你显然已经展示出随着时间的推移,无论模型改进的方式如何,你都有相应的产品来支持。我相信这在很大程度上源于你与客户的深度绑定。我可以想象有些人会说:“构建强化学习环境与早期的 surge 行为截然不同。他们有多大权利去做这件事,而不是那 30 个新进入者?”我很好奇这在实地是如何感受的。
Well, and what I'm struck by is obviously it's not clear to anyone how long RL environments will be the flavor of the day to improve models. Maybe that persists for a while, maybe it doesn't. But what you've clearly shown the ability to do over time is whatever the way models are being improved, you have an offering that helps support that. And I'm sure a lot of that comes from being deeply embedded in your customers. I could imagine some folks saying, "Look, building an RL environment is so different than the first few acts of surge. To what extent do they have the right to go do that versus the 30 new entrants?" I'm curious how on the ground that's felt.
如果你考虑我们的环境涉及什么,它们涉及三件事。第一,我们的环境只是训练以启用 AGI 所需数据的下一个迭代。这符合我们的基本论点:我们想要创建任何需要的数据来实现这一点。所以这是很大一部分。第二,我们的环境需要大量的工具——创建所有这些工具、运行模型、测量模型、分析模型等等的工具。这与基于人类反馈的强化学习(RLHF)也需要工具没有什么不同。这是一种不同类型的工具,但所有这些部分都是一样的。即使在我们做基于人类反馈的强化学习(RLHF)时,你也需要大量工具来确保你能分析模型在对话中的进展,理解成功和失败,并确保创建提示和评估响应的 surge 工作者能够以高质量和多样化的方式完成所有这些事情。所以显然你需要大量工具来支持这一切。这与我们领域中的许多其他公司非常不同,它们本质上只是人员配备机构,历史上没有构建任何技术。但我们始终首先是一家技术公司。所以这只是另一种类型的工具,就像任何技术公司构建工具一样。第三,创建强化学习环境完全是为了获得非常丰富、复杂、创造性的数据,除了使用人类之外别无他法。例如,想想 SWE-bench:什么是 SWE-bench?它是由人类创建的 PR 集合,而且即使 SWE-bench 本身也不够干净。
If you think about what our environments involve, they involve three things. One is that our environments are just the next iteration of the data needed to train to enable AGI. That fits with our fundamental thesis: we want to create whatever data is needed to enable that. So that's a big piece. Second, our environments require a lot of tooling—tooling to create all these tools, to run the models, to measure them, to analyze them, and so on. That's not any different from the fact that for RLHF you needed tooling as well. It's a different kind of tooling, but all those pieces are the same. Even when we've done RLHF, you need a lot of tooling to ensure you can analyze the models as they progress through the conversation, understand the wins and fails, and make sure the surge workers creating prompts and evaluating responses can do all these things in a high quality and diverse way. So you obviously need a lot of tooling to support all this. This is very different from a lot of other companies in our space that are essentially just staffing agencies and historically haven't built any technology. But we've always been a technology company first and foremost. So it's just another type of tooling, in the same way that any technology company builds tooling. And the third piece is that creating RL environments is all about getting really rich, complex, creative data, and there's just no other way to do it beyond using humans. For example, think about SWE-bench: what is SWE-bench? It's a collection of PRs created by humans, and even SWE-bench itself wasn't quite clean enough.
所以人们需要构建 Sweet Bench 验证,也就是人们基本上以各种方式处理问题、评估和清理。同样,我们从根本上相信,创建我们的环境是一个人类数据问题,只是需要大量的技术。我认为有些人可能认为质量就等同于资历,对吧?他们会说,如果我有某个领域的博士学位,做标注或者花时间评估某件事,那当然是高质量的。但也许帮我们理解一个例子,就是实际上有一个人看起来资历很好,能改进模型,但在实践中却做不到。
So people need to build Sweet Bench verified and that's, you know, people basically taking problems and evaluating them and cleaning them up in various ways. And so just in the same way, I think we fundamentally believe that creating our environments is a human data problem that just requires a lot of technology. I think some people probably just assume quality is synonymous with credentials, right? And they're like, well, if I have a PhD in something, labeling or spending time on evaluating something, of course that's high quality. But maybe help us understand an example of something where you actually have someone that seems credentialed on paper to improve models, but it's not happening in practice.
我一直喜欢举的一个例子是海明威。海明威没有博士学位,我甚至不知道他有没有读完大学。我们寻找的是世界上在每个技能上最优秀的人,不管他们的资历如何。想想谷歌雇人的方式,对吧?谷歌不是只看简历上的学校或学位。在谷歌这样的公司,晋升靠的是你实际做的工作。我们平台的一个有趣之处是,我们有一个技术平台,查看他们创建的所有数据并进行衡量。我们每天收集数百万个关于工人的信号。我们看到他们生产的数据类型,这几乎是你所能想象的最精英化的制度。所以,不是因为你碰巧有哈佛的学位就能晋升,不,我们会实际衡量你所做的工作,并据此提升你。当然,我们平台上有大量哈佛学生和博士。我认为我们可能是世界上最大的博士来源。但这还不够,原因有两个。第一,即使考虑程序员,如果你是一个麻省理工毕业生,有计算机科学学位,而且你真的非常优秀,你可能不会真正去创建好的数据来训练这些模型,反而可能试图欺骗系统。你是一个很好的程序员,你对红队系统着迷,对对抗性攻击着迷,你会试图找到一种方法……另一个原因是,即使回想我在麻省理工的同学,或者我为 First 面试过的麻省理工毕业生,说实话,一半的人甚至不会编程。在教科书上读到某件事和在实际中执行它之间有很大的区别。这又回到了我之前说的基准测试表现与现实世界表现的区别。前沿实验室的一个大问题是,它们的模型过于教科书式智能,而不是拥有在现实世界中做事的街头智慧。
An example I always love to give is take Hemingway. Hemingway didn't have a PhD. I don't even know if he completed college. And yeah, I mean, what we're looking for is the greatest people in the world at every skill regardless of their credential. So even just think about who works at Google, right? Google doesn't just hire people based on what school they have on a resume or what degree they have. And that's not how you progress at a company like Google or any other company. The way you progress is based on the actual work that you do. And one of the interesting things about us as a platform is that we have a technology platform that looks at all the data they're creating and then measures it. So we gather millions of signals on our workers every day. We see the types of data they're producing, and so it's almost like the most meritocratic thing possible that you can imagine. So as opposed to, okay, you progress because you happen to have a degree from Harvard, no, we are going to actually measure what you do and advance you based on that. And sure, we have a ton of Harvard students on our platform. We have a ton of PhDs on our platform. I think we're probably the biggest source of PhDs in the world. But that just isn't sufficient. And it's not sufficient for two reasons. One is, even if you think about coders, if you're an MIT grad who has a computer science degree and you're actually really, really good, what you're probably not going to do is actually try to create really good data to train these models. Instead, you're probably just going to try to cheat the system, right? Like, okay, you're a really good coder, you're fascinated by red teaming systems, you're fascinated by adversarial attacks. What you're going to try to do is find a way to... And another part of it is, even if I think about all of the people who were in my class at MIT, or the number of people that I've interviewed from MIT for First itself, honestly, half of them can't even code. There's a very, very big difference between reading about something in a textbook and then having the street smarts to execute it. This almost ties back to what I was saying about performance on benchmarks versus performance in the real world. A big problem that the frontier labs have had is that their models are almost too textbook intelligent instead of having the street smarts to do things in the real world.
你之前提到,今天有一个实验室让你参加他们的内部大型会议,这显然说明了你与这些团队的合作有多紧密。几个月前,Meta 决定收购他们密切合作的一个供应商,也就是 Scale。鉴于这些关系的密切程度,我知道其中一部分也是为了引进人才,但我很好奇你的反应,以及你是否认为随着时间的推移,考虑到关系的密切程度,这些事情是自然发生的?
You know, you alluded to earlier, I think you mentioned one of the labs is having you be at their internal big conference today, and obviously it speaks to just how closely you work with those teams. You know, obviously you had, I guess, a few months ago, Meta decide to buy one of the vendors they work with really closely, and Scale. And given the proximity of those relationships, I know part of that was also bringing some talent into the organization, but I'm curious about your reaction to that and do you think over time some of those things are natural to happen given just the proximity of relationships?
我认为 Scale 的收购对我们来说实际上很棒,因为在那之前,我们一直很低调。所有顶尖的研究人员和实验室都已经知道我们。他们知道我们拥有迄今为止最高的质量,也知道我们是最大、最快的。但 AI 领域发展如此之快,每天都有越来越多的人进入这个领域。所以 Scale 的收购基本上让我们成为焦点,这对我们的扩张非常有帮助。一夜之间,我们从所有这些新团队那里获得了大量新需求,这对我们非常有利。
I think the Scale acquisition was actually amazing for us because up until then, we'd been pretty under the radar. So, I think all of the top researchers, all of the labs already knew about us. So, they knew that we had the highest quality by far. And they knew that we were the biggest and the fastest. But the AI field has just been growing so large that more and more people are entering the field every single day. And so basically the Scale acquisition just put a spotlight on us that was really, really helpful for expanding. I think we just got so much new demand from all these new teams overnight, and so that was really beneficial for us.
你提到了这些公司吸引的人才文化差异,以及人们优化目标的不同。我想知道这在你们为 Search 所做的不同决策中是如何体现的,以及这些业务随着时间的推移可能走上不同的轨迹。
You kind of mentioned this difference in maybe the culture of folks that are attracted to these types of companies and the things that people are optimizing for. I'm wondering how that manifests itself in maybe the different decisions you've made for Search today and also maybe the different trajectories that these businesses go in over time.
好问题。我认为这在很多方面塑造了我们,其中最主要的一个是招聘。我们试图建立的文化和招聘的人,是那些内心像研究员、从根本上关心数据和 AI 的人。当我们思考作为一家公司要构建什么时,我们更像一个研究实验室,而不是另一个追逐金钱、炒作和估值的硅谷初创公司。一个具体的体现是,我们招聘那些对研究、数据和实现 AGI 有根本兴趣的人,而不是那种由增长黑客和为了增加收入不择手段的人所代表的硅谷类型。想想那些人的动机,他们通常会试图卖给你你可能不需要、他们也不认为会真正改进你模型的东西。本质上,他们只是像销售员一样,而不是深入挖掘你的需求、你的模型问题,确保你理解所有衡量模型改进的方式,而不是卖给你蛇油。
Good question. So I think it's shaped us in a bunch of ways, and maybe one of the most prevalent ones is in hiring. If I think about the culture that we're trying to set and the type of people that we're trying to hire, it's people who are kind of like researchers at heart and people who just fundamentally care about data and AI. Like when I think about us as a company and what we are trying to build, I think of ourselves as a lot more like a research lab than just another Silicon Valley startup that's trying to chase money and hype and valuations. And one of the concrete ways that manifests is again if you think about this idea of hiring people who are fundamentally interested in research and data and enabling AGI, as opposed to this type of Silicon Valley person who's kind of embodied by growth hackers and people who are just doing whatever it takes to increase your revenue. If you think about the incentives of those types of people, what they will often do is they'll basically try to sell you things that you may not need that they don't think will actually improve your models. I mean, essentially, they'll just try to act like salesmen as opposed to digging deep into what you need, digging deep into the problems with your models, trying to make sure that you understand all the different ways that you should be measuring your models to make sure that they're actually improving, as opposed to just selling you snake oil.
说到模型进步,你如何描述模型从这里变得更好的路径?你觉得大多数顶级实验室之间是否有共识,还是你看到了一些相当不同的方法?
Speaking of model progress, how do you articulate the path to models getting better from here? Does it feel like there's consensus among most top labs, or do you see some pretty divergent approaches?
是的,我认为在所有训练范式上,分歧比我预期的要多得多。几乎感觉每个前沿实验室都有自己的看法。有时这些看法截然不同,有时只是彼此略有变化,但分歧比我预期的要多得多。
Yeah, I think there's been a lot more divergence than I expected for all the training paradigms out there. It almost feels like every frontier lab has their own take on it. Sometimes those takes are wildly different. Sometimes they're just slight variations on each other, but there's a lot more divergence than what I expected.
这种分歧存在的一些关键方面是什么?
What are some of the key vectors where that divergence exists?
我认为在高层次上,公司有两种分歧方式。一是他们选择目标函数——基本上是什么类型的训练算法,什么类型的训练数据。我不能说得太多,但我认为所有前沿实验室之间一个被低估的差异是他们在优化什么和关注什么上的选择。我可以举几个例子。一个例子是我之前提到的 AlpacaEval。我觉得有趣的是,一些前沿实验室完全选择不去关注它。我认为那些前沿实验室做得更好,因为那些不得不关注它的前沿实验室——那些实验室的研究人员经常告诉我,他们讨厌 AlpacaEval。他们理解优化 AlpacaEval 会导致负面进步的所有方式。它会导致模型产生幻觉,因为 AlpacaEval 用户不在乎幻觉。相反,他们喜欢幻觉,因为当你的模型产生幻觉时,它听起来狂野、疯狂,有趣且诱人,就像小报一样。所以,一些前沿实验室有勇气和基本论点,即他们相信通往 AGI 的道路,而不是觉得需要追逐公众关注和炒作得很厉害的公开排行榜——我认为某些前沿实验室因为对自己的目标有核心信念而感到自由和勇气不去关注它——这确实塑造了很多模型进步。然后另一个有趣的分歧是前沿实验室试图优化的目标选择。我认为你可以在 OpenAI 和 Anthropic 之间看到最明显的区别。想想 OpenAI,他们现在在优化什么?他们似乎更倾向于优化用户参与度——很长的会话时长或每日用户数量——而像 Anthropic 这样的公司可能更倾向于优化生产力,以及你能提取多少价值,几乎像 GDP 或通过交互模型节省的时间。我认为这塑造了他们构建的产品类型、吸引的人才类型以及模型的能力。所以我认为这确实以我们开始看到的方式塑造了他们。
I think at a high level there are two ways in which the companies diverge. One is once they choose their objective function—basically what type of training algorithm, what type of training data they are going to gather. I can't speak too much about that, but I think an underestimated difference between all the frontier labs is their choice in what they optimize for and what they pay attention to. So I can give a couple examples. One example is I've mentioned AlpacaEval earlier. And I think one of the fascinating things to me is that some frontier labs have just chosen not to pay attention to it at all. And I think those frontier labs have done better because the frontier labs who have had to pay attention—what researchers those labs have often told me is that they hate AlpacaEval. They understand all the ways in which optimizing for AlpacaEval will lead to negative progress. It will lead to models that hallucinate because AlpacaEval users don't care about hallucinations. If anything, they love them because when your model hallucinates, it just makes it sound wild and crazy and kind of fun and enticing, again like a tabloid. And so the fact that frontier labs like some frontier labs have had the fortitude and the underlying thesis, the underlying belief in what they see as a path to AGI, as opposed to feeling like they need to chase publicity and chase a very hyped up and publicly visible leaderboard—I think just the fact that certain frontier labs have felt a freedom and fortitude not to pay attention to that because they had such a core belief in what they're going for instead—I think that has really shaped a lot of model progress. And then I think another interesting divergence is again in the choice of objective that the frontier labs are trying to optimize for. And I think you can see this clearest difference between OpenAI and Anthropic. If you think about OpenAI, what are they optimizing for now? It's almost like they are leaning more towards optimizing for user engagement—really long sessions or amount of daily users—as opposed to a company like Anthropic who might be optimizing for something more akin to productivity and how much value you can extract, how much almost like GDP or how much productivity or time savings you can extract by interacting with the model. And I think that shapes the types of products they build, the types of people they attract, the capabilities of their models. So I think it just really shapes them in ways we're starting to see.
这是一个有趣的观点,因为它引出了一个更大的问题:迄今为止的模型进步都归结为一个能做很多不同事情的核心大模型。现在感觉你可以想象模型围绕消费者参与度优化,然后是一个生产力模型。你认为会有一个模型能够根据优化目标进行上下文切换,还是我们最终会进入一个可能需要企业模型甚至按行业划分的模型(如金融、法律或其他领域)的世界?
It's an interesting point because it feeds into this larger question of model progress to date has all come back to one core large model that can do lots of different things. Now it feels like you could imagine models optimizing around consumer engagement and then a productivity model. Do you think there'll be one model that can context switch across whatever is being optimized for, or do we end up in a world where you probably want this enterprise model or even per industry models for finance, legal, or other things?
是的,所以我认为这是我想法分歧很大的一个地方。我过去认为基本上会有一个模型统治一切——你有一个超级智能模型,它应该能够上下文切换并适应你想让它做的任何事情。但实际上,在过去一年里,我开始意识到,几乎每家公司都应该有一个论点,即世界是如此丰富。永远不会有放之四海而皆准的解决方案。相反,每家公司、每个实验室或每个人工智能都需要有一个基本论点,关于什么在现实世界中有用,以及哪种人工智能最能服务人类。这个论点将塑造模型的行为。当然,两个模型可以同样智能,但它们会有不同的个性,对特定问题的回答有不同的偏见,与你交谈的方式也不同,等等。就像如果谷歌要构建一个社交媒体平台,它会与 Facebook 构建社交媒体平台的方式非常不同。如果 Facebook 要构建一个搜索引擎,它会与谷歌构建搜索引擎的方式截然不同。所以没有绝对的对错。只是不同的公司、不同的人对什么对世界有用和有益有不同的基本信念。所以我认为同样的事情会发生在人工智能上。
Yeah, so I think this is one place where my thinking has diverged a lot. I used to think that there would essentially be one model to rule them all—you have some super intelligent model and it should be able to context switch and adopt whatever you want it to do. But actually, I think over the past year I've started to realize that it's almost like every company should have a thesis on like the world is just so rich. There's never going to be a one-size-fits-all solution. Instead, every company or every lab or every AI needs to have a thesis underlying it of what will be useful in the real world and what kinds of AI will best serve people. That thesis will shape how the model behaves. Sure, two models can be just as intelligent as each other, but they'll have different personalities, different biases for how they answer particular questions, different ways in which they converse with you, and so on. And so just in the same way that if Google were to build a social media platform, it's going to be very different from the way Facebook built social media platforms. If Facebook were to build a search engine, it would have a very different take from how Google would build a search engine. So there's no right or wrong answer per se. It's just that different companies, different people have different fundamental beliefs in the things that are useful and good for the world. And so I think the same thing will happen with AI.
那么这对你关于应该有多少人构建模型的看法意味着什么?你认为随着时间的推移我们会看到更多的模型参与者,还是从现有的参与者中整合?
So what does that mean for your take on how many people should be building models? Do you think we'll see more model players over time or consolidation from the folks we do have today?
我绝对认为应该有更多的人构建人工智能,因为关于这一点,我认为需要许多许多不同类型的论点来说明什么样的人工智能对世界有用,而且我认为还没有人弄清楚这一点。
I definitely think that there should be more people building AI because to that point, I think that many, many different types of thesis are needed on what kinds of AI will be useful for the world, and I just don't think anybody's figured that out yet.
公司应该训练自己的模型吗,比如一个非常大的金融公司或一个非常大的医疗组织?
Should companies be training their own models, like a really large finance firm or a really large healthcare organization?
所以我确实认为,最终每家公司都应该训练自己的模型,因为这些模型对世界如此重要。你最终会希望将它们部署到 99.9999% 的用例中,对吧?如果你仅仅依赖前沿实验室的模型,它们优化的方向可能并不是你想要的。再说一次,因为 AI 将如此重要,你希望把它部署到各处,我认为如果你想获得最佳价值和最佳性能,是的,你应该自己训练。
So I definitely think that eventually every company should be training their own models, and that's because these models will be so important to the world. You'll want to eventually deploy them to like 99.9999% of use cases, right? And if you simply rely on models from the frontier labs, what they're optimizing for may not be what you're optimizing for. Just again, because AI will be so important and you want to deploy it everywhere, I think if you want to get the best value, best performance possible, yeah, you should be training your own.
你能通过一些好的提示词或轻量微调来实现这种优化吗?还是你认为实际上需要从头开始构建?
Can you achieve that optimization through just some good prompting or some light fine-tuning, or do you think it actually requires building somewhat from scratch?
所以,我认为这又回到了我之前说的:关于 AI 应该如何服务你的客户,以及你想构建什么样的 AI,你需要有一个基本论点。如果你有一个强有力的论点——这几乎就像产品论点——而不是仅仅构建一个商品化产品,比如我们相信你对模型的行为有独特的见解,那么是的,我认为这完全合理。显然,我们会好奇这样做的成本如何随时间变化,因为做出这种投资或持有独特见解的权衡——今天要达到最先进水平相对昂贵,但我想很多人可以带着论点构建这些有主见的模型,即使落后最先进水平 6 到 12 个月,也完全可以接受。
So, I think this again goes back to what I said earlier about having a fundamental thesis on how AI should serve your customers and what types of AI you want to build. So if you have a strong thesis — it's almost like having a product thesis — but if you have a strong thesis, as opposed to just building some commodity product, like we believe that you have some unique take on how they should behave, then yes, I think it makes absolute sense. Obviously, we'll be curious to see how the cost of doing that changes over time, because certainly the trade-off of making that investment or having that unique take — today it's kind of relatively prohibitively expensive to get to the state-of-the-art, but I imagine a lot of people could build these opinionated models with a thesis and still be 6 to 12 months behind state-of-the-art and be okay.
是的,没错。就像我认为公司现在还没有完全准备好,考虑到——我是说公司自身和 AI 的状态——但随着 AI 越来越好,我认为它会越来越重要。
Yeah. Yeah. Exactly. Like I don't think companies are quite ready right now given the state of — I mean both the companies themselves in a state of AI — but as AI gets better and better, I think it will be increasingly important.
显然,模型在编程和这些容易验证的领域进步很大。你知道,曾几何时,你用 ChatGPT,一个新模型出来,你会明显感觉到模型变好了。我不知道过去 3 到 6 个月是否还是这样。你显然身处其中。你觉得现在模型在编程之外还在变好吗?
Obviously it's very clear models are getting way better at coding and these easily verifiable domains. You know, I think there was a time where you'd use ChatGPT and a new model would come out and it would be blindingly obvious that the models had gotten better. I don't know if I would necessarily say that's been the case over the last 3 to 6 months. You're obviously on the inside of this stuff. Do you feel like models are still getting better outside of coding right now?
是的,我确实认为它们在变好。部分原因是我运行的所有评估都显示持续进步,但同样真实的是,就在前几天,我开始更多地用 Claude 来写作,我震惊于它现在比几个月前好太多了。
Yeah, I definitely think they do. And in part that's through all the evaluations that I run where we see this constant progress, but it's also true that I think just the other day I started using Claude a lot more for writing in particular and I was just shocked by how much better it was today compared to a couple months ago.
我确实想谈谈正在构建的多模态模型,无论是视频、机器人还是生物领域的工作。我相信你考虑过这些。这对你有多大吸引力?是类似的一系列问题,还是你如何描述其中的异同?
I do want to make sure I hit on some of the multimodal models that are being built, whether it's video, robotics, stuff being done in bio. I'm sure you've thought about some of this stuff. To what extent is that interesting to you? Is it kind of like a similar set of problems or how would you characterize what's similar and different there?
是的,我认为所有这些模态都让我着迷。我认为人们没有意识到的一点是,我们实际上已经在所有这些领域大量工作了。比如,我认为我们目前 50% 甚至更多的工作实际上是在纯文本之外的领域。
Yeah, so I think all these modalities are fascinating to me. I think one of the things that people don't realize is that we actually work very heavily across all of these spaces already. Like, I think maybe 50% or more of our work today is actually in domains outside of pure text.
所以我确实觉得这很迷人。再说一次,如果你想想我们的论点和我们试图实现的目标,那就是我们想不惜一切代价实现 AGI。是的,如果我们想要在现实世界中有用的 AI,它需要理解所有这些能力,需要跨越所有这些领域运作。所以我们只想不惜一切代价实现这一点。在视频语境中,质量意味着什么?我当然能想象文本用例,但对于视频,这对你们来说意味着什么?
So I actually do think it's fascinating. And again, if you think about our thesis and what we're trying to enable, it's this idea of we just want to enable AGI no matter what it takes. And yeah, if we want AI that is useful out there in the real world, it needs to understand all these capabilities, needs to operate across all these domains. So we just want to do whatever it takes to make that happen. What's like quality mean in the video context? I could certainly imagine text use cases, but for video, what's that mean to you guys?
是的,我认为人们常常低估质量的重要性,即使是在那些被认为出奇简单的模态中。举个例子,即使在你创建的提示词类型中,你也需要大量的创造力和技术来确保你探索了空间的完整分布。我认为人们常常低估这一点,因为他们认为人们可以凭空创建提示词,覆盖模型能力的完整分布,而实际上这极其困难。我们的一个重要目标就是经常尝试教导人们,或确保我们的客户理解,当你考虑质量时,你需要超越机械的指令遵循和机械的正确性,去思考提示词本身的所有其他含义。所以,是的,我认为人们常常低估质量意味着什么。
Yeah, I think people often underestimate how important quality is even across these modalities that people think are surprisingly simple. So, one example is even in the types of prompts that you're creating, you need a lot of creativity and kind of technology to make sure that you're exploring the full distribution of the space. I think people often underestimate that because I think they just think that people can create prompts out of thin air that target the full distribution of a model's capabilities when it's actually sparkling difficult. One of our big goals is that we often try to teach people or try to make sure our customers understand that when you think about quality, you need to go beyond robotic instruction following and robotic correctness to think about all these other implications of the prompt itself. So yeah, I think people often underestimate what quality means.
是的。我的意思是,是什么让视频评估器或视频本身质量更高或更低?当你想到文本模型在 Elo 评分中遇到的问题时,我想视频方面也有类似的陷阱。你对此有什么了解?
Yeah. I mean what makes like either a video evaluator or like video itself like higher or lower quality? As you think about maybe the problems that text models have run into with Elo ratings, I imagine there are similar traps on the video side. What have you learned around that?
是的,所以我认为这又归结为我之前提到的品味和精致度的概念。就像,好吧,当然。你让斯科塞斯拍一个视频,做一个关于鱼的视频。然后你让一个高中艺术毕业生做同样的事,就在大急流城拍。当然,两个人都能拍出关于鱼的电影,但斯科塞斯的电影可能会好得多。这就是品味、精致度、创造力和超越常规的用武之地。因为当你思考你想从模型中得到什么时,不仅仅是字面遵循指令、做你所说的事。而是创造一些让你惊叹的东西,一些富有想象力和创造力、提升标准的东西。这就是我们努力的方向。
Yeah, so again I think it boils down to this notion of taste and sophistication that I've mentioned before. It's like, okay, sure. You ask Scorsese to film a video, create a video about a fish. Then you ask your high school arts graduate to do the same thing, just pull over Grand Rapids. Sure, both people can create a film about a fish, but Scorsese is probably going to have a much better film about the fish. And that's where that notion of taste and sophistication and creativity and just going above and beyond comes in. Because when you think about what you want from models, it's not just the ability to literally follow your instructions and do whatever you say. It's to craft something that will blow your mind, something that feels imaginative and creative and raises the bar. That's what we're striving to do.
然后像机器人和生物这些领域,需要硬件组件来收集数据。
And then like robotics and bio in these spaces that require like a hardware component for data collection.
你认为这些是像你们这样的公司的自然延伸,还是你们可能已经在做了,或者这完全是另一批公司会冒出来做的事情?
Do you think those are natural extensions for companies like you guys or maybe you're already working in them or is that like a totally separate set of companies that might pop up and do that?
我的想法是,我们愿意不惜一切代价去获取那些能加速 AGI 的数据。有时这需要构建新工具,有时需要购买新硬件设备,有时需要拓展到我们所在的任何领域。我们是一家科技公司,会尽一切努力实现目标,而不是一家受局限的公司。我们刚刚转型进入这个领域,并没有只考虑短期。与其他一些公司不同,我们更长远地思考实现这一切所需的一切。
The way I think about it is we want to do whatever it takes to enable the data that is going to help accelerate AGI. Sometimes that involves building new tools, sometimes buying new hardware equipment, sometimes expanding to whatever space we are in. We're a technology company. We're going to do whatever it takes to make that happen, as opposed to being a narrowly constrained company. We just pivoted into this area and we're not really thinking short-term. In contrast to some other companies, we are thinking longer term about everything that's needed to achieve all these things.
太棒了。我们总是喜欢在采访结束时来一轮快问快答,听听你对一些标准话题的看法。我们可能已经聊过几次了,但我很好奇,过去一年里你在 AI 方面有什么看法发生了改变?
Amazing. Well, we always like to end our interviews with just a quick fire round where we get your take on a standard set of topics. So maybe we've hit on this actually a few times, but I'm curious one thing that you've changed your mind on in AI in the last year.
最大的变化可能是,我以前认为会有一个模型统治一切,但现在我确实看到了这些不同的产品观点、AI 观点将如何塑造未来的每一个 AI。
Probably the biggest thing is this idea that I used to think that there would be one model to rule them all, and now I actually do see how these different product opinions, AI opinions, will shape up every AI going forward.
显然,Serge 似乎取得了一系列令人难以置信的成功,你建立了一家了不起的公司,但我很好奇,回顾过去,你在创业过程中犯过最大的错误是什么?
Obviously it seems like Serge has just been a series of incredible wins and you built an amazing company, but I'm curious in reflecting back what's the biggest mistake you've made in building the business?
我的背景一直是研究和数据科学,所以我以前很喜欢发表文章和写博客。我喜欢分享我们所有的见解,早期我们确实这么做了。但后来我太忙了,就没再继续,我很怀念那种向世界传授知识、分享我们对行业的看法以及为了确保走在正确道路上需要做出什么改变的感觉。所以我认为最大的错误是过去两三年我们停止了大量发表。我现在希望能纠正这一点。
So my background has always been in research and data science, and so I used to love publishing and blogging. I used to love sharing all of our insights, and we did that early on. Then I somehow just got too busy to do it, and I really missed that idea of teaching the world and sharing our viewpoints on the industry and what needs to change or happen to make sure we're on a good path. So I think the biggest mistake is that we kind of stopped publishing as much in the past two or three years. I'm hoping to fix that now.
如果你有一周假期,可以坐下来写一篇很长的文章,你会写什么?或者说,你目前最关注的是什么?
If you were to get a week of vacation where you could just sit back and write a really long piece, what would you write about or what's kind of most top of mind for you?
我最关注的是目标函数这个概念,以及每个前沿实验室都在优化什么。我认为这非常微妙,而且影响深远。你是在优化参与度?优化有用性?优化用户数量?还是优化 GDP?无论是什么,我认为这个概念对整个行业和 AI 都有非常深远的影响。
So the most top of mind thing to me really is this concept of objective functions and what every frontier lab is optimizing for. I think it's surprisingly subtle and has surprisingly far-reaching consequences. Are you optimizing for engagement? Are you optimizing for usefulness? Are you optimizing for number of users? Are you optimizing for GDP? Whatever it is, I think that concept has very far-reaching consequences for the industry and for AI at large.
如果你在运营一个实验室,你会优化什么?
What would you optimize for if you were running one of the labs?
我会优化这样一个概念:一个月后,你会为这次与模型的互动感到高兴吗?它是否在某种程度上改变了你的生活?这样的时刻越多越好。它可能改变你的生活,比如你问关于度假的事,它向你介绍了一个你从未想过的新地点;或者你有一个医疗问题,不太知道怎么表达,但 AI 偶然注意到了一些东西,教会了你一些你本来不会发现的事情。我认为我们现在非常需要更多这样的东西。
So I would optimize for this notion of a month later, would you be happy that you had this interaction with this model? Would it have almost changed your life in some way? The more moments they can get like that, and could it change your life because maybe you were asking about a vacation and it introduced you to this new location that you've never thought about before, or maybe you had a medical question and you didn't quite know how to phrase it, but then the AI serendipitously noticed something and taught you something that you wouldn't have figured out otherwise. I think that is something we need a lot more of right now.
我注意到,我们在 AI 中提出的这些问题,很多都引出了我们社会一直面临的挑战,对吧?你之前提到了 SAT 的挑战,那当然是一种不完美的衡量智力的方式,但我们从未真正找到更好的方法。同样,关于技术应该优化什么、我们应该努力改善人们生活的哪些方面,也有很多讨论。这是一个很难回答的问题,但随着这些模型越来越好,并且会在我们优化的任何目标上爬山,这个问题显然变得越来越重要。
I'm struck by how much of these questions we're asking in AI just bring up challenges that we've always had as a society, right? I mean you were alluding to the SAT challenge earlier and that's certainly an imperfect way of measuring intelligence, but we've never really found much better ways. Similarly, there's been lots of talk about what technology should be optimized on and what we should be trying to improve in people's lives. It's a hard question to answer, but obviously an ever important one as these models do get better and are going to hill climb on whatever it is we are optimizing for.
是的,完全正确。我认为我对 AI 的很多思考,以及 AI 带来的担忧和后果,都与社交媒体有相似之处。
Yeah. Exactly. I think a lot of the way I think about AI, and maybe the worries and consequences of AI, is analogous to the parallels with social media.
完全同意。Vlad,这是一次非常精彩的对话。我想把最后一句话留给你。我们的听众可以去哪里了解更多关于你和 Surge 的信息?还有别的吗?话筒交给你,你想把大家引向哪里都行。
Totally. Vlad, this has been a fascinating conversation. I want to make sure I leave the last word to you. Where can our listeners go to learn more about you about Surge? Anything else? The mic is yours. Wherever you want to point folks.
是的,我强烈推荐我们的博客。我们开始更多地写博客,分享更多的见解和分析。所以一定要去看看。
Yeah, so I would definitely suggest our blog. We're starting to blog a lot more, starting to share a lot more insights and analyses. So definitely check that out.
太棒了。非常感谢,这非常有趣。
Amazing. Well, thanks so much. This was a ton of fun.
非常感谢。
Thanks much.