Applied Compute CEO Yash 谈强化学习的边界、新型 AI 超大规模服务商,以及为何后训练决定推理胜负

Applied Compute CEO Yash on the Limits of RL, the New AI Hyperscaler, and Why Post-Training Wins Inference

亚什·帕蒂尔 Yash Patil · Unsupervised Learning · 2026-10-06 · 约 57 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Applied Compute CEO Yash 阐释为何“拥有自己的智能”关乎灵活性与控制权,以及后训练如何成为赢得推理的关键。

Applied Compute CEO Yash explains why owning your own intelligence is about flexibility and control, and how post-training is becoming the key to winning inference.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 34)

全文 · Full transcript(中英对照)

介绍与嘉宾背景 Introduction and Guest Background

Host

你如何描述我们今天的处境?

How do you characterize where we are today?

Yash

我们在强化学习中所拥有的本质上是一台爬山机器。最难的部分实际上是定义要爬的山,这就是为什么评估对公司来说非常重要,要专注于构建并仔细保护。你的员工不是可替代的。你不会愿意他们去另一家公司工作。这些模型也是如此。拥有自己的智能并训练自己的模型在当今的 AI 时代精神中非常流行。关于所有这些话题,可能没有比来自 Applied Compute 的 Yash 更合适的人选了。Applied Compute 正在与一些最前沿的 AI 应用合作,处理非常有趣的后训练示例。Yash 和我能够讨论他对当今开放模型一系列顶级问题的看法,包括何时真正有意义去托管和训练这些模型,如何思考未来,当这些事情有意义时,尤其是随着实验室随着时间的推移从不同领域获取越来越多的数据。我们还讨论了何时有意义去利用优化 versus 实际进行后训练,以及 Yash 对当今生态系统中许多非常有趣的事情的看法。总之,与这个领域的一位杰出人才进行了一次引人入胜的对话。我希望大家能像我一样享受它。事不宜迟,有请 Josh。

What we have with RL is essentially a hill climbing machine. The hardest part is actually defining the hill to climb, which is why eval is really important for companies to focus on building and safeguard carefully. Your employees are not fungible. You would not be comfortable with them going to another company and doing work there. Same thing with these models. Owning your own intelligence and training your own models is incredibly in the AI zeitgeist these days. There's probably no better person to speak with about all these topics than Yash from Applied Compute. Applied Compute is working with a bunch of the most cutting-edge AI applications on really interesting post-training examples. Yash and I were able to talk about his perspective on a bunch of the top questions around open models today, including when it actually makes sense to host and train these models, how to think about the future, when this stuff makes sense, especially as the labs get more and more data from different domains over time. We also talked about when it makes sense to harness optimize versus actually go and post-train, as well as just get Yash's thoughts on a lot of the really interesting things happening in the ecosystem today. All in all, just a fascinating conversation with a brilliant mind in the space. I hope folks enjoy it as much as I did. Without further ado, here's Josh.

Host

非常感谢你来参加播客。

Well, thanks so much for coming on the podcast.

Yash

谢谢邀请我。是的。

Thank you for having me. Yeah.

Host

期待这一期已经有一段时间了。

Been looking forward to this one for a while.

Yash

是的。是的。不,这是很久以来一直期待的事情。很高兴我们终于能亲自做这件事。

Yeah. Yeah. No, it's been a long time coming. It's good we can finally do it in person as well.

Host

是的。在我们这个小壁橱般的工作室里,面对面总是更好。

Yes. Always better in person in our little closet of a studio here.

Yash

是的。这很棒。我的意思是,它非常舒适,就是为这种场景建造的。

Yeah. This is awesome. I mean, it's very cozy, built exactly for this setting.

Host

是的,绝对不是改造的图书馆,而是专为播客建造的。

Yeah, definitely not a repurposed library, purpose-built for podcast.

Yash

所以谢谢邀请我们。是的。

So thanks for having us. Yeah.

Host

当然。我觉得我们可以从很多地方开始,但我认为对听众来说最有趣的是,拥有自己的智能似乎成了时代精神。我不知道转折点是什么时候发生的,大概三四个月前,但它确实在讨论中,我觉得关于依赖前沿实验室的风险,已经有很多不同版本被阐述。有人说,你永远不知道他们什么时候会撤走 API,当他们决定某件事是安全风险时,就不给你访问权限。还有人说,他们会构建产品并与你竞争。如果你使用他们的模型,他们会在你的数据上训练。有各种各样的论点。我很好奇,当你思考这个问题时,支持这一点的理由。哪些论点引起了你的共鸣?哪些你认为可能有点被夸大了?

Of course. Well, I feel like there's a bunch of places we could start, but where I thought would be most interesting for our listeners is just owning your own intelligence feels like it's in the zeitgeist. I don't know when the tipping point happened, you know, 3-4 months ago, but it is certainly in the conversation and I feel like there's been a bunch of different versions articulated of the risk of depending on Frontier Labs. There's been people saying, you know, you'll never know when they'll pull their APIs away, when they decide something's a safety risk and they don't give you access. There's been, you know, they'll build products and compete with you. They'll be if you use their models, they'll train on your data. There's all sorts of different arguments. And I'm curious when you think about this, the case for this. Which of these arguments resonate for you? Which of them do you think maybe are a little overblown?

掌控智能:灵活与控制 Owning Your Intelligence: Flexibility and Control

Yash

是的。是的。所以,如你所知,我的背景是我实际上在开放实验室工作过。我认为,当谈到拥有自己的智能时,风险甚至不是风险,只是我思考的事情,并不是那些假设实验室是恶意的,比如他们在窃取你的数据之类的。我认为这些公司由伟大的人组成。有合同义务。实际上,在保护数据等方面有很多工程投入。但我确实认为,拥有自己的智能的想法实际上是关于灵活性和控制。所以你说,哦,我不知道这个拥有自己智能的叙事是什么时候出现的。我认为它真正出现是在开放模型开始变好,你实际上可以用它们做事的时候。但你获得的控制实际上可以比这更深,对吧?比如这些模型在哪里运行,你如何部署它们,你优化它们是为了成本、延迟还是特定领域,你获得的控制越多,这就是我们所说的拥有自己的智能。所以如果你想拥有自己的智能,并且希望能够训练这些模型在你的数据上执行任务,你实际上需要访问权重。你需要有基础设施来推动训练它们、推理它们、在它们之间路由。你知道,我确实认为过度依赖单一模型供应商存在风险。幸运的是,我认为实验室正在尽力服务尽可能多的人。但有时模型会从不同产品中被撤下。这实际上发生过不止一次,特别是在编码领域。所以,我认为随着这些实验室开始更多地进入特定垂直领域,不清楚他们是否会继续服务那些直接竞争的公司。因此,最安全、最持久的事情,你可以做并建立优势,我认为是投资于全栈,能够制造为你的产品提供动力的模型,这些模型是智能的,做你想做的事情,并且不对任何一组模型或公司有极端的供应商依赖。

Yeah. Yeah. So, as you know, my background is I actually worked at open labs. And I think the risks or not even risks, just the things that I think about when it comes to owning your intelligence aren't more the ones that presume that the labs are malicious or something like that where they're stealing your data or whatnot. I think these companies are made up of great people. There are contractual obligations. There's actual, you know, there's a lot of engineering that goes in protecting data and whatnot. But I do think that the idea of owning your intelligence really is about flexibility and control. And so you said, oh, I don't know when this own your intelligence narrative popped up. I think it really popped up when the open model started getting good. And you could actually go and do things with them. But the control you get, it can actually be a lot deeper than that, right? Like where these models run, how you deploy them, what you optimize them, whether it's cost or latency or specific domains, the more control you get, that's what we're talking about when we say own your intelligence. So if you want to own your intelligence and you want to be able to train these models to do tasks on your data, you actually need to have access to the weights. You need to have the infrastructure to push train them, to inference them, to route between them. You know, I do think there are risks about being super dependent on a single model vendor. Luckily I think the labs are doing their best to sort of serve as many people as possible. But there are times when models can get pulled from different products. That's happened actually more than once, particularly in the coding space. And so, I think as these labs start to move more into specific verticals, it's not clear whether they'll continue serving companies that are directly competing against. And so the safest, most durable thing that you can go and do and the advantage you can build up I think is investing in the full stack to be able to make models that are powering your products that are intelligent that are doing the things that you want and not having extreme vendor dependence on any one set of models or companies.

多模型未来与OpenAI携手Baseten Multi-Model Future and OpenAI's Partnership with Baseten

Host

是的。感觉这个叙事真的赢了。我的意思是,我认为也许最有趣的体现之一是,甚至昨天 OpenAI 宣布与 Baseten 的合作,对吧?你可以使用 OpenAI 的积分,几乎通过 Baseten 在开放模型上消耗它们。

Yeah. And it feels like this narrative has really won out. I mean, I think maybe one of the most interesting manifestations of this is you even yesterday OpenAI announcing this partnership with Baseten, right? Where you can kind of use OpenAI credits almost burned them down on open models through base 10.

Yash

所以,我认为首先,我们是 B 10 团队的大力崇拜者。我认为 Tuin 和 Danny 以及所有那些人,他们真的是一家了不起的公司。但是的,我认为这是一个明确的信号,嘿,这是一个多模型未来。客户想要选择。他们想要灵活性。我认为在 OpenAI 方面,这是一个明确的动向,嘿,我们想更像一个平台。我们希望人们能够从他们在不同产品上的承诺中提取。我们想成为企业使用 AI 的网关。我认为这对开源模型来说是一个巨大的胜利。你知道,很多人想使用这些东西。现在有一个网关,你可以去使用各种专有或开源模型,我认为这是朝着这个多模型未来迈出的一大步。

So, I think first of all, we're big admirers of the B 10 team. And I think that Tuin and Danny and all those folks, they're really an incredible company. But yeah, I think this was a clear signal that, hey, it's a multi-model future. Customers want choice. They want flexibility. I think on the OpenAI side, it's a clear motion towards, hey, we want to be more of a platform. We want people to be able to draw down from their commits on different products. We want to be the gateway to enterprises using AI. And I think it was a big win for open source models. You know, a lot of people want to use this stuff. The fact that there is a sort of gateway now that you can go and use a variety of models proprietary or open source I think it's a big step towards this multi-model future.

转向帕累托曲线 The Shift to Pareto Curves

Host

不知道你有没有注意到,几个月前的叙事绝对是:嘿,我们要讨论、推广和分析最强、最大、最厉害的模型。而现在更多是在谈帕累托曲线。对,存在成本与性能的权衡。你不需要为每个任务都用一个巨型模型,实际上你能获得更多选择和灵活性。

I don't know if you noticed this, but the narrative a couple months ago was definitely like, hey, we are going to talk about and promote and analyze the strongest, biggest, baddest model out there. And now it's like much more talking about a Pareto curve. Right. There's some cost-performance trade-off. You don't need a mega model for every single task and you actually get more choice and flexibility.

Yash

我认为开放模型把这一点推向了更极端的程度:你可以通过选择不同家族、不同规模的不同开放模型,沿着帕累托曲线滑动;但我们的核心论点是,你可以把这条帕累托曲线向外推,你可以训练模型在特定事情上做得更好,而且实际上能以更低的成本获得更好的模型。所以所有这些,我认为都是人们想要更多灵活性的普遍趋势。他们对在生产用例中部署 AI 已经变得更自在了。他们做了大量探索,现在情况是:嘿,这个 AI 东西真的在起作用了。让我们去优化它。

I think open models take that to an even more extreme where it's like hey you can slide along the Pareto curve by picking different open models of different sizes from different families but also in you know what our core thesis is is you can move that Pareto curve out you can train models to be better at particular things and and you can actually get better models for cheaper cost and so all of these things I think are like um you know sort of the general trend of people wanting more flexibility they've gotten more comfortable with uh deploying AI and production use cases. They've done a lot of exploration and now it's like hey we're this AI thing is really working. Uh let's go and optimize it.

OpenAI的中立市场策略 OpenAI's Neutral Marketplace Play

Host

你觉得那个开放 IBS 10 的事情会成功吗?我的意思是,这挺有意思的,他们是否处于一个位置,能成为企业的中立伙伴,就像:你知道,我们会给你最好的模型,不管它是哪个。

Do you think the like opening IBS 10 thing will work? I mean it's an interesting like are they in a position to be this like neutral partner to enterprises of like you know we'll give you whatever model is the best.

Yash

这也是我们很兴奋能参与的事情。棘手的地方在于,比如对亚马逊来说,做一个市场就非常合理,因为你有 Bedrock,你可以运行 Anthropic 的模型,你可以运行 OpenAI 的模型,你可以运行开放模型。那么,OpenAI 会不会把它推到极致,让你也能运行 Anthropic 的模型和 Google 的模型?我不确定。

It's something that we're also excited about participating in. Uh the tricky thing is right like it makes a lot of sense for an Amazon for example to be uh a marketplace because you have bedrock and you can run anthropic models, you run open AI models, you can run open models. Now like is openi going to take it to the extreme where you can run anthropic models and ji models? I I'm not sure.

后训练降本提速 Post-Training for Cost and Speed

Host

你提到后训练能够把帕累托前沿向外推,看起来这方面有很多非常有趣的例子。你知道,最明显的是很多应用构建者已经意识到:哇,当我在大规模下做类似的工作负载时,推理账单真的会飙升。或者,嘿,当体验更快时,我的客户体验会好得多。完全同意。所以我觉得我听说过的很多早期后训练用例,都是关于成本和速度,感觉它们是非常有价值的东西。

You mentioned kind of this pushing out of the Pareto frontier that post training enables and it seems you know like there's been a ton of of really interesting examples of this. you know, it it feels like most notably a lot of application builders have realized, wow, when I do similar workloads at massive scale, like that inference bill really runs up. Um, or you know, hey, uh, this experience for my customers is so much better when it's when it's faster. Totally. And so I feel like a lot of the, you know, initial post training use I heard about were like, you know, cost and speed feel like tremendously valuable things.

Yash

完全同意。

Totally.

Host

能力方面。我很好奇你能不能稍微谈谈,你知道,它们确实在把帕累托前沿向外推,对于同样规模的模型,你大概能获得更好的能力。你怎么看当前能够通过后训练在能力上击败前沿模型、比如最好的巨型模型的状态?

Capabilities. I'm curious if you could talk a little bit about like you know they're certainly pushing out a Pareto frontier of of you could get for the same size model probably better capabilities. How do you think about like the current state of being able to post train to beat frontier models like the best mega models at at like you know capabilities?

Yash

我认为我们看到的很有意思的一点是,显然前沿模型在基础能力上领先于开源模型。但你能对开放权重模型进行后训练,这意味着如果你有差异化的数据,一些你专有的、处于模型分布之外的数据,那时你实际上也能把这些模型的能力推得更高。我们的一般看法是,你的数据越处于分布之外,后训练就越可能真正给你带来能力提升。但同样,一切都关乎那条帕累托曲线,对吧?所以如果我能拿一个更便宜的模型,把它训练到与前沿模型相当,但成本好得多,你知道,有时你就能服务更多用户,为此你的数据必须确实相当差异化,但大多数人实际上做后训练,基本上是为了优化性价比。

So I think the the interesting thing that we're seeing is you know there's obviously frontier models that are are ahead of the open source models on base capabilities. Um but the fact that you're able to post-train the open weight models means if you have differentiated data something proprietary to you that's out of distribution of the models that's when you can actually push these models to be better on capabilities as well. Our general take has been like the more out of distribution your data is the more likely that post training is actually going to give you that capabilities lift. Um but also again it's all about that Pareto curve right? So if I'm able to take a cheaper model and train it to be parody with the frontier but at much better cost um you know sometimes you're able to serve more users for that your data has to actually be quite differentiated but most people are actually postraining um to basically optimize uh price performance

随时间变化的分布外数据 Out-of-Distribution Data Over Time

Host

在能力方面,你知道,我觉得一个有趣的问题是:哪些数据在多长时间内仍然处于分布之外?对吧,显然,你知道,我觉得实验室们正聚焦在编程上,所以,也许他们显然也在其他领域做工作,但可能不像几年后那么专注。我很好奇你怎么看,比如,我不知道,五年后。你知道,Applied Compute 合作的多少企业会觉得:哦,我们的数据仍然处于模型所做的事情的分布之外。

on the capability side you know I think it's an interesting question of like what data remains out of distribution over what period of time right like obviously you know I feel like the labs are laser focused on coding you know and so you know, maybe they're obviously doing work in other domains, but maybe not as as focused as they might in in a few years from now. I'm curious how you think about like, you know, I don't know, five years from now. Uh, uh, you know, how many businesses that applied comput works with feel like, oh, our data is still out of distribution from like what the models did.

Yash

是的。所以,我认为我在这里的诚实回答是,我们训练模型的方式将会改变,对吧?我们一直处于离线强化学习的范式里:你去从数据供应商公司采购高质量数据集,你去专门对它们做爬山优化,然后得到一个在那个领域非常尖峰的模型。强化学习就是从经验中学习。我认为很多在线强化学习方法实际上是在利用生产推理来改进你的模型,这实际上将成为大多数企业把它们的判断专长编码化的方式:基本上,它们越使用模型,模型就变得越好。但一般来说,真正会处于分布之外的专有数据,是公司内部产生的东西。所以,你知道,这方面最极端的例子是——向 Liam 致敬,他曾经是我的经理,Liam Fettis,极其处于分布之外,因为他们在制造那些数据,他们实际上在真实世界里做实验、获取奖励信号之类的。但我认为在企业里,那看起来像是判断轨迹。你知道,一家公司的运作方式可能与另一家公司非常不同,因为你知道,它们有不同的风险阈值,有不同的运营模式,有不同的专长、历史数据,所有这些东西,而这些都会影响它们所做的决策。我们思考分布外数据的方式,也许是我们把它放在离线强化学习的语境里,就像:哦,必须有来自某个数学问题或某种基于评分标准的强化学习问题之类的高度可验证的信号。有时候你会通过配置来优化 AI 系统。所以你会去给模型写提示。你会把东西放进上下文里。而有时候,实际上优化策略要合理得多。也就是它如何做判断、如何做推理,而这些实际上更难编码。这就是我们将在五年后看到的地方,我认为很多公司基本上会有系统来收集专家的输入,然后把它转化为更好的策略。是的。

Yeah. So, so I think like my my my honest answer here is that the way we're going to be training models is going to change, right? So we've done you know we we're in this paradigm of like offline RL you go and you procure high quality data sets from uh data vendor companies you go and hill climb those specifically and you get a model that's very spiky in that domain RL is uh reinforcement learning it's learning from experience um I think a lot of the online RL methodologies are actually using production inference to go and improve your models that is is is actually going to be like how most enterprises codify their judgment expertise by basically the more they use the model the better that it gets. Um but generally right like the the the proprietary data that will be out of distribution is things that are produced inside of the company right so like you know the very extreme example of this right is uh shout out to to Liam who used to be my my manager Liam Fettis extremely out of distribution right because they're making that data they're actually going in the real world um running experiments getting reward signal things like that but I think in the enterprise what that looks like is like judgment traces Right. Um you know uh the way one company operates can be very different than the the way another company operates because you know they have different uh thresholds for risk, they have a different operating model, they have different expertise, historical data, all that sort of stuff and that contributes to the decisions that they make. The way we think about out ofdistribution data is um maybe like we're putting it in the context of offline RL where it's like oh there has to be this like highly verifiable signal from like some math problem or some um you know rubric based RL problem or something like that. There are times when you're going to go and optimize an AI system by configuration. So you're going to go and prompt the model. You're going to put things in context and there are times where you're actually it makes a lot more sense to optimize the policy. So how it does judgment, how it does reasoning and these things are actually like harder to codify. Um and that's where we'll see you know in five years from now I think uh you know a lot of companies will basically have systems to basically gather uh input from experts and then turn that into better policies. Yeah.

将组织判断力编码化 Codifying Organizational Judgment

Host

所以,你知道,随着这些模型被使用,它们会编码一个组织做决策的方式,然后你用它来创建、来改进模型。

So you know as these models get used kind of codify a bunch of the like the the way that an organization makes decisions and you kind of use that to create you know to to improve models.

CEO在研究与企业间的角色 Data Efficiency and Continual Learning

Host

生成的数据足够做这件事吗?我觉得做很多这类事情都需要大量数据。

Is there enough data that's generated to do that? I guess it feels like to do a lot of these things requires a large amount of data.

Yash

我认为部分原因在于这是一个全新的领域,对吧?这里正在进行新的研究。一切都关乎数据效率和从稀疏奖励中学习。我们有预训练,它数据效率不高但泛化性很强。然后我们有监督微调——这是对无监督学习的双关,对吧,播客的名字——但监督微调数据效率更高,但仍然不是很好。然后有带可验证奖励的强化学习,你通过可重放环境和某种验证器来传播强化学习信号,用数据数量换取质量。这仍然相当糟糕,对吧?仍然不是非常数据高效。对你问题的回答是,将会有更数据高效的训练方法。这就是在线强化学习或同策略自蒸馏或一些这类技术或 SAO 等许多新方法正在推进的方向。我基本上认为,阻碍持续学习的是从稀疏奖励中进行极其数据高效的训练,这是一个未解决的问题。

I think part of this is a green field space, right? There's new research happening here. It's all about data efficiency and learning from sparse rewards. So we had pre-training, which is not that data efficient but very generalizable. Then we had supervised fine-tuning—pun on unsupervised learning, right, the name of the podcast—but supervised fine-tuning is more data efficient but still not very good. Then there's RL with verifiable rewards where you trade data quantity for quality with replayable environments and some sort of verifier to propagate the RL signal. It's still pretty bad, right? It's still not very data efficient. The answer to your question is there will be more data efficient methods of training. That's what a lot of these newer methodologies in online RL or on-policy self-distillation or some of these techniques or SAO—these types of things that people are pushing on. I basically think what's blocking continual learning is extremely data efficient training from sparse rewards, which is an unsolved problem.

Host

是的。

Yeah.

Yash

是的。

Yeah.

非静态权重与反馈循环 CEO's Role in Research vs. Enterprise Focus

Host

作为 CEO,你如何看待在应用计算背景下你应该在多大程度上致力于解决这个问题,还是说,嘿,只要真正理解企业问题,随着研究进展,你可以取用一些成果?

How do you think about as a CEO the extent to which you should be working on that problem within the applied compute context, or is it like, hey, just understand the enterprise problems really well and as that research progresses you can take pieces?

Yash

是的。所以我们现在实际上正在与客户一起做这件事。这种在线训练,我们能够利用大型推理部署,人们服务数万、数十万用户,然后使用这些生产数据来改进模型。我们并不拘泥于任何特定类型的训练。强化学习和这些在线方法目前是最好的,但这些会继续发展和改进。这就是为什么我们团队很大一部分实际上是研究系统团队,与真实客户一起在生产中尝试许多新方法。更新模型和训练模型有巨大潜力。而且,我相信任何玩过这些工具的人都意识到——我不知道你是否做过一些非常棒的提示工程或创建一些技能,或者有其他获取知识的方式。我认为有时可能会有一种感觉,研究人员可能有点像拿着锤子找钉子,比如,哦我们可以用模型来做,还有其他方法。你有多确信长期解决方案涉及组织内许多不同种类的模型,而不是以其他方式拼接东西?

Yeah. So we're actually doing this today with customers. This sort of online training where we're able to take advantage of large inference deployments where people are serving tens of thousands, hundreds of thousands of users, and then use that production data to improve the models. We aren't married to any specific type of training. RL and these online methodologies are the best things right now, but these will continue to evolve and get better. That's why a large portion of our team is actually a research systems team trying a lot of new methodologies in production with real customers. There's a ton of potential in updating models and training models. And also, I'm sure anyone that's played around with these tools has realized—I don't know if you do some really great prompt engineering or create some skills or there's other ways to take knowledge. And I think there's maybe a feeling sometimes that researchers can be a little bit of a hammer looking for a nail, like, oh we can do it with the model, and there's other ways to do it. How confident are you that the long-term solution involves many different kinds of models of an organization versus stitching together things in other ways?

真实数据与在线训练案例 Non-Static Weights and Feedback Loops

Yash

是的,我认为我们相当确信未来看起来像是一组非静态的权重。这是我们深信不疑的。将会有某种从现实世界回到模型的反馈循环。这需要对每个任务都做吗?

Yeah, I think we are pretty confident that the future looks like a non-static set of weights. That is something that we deeply believe in. There is going to be some sort of feedback loop from the real world back into the model. Does that need to be done for every single task?

Host

在编程领域有一些很好的例子,但你觉得在哪些方面反馈循环有意义,哪些方面可能没有?

There's been some good examples of that in the coding world, but what do you feel like above the line of where that feedback loop makes sense and maybe what's below it?

Yash

基本上是那些随时间变化的事情。你通常想要某种闭环,让新信息回到模型中。但我也认为,从基础模型开箱即用获得的智能范围正在扩大,到了某个时候,这对于许多手动、面向流程的任务来说就足够了。需要迭代训练的是那些更面向判断的事情。

Basically things that change with temporality. You often want some sort of closed loop where new information is making its way back into the model. But also I just think there's this expanding umbrella of intelligence that you're getting from the base models out of the box, and at some point that is going to be sufficient for a lot of these manual, process-oriented tasks. The things that will require sort of iterative training are the things that are more judgment-oriented.

降低TCO与持续训练 Real-World Data and Online Training Examples

Host

在你可以谈论的范围内,有哪些例子能想到,关于如何利用这些现实世界数据来改进?

To the extent you can talk about it, examples that come to mind of ways of taking this real-world data and using it to improve?

Yash

是的,是的,基本上我们今天合作的一个客户案例是他们有一个助手产品,很多用户使用。他们让这个助手去完成不同的任务,然后反馈模型做错了什么或做对了什么。所以,与其——构建某种巨大的技能文件或上下文之类的东西非常不切实际。所以我们实际上能够基于直接的用户反馈,使用一些在线训练方法来偏置模型,我们能够看到用户满意度、完成率的真正提升,工具调用失败减少等等。要回答的主要问题是做这件事的总拥有成本是多少?后训练有多贵?这就是我们正在构建的很多东西:我们如何真正降低强化学习模型的成本?我们如何能够自动和在线地做到这一点,而不需要你去构建完全可重放的手动调整的强化学习环境等等?所以我们看到的是,后训练和实际调整权重的成本正在下降。你仍然最好去做一堆上下文优化和框架优化。但一旦你达到一个点,你已经从你构建的框架和工具中榨取了很多,你实际上仍然可以通过优化策略来更好地使用这些工具,继续榨取很多。所以我们合作的一些公司,他们有非常专业的框架和自定义工具等等。比如,我们正在与一家芯片公司合作。他们实际上构建了一堆软件,通用模型并不真正知道如何很好地使用,去做 ETL 验证之类的事情。所以事实证明,当你应用这种强化学习的金发姑娘算法时,你实际上可以将模型推向更好的性能。

Yeah, yeah, so basically one of the examples of customer work we're working with today is they essentially have some sort of assistant product that a lot of users use. They ask this assistant to go and accomplish different tasks and then give feedback on what the model did wrong or right. And so instead of—it's pretty impractical to go and build some sort of giant skills file or context or something like that. So we're actually able to bias the model based off of the direct user feedback using some of these online training methodologies, and we're able to see real upticks in user satisfaction, completion rates, decreased tool call failures, things like that. The main question to answer is what's the TCO of doing this? How expensive is it to post-train? And that's a lot of the stuff that we're going and building: how do we actually decrease the cost to go and RL model? How can we actually do this automatically and online without you having to go and build hand-tuned RL environments that are fully replayable and things like that? And so what we see happening is the cost to post-train and actually adjust the weights is falling. You're still better off going and doing a bunch of context optimization and harness optimization. But once you reach a point where you've sort of squeezed a lot out of the harness that you've built and the tools that you've made, you actually can still continue to squeeze a lot after out of optimizing the policy to use those tools better. So some of the companies that we work with, they have really specialized harnesses with custom tools and things like that. Like, we're working with a company in chips. They've built a bunch of software actually that the general models don't really know how to use well to go and do like ETL verification and things like that. So it turns out when you apply this Goldilocks algorithm of RL, you can actually push the models to a lot better performance.

每家公司拥有自己的模型 Lowering TCO and Continuous Training

Host

你谈到了降低总拥有成本(TCO)并实现更持续训练的努力。那里最难解决的问题是什么,我们在这个旅程中处于什么位置?

You talked about this effort to lower the TCO and enable more continuous training. What are the hardest problems to solve there, and where are we on that journey?

Yash

是的。强化学习和基础设施方面最难解决的问题是将总拥有成本(TCO)降到可以进一步产品化的程度。我的联合创始人 Lyndon,我们的首席架构师,是这方面的奇才。他是我见过的最聪明的人之一。你会惊讶于 GPU 的利用率有多低。如果你在全球范围内聚合资源,可以从现有硬件中挤出更多性能,从而大幅降低成本。你可以投资于更数据高效的算法,这样训练时间更短,算力消耗更少。在系统侧和算法侧,是目前在可用数据等方面最能挤出效益的地方。有合成数据生成技术,可以取一定量的 token 并将其扩展成更大的数据集,从中提取学习信号。但坦率地说,大部分成本优化和削减将发生在算法侧和系统侧。

Yeah. The hardest problems to solve with RL and infra are to bring the TCO down to the point where you can productize it even more. My co-founder Lyndon, who's our chief architect, is a wizard with this. He's one of the smartest people I've met. You'd be surprised how underutilized GPUs are. If you aggregate things globally, there is a lot more that can be squeezed out of the available hardware to lower these costs quite drastically. You can invest a lot in algorithms that are more data efficient, so you need to train for less long and use less compute. The system side and the algorithm side are where you can squeeze the most right now in terms of data available. There are synthetic data generation techniques to take some amount of tokens and explode it into something much larger and pull learning signal out of that. But most of the cost optimization and cutting down costs is going to happen on the algo side and the system side.

行业差异与模型整合 Every Company Having Its Own Model

Host

Satia 公开表示,世界上的模型数量应该和公司数量一样多。这显然对你们来说是一个很好的卖点。这个愿景引起你的共鸣吗?你有多相信这个愿景,即每家公司都是独特的雪花,拥有自己独特的判断力,需要自己的模型?

Satia came out and said that there should be as many models in the world as there are firms in the world. Which obviously was a good selling point for you all. Is that a vision that resonates with you? How much do you believe in this vision that every single company is its own special snowflake with its own unique judgment that requires its own model?

Yash

是的,我认为原则上,这是否意味着每家公司都需要预训练或中期训练自己的模型,我并不一定认为如此。但我确实认为每家公司内部都有值得捕捉的闭环反馈系统。现在有很多低垂的果实。所以人们从 harness 优化和上下文优化开始。但不同公司是独特的,它们运营方式不同,有不同判断结构,并不相同,我们从根本上同意这一点,这也是我们整个公司的核心论点。如果每个人都使用相同的模型,伟大的模型为所有人设定了下限,但如何优化这些模型并构建真正出色的 AI 系统,那将决定上限。所以人们可以构建的产品确实存在差异,既体现在他们所做的 harness 和上下文工程上,也体现在他们放在这些产品背后的模型上。

Yeah, I think in principle, whether this means that every company needs to pre-train a model or mid-train a model of their own, I don't necessarily think that's the case. But I do think there are closed loop feedback systems inside of each company that are worth capturing. Right now there's just a lot of low-hanging fruit. So people are starting there with harness optimization and context optimization. But the idea that different companies are unique and they operate differently and they have different judgment structures and are not the same, we fundamentally agree with, and that's kind of the thesis of our whole company. If everybody's using the same model, great models sort of set the floor for everybody, but then how you optimize those models and build really amazing AI systems, that's what's going to set the ceiling. So there actually is a difference in the products that people can build, both on the harness and context engineering that they do, but also the models that they put behind those products.

非可验证领域的RL Industry Differences and Model Consolidation

Host

当然,你把它推向一个极端,比如制药公司。制药公司的数据彼此差异很大。这很合理——它们专注于一个疾病领域,生成大量实验数据,所以当然应该有自己的模型。如果拿银行来说,一家银行和另一家银行有多大不同?当然,可能有一些不同的判断决策,但我想知道,如果一个前沿模型正在对大量非常好的通用金融数据进行后训练,那么有多少独特的东西是……

Certainly, you take it to one extreme, like a pharma company. Pharma companies have such different data from each other. It makes so much sense—they focus on one disease area, generate a lot of experimental data, so of course they should have their own model. If you take banks, how different is one bank from another bank? Certainly maybe there's little judgment calls that are different, but I wonder if a frontier lab is post-training on a bunch of really good general finance data, how much uniquely unique things is a...

Yash

是的,我认为市场上肯定有一些东西会从整合中受益。这就是为什么会有 rollup,对吧?你拿会计师事务所之类的来说——它们由于业务性质而分散。你有一个会计师和一个客户,他们是一对一匹配之类的,所以把会计师工作中的琐碎劳动拿出来放到模型中是有道理的。这类事情,我认为你会看到不同公司运营方式的整合。但像银行,我认为每家公司做出的一系列决策和权衡都与其他公司不同。

Yeah, I think there are certainly things in the market that will benefit from consolidation. That's why you have rollups, right? You take accounting firms or something like that—they're fragmented just by the nature of the business. You have an accountant and a client, and they're one-to-one match or something like that, so it makes sense to take the menial labor out of doing the work of an accountant and put that into a model. Those types of things, I think you will see consolidation when it comes to how different companies operate. Like banks, though, I think there is a whole litany of decisions and trade-offs that each company makes that is different than the other.

Host

是的。

Yeah.

Yash

所以它们不太像商品,你可以用这个或那个。这就是为什么银行的表现分布很广,有的做得好,有的做得差。我认为实际上有竞争的空间。

So they're less of a commodity where you can use one or the other. That's why you have such a wide distribution of banks that are doing well and banks that are doing poorly. I think there actually is room to compete.

Host

看到这个会很有趣,对吧?本质上,我们将通过拥有自己的模型相对于使用现成的平均数据标注器有多大益处,来了解这些行业和公司实际上有多不同,或者做事方式有多不同。

It'll be fascinating to see, right? Essentially we'll learn the extent to which a lot of these industries and companies are actually particularly different or have different ways of doing things by the extent to which having your own model is actually beneficial relative to using whatever the off-the-shelf average data labelers.

Yash

完全同意。我认为这实际上体现的方式并不总是能力,对吧?我们对正在构建的东西感到非常兴奋的一个原因是我们正处于供应紧缩中。算力有限。每个人都感受到了这一点。所以你能从模型中挤出的越多,从你实际可用的算力中挤出的越多,你能做的事情就越多。所以基本上这就是为什么我们认为拥有优化模型并使其更高效的基础设施,即使你是在推动帕累托曲线外移,而不一定是在推动能力前沿,这种差异也是一种机会或优势。如果我能比竞争对手便宜 10 倍地做某事,那就是差异化,那就是我可以在竞争格局中战胜他们的东西。

Totally. And I think one of the ways that this actually manifests is not always capabilities, right? One of the reasons why we feel very excited about what we're building right now is we are in a supply crunch. There is limited compute. Everyone's feeling this. So the more you can squeeze out of your models and the more you can squeeze out of the compute that you actually have available to you, the more stuff you can do. So basically that's why we think having the infrastructure to optimize models and make them more efficient, even if you are pushing the Pareto curve out and you're not necessarily pushing the frontier on capabilities, that delta is some sort of opportunity or advantage. If I can do something 10x cheaper than my competitor, that is differentiation, and that is something that I can actually win against them on the competitive landscape.

RL与可验证领域 RL in Non-Verifiable Domains

Host

我想也许退一步,谈谈那些正处于决定是否要开始构建自己模型边缘的人。显然你谈过应该先尝试很多 harness 优化之类的东西,但你知道,我认为很多人正在试图弄清楚强化学习在这些不可验证的领域效果如何。你如何描述我们今天所处的位置?

I guess maybe zooming out and talking about people that are at the precipice of deciding whether they should embark on their own models. Obviously you've talked about you should try a lot of the harness optimization stuff first and a lot of other things, but as you know, I think a lot of folks are trying to figure out how well does RL work in these non-verifiable domains. How do you kind of characterize where we are today?

Yash

是的。

Yeah.

合格业务与推理焦点 RL and Verifiable Domains

Yash

所以我想说的是,强化学习起步时,或者说这些项目起步时,内部训练的模型都是在高度可验证的领域,比如数学和编程。部分是因为它们高度可验证,部分是因为实验室里每个人都喜欢数学、喜欢代码,但更多是前者。我认为,如果你把非可验证领域转化成某种代理可验证领域,在上面爬山是出奇地容易。所以基于评分标准的强化学习之类的方法实际上效果很好。所以我想说,我们目前的状态是,如果它不是只有一个确定答案的东西,基本上给它一个专家答案并据此评分,就是一个相当好的代理。所以这实际上就是为什么不同公司有不同的专家,所以针对两个不同的群体和两家不同的公司进行优化,实际上会给你非常不同的模型。

So what I would say is that RL started, or the way these projects started, and the internal models that were being trained were on highly verifiable domains like math and coding. Partially because they're highly verifiable, partially because everyone at the labs loves math, loves code, but it was more of the former. I think it is surprisingly easy to hill climb on non-verifiable domains if you turn them into some proxy verifiable domain. So rubric-based RL and this sort of stuff actually works quite well. So I would say the state where we're at right now is that if it's not something that has one definitive answer, basically giving it an expert answer and grading against that serves as quite a good proxy. So that's actually why different companies have different experts, and so optimizing to two different groups and two different companies can actually give you very different models.

Host

所以是为了让你所有前员工都不参与数据标注?

So to keep all your ex-employees off of the data labeling?

Yash

没错。是的。不,我认为实际上我们通过强化学习得到的就是一台爬山机器。最难的部分实际上是定义要爬的山,这就是为什么我认为评估对公司来说是一件非常重要的事情,既要专注于构建,也要相当仔细地保护。如果你考虑竞争之类的事情,如果你看每一个公开基准,它都会被刷爆,对吧?所以如果你有非常特定于你业务运作方式的东西,并且你想让你的模型在这方面非常擅长,那么不告诉别人好坏的标准实际上对你有利,这样别人就无法让他们的东西在这方面变得非常擅长。所以这就是为什么我认为实际上保留你的评估和你的员工等等是非常重要的。就像你的员工不是可互换的。你不会愿意他们去另一家公司在那里工作。这些模型也是一样。

Exactly. Yeah. Yeah. No, I think it's actually what we have with RL is we essentially have a hill climbing machine. The hardest part is actually defining the hill to climb, which is why eval I think is a really important thing for companies to a) focus on building and b) actually safeguard pretty carefully. If you are thinking about competition and stuff like that, if you look at every public benchmark, it's going to get benchmaxed, right? So if you have things that are very particular to how your business operates and you want to make your models really good at that, it's actually to your advantage not to tell everybody else what good and bad looks like so people can go and make their things really good at that. So that's why I think it is really important to actually keep your evals and your employees and all that sort of stuff. Like you wouldn't—your employees are not fungible. You would not be comfortable with them going to another company and doing work there. Same thing with these models.

Host

我觉得如果你是一家领先的应用公司,就会有一种张力,显然你想扩大你能提供给客户的能力范围,所以去实验室告诉他们你的评估是有帮助的,而另一方面,很多这些来之不易的洞见就像——

I feel like there's this tension if you're a leading app company where obviously you want to just expand the set of capabilities you can deliver to your customers, and so going to labs and telling them your eval is helpful, and then the other side it's like a lot of these hard-earned insights like—

Yash

完全正确,这是一个艰难的处境,这就是为什么我认为投资于开放模型基础设施、投资于多模型的未来是每个公司绝对都应该做的事情,因为是的,你有点进退两难。

Totally, and it's a tough position, which is why I think investing in open model infrastructure, investing in a multi-model future is something absolutely every company should be doing, because yeah, you're kind of caught between a rock and a hard place.

Host

是的。

Yeah.

Yash

是的。

Yeah.

推理业务与市场动态 Qualifying Business and Inference Focus

Host

我想我很想转到 Black Compute 这家公司。我在想——我相信肯定有很多企业会觉得,能和你们合作、让你们卷起袖子帮忙解决各种问题,那真是太好了。在这个阶段,你们如何筛选业务,确定和谁合作、不和谁合作?

I guess I would love to shift to Black Compute, the company. And I'm wondering—I'm sure that there is no shortage of enterprises that would be like, it would be pretty sweet to work with you guys and get you rolling up your sleeves and helping whatever the set of problems is. How do you qualify business and figure out who you work with and who you don't at this stage?

Yash

好问题。所以我们基本上把后训练以及实际训练这些定制模型视为差异化因素,以及让业务具有粘性、并在现成模型之上提供更多价值的部分。但实际上,推理才是大部分价值捕获发生的地方。对吧?如果你看花在模型上的钱,绝大多数都用于实际服务和使用的模型,而不是训练模型本身。所以当我们与客户合作时,我们通常会追求非常高价值的用例,要么是因为我们将提供的模型能力提升极其有价值——比如你之前描述的那些公司,像制药公司、网络安全公司或芯片公司之类的。另一种是推理工作负载真的很大,对吧?所以基本上是将某种效率或优化收益或质量改进分摊到非常大的用户群或数十亿、数万亿个 token 上。这实际上也非常有价值。所以我们实际上寻找的是大型推理工作负载。这些通常是我们最好的客户,或者是非常分布外的数据,人们想要训练那些只对他们独特的东西。

Great question. So we basically see post-training and actually training these custom models as the differentiation and the part that makes the business sticky and actually provides more value on top of just the off-the-shelf models. But really inference is actually where most of the value capture happens. Right? If you look at the dollar spent on models, the vast majority goes to actually serving and using the models than training the models by themselves. So when we work with customers, we typically go after very high value use cases, and that's either because the capability increase in the model that we're going to provide is extremely valuable—so some of these companies like that you were describing before, like a pharma company or cyber security company or chip company or something like that. The other is if the inference workload is really big, right? So basically amortizing some sort of efficiency or optimization gain or quality improvement distributed over a very large user base or billions or trillions of tokens. That's actually super valuable as well. So we actually look for large inference workloads. These are typically the best customers for us, or very out-of-distribution data where people are wanting to train on things that are kind of only unique to them.

训练-推理协同设计 Inference Business and Market Dynamics

Host

你提到了推理方面。显然,如今任何基础设施公司都感觉,拥有推理业务是相当不错的。我觉得对此有不同的看法。比如,构建一个真正好的推理引擎,这个产品有多难?

You kind of mentioned the inference side of things. Obviously it feels like the temptation for any infrastructure company these days is like, having an inference business is pretty nice. I feel like there's different schools of thought on this. Like how difficult of a product is that to build, like a really good inference engine?

Yash

所以我想,也许退一步说,为什么我们认为——首先,我们认为推理将成为世界上最大的市场之一,如果不是最大的话。很多人都说过这一点;这不是什么独特的见解。我的意思是,如果你相信 AI 会做很多——

So I think maybe taking a step back, why we think—so first of all, we think inference is going to be one of the biggest markets in the world, if not the biggest. Tons of people have said this; it's not a very unique take. I mean, it's kind of obvious if you believe in AI doing a lot of—

Host

如果没有的话,就不会有很多训练资金——

Not going to get a lot of training dollars if there's not—

Yash

没错。所以我们的观点是,最好的后训练实际上会赢得推理。所以如果你基本上——我们看到的趋势是,公司会从前沿模型开始,他们会做大量实验,然后他们的工作负载会变得成熟,他们会扩展那个工作负载,然后他们会考虑开放模型,以及如何成本优化、性能优化,基本上把帕累托曲线往外推。有趣的是,最大规模的工作负载,也就是你花钱最多的最大工作负载,实际上是你最想先去进行后训练的地方,对吧?因为那实际上是最能获得收益的地方,而且在所有其他用例之间有点像幂律分布。有趣的是,你可以同时协同优化你的推理和训练。所以基本上,我们训练模型的方式实际上会影响我们如何为客户设置推理部署。

Exactly. And so what our view is is that the most post-training actually wins inference. So if you are basically—what the trend we've been seeing is that companies will start off with frontier models, they'll do a lot of experimentation, then their workload will become mature, they'll be scaling that workload, then they'll think about open models and how to sort of cost optimize, perform, basically push the Pareto curve out. And the interesting thing is that the most at-scale workloads, so the largest workloads where you're spending the most money, is actually where you're going to want to go and post-train first, right? Because that's actually where the most gains are to be had, and it's kind of power law between all the other use cases. And the interesting thing you can do is you can both co-optimize your inference and your training. So basically the way we train models actually influences how we go and set up inference deployments for customers.

横向扩展与倒金字塔 Training-Inference Co-design

Yash

如果它是一个重度使用工具的模型,我们训练它做大量工具调用并行化之类的事情。

If it's a heavy tool-using model and we train it to do a bunch of tool call parallelization, that kind of stuff.

Host

你能多讲讲吗?这在实践中到底意味着什么?这如何影响你做出的推理选择?

Can you say more about that? Like what does that actually mean in practice? How does that influence the inference choices you make?

Yash

是的,是的,这是个好问题。基本上,如果我们做一个长时程的工具使用智能体,我们设置推理部署或推理引擎的方式与做重度解码型工作负载不同,对吧?所以你可以设置这些解耦的配置,使用多种不同类型的芯片分别进行预填充和解码。因此,对模型训练方式的影响实际上可以影响我们如何为超大规模工作负载设置推理部署。

Yeah, yeah, so it's a great question. So basically, if we are doing a long horizon tool-using agent, the way we set up the inference deployment or the inference engine is different than if we were doing a heavy decode-heavy workload, right? So you can set up these disaggregated setups where you're using multiple different types of chips for prefill and decode. And so, having influence on how the models are trained actually can influence how we want to set up the inference deployments for very large scale workloads.

Host

这太有趣了,我真的很想深入探讨,因为我觉得这个问题已经存在一段时间了:创建模型在多大程度上能帮助你更好地运行推理。它最初出现在很多早期开源模型的宣传中,对吧?当时的想法是,嘿,我们训练这些模型,所以我们能在它们上更有效地运行推理。但至少早期实际的发展是,像 Fireworks 这样的公司会说,谢谢你训练了一个很棒的开源模型,现在我们将是最有效率地运行它的人。

It's such an interesting, I'm really curious to dig into this because I feel like there's been this question for a while of how much does creating the model help you run better inference. And it originally showed up in a lot of the original open source model pitches, right? Where the idea was, hey, we train these models so we're able to run inference way more effectively on them. And then the way it played out, at least in the early days, was actually someone like Fireworks would be like, thank you for training a great open source model, now we'll be the ones to run that as efficiently as possible.

Yash

是的。所以我认为发生的情况是,基本上,你用于训练和推理的内核,显然希望它们之间匹配。我们在训练方面做了大量工作来优化训练速度、训练效率,而很多工作可以转化到推理方面。我还想指出另一件非常有趣的事情,这和我们刚才讨论的有点无关。我认为经常被忽视的一点是,你可以做一堆推理优化之类的事情,如果你想将推理费用削减 10%。你可以做一堆推理优化来降低成本 10%。或者我认为这些模型在很多事情上实际上相当 token 低效。所以我们一直在和客户说的是,嘿,我们可以通过让你的模型 token 效率提高 10% 来削减你的 token 费用,同时在你关心的某种 EVL 上保持相同的性能。所以我认为我们能够告诉客户的一件事是,嘿,这些事情是同一枚硬币的两面。如果你和我们合作,你实际上可以同时影响两者,这就是为什么我们试图成为训练和推理的集中式提供商。

Yeah. So I think what's happened is, basically, the kernels that you use for training and inference, you obviously want matches between them. We do a lot of work on the training side to optimize how fast we're able to train these things, how efficiently we're able to train these things, and a lot of that work translates over to the inference side. I think also one other really interesting thing to point out, this is kind of unrelated to what we were just talking about. One thing I think is really often overlooked is you can do a bunch of inference optimization and things like that if you want to cut your inference bill by 10%. You could do a bunch of inference optimization to cut cost by 10%. Or I think these models are actually quite token inefficient for a lot of things. So what we've been doing with customers is say, hey, we can actually cut your token bill by making your model 10% more token efficient, preserving the same performance against some sort of EVL that you care about. And so I think one of the things that we've been able to tell customers is, hey, these things are two sides of the same coin. You can actually influence both of them if you do it with us, which is why we're trying to be the centralized provider for both training and inference.

Host

自己进行训练能在多大程度上让你对推理工作负载有更深入的了解,这将很有趣。

It'll be interesting to see the extent to which doing the training yourself gives you even more insight into the inference workloads.

Yash

回到你最初关于推理产品以及它们多容易启动的问题。我认为我们看到的主要是要有容量,对吧?所以你需要实际的 GPU 来服务模型等等。像 VLM 和 SG Lang 这样的开源项目是很好的起点。事实上,我认为几乎所有的推理云都已经淘汰了很多他们定制的推理引擎,转而采用这些推理引擎之一,进行工作负载调优和性能调优,也许还会加入几个自定义内核。所以我认为 Dylan Patel 在 Dwarkesh 的播客上,他实际上谈到了这有多容易,基本上,如果你有一些 GPU,你放一个 VLM 上去,然后把它扔到 open router 上。你就有了一种生产推理产品。现在,我不想过于轻描淡写。我认为需要大量的优化和艰苦的工作才能使这些推理部署更高效、更快的 token 速度,尤其是在你达到规模时,这很重要,对吧?因为成千上万芯片上节省 1% 实际上会累积起来。但就实际启动生产质量的模型服务而言,在训练和推理中,这绝对是更容易的事情。

And I think, back to your original question about inference offerings and how easy they are to spin up. I think what we've seen is the main thing to have is capacity, right? So you need the actual GPUs to be able to serve the models and things like that. And the open-source projects like VLM and SG Lang are great starting points. In fact, I think almost all of the inference clouds have sunset a lot of their custom inference engines in preference for taking one of these inference engines and doing the workload tuning and performance tuning and maybe having a couple custom kernels in there. So I think Dylan Patel was on like Dwarkesh's podcast and he was actually talking about how easy it is to basically, if you have some GPUs, you put like a VLM on it and you throw it up on an open router. You kind of have a production inference offering. Now, I don't want to trivialize it too much. I think there is a lot of optimization and hard work that goes into making these inference deployments a lot more efficient, faster token speeds, and that matters especially as you hit scale, right? Because 1% of savings across tons and tons of chips actually adds up. But in terms of actually getting production quality model serving going, it's definitely the easier thing to do out of training and inference.

基础设施战略与焦点 Horizontal Expansion and Inverted Pyramid

Host

是的,是的。这很有趣。我觉得作为基础设施 CEO,你一定感觉到有很多公司从不同的地方起步,并且有一种愿望或信念,即随着时间的推移,它们中的许多会变得相当横向,并以各种方式相互侵入。这在业务过程中是如何演变的,你今天如何看待它?

Yeah, yeah. It's interesting. I feel like as an infrastructure CEO, you must feel like there's a bunch of companies starting in different places and there's this desire or this kind of belief that over time a lot of them will go pretty horizontal and kind of intrude on each other in all sorts of ways. How has that evolved over the course of the business and how do you think about it today?

Yash

是的,这是个好问题。所以我们的大规模愿景是,我们认为需要构建一个新的 AI 超大规模云服务商。就像在云时代,你有 AWS、GCP 和 Azure,他们利用 CPU 算力,在上面构建了非常棒的灵活算力产品,比如网络、存储、数据库、数据湖以及用于构建应用程序的东西等等。现在你有了这个新的计算基础,GPU。你有了这个奇怪的东西,即称为模型的一组权重和参数。事实证明,你可以用推理引擎来服务它们,这是第一个在 AI 软件栈中实现真正粘性的软件,你可以提供这些有价值的智能 token。我们的观点是,在那个软件层中,除了推理之外,实际上还有大量的东西。我们的大赌注是,嘿,后训练是栈中被严重低估的部分。我如何将 GPU 算力转化为更智能的 token?那就是训练。我如何服务这些 token?那就是推理。我如何选择正确的 token 来服务?那就是路由。但还有大量的东西,比如安全监控和身份智能体沙箱等等。所以当我们思考我们正在构建的东西时,我们实际上希望非常横向,但我们实际上想从倒金字塔的角度来思考。在金字塔底部是很少有人能做的事情,在金字塔顶部是很多人能做的事情。所以在金字塔的最底部,我们认为那是模型训练。然后在它之上是推理。

Yeah, it's a great question. And so our large scale vision, so what we're actually building is we think there's a new AI hyperscaler to be built. Like in the cloud era you had AWS and GCP and Azure, they took CPU compute, they built really awesome flexible compute products on top of it, like for networking and storage and databases and data lakes and stuff for building applications, all that. Now you have this new computer substrate, GPUs. You have this weird thing which are this collection of weights and parameters called models. It turns out you can serve them with an inference engine, which is the first piece of software that has achieved real stickiness inside of the AI software stack, and you can serve these valuable intelligent tokens. Our view is that there's a ton of stuff in that software layer actually besides just inference. Our big bet was, hey, post-training is a massively underlooked part of the stack. How do I take GPU compute and turn that into more intelligent tokens? That's training. How do I serve those tokens? That's inference. How do I pick the right tokens to serve? That's routing. But there's also a ton of stuff like security monitoring and identity agent sandboxing, all this sort of stuff. So when we think about what we're building, we actually want to be very horizontal, but we actually want to start, we think of things as an inverted pyramid. There's things that very few people can do at the bottom of the pyramid and things a lot of people can do at the top of the pyramid. So at the very bottom of the pyramid, that's where we think model training is. And then above that is inference.

构建RL环境 Infrastructure Strategy and Focus

Yash

在那之上是路由。再之上是上下文和 harness 以及所有这类东西。所以我们的策略一直聚焦在那些极具杠杆效应、极高 ROI、非常难做的事情上,你需要构建超级稳健的基础设施。你需要一支真正拥有并运营的 GPU 集群,诸如此类。所以作为一家基础设施公司,我们的目标是随着成熟逐步拓宽产品线,你知道我们现在才成立 16 个月。所以第一步是从训练走向推理。但我们的路线图上有很多令人兴奋的事情,以及我们想要去构建的东西。

Above that is routing. Above that is context and harness and all that sort of stuff. And so our strategy has really been focused on the things that are extremely high leverage and extremely high ROI that are very hard to do that you need to build super robust infrastructure. You need a GPU fleet that you actually own and operate and all this sort of stuff. And so our goal is to as an infrastructure company widen the offering as we mature and you know we're 16 months old right now. So we're the first step was going from training into inference. But we have a lot of exciting things on the road map and things that we want to go and build.

Host

有趣的是,我想象当你们 80% 的潜在客户来找你时,你会觉得这大概可以通过一个好的 harness 和一些数据工作之类来解决。所以你们在 harness 优化方面最终会做些什么?我想现在很多人都在思考这个问题。

What's so interesting is that I imagine when like 80% of your pipeline approaches you you're like this could probably be solved by like you know a good harness and like some data work or something. And so what you know uh what do you end up doing on like the harness optimization side like I imagine a lot of folks are thinking through this right now.

Yash

好问题。我认为我们现在真的在追求专注。所以我们实际上,你知道,我们合作的客户都有成熟的 AI 产品。他们已经建立了出色的产品团队,并且正在投资创建非常棒的 harness。所以我们目前的运营模式实际上是我们不专注于 harness 层或上下文层。我们专注于那里的模型,而且我们不一定做从零到一的智能体构建,我们优化成熟产品背后的模型。

Great question. I think we are really going for focus right now. So we actually you know the customers we work with have mature AI products. They have already like they've built amazing product teams and they are investing in creating really great harnesses. And so our current operating model has actually been we're not focused on the harness layer of the context layer. We're focused on the model there and um you know we we don't necessarily do the zero to one agent builds um we we optimize the models behind mature products.

竞争动态与模型访问 Building RL Environments

Host

有道理。我想就今天构建强化学习环境而言,也许我们的听众很想听听实际上,哪怕只是一个说明性的例子,构建其中一个环境到底涉及什么,以及你如何看待,我确信有 500 个烦人的部分,那些需要改进或你希望改进的部分。

Makes sense. And I guess in terms of like the building the RL environments today maybe like I'm sure our listeners would love to hear like what is actually you even just take an illustrative example like what is actually involved in like building one of these and you know how do you think about like I'm sure there's 500 annoying parts of it like the parts that need to get better to or that you hope get better.

Yash

好消息是,就像我之前说的,强化学习就像爬山机,对吧?所以,实际上你构建训练数据的方式和你构建评估的方式是一样的。所以,希望这能帮助人们概念化,构建优秀评估需要什么,就是,嘿,你必须要有正确的任务,你必须要有正确的环境,你必须要有正确的验证器,对吧?这些是进行强化学习时的三个主要方面。问题是什么?模型可以访问的工具和服务的范围是什么?然后我如何判断它是好是坏?举个例子,我们做了这个 Harvey review table 的自定义模型,我们训练的。我们发布了一个案例研究。但为此,我们基本上与 Harvey 团队合作,能够收集一堆专家评分标准和答案,还用合成数据补充,并让模型访问一些工具。能够说嘿,这是你应该去提取所有信息的任务,用真正的强化学习根据专家答案评分,这就是我们如何创建任务,环境层是模型在文档库中可以访问的工具,然后验证器是根据专家答案评分。

The nice thing is like I said before, RL is just like hill climbing machine, right? So, um it actually is the same way you're going to go and build your training data is actually the same way you're building your eval. Um so, uh hopefully, you know, that can help folks like kind of conceptualize, you know, what is go what goes into building great eval is like, hey, you have to have the right uh you have to have the right task um and you have to have the right environment and you have to have the right verifier, right? Those are kind of the three main things when you're going to do RL. What is the problem? What is the uh scope of tools and services that the model has access to? And then how am I going to judge whether it's good or bad? So take for example um you know we did this uh um uh uh Harvey review table um custom model that we we train. We put out a case study about it. But for that right, we basically um had worked with uh you know Harvey Harvey the Harvey team um were able to cor collect a bunch of expert uh rubrics and answers also supplemented this with like synthetic data and gave the model access to some tools. um were able to say hey here's the task that you're supposed to go you know extract all this information out uh graded it against the expert answers with uh you know real request RL and that's how we were able to create the task the environment tier was the tools that the model has access to in the document base and then the verifier was the grading against the expert answers

Host

最难的部分是不是就是让人们创建这些评分标准,或者创建你知道的现实文档,这些会存在于这些东西中?

Is the hardest part of that just like actually getting people to create these rubrics or like you know create you know what are realistic documents that would live in these things

Yash

是的,所以我们与所有客户紧密合作,因为他们是领域的专家,对吧?我们真正带来的是基础设施、训练团队和推理,你知道,生产服务等等。所以这非常非常协作。我们经常可以在合成数据方面做很多工作,如果是一些,比如说你可以拿某种文档反向翻译成问答对,例如,这是一个非常具体的例子,我们可以合成地做这些事情,但当涉及真正的专家和真正的律师时,我们实际上想直接与团队合作。

Yeah so we we we work really closely with um all of our customers ers um because they are the experts in the domain right uh we we really come in with the infrastructure and the training team and the inference uh you know production serving and stuff so you know this was like pretty pretty pretty collaborative we oftent times can do a lot of work on the synthetic side um if it's like uh some take say you can take some sort of document and back translate it into uh you know question answer pairs for example that's very concrete example um we can do that stuff synthetically, but when it involves like real experts and real lawyers, we actually want to work with the teams directly.

快问快答 Competitive Dynamics and Model Access

Host

是的。我确信,你知道,整个论点似乎是这类工作会比律师数据标注公司得到的质量高得多。我想这可能已经是遥远的记忆了,但比如 OpenAI 关闭 Cursor,你预期这类事情会更多发生吗?比如 Anthropic 和 Windsor?是的。你预期这类事情会更多发生还是?

Yeah. And I'm sure like you know the the entire thesis of this seems to be like that type of work is going to be way higher quality than like whatever lawyers data labeling company gets on and tries to like you know create I think you know probably distant memory at this point but like OpenAI shutting off cursor like do you expect that kind of stuff windf right anthropic with Windsor? Yeah. Do you expect that kind of stuff to happen more or like

Yash

嗯,我的意思是它发生了两次。我认为竞争动态是真实的,他们的激励也是真实的。所以特别是当公司开始更多地进入应用层,并与他们模型的最大消费者构建自己的竞争产品时。如果发生更多,我不会感到惊讶。

um I mean it happened twice. I think competitive dynamics are real and like their incentives are real. So especially as companies start to move more into the application layer and build their own competitive products with their largest consumers of their models. Um I wouldn't be surprised if it happened more.

成本优化与模型选择 Quick Fire Round

Host

我总是喜欢有一个快速问答环节,我基本上以非结构化的方式塞进所有我之前没有涉及到的内容。所以也许开始吧。

I always like to have a quick fire round where I basically shove in like everything that I uh in an unstructured way that I haven't uh gotten to before. And so maybe to to to start um

Host

强化学习会泛化吗?

is RL going to generalize

Yash

在某些方面?在某些方面它确实在泛化,比如模型的能力,你知道,进行错误恢复或使用工具和子智能体之类的事情。但我认为如果你实际经验性地运行这些实验,你可以在数学数据集上训练,看看有多少转移到法律领域或其他类似的东西。转移并不真正存在。当然不像预训练那样的泛化程度。所以我认为它会渗透到相邻领域,就像智能是某种球体或圆形。它有锯齿状的点。点更接近的地方,你会看到更多的泛化。但当然没有像预训练那样的泛化。

in some ways? In some ways it is generalizing right like uh models ability models abilities to u you know do error recovery or um you know use tools and sub aents and things like that. Uh but I think like if you actually empirically you can like run these experiments, you can train on a math data set and see how much that translates to uh something in the legal domain or something like that. The transfer is is not really there. Certainly not to the extent of like uh the generalization that pre-training has. Um so I think it will it will bleed into adjacent areas like intelligence is some sort of sphere um or circle. uh it's got jagged points on it. The points that are closer together, you'll see more generalization there. Um but certainly does not have the same generalization as like pre-training or something like that.

Host

是的。

Yeah.

Host

嗯,在过去一年里,你在 AI 方面改变了什么想法?

Um what have you changed your mind on in AI in the last year?

Yash

是的。我的意思是,我想有一件事,也许我没有改变想法,但我没有完全意识到 Jeban 悖论有多真实。

Yeah. I mean, I guess one thing that I didn't maybe I didn't change my mind on it, but I didn't fully appreciate is how real uh Jeban's paradox is.

Host

有没有什么特别的事情你看到了?

Is there like a particular thing you saw that

Yash

是的。

Yeah.

算力预算与商业策略 Cost Optimization and Model Selection

Yash

一个有趣的现象是,如果你看一些 OpenRouter 的统计数据,每当某个模型降价,使用量就会激增。所以我真的认为我们正在走向一个极度成本优化的世界。我们感受到了压力,就像现在每个人都在感受压力一样,要尽可能提高效率,降低 AI 账单。我没想到成本优化会成为一个如此重要的驱动因素,甚至超过能力。我们已经到了一个点,我敢说大多数开源模型几乎能做你想让它们做的一切。当然,还有很多前沿的东西,我不想一概而论。前沿模型很棒,但我没想到人们会如此关心真正的成本优化,以及为合适的任务使用合适的模型。

One interesting thing is that if you look at some of the OpenRouter statistics, whenever there's a price cut for a model, usage just spikes. So I really think we're moving to a world where things are extremely cost optimized. We're feeling the pressure, like everyone's feeling the pressure right now, to make things as efficient as possible, lower the AI bill. I don't think I fully appreciated how much of a driving factor that would be over something like capabilities. We are reaching a point where I would say most of the open source models can kind of do everything that you might want them to do. There are obviously a lot of frontier things, and I don't want that to be a blanket statement. Frontier models are amazing, but I didn't expect how much people would care about really cost optimizing and using the right model for the right task.

Host

是的。这很有趣,因为显然对于某些任务,人们使用前沿模型是因为能获得很多价值。但如果有一种方法能在用例上便宜得多……

Yeah. Which is interesting because obviously for some of these tasks, people were using frontier models because you got a lot of value. But if there is a way to do it much cheaper on the use case...

Yash

是的,一些边际价值回报。到了某个时候,人们会进行实验,尝试一堆东西,然后整合到正确的模型上。

Yeah, some marginal return of value. And at some point people are going to do the experimentation, try a bunch of things, and then consolidate on the right model.

Host

当像 Jev 这样的东西出现时,在 Applied Compute 会怎样?你们是怎么考虑的?

When something like Jev comes out, how does that work at Applied Compute? How have you guys thought about that?

Yash

我认为这是一个非常酷的模型。我基本上认为可能会有很多类似的东西出现。它开启了一整套新的用例。我的意思是,名字里就体现了——这就是为什么他们叫它 Jev,对吧?便宜、快速、为特定任务集设计。这就是我们在 Applied Compute 构建许多东西时的思考方式。所以像 Jev 这样的模型,我们正在研究:如何将其集成到训练过程中,作为对 rollout 进行大规模十亿 token 级别的分类,或者我们以前因为成本过高而无法做的事情。所以是的,当这样的东西出现时,我认为非常令人兴奋,因为它基本上扩展了我们可以构建的东西的数量。这有点像某种矩阵,对吧?所以它是高度倍增的——这些东西实际上会叠加。Jev 与其说是替代品,不如说是我们在训练和推理方面以前无法做的事情的推动者,比如用于可观测性之类的。

I think it's a super cool model. I basically think there are probably going to be a bunch of things like that coming out. It enables a whole set of new use cases. I mean, it's in the name—that's why they called it Jev, right? Things that are cheap, fast, designed for a specific set of tasks. That's how we've been thinking about a lot of things that we're building at Applied Compute. So models like Jev, we're working on: what does this look like to integrate into the training process as sort of massive billion-token scale classification on rollouts, or things that we couldn't do before that were cost prohibitive. So yeah, when things like that come out, I think it's super exciting because it basically just expands the amount of things that we can go and build. And it's some sort of matrix, right? So it's highly multiplicative—these things actually end up stacking. Jev is much less of a replacement for stuff rather than an enabler for things that we weren't able to do on both training and on inference, say for observability or something like that.

Host

是的。只是因为在那样的成本水平上,你……

Yeah. Just because at that kind of cost level you...

Yash

是的,在那个成本或价格点上,它基本上开启了一整套新的事情,我们可以实际去做以前无法做的事情,即使使用最小最便宜的 Quen 模型之类的也不行。

Yeah, at that cost or that price point, it basically enables a whole new set of things that we can actually do that we couldn't have done before, even with like the smallest cheapest Quen model or something like that.

Host

是的。你有没有一个你见过的生动的例子?

Yeah. Is there like a visceral example for you of that that you've seen?

Yash

所以我们实际上在一篇博客文章中发布了这个。我们基本上使用 Jev 来做大规模的 trace 分析,用于在线训练。

So we actually put this out in a blog post. We basically used Jev to do essentially massive trace analysis for online training.

AI编程与工程建议 Budgeting Compute and Business Strategy

Host

我想你正处于指数级增长的斜坡上。那么你到底如何预算你想要的算力,因为你可能正在进入一些……

I imagine you're on this exponential slope of growth. So how in the world do you budget for the amount of compute you want, because you're probably entering into some of these...

Yash

所以我认为我们的思考方式是,我们试图真正打造一个非常有粘性、差异化的业务。所以我们不会追逐每一个可能的机会——比如,哦,有人需要推理之类的。我们非常相信这个论点:人们会使用定制模型,他们会想要投资于我们用于后训练的基础设施。所以我们实际上不是看所有事情的总需求,而是看:我们可以与哪些最具战略意义的公司合作,我们可以与他们一起成长,他们也是合适的人选,基本上能够证明这个论点:定制模型比某种现成的模型更有粘性、更有价值,而且 ROI 是持久的。

So I think the way we think about it is we're trying to actually make a very sticky, differentiated business. So we aren't going after every single opportunity that we could—like, oh, someone needs inference or something like that. We are very convinced of this thesis that people will use custom models, they will want to invest in the infrastructure that we use to post-train. So we actually are, instead of looking at the total demand for everything, we actually are like: what are the most strategic companies that we can go and work with, that we can grow with them, and they're also the right folks to basically go and prove out this thesis that custom models are much stickier and valuable, and the ROI is persistent besides just some sort of off-the-shelf model.

AI原生招聘超越工程 AI Coding and Engineering Advice

Host

我想谈两件我听你说过的事情。一件是你说过一些我觉得非常有趣的话:最好的工程师都是在 AI 存在之前就学会了编程。我想知道你能否稍微展开一下。现在有很多 16 岁的孩子,他们编程,是这些 AI 工具的原住民——他们只通过 AI 工具学习编程。你如何考虑给人们建议……

Two things I wanted to hit on that I've heard you say. One is you said something I thought was really interesting: that the best engineers all learned to code before AI existed. I wonder if you could expand on that a little bit. There are all these 16-year-olds now that are coding and are native to these AI tools—they've only learned to code with these AI tools. How do you think about advising folks on...

Yash

是的,不,这是一个很好的问题。澄清一下,我不反对 AI 编程。实际上,

Yeah, no, it's a great question. To be clear, I'm not against AI coding. Actually,

Host

我想你用了相当多的 AI。

I'd imagine you use a fair amount of AI.

Yash

没错。而且当我在 OpenAI 时,我主要做的是 Codex。那实际上是我花所有时间做的事情。所以我对 AI 编程非常兴奋。我认为将你的思考委托给编程智能体,与使用它们来执行你想到的某种计划和设计,是有区别的。所以我想说的是,我实际上认为人们最应该关注的是系统设计。来自一家基础设施公司的人这么说可能有点好笑,但你在构建可扩展系统时面临的问题类型,真的迫使你思考你在做什么,以及为什么你要这样设计和构建。我认为模型变得非常好了。所以你不需要知道确切的语法之类的。把框架和设计做好比任何实际编码都重要得多。所以我绝对是 AI 编程的忠实粉丝。只是你不能把思考委托出去,基本上依赖它作为拐杖,因为那样你会得到大量你无法解释的垃圾代码库。所以可能最好的事情是学会解释你的代码。这就是我们面试时考察的——我们会在现场面试时让某人尽情发挥,但然后我们会问大量问题:你为什么这样构建?你考虑了哪些权衡?如果他们说是 Claude 做的,或者 Codex 做的,那不是我们特别兴奋的事情。

Exactly. And also when I was at OpenAI, the main thing I worked on was Codex. That was actually the thing I was spending all my time on. So I'm very excited about AI coding. I think there's a difference between delegating your thinking to coding agents versus using them to go and execute on some sort of plan and design that you've thought of. So what I would say is I actually think the best thing for people to focus on is system design. Funny coming from an infra company or something like that, but the types of problems you face in building scalable systems really force you to think about the things that you're doing and why you're designing and structuring things that way. I think the models are getting really good. So you don't need to know the exact syntax and all this sort of stuff. It's much more important to get the frameworks and design right than any of the actual coding. So I definitely am a big fan of AI coding. It's just you can't delegate your thinking away and basically rely on it as a crutch, because then you get these massive slop code bases that you can't actually explain. So probably the best thing is learn how to explain your code. That's what we interview for—we'll let someone go ham when they're on their on-site with us, but then we'll ask a ton of questions about why did you build this this way? What were the trade-offs you considered? And if they said, oh, Claude did it, or Codex did it, that's not something that we're particularly excited about.

Host

是的。有趣。

Yeah. Interesting.

什么让RL环境公司更优 AI-native hiring beyond engineering

Host

显然,大家都在向最 AI 原生的公司学习它们是怎么做的。你们在工程岗的面试上是怎么考虑的——我相信你们也招其他岗位,而且我相信你希望这些人也相对 AI 原生、用得熟练。对于工程以外的其他类型岗位,你在面试或流程调整上是怎么考虑的?

Obviously everyone's trying to learn from the most AI-native companies on how they do things. How have you thought about interviewing on the engineering side — I'm sure you hire other roles too, and I'm sure you want people to be relatively AI-native and fluent. How have you thought about interviewing or changing processes around that for other types of roles outside of engineering?

Yash

哦,好问题。对,我觉得同样的原则也适用。比如财务岗或商务拓展岗之类的,人们常常会用 AI 工具去搭一个财务模型之类的,如果他们说不清为什么这样搭模型,那基本上就是我们要考察的点。其实我们的观点是,人们应该用一切可用的工具,但他们必须对那些工具替他们做的一切有完整的理解。所以我们的面试流程更少约束,但在审核、或者说评估他们对产出结果的理解上,非常讲究。

Oh, great question. Yeah, I think the same principles apply. So for finance roles or bizdev roles or things like that, people will often use AI tools to go and build a financial model or something like that, and if they can't explain why they built the model that way, that's basically what we're looking for. Really, our view is people should use whatever tools are available to them, but they should have a complete understanding of everything those tools did for them. So our interview process is much more unconstrained, but then very particular in how we audit or essentially assess understanding of the outputs that were produced.

前银行家与内部数据的差距 What makes an RL environment company better

Host

在今天所有的环境公司里,是什么让一家强化学习环境公司比另一家更强?

Among all the environment companies today, what makes an RL environment company better than another?

Yash

我觉得真的是人的部分——就是能触达的人。

I think it really is the human part — like people are able to get to.

Host

对,没错。专家网络。

Yeah, exactly. The expert networks.

Yash

比如 Mercor 和 Handshake 这类公司,它们有非常厉害的专家网络。其实我在 OpenAI 的时候,我们和 Mercor 合作过——我想他们现在显然还在合作——我们和他们在一些最早的智能体产品上合作得非常紧密,比如 deep research 和 Codex。那里你能拿到全世界最顶尖的编程、数学之类的人才,而且仍然有信号需要流入这些模型。所以模型是数据的函数。数据质量越高,你的模型就越好。所以我会说,差异化的地方大概就是:能非常快地在广大的网络里发起一场活动,然后基本上把它收窄到最高质量的数据点。因为你在训练的时候,其实用不了那么多数据,对吧?你用的是几百个数据点,每个任务大概几百到几千个。所以如果你的质量保证很低,那对模型的最终表现会有实质性的影响。

So like Mercor and Handshake and these types of companies, they have incredible expert networks. Actually, when I was at OpenAI, we worked with Mercor — and I think they obviously still do — but we worked with them very closely for some of these first agent products like deep research and Codex. And that is where you are getting the best of the best in the world for coding or math or stuff like that, and there's still signal that needs to flow into these models. So models are a function of data. The higher quality of the data, the better your model. So that's what I would say probably differentiates — the ability to very quickly run a campaign across a wide network and basically narrow that down to the highest quality data points. Because when you're going into training, you're not actually using that much data, right? You're using a few hundred data points, all of those each a few hundred to a couple thousand per task or something like that. And so if your quality assurance is pretty low, that has a material effect on the model's final.

奖励黑客的最佳案例 The delta between ex-bankers and in-house data

Host

感觉这里面很多东西——公司专属模型的能力,会取决于这个差值:一边是 Mercor 能招到的前银行家、能搞清楚任务是什么,另一边是摩根大通内部的银行家、他们手上的任务、他们写的邮件。我不知道这两者之间的差距最后会有多大。

It feels like a lot of this — the power of company-specific models will be determined by this delta between the ex-bankers that Mercor can recruit and figure out what the tasks are, versus the bankers within JP Morgan and the stuff, the tasks they have, the emails they create. I don't know what the gap ends up being between those two.

Yash

对。我觉得当银行家离开时——就拿你的例子来说——当银行家离开时,他们拿不到很多长期积累起来的专有数据。他们显然在做大量的推理和判断之类的事情,但这也深受他们查看、并据以开展工作的那些数据的影响。不过对,我觉得这确实是一种 alpha 的来源,会渗透到企业的边界。我的意思是,这一直都是这样,对吧?这就是为什么人们会有竞业禁止之类的东西,对吧?

Yeah. Well, I think when bankers leave — take your example — when bankers leave, they don't get access to a lot of the proprietary data that's been built up over time. They're obviously doing lots of reasoning and judgment and stuff like that, but it's also very heavily informed by the data that they look at and that they do their work on. But yeah, I think it's a real source of alpha permeating the boundary of an enterprise. I mean, it's always been the case, right? That's why people have non-competes and things like that, right?

Host

这几乎像是——对,显然你见过这种趋势:一家公司倒闭了,人们为了数据把它买下来。就像如果我是银行,而另一家银行倒闭了,我可不希望某个实验室去收购它之类的。我不知道在这种情况下他们到底能拿什么数据来训练。

It almost seems like — yeah, obviously you've seen this trend of a company goes under and people buy it for the data. It's like if I was a bank and a bank went under, I would not want one of the labs to go acquire whatever. I don't know what data they can actually train on in this.

Yash

不。没错。对。我觉得这几乎算是这些公司为什么有价值的一个论据,对吧?那些数据有价值,所以那家公司就有价值,因此你越能把东西保持为专有,就越好。

No. Exactly. Right. I think it's almost a sort of argument for why these companies are valuable, right? That data is valuable. So that company is valuable and therefore the more you can keep things proprietary to you, the better.

Host

对。看看那条边界最后会落在哪里,会非常有意思。像黑猩猩、生物学——那感觉是很难流出去的。至于实验室最终会不会搞懂经营一家银行的基本功,或者别的什么,那会——

Yeah. And it'll be so interesting to see what that boundary ends up being. It's like clearly chimps, biology — that feels like it's hard to get out there. Whether the labs eventually figure out the basics of running a bank or something else will be —

Yash

对。我觉得这里面有太多——我的意思是,经营一家银行就像千层蛋糕,对吧?我觉得不清楚一个单一模型、或者一个 ASI 模型之类的,能不能一上来就做到。我觉得人很重要——他们积累了多年做这些工作的专业能力。把这些编码进模型其实相当困难。但没错,我觉得不同公司里会有很大一部分岗位肯定会被自动化掉。

Yeah. Well, I think there's just a lot of — I mean, running a bank is like a thousand layer cake, right? I think it's unclear whether a single model or an ASI model or something like that is just going to be able to do that out of the box. I think people are important — they've built up years of expertise doing these jobs. Codifying those into models is actually quite difficult. But yeah, I think there will be large portions of roles at different companies that will get automated away for sure.

结束语 Best example of reward hacking

Host

对。非常有意思。我在想该怎么收尾,有一件事我觉得会很有意思——因为你整天都在做强化学习——很多情况下,当你把这些东西搭起来时,模型有时会做出一些疯狂的奖励黑客行为。我很好奇你见过的最好的例子是什么?你不用提具体客户,就是某个部署里,也许某次运行没有完全如你所愿。

Yeah. Super interesting. Well, I was thinking about where to end, and one thing I thought would be interesting is — because you're doing RL all day — a lot of things like models do some crazy reward hacking sometimes when you set these things up. I'm curious what's the best example you've seen? You don't need to mention the specific customer, but in some deployment where maybe some run didn't go quite as you expected.

Yash

对。对。对。不,不,我的意思是,甚至在 OpenAI 的时候,就有各种有趣的方式显示这些模型真的很聪明。它们会找到——就像水一样,它们会找到阻力最小的路径。所以我觉得对我们来说,比如在和一些做编程的公司合作时,我们能在一些软件包里找到一些漏洞利用点,模型就用它们来在任务上轻松拿到奖励——你知道,我觉得公开的那些东西里有比这糟糕得多的。我觉得老实说,这就是为什么我们看到网络安全公司对训练强化学习模型之类的事情有这么大的兴趣,因为这几乎是完美的、可以直接施加优化压力的问题,因为这些模型是非常好的编程先验,而且进攻性网络能力相当强。

Yeah. Yeah. Yeah. No, no, I mean, even from OpenAI days, there are all these interesting ways that these models are really clever. They figure out — it's like water, they find the path of least resistance. And so I think for us, with some of the companies we work with around coding, for example, we've been able to find some exploits in packages that the model just used to sort of get easy reward on the tasks, which is — you know, I think there are much worse things that are public out there. I think that's honestly why we're seeing so much interest from cyber companies in training RL models and things like that, because it's almost the perfect problem to go and just apply optimization pressure to, because these models are really great coding priors, and the offensive cyber capabilities are quite good.

播客结尾与致谢 Closing

Host

好,这是一次非常精彩的对话。非常感谢你来上播客,抽出时间。

Well, this has been a fascinating conversation. Really appreciate you coming on the podcast and taking the time.

Yash

谢谢邀请我。很棒。非常愉快。

Thanks for having me. This is great. It's been awesome.

Podcast Outro and Thanks Podcast Outro and Thanks

Host

我是 Jacob Efron,这里是 Unsupervised Learning,一档播客节目,我在节目里与 AI 领域最聪明的人对话,向他们抛出大量问题,聊模型正在发生什么,以及这对世界上的企业意味着什么。我希望大家能看出来,我做这件事非常开心。这是我利用夜晚和周末做的一个项目,除此之外我白天还在 Redpoint 做投资人。但我们能请到这些了不起的嘉宾,真正靠的是像你们这样的听众订阅这档播客、把它分享给朋友。说到底,正是这些让整件事得以运转。所以,请考虑这么做。非常感谢你们的支持和收听。我们下期见。

I'm Jacob Efron and this has been Unsupervised Learning, a podcast where I get to talk to the smartest people in AI and ask them tons of questions about what's happening with models and what it means for businesses in the world. As I hope is clear, I have a ton of fun doing this. It's a nights and weekends project in addition to my day job as an investor at Redpoint. But our ability to get these incredible guests on really comes from folks like you subscribing to the podcast, sharing it with friends. It's really what ultimately makes this whole thing work. And so, please consider doing that. And thank you so much for your support and listening. We'll see you next episode.

互动版:逐字朗读 + 针对本期提问 →