Owning Intelligence: Post-Training for Durable AI Businesses
打开互动全文版(中英对照 + 朗读 + 问答)→Firework CEO Lynn 解释了为何后训练是公司将独特判断和数据嵌入 AI 模型、超越现成 API 的关键。
Lynn, CEO of Firework, explains why post-training is key for companies to embed their unique judgment and data into AI models, moving beyond off-the-shelf APIs.
好的,接下来我们有请 Lynn,Brendan 的密友兼合作伙伴。Lynn,先举手示意一下,这里谁做后训练?好的,房间里有不少人。那谁在用 Firework?好的,也有不少人。所以,Lynn,你在观众里有一些朋友。那么,对于接下来的这场演讲,我们要做的是聚焦后训练。Lynn,和 Brendan 的演讲设置类似。我们会用 15 分钟讲如何应对后训练的问题,然后用 15 分钟左右做问答。Lynn,非常高兴你能来,感谢你的参与。
Okay, up next we have Lynn, close friend and collaborator of Brendan's. Lynn, actually show of hands, who here does post training? Okay, good amount of the room. And who here uses Firework? Okay, good amount of the room, too. So, Lynn, you have some friendlies in the audience. And so for this next talk, what we're going to do is we're going to focus on post training. Lynn, similar setup to Brendan's talk. We're going to do 15 minutes on how to approach the problem of post training, and then 15 minutes or so for Q&A. And Lynn, really delighted to have you here. Thank you for joining us.
谢谢邀请。大家好,早上好。我是 Lynn,Firework 的 CEO 兼联合创始人。那么,随着幻灯片展示,今天我会稍微深入探讨后训练,以及为什么后训练可能与你建立自己的业务非常相关。首先,稍微像之前一样,我认为今年我们已经看到……首先,我是 Firework 的,专注于智能平台。在我们之上构建了无数应用,从初创公司到数字原生企业和企业级客户。所以,我们……我工作中有趣的部分是,我们能看到很多模式。人们在我们的平台上构建什么创新,他们面临什么挑战,以及开发者们正在构建什么趋势。
Thanks for having me. Hi everyone, good morning. I'm Lynn, I'm CEO and co-founder of Firework. So, as we bring up the slide, today I'm going to do a little bit of a deep dive on post training and why post training could be very relevant to you building your own business. So, first of all, a little bit like prior, I think this year we have seen... So, first of all, I'm a Firework, specializing in intelligence platform. There are tons and tons of applications built on top of us, all the way from startups to digital native and enterprises. So, we get... the fun part of my job is we get to see a lot of patterns. What are the innovations people are building on top of us, and how they are... what challenges they are facing, and what trends people, developers, are building on top of us.
所以,尤其是在过去一年里,软件开发和应用程序开发在某种程度上被颠覆了,因为过去,你有一个好主意,想实现它,扩大生产,需要一支由数十名非常强的产品工程师、产品经理组成的团队,合作多个季度才能交付。而现在,一个人,几周时间,甚至不懂写一行代码,就能做到。因此,在时间线和深度专业知识方面的资源需求大幅缩减,这正在改变应用领域的竞争格局,也正在促使人们从在现成的黑盒 API 之上构建,转向构建更深的模型,从而建立更持久的企业。
So, one of the things, especially in the past year, software development and application development has been somewhat disrupted because it goes... in the past, you have a good idea, you want to implement it, scaling production, it requires a team of tens of very strong product engineers, PMs working together, multiple quarters to deliver that. And right now, with one person, a few weeks, without understanding how to write a single line of code, you can do that. So, that collapsing of resource required both in terms of timeline and deep expertise is shifting how competitive the application space is, and it's shifting people from thinking about building on top of off-the-shelf black box API to build a much deeper model. So that they can build a much more durable business.
正如你们所见,尤其是在过去一周,关于开放模型和封闭模型之间有很多讨论,以及对开放模型的各种支持和集结。这种联盟的深度是因为我们相信……我的意思是,这个行业要深得多,因为看看整个行业,有这么多公司,对吧?有这么多公司,你们都在建立自己的公司。每家公司存在都有原因,因为它们专注于以特殊方式解决一个独特的问题。这意味着它们将自己的判断、品味、决心和信念注入产品中,这就是公司存在的原因。今天,如果你在现成的 API 之上构建,那么你真的需要考虑如何保持那种特殊的品味、判断和独特之处。我们相信,每家公司建立持久业务的一种方法,是真正将你的判断、品味和对客户的深刻理解融入你构建的智能中,而不仅仅是现成的 API。所以,这就是行业正在发生的事情、我们看到的趋势,以及为什么后训练可能与你非常非常相关的背景。
As you have also seen a lot of discussion, especially in the past week, between open model and closed model and all the rallying and support across the open model. The depth of that alliance is because we believe... I mean, the industry is much deeper because look at the whole entire industry, there are so many companies, right? There are so many companies, all of you are building your own company. Every single company exists for a reason because they focus on solving a unique problem in a special way. And that means they carry their own judgment, taste, and determination, conviction into that product, and that's why company exists. Today, if you build on top of off-the-shelf API, then you really need to think about how you keep that special taste, judgment, and unique part forward. And we believe one approach for every company to build a durable business is to actually bake your judgment, taste, and customer deep understanding into the intelligence you build on top of, instead of just an off-the-shelf API. So, that's kind of a little bit of context of what's happening in industry, what we are seeing, and why post training could be very, very relevant to you.
好的。所以,你可能听说过很多关于拥有智能而不是租用智能的说法。那么,拥有你自己的智能意味着什么?实际上它意味着很多事情。首先,它从数据开始。智能是数据的衍生物。显然,基础实验室,我们使用的所有基础模型,都是建立在公共数据和已解决常见任务的标注数据之上的,但你们所有人都在解决一个特定的任务。这就是为什么你们在建立业务,建立公司,并且能够策划高质量的生产数据,甚至生成合成数据来丰富你的生产数据,是拥有你自己智能的一步。
Okay. So, you probably heard a lot about owning intelligence, not rent. And what does owning your own intelligence mean? It actually means many things. So, first of all, it starts from data. Intelligence is derivative of data. And obviously the foundation labs, all the foundation models we're using, are building on top of public data and label data that has solved common tasks, but all of you are solving a specific task. That's why you're building a business, you're building a company, and being able to curate production data with high quality and even generate synthetic data to enrich your production data is one step of owning your own intelligence.
然后,在你有了数据之后,你会开始利用这些数据转化为模型,在现有模型之上构建并拥有权重。有一系列技术可以用来实现这一点。这些技术针对你可能遇到的不同类型的问题而定制。这些技术也可以相互配合,帮助你达到最终目标。然后,在你拥有一个属于你自己的优秀模型,很好地解决你的特定问题之后,你需要致力于服务它。你可能首先会做一些 A/B 测试,以确保它真正推动你的产品指标,然后回到这个循环中。
And then after you have data, you will start to use that data to turn into a model, build on top of existing model and own the weights. And there are a collection of techniques you can use to get there. Those techniques are tailored to solve different kinds of problems you could possibly have. And those techniques can also interoperate with each other for you to reach your final goal. And then after you have a great model belonging to yourself, solving your specific problem really well, and then you need to work on serving it. You probably will first do some A/B testing to make sure it really moves the needle for your product metrics, and then go back in this loop.
显然,我认为你不应该立刻跳入后训练。你会经历不同的阶段。首先,提示。每个人都从提示开始,按原样使用模型。这是少样本。如果快速的话,你可以用它来测试你的想法。然后你使用 RAG 将 AI 的使用扎根于你自己的数据,并从中开始做大量的上下文工程。这些是从几分钟的交互到几小时的交互。然后你首先进展到,嘿,我有一些数据。我想看看我的数据如何反映并让模型更好地与我的产品配合。所以你会开始做监督微调。这会花你几个小时,到,嘿,我希望模型反映个性化的品味。这是我产品非常独特的选择。因此,你想开始使用从用户交互中收集的偏好信息,无论是点赞、点踩,这类信息,并帮助模型学习你的产品品味。
Obviously, I don't think you should jump into post training right away. And there are different phases you will go into. First, prompt. Everyone starts from prompt, use the model as is. It's few shots. If quickly, you can use that to test your ideas. And then you use RAG to ground the usage of AI into your own data, and you can start to do a lot of context engineering from that. Those are from minutes of interaction to hours of interaction. And then you first progress into, hey, I have some data. I want to see how my data is going to reflect and make the model work better with my product. So you will start to do supervised fine-tuning. That will take you a few hours to, hey, the model I want to reflect in personalized taste. It is very unique choice of my products. So therefore you want to start to use preference information where you collect from user interaction, whether thumbs-up, thumbs-down, a lot of those kind of information, and help the model learn your product taste.
最后,你想构建一个模型,以理解并承载你在该领域的专业知识。无论这些专业知识是跨法律、金融、医疗保健、客户支持、招聘、营销、销售,等等。即使在一个行业垂直领域,也有如此多的子领域。所以所有这些对于你正在构建的产品来说都是独特和特别的。所以通常,你可能听说过很多关于强化学习的内容,那实际上是为了构建专业性。所以这个进展非常类似于我们人类随时间学习知识的方式。例如,我们实际上通过阅读文献学习很多知识,对吧?文献中会说,嘿,什么是正确的,什么是不正确的?所以这非常类似于监督微调。而随着我们的成长,我们发展出自己的品味和判断,关于我们想要如何以特定方式处理问题,那就是偏好或 DPO。
And finally, you want to build a model towards understanding and carrying your expertise in that domain. Whether that expertise is across legal, finance, healthcare, customer support, recruiting, marketing, sales, you name it. Even in one industry vertical, there are so many subdomains. So all of that is unique and special towards the product you're building. So usually, you probably heard a lot about reinforcement learning, and that is to actually build towards a specialty. So this progression is very similar to how we human beings learn knowledge over time. For example, we actually learn a lot of knowledge by reading literature, right? In the literature it will say, hey, what is correct, what is not correct? So this is very similar to supervised fine-tuning. And as we grow, we develop our own taste and judgment of how we want to conduct the specific way we want to approach a problem, and that is preference or DPO.
随着时间推移,我们会在某个领域变得非常擅长,无论是会计——我想当会计师,我就会真正懂得如何深入财务数据——还是我想当牙医,你会学习牙科操作的细节。所以,这一切和我们人类获取知识的方式非常相似。而且,有趣的是,不同的技术用于解决不同类型的问题。例如,如果模型不知道某个事实,而事实又是非常动态的——你产品中的数据事实非常动态——那么通常你会用 RAG 来解决这个问题。然而,如果你的模型输出行为或结构不对,那么你会整理数据并进行监督微调来纠正。如果你的模型回答质量是个人化的,或者特定于你产品的品味,那么你会用偏好调优。或者如果模型在你试图解决的特定问题上很弱,那么你会用强化学习。如果模型的最终结果太慢或太贵,无法在生产环境中服务,那么你会用蒸馏,让教师教学生——一个更小的学生模型——这样它就能更高效或更经济。显然,蒸馏对 LLM、VLM 或图像生成模型也意味着不同的东西。如果你感兴趣,我们可以进一步讨论。
And over the course of time, we learn to be really good at a certain area, whether it's accounting—I want to be an accountant, and I would really know how to build into the financial data—or I want to be a dentist, and you learn the details of how to operate with dentistry. So, all this is very similar to how we human beings acquire knowledge. And also, very interestingly, different techniques are there to solve different kinds of problems. For example, if the model doesn't know a fact, and the fact is actually very dynamic—the facts of your data in your product is very dynamic—then typically you use RAG to solve that problem. However, if your model's output behavior or structure is off, then you curate data and do supervised fine-tuning to correct that. If your model's answer quality is personal or specific to the taste of your product, then you use preference tuning. Or if the model is quite weak on the special problem you're trying to solve, then you use reinforcement learning. And if the end result of the model is too slow or too expensive for you to serve in production, then you use distillation to let the teacher teach the student—a much smaller student model—so it can be more performing or economical. Obviously, distillation also means different things for LLM, VLM, or image generation models. If you are interested, we can talk more about that.
所以,团队有很多不同的方式开始尝试这些技术,他们可能对体验不满意,因为他们会烧钱和时间,但可能达不到理想结果。所以,这里有一些你可能感到沮丧的领域。例如,当涉及到数据——数据是调优的本质——数量不是最重要的。实际上,质量才是最重要的。但有时仅仅向训练过程投入海量数据可能不会带来好结果。所以,你真的需要控制数据质量。而且,通常谁是数据质量的最佳评判者?实际上是你的产品团队。所以,这就是我们看到融合的地方——在 GenAI 之前,产品团队和研究团队或机器学习团队是分开的组织、分开的团队。他们携手合作来推动事情。而如今很多时候,当人们后训练 GenAI 模型时,我们看到产品团队需要判断数据质量并深入参与过程的融合,这样他们才能确保最佳结果。
So, there are many different ways teams start to try these technologies, and they may not be happy with the experience because they can burn money and time, but may not reach their ideal results. So, here are the areas that you can possibly feel frustrated. For example, when it comes to data—data is the essence of tuning—the quantity is not the most important. Actually, quality is the most important. But sometimes just by throwing tons and tons of data into a training process may not lead to a great outcome. So, you really need to control the data quality. And usually, who's the best judge of data quality? It's actually your product team. So, that's where we see the convergence of—before GenAI, product team and the research team or ML team, they're separate organizations, separate teams. And they work hand-in-hand to make things happen. And a lot of times these days, when people post-train GenAI models, we see the convergence of the product team needing to make judgment calls on data quality and starting to be deeply involved in the process, so they can ensure the best result.
第二点是你会有评估。我知道你们都很忙,而且为了快速发布,很多评估都是凭感觉评估。
And the second is you're going to have evals. I know all of you are very busy, and part of launching quickly, a lot of evals is vibe evaling.
创始人真的会看结果,然后觉得,嘿,对不对?
And the founders really look at the result and feel, hey, is it right or not?
实际上,这就是判断。这是你对最终结果做出判断,决定它好不好。然后你把那个判断转化为系统化的评估。所以,这和传统软件开发没什么不同,你有单元测试、集成测试来确保质量。同样,如果你考虑做后训练,然后有办法构建评估,并把你自己的判断构建成可重复的过程,这极其重要。
Actually, this is judgment. This is you put your judgment of the end result and decide whether it's good or not. And that judgment you convert into systematic evaluation. So, this is no different from traditional software development where you have your unit test, integration test to ensure quality. Similarly, if you think about doing post-training and then have a way to build the eval and build your own judgment into a repeatable process, it's extremely important.
然后,当你做强化学习时,可能会有草率的强化学习环境,你构建一个模拟,但模拟并没有真正反映现实,然后模型可能会在糟糕的模拟上爬山,它还可能做奖励黑客和各种奇怪的事情。所以,关于奖励黑客的一个有趣故事是,我们一直要求一个模型生成最小化编译错误的代码。好的。那么你猜模型做了什么?模型生成了零行代码。好的,没有编译错误,但这绝对不是你想要的结果。所以,这是奖励黑客的一个例子;这很常见,因为模型非常聪明,它会用各种不同的方式来实现你的目标,但可能不是你想要的。所以,注意所有这些细节,尽量不要让模型比你更聪明。
And then there could be when you do RL, there could be sloppy RL environments where you build a simulation and the simulation is not really reflecting reality, and then the model can hill-climb on a bad simulation, and it could possibly do reward hacking and all kinds of weird stuff. So, a fun story about reward hacking is we have been asking a model to generate code that minimizes compilation errors. Okay. So, guess what the model did? The model generated zero lines of code. Okay, there's no compilation error, but that's absolutely not what you want. So, this is one example of reward hacking; it's very common because the model is very smart, it'll try in all different ways to get your goal, but it may not be what you want. So, pay attention to all these details and try not to let the model outsmart you.
显然,在你的实验之间——把你的开发过程看作你实际上在做大量实验,从训练模型开始——训练模型并不是你实验的终点。最终的评判标准是产品指标是否在变化。所以,你需要把最终模型带到服务层,做 A/B 测试。而这个过渡极其重要,因为如果你从一个训练栈迁移到服务栈而没有对齐这两者,质量可能会下降。因为想想模型——它有大量的计算、数学和矩阵乘法。而做矩阵乘法和计算的方式,如果你使用不同的库、不同的数值和不同的优化,会导致不同的结果。因此,训练的结果可能无法复现,甚至在服务过程中你会失去精度。所以,这种对齐非常重要。我可以给你更多这方面的例子。还有几个你会遇到的其他挑战,很高兴离线和你详细讨论。
Obviously, between your experimentation—think about your development process as you're actually doing a lot of experimentation from training the model—training the model is not the end of your experiment. The final judge is whether product metrics are moving or not. So, you need to bring the final model into the serving tier and do A/B testing. And that transition is extremely important because the quality can drop if you move from one training stack to a serving stack without aligning these two. Because think about the model—it's tons and tons of calculations, math, and matrix multiplication. And the way to do matrix multiplication and calculation, if you use different libraries, different numerics, and different optimization, it will lead to different results. And therefore, the end result of training may not be replicated or even you lose precision during serving. So, that alignment is very important. I can give you more examples of that. And there are a few other challenges you will run into, and happy to talk with you more about in detail offline.
所以,我们合作过很多后训练的先驱。我会说 Cursor 是其中之一。他们从去年年初就开始登上这趟列车。原因有很多。一是他们真的想掌控模型供应的命运。显然,他们有很多用户参与。他们理解客户偏好,有大量数据。所以,这成为他们旅程的开始。他们做了相当深入的从中间训练到后训练。你听到他们每隔几个月就宣布更新模型。结果很棒。他们的灵感是在前沿质量上竞争。非常大胆的抱负。他们正在实现。所以,Composer 2、Composer 2.5 的结果非常令人印象深刻,与——他们总是试图与前沿实验室的质量持平或超越。所以,他们是其中一个例子,一个非常生动的例子,只要保持专注并拥有正确的工具,Cursor 在我们之上构建,这是完全可行的。
So, there have been many pioneers we worked with in post-training. I would say Cursor is one of the few. They have started onboarding onto this train from the beginning of last year. There are multiple reasons. One is they really want to control their destiny of the supply of the model. And obviously, they have a lot of user engagement. They understand customer preferences, a lot of data. So, that becomes the beginning of their journey. And they do pretty deep mid-training to post-training. And you have heard them announce continuously newer models every few months. And the result is great. Their inspiration is to compete at the frontier quality. Very bold aspiration. And they're getting there. So, very impressive results from Composer 2, Composer 2.5, on par with—they're always trying to be on par or beat the frontier labs' quality at the same time. So, they are one of the examples, a very vibrant example, that it's completely doable as long as you stay focused and have the right tools and Cursor builds on top of us.
还有来自医疗保健领域的各种例子。
There's another huge variety of examples from healthcare.
例如,Doximity 就是一个例子,他们做临床 AI,基本上让医生提出深入的医学问题,将症状与药物和副作用匹配,并对医学领域的一切进行全面的研究。他们实际上也在 Fireworks 上训练,并且他们在斯坦福哈佛临床安全基准测试中名列前茅。我们为这一结果感到非常自豪,他们也在继续这个训练循环。Factory 是另一家红杉资本的公司。他们在 Fireworks 之上构建,特别关注编码的安全部分。这是一个非常困难的话题,因为安全不容忍高容错。你必须做对。他们调优的模型也在安全领域的基准测试中名列前茅。所以你可以看到所有这些例子都是一种展示。所有这些公司都在以特殊的方式解决非常独特的问题,他们通过基于开放模型构建,能够用自己的数据应对自己的挑战。这些就是细节。
For example, Doximity is one example where they do clinic AI where they basically let doctors ask deep medical questions matching symptoms to medication and side effects and have well-rounded research around everything in medical space. They actually also train on Fireworks and they top a very important benchmark which is Stanford Harvard clinic safety benchmark. And we are very proud of that result and they keep working on this training loop. Factory, that's another Sequoia company. They build on top of Fireworks and especially focus on security part of the coding. That is a very hard topic because security is not high tolerance. You need to get it right. They tune a model that also top the kind of benchmarking security area. So you can see all these examples is a demonstration. They solve all the companies are solving a very unique problem in a special way and they have been able to kind of build their own challenges use their own data by building on top of open model. And those are the details.
显然,如你所知,去年 2025 年是编码之年。在 Fireworks,我们支持所有基于我们构建的编码公司。在从中间训练到后训练中取得了许多成功,产生了非常强大的模型。如你所见,这些基准测试令人印象深刻,所有这些编码公司都受到启发,要在编码领域赶上或超越前沿实验室,但这不仅仅是编码。今年,我们开始看到各种不同协作空间的有趣发展,从通用协作到专业领域特定的协作,正如我提到的。有法律、金融、营销、招聘、销售、客户支持等协作空间。它们都开始进行后训练,并拥有自己的智能来驱动产品。这里,Jen Spark 是一个通用的协作应用,他们为专业人士构建深度研究和幻灯片生成。如你所见,他们与前沿模型竞争,略胜一筹,但成本显著更低。所以我们说的是 5 到 10 倍的成本降低。因此,作为初创公司,一旦你遇到这样的问题,你可以迅速扩展成一个可行的业务。我想另一个医疗保健的例子,医疗保健数据或用例在前沿实验室模型中没有被很好地捕捉。所以你可以看到调优模型的质量显著优于最先进的封闭模型。
Obviously, as you know, last year 2025 is the year of coding. There Fireworks we support all the coding companies build on top of us. A lot of success in mid-train to post-train very strong models. As you can see those benchmarks are very impressive and all these coding companies are inspired to get on par or beat Frontier Labs in the coding space, but it's not coding. And we start This is the year we start to see very interesting development in all different kind of co-work space from general purpose co-work to specialized domain specific co-work as I mentioned. There are like legal, finance, marketing, recruiting, sales, customer support co-work space. They are all starting to post train and own their own intelligence power their product. Here, Jen Spark is one of the generic co-work application and they build deep research for professionals and slide generation. As you see, this they compete with a frontier model and it's slightly better, but the cost is significantly lower. So here we're talking about five to 10 10 times cost reduction. So as a startup, you can once you hit a problem like if it you can scale quickly into a doable business. And I guess another health care example and health care data or use case is not well captured in in frontier labs model. So you can see the quality of tuned model is significantly better than the state of art closed model.
嗯,我们有多种不同类型的开发者基于 Fireworks 进行后训练。那么这些类型是什么呢?我们看到了许多前沿智能体构建者。前沿智能体构建者通常构建定制的、个性化的框架。他们不使用常见的框架。为了将他们的框架与他们的系统以及他们想要构建的各种工具进行协同优化,通常需要后训练。嗯,所以,我们看到了很多可重复的成功案例。显然,我们也看到很多大公司进行后训练,因为有趣的是,这些现有企业拥有巨大的流量。对他们来说,向所有客户全面部署 AI 功能是一个巨大的成本负担。巨大的成本负担。嗯,所以,如果我们说要小心不要规模扩张到破产,这不仅仅是针对初创公司。实际上也是针对现有企业。他们的 CFO 真的因为成本而阻止 AI 功能的发布。而后训练是解决这个问题的方法。嗯,我们还看到了很多专业模型运营商,这些可能是新的实验室,很多前沿模型开发者也会进行大量的后训练。
Well, we have many different type of developers build on top of fireworks to do post training. So what what are these types? So we have seen many frontier agent builder. So frontier agent builder, they usually build bespoke customized harness. They don't use common harnesses. And to co-optimize their harness with their systems and all different kind of tools they want to build on top of often time requires post training. Uh so, we have we have seen a lot of repeatable success on doing that. And obviously, we have seen a lot of big companies put to post training because interestingly, those incumbents have huge amount of traffic. For them to deploy AI features across the board to all their customers is a huge cost burden. Huge cost burden. Uh so, if we say we we need to be careful not scale into bankruptcy, it's not just for startups. Uh it's actually for incumbents. They literally their CFO is blocking their AI feature launch because of the cost. And the post training is way to remediate that. Uh and we have seen a lot of specialized model operator, uh those could be uh new uh new labs and a lot of kind of cutting-edge model develops also um do significant post training.
所以,嗯,这是我们反复听到的一句话:在产品市场匹配之后,后训练成为许多许多公司构建专业智能的载体。我们确实相信,我们也看到了趋势,未来可能会有数百万个专业模型。每个用例一个应用。会是数百万个。我就是这么认为的。嗯,显然,我们也看到了将数据转化为模型的不同方式。嗯,奖励信号以及如何构建这些奖励函数是故事中非常重要的一部分。嗯,所以,我们的观察是,不,要尽早开始,尽早开始,然后开始进行更多的迭代实验,并获得实践经验,我们看到了人们如此迅速地加入。实际上,并没有很多人感觉到的障碍,因为整个奖励工程与软件工程非常相似。逻辑推理和思维方式非常非常相似。所以,是的,动手实践,开始测试吧。酷。我想,是的,这些是关键要点,嗯,我想我们想承认,在这个领域,你现在拥有不同的专长或不同的知识。
So, uh this is a quote we have heard repeatedly uh that after part of market fit, post training becomes the vehicle to for many, many companies to build specialized intelligence. We do believe we do see the trend that in the future there could be millions of specialized models. One application per use case. It it it it will be millions. That's how uh I believe in that. Um And obviously, there are um we have seen kind of different way to convert the data into into model. Um and the the signal from the reward and how to build those reward function is very important um part of the story. Um so, our observation is no, start early, start early, and the start to do experimentation a lot more iteratively and they get hands-on experience and we we have seen people on board so quickly. There there's actually not the barrier as many people feel because this this whole reward engineering is very similar to software engineering. The the the logical reasoning and mindset is very very similar. So so yeah so get your hands dirty and and start to kind of test it out. Cool. I think Yeah so those are the kind of key takeaways and uh I think we we we want to acknowledge there are different specialties or different knowledges you have right now in this space.
例如,我们与 Cursor 和 Cognition 这样的公司合作。他们有深入的研究人员。他们希望尽可能控制每一个旋钮,以获得极致的结果。我们给他们最低层的 API。这意味着我们只提供 RL 回滚,他们可以直接交互,并完全控制训练器。这给了他们最好的结果。但许多公司有产品工程师或机器学习工程师。你对后训练有足够的信息和知识。你想要一些控制,但你不想要最低层的控制。所以这就是我们的训练 SDK 与你们所有人紧密迭代,让你们上手的地方,我们也乐意部署我们自己的研究人员。这就是我们会看到,嗯,动手帮助,比如训练你的团队,让他们上手。所以,这些就是我们看到的不同团队想要参与并熟悉整个过程的多种方式。酷。嗯,我在这里暂停一下,看看你们是否有问题。
For example, we work with the companies like Cursor and Cognition. They have deep researchers. They want to control every single knob as much as possible to get extreme results. We give them the lowest API. That means we just have RL rollout that they can directly interact with and they fully control the trainer. That gives them the best result. But many of the company you have product engineers or machine learning engineers. You have good enough information and knowledge about post training. You want to have some control but you do not want to have the lowest level control. So that's where we have our training SDK iterate very closely with all of you to get to get you going and we also are happy to kind of deploy our own researchers. That's where we will see, um, help hands-on like training your team to to get on board. So, those are the kind of different way we see different teams want to engage and and get familiar with the whole entire process. Cool. Uh, I'm going to pause here and see if you have any questions.
我来插一句。嗯,如果人们已经有数据和奖励信号,也许来自他们的产品,他们是最成功的。你见过的哪些奖励信号的最佳例子,能够非常快速且成功地训练?
I'll jump um People are most most successful if they have data and reward signal already, maybe from their product. What are the what are the best examples of reward signals you've seen with very successfully trained very quickly?
所以,嗯,通常你可以把奖励想成代码。你用代码写奖励。嗯,通常想想奖励的评分标准。嗯,想想你想从多个维度对结果进行评分。嗯,例如,如果你在构建一个招聘智能体,嗯,然后你可以想想,“嘿,我们如何评估候选人筛选?”嗯,不同的公司,我保证你有不同的标准。嗯,例如,我们想评估候选人的能力。他们是否真的渴望?他们不接受“不”作为答案。他们会打破壁垒。这是一个指标。第二个是他们在构建事物和取得进展方面非常快。等等。所以,然后你有一个分数混合来合并这些。所以,想想不同维度的评分标准作为奖励。
So, um, usually the rewards you can think about rewards actually is code. You write rewards in code. And uh, usually think about rubrics of rewards. Um, and think about uh, you want to grade the result in multiple dimensions. Uh, for example, if you are building a recruiting agent, uh, and then you can think about, "Hey, how do we evaluate candidate selection?" Um, and a different company, I guarantee you you have different criterias. Uh, for example, we want to grade a candidate um, um, aptitude. Are they really hungry? They do not take no as an answer. They they will break down walls. And that's one matrix. The second is they're really fast in kind of um, building things and making progress. And and so on. So, then you have a blend of score to blend to merge this. So, so think about reward in different dimensions of rubrics.
然后不同公司有不同的融合方式。那是你独特的部分和秘密配方。哦,抱歉。呃,讲得很好,Lin。我是 Vercel 的 Raj。我很好奇,你提到了过早开始的陷阱,但根据你在 Vercel、Vercel cognition 或其他公司看到的,他们什么时候真正开始考虑后训练?是他们觉得准备好了?还是成本太高了?因为他们在基准测试上击败了前沿模型,这是个很好的信号,说明你能打败前沿模型,但这是主要的动机吗?那 2% 的提升值得投入成本或投资到整个框架吗?他们怎么考虑投资回报率?你知道,闭源模型更贵,但后训练也有金钱、投资和后续维护的截点。所以我很好奇,他们是在成本成为问题时开始,还是当使用量激增、他们真的必须提前一两年规划时?
And then different companies have a different way to blend those. And that's your unique part and secret sauce. Oh, sorry. Uh, great talk, Lin. Raj from Vercel. I'm curious about I think you mentioned what are the pitfalls of starting too early, but based on what you're seeing from Vercel, Vercel cognition, or other companies, when do they actually start thinking about post-training? Is it when they feel ready? Is it when the cost is now too much? Because the benchmarks where they're beating it, that's a great signal that yes, you can beat the frontier models, but is that the primary motivator? Like that is that 2% bump worth the cost or investment into this whole framework? And how do they think about ROI? You know, close models are more expensive, but then post-training has some intercept of, you know, money and investment and maintenance going forward. So, I'm curious of like, are they approaching when cost becomes a concern or is it when the usage spiking up and they really have to now plan for a year or two in advance?
是的,这是个很好的问题。通常,在 AI 领域,要把产品市场匹配和业务扩张看作两个阶段。在 SaaS 时代,这是一个概念。你实现了产品市场匹配,就扩张,尽可能多地扩张。但现在我们看到了分化。产品市场匹配并不真正意味着你有一个持久可扩展的业务。通常,我们看到公司采取的策略是,先通过在前沿实验室模型之上构建来专注于产品市场匹配,因为你不需要担心任何事。所以只管花钱,实现产品市场匹配。另一个非常重要的事情是,只有在你实现产品市场匹配之后,从产品表面收集的数据才真正有意义。你也会开始从产品表面收集到大量高质量数据。这成为你开始拥有自己智能的燃料。所以我们认为这是一个非常强烈的信号,因为一旦你实现了产品市场匹配,你真的开始考虑扩张业务时,你会考虑两件事。一是继续保持你的竞争优势。二是建立一个持久的业务,让你的收入和成本处于健康状态。所以后训练成为一个非常有吸引力的解决方案,因为后训练允许你基本上将你独特的品味编码进一个没人能偷走的模型里。因为克隆和复制应用程序本身非常容易,正如你们所知,对吧?从编码智能体来说非常容易。所以从截图,砰,生成相同甚至更好的应用。所以很可怕的是,应用本身作为护城河正在被削弱。而你将反映产品参与度和所有深层知识的数据烘焙进你的模型,是保留这一点的方式。第二,你可以后训练一个模型,把成本降低 5 到 10 倍。这意味着你可以用同样的预算支持 5 到 10 倍的流量。你的单位扩张经济性要好得多,所以你避免了扩张到破产。所以我们认为这是背后原因和时机背后的两个令人信服的故事。
Yeah, this is an excellent question. So usually, think about in the AI building product-market fit and scaling the business as actually two phases. In the SaaS time, it's one concept. You hit product-market fit, just scale. You've got to scale as much as possible. But now we see a bifurcation. Product-market fit doesn't really mean you have a durable scalable business. And typically, we see companies deploy the strategy of focus on product-market fit first by building on top of frontier labs' models, because you don't need to worry about anything. So just spend your money and hit product-market fit. And the other very important thing is only after you hit product-market fit, the data you collect from product surface area are really meaningful. And you will also get the volume of high-quality data starting to collect from product surface area. That became the fuel of you start to own your intelligence. So we see that as a very strong indication, because once you hit product-market fit, you really think about starting to scale a business, you think about two things. One is continue to keep your competitive edge. Two is build a durable business, so your revenue and your costs are in a healthy state. So then post-training becomes a very appealing solution, because post-training allows you to basically encode codify your unique taste into a model that no one can steal from. Because it's very easy to clone and copy application as is, as you all know, right? From coding agent it's very easy. So from screenshot, boom, generate the same or even better app. So that's very scary that application itself as a moat is being reduced. And you bake your data reflecting the product engagement and all the deep knowledge into your model as a way to preserve that. And second, you can post-train a model that brings down the cost five to ten times. And that means you can support five to ten times much higher traffic with the same budget. And your unit economics of scaling is so much better, so you avoid scaling into bankruptcy. So we see those as two compelling stories behind the reason and timing.
谢谢你,Lin。
Thank you, Lin.
非常感谢。
Thanks a lot.