Sovereign AI: Building the UK's First LLM on a Startup Budget
打开互动全文版(中英对照 + 朗读 + 问答)→Cosine 首席执行官讲述如何获得政府支持,打造英国首个主权大语言模型,克服算力限制,与数十亿美元的项目竞争。
Cosine's CEO discusses how they secured government backing to build the UK's first sovereign LLM, overcoming compute constraints and competing with billion-dollar efforts.
Alistair,很高兴见到你,老兄。我们在伦敦,按惯例得说声“你好,Giza”。
Alistair, it's great to meet you, mate. We are here in London where it's customary to say hello, Giza.
你好,Geyser。
Hello, Geyser.
那么,我们在伦敦的哪个位置?
So, where in London are we?
我们现在在 Hawkton,也就是肖尔迪奇。离 Cosine 起步的我的公寓大约半英里,就在那边的 Hawkton 广场。所以我们没走多远,但自那以后已经扩张了不少。
We are in Hawkton right now. So, we're in Shoreditch. We're about half a mile away from where Cosine started in my apartment, which was in Hawkton Square, just over there. So we haven't come very far, but we have expanded a fair bit since then.
Cosine 是做什么的?
And what is Cosine?
Cosine 是一家位于英国的前沿实验室。大约 3 个月前,我们专门为高度监管和高安全环境构建了顶尖的编码智能体。比如金融服务、保险、国防等领域。但最近,我们获得了构建英国首个主权大语言模型(LLM)的授权,这是一个更宏大的愿景,规模远超以往,但能参与其中非常令人兴奋。
Cosine is a frontier lab based here in the UK. Prior to about 3 months ago, we built best-in-class coding agents specifically for highly regulated and high-side environments. So think things like financial services, insurance, defense, and so on. But more recently, we have obtained the mandate to build the UK's first sort of sovereign LLM, which is a much more ambitious vision and something which is on a scale much larger than we've done before, but it's very exciting to be working on it.
跟我讲讲主权 AI 这部分吧。顺便说一句,我读了《每日电讯报》上关于你的文章。当 Fable 突然被封禁时,我和很多人一样感到震惊。因为出口管制,现在每个人突然都在思考主权 AI。所以给我讲讲这个故事。
So tell me about that sovereign AI piece. I should say by the way I read the article about you in the Telegraph. And I was horrified as many were when Fable suddenly got banned. Because of this export control and now everyone suddenly is thinking about sovereign AI. So tell me the story.
这跟好几件事有关。它跟 Cosine 的背景故事有关,我们之所以有幸能从事主权 AI,是因为我们在模型训练、模型构建以及所需的基础设施、算法、数据和人才方面拥有大量专业知识。我们已经做了相当长一段时间,所以组织内部已经具备了所有这些条件。大概 9 到 10 周前,我们被纳入政府的“主权 AI 部门”,或者说得到了他们的支持。这是他们为了增加英国本土主权 AI 初创公司数量而推出的举措。对我们来说,实际表现是获得了布里斯托尔 Isambard 超级计算机集群的算力分配。老实说,这让我们甚至有能力去追求这样的雄心。对于像我们这样规模或更小的初创公司来说,最大的障碍之一就是算力。如果你筹集了 5000 万到 1 亿美元,很大一部分都会花在这个项目的算力上。从主权 AI 部门获得分配意义重大,因为它真正让这件事成为可能。我们仍然会使用一些私有算力作为补充,但基本上所有工作都会在 Isambard 上完成,这非常酷。老实说,今年 1 月我们刚开始时,这完全不在我的预期之内。我们仍然在做传统的编码智能体业务和已有的模型,但我们已经能够将那个愿景推向极致,否则我们无法做到。
It ties into a bunch of different things. It ties into the backstory of Cosine and one of the reasons we're fortunately placed to be able to do sovereign AI is that we have a lot of expertise around model training, model building, all of the infra, algorithms, data, people that you need to do that kind of thing. And we've been doing that for some time. So we already had all of those things in the organization. Then probably 9 or 10 weeks ago, we were inducted into the government's sovereign AI unit or backed by them I should say. That is something that they have put out to increase the number of sovereign AI initiative companies that are being built in the UK. What that looks like in practice for us is an allocation of compute on the Isambard supercomputer cluster out in Bristol. That has allowed us to, honestly, it's one of the things that's unlocked our ability to even have the ambition to do something like this. Fundamentally, one of the biggest blockers for a startup of our size or smaller to approach work like this is the compute. If you raised 50 to 100 million bucks, a good chunk of that would go on compute on doing a project like this. To have allocation come from the sovereign AI unit is huge because it genuinely does enable it. We still use some private compute on the side, but fundamentally all of it will be done on Isambard, which is super cool. To be honest with you, it was not something that was on my bingo card in January at the beginning of the year when we started out. We still do our conventional business of coding agents and the models that we already have built. But we've been able to take that vision and really take it to the extreme in a way that we wouldn't have been able to otherwise.
我不是在开玩笑,但价值百万美元的问题是——实际上还不止。你看,美国那边的人可能拥有数千亿美元的规模。Mistral 大概是 140 亿,单到双位数十亿。差不多这样。你怎么能用数百万美元做到他们用数十亿美元做的事?
So, I'm not being funny, but the million-dollar question is, well, you actually more than that. So, you know, folks over there in the US, they've got probably on the order of hundreds of billions. You've got Mistral, which is on the order of, let's say 14, single to double digit billions. Something like that. How can you do in millions what they are doing with billions?
这是个非常合理的问题,可能也是我被问得最多的一个问题。我们 Cosine 不是一家推理公司。这跟我们产品的部署方式和销售方式有关。对于你的观众,我可能需要提供一些背景。由于我们主要部署在高度安全、高安全级别的环境中,现在绝大多数时候,我们并不自己托管模型。客户不会访问 cosine.com/api/v1 然后调用我们的聊天补全接口。他们要么拿走我们提供的模型权重,部署在自己的 GPU 上——这是我们最隔离、最安全的部署方式;要么在他们已经使用的超大规模云服务商那里租用 GPU,比如 Azure 或 AWS 等,然后在那里运行模型。这对 Cosine 的实际意义是,我们授权我们构建的技术,而不是通过 token 等盈利。所有这些都跟你的问题相关:我们不必像美国公司那样在推理用的数据中心上花大钱。当然,这并不意味着训练不需要大量算力——显然需要,美国公司的大量基础设施也用于训练。但我认为,像 Anthropic 最近陷入困境以及他们与 Colossus 集群等达成交易的最大原因之一,是推理而非训练。所以,如果你不做推理部分,所需的资源会少得多。幸运的是,我们的产品销售方式意味着我们真的不需要做推理。
It's a very fair question and one that I probably get more than anything else. We at Cosine are not an inference company. That ties into the kinds of deployments and the way that we sell our product. For your viewers, I should probably give a bit of background. Given the fact that we predominantly deploy into highly secure, high-sided environments, most of the time now, nearly all of the time, we are not hosting the model ourselves. A customer isn't hitting cosine.com/api/v1 and then hitting a chat completions endpoint from us. They are either taking the model weights that we give to them and deploying them on their own GPUs. We have a lot of that—that is the most air-gapped, most secure deployment we do. Or they are renting GPUs in some hyperscaler cloud that they're already a part of, whether it be Azure or AWS or whatever, and then they'll run the model there. What that means in practice for Cosine is that we license the technology that we build. We don't actually make a margin on tokens or anything like that. All of this ties into your question, meaning we don't have to spend a lot of the money that the Americans are having to spend on data centers for inference purposes. Now that's not to say that you don't also need a huge amount of compute for training. Obviously you do, and a huge amount of the infrastructure they have in the US will also be used for training. But I think that one of the biggest reasons you've seen people like Anthropic struggle recently and the reason they've signed the deals they have with the Colossus cluster and so on is because of inference and not because of training. So you do need significantly less resource if you're not going to do the inference bit. And we're fortunate in that the way that we sell the product means we don't really have to.
另一方面,我们在如何实现这一目标上采取了一些有趣的研究方法,比如模型的架构设计、训练方式以及一些算法层面的东西。总的来说,我们确实有可信的机会完成整个流程,包括持续预训练、中期训练和后训练等所有环节。但坦白说,容错空间和回旋余地并不大。我们在一些地方不得不做出权衡决策,这包括强化学习的范围。我总希望进行更大规模的强化学习运行、更多的生成次数,以及在强化学习过程中投入更多的推理时算力以获得更多多样性。但如果我们没有 10 倍的算力,就无法做到那么多,对吧?所以存在权衡。但基本上,考虑到我们项目的规划范围——我们稍后可以深入细节——我认为这是可行的。在这个非常狭窄的范围内,英国一些最大的公司都在向我们提供用例和他们对模型能力的期望,这样我们就可以训练一个真正为他们服务的模型。我认为这可能是该领域尚未充分探索的反馈循环。
And then on the other side of that, we are taking some interesting research approaches in terms of how you pull something like this off. We can talk more about that in a minute, I'm sure, in terms of how we are architecting the model, how we're training it, some algorithmic stuff. All of that's to say that we do have a credible shot at pulling off the full run, including the continued pre-training, the mid training, the post-training, all of those bits. But to be completely transparent with you, there isn't a huge amount of room for error or a huge amount of wiggle room. There are obvious places where we have had to make trade-off decisions. That includes the scope of RL. I'd always like to do larger RL runs, more generations, more inference time compute during the RL process to get more variety. We can't do as much of that as we would like to if we had 10 times more compute, for instance, right? So there are trade-offs. But fundamentally, I think given the way that we've scoped the project, and we can get into more detail I'm sure, but given the way that we've scoped the project, I think it is viable. In that very narrow scope, we have some of the largest companies in the UK all feeding use cases and their desires for what they want the model to be able to do directly into us, so that we can train a model that's really for them. I think that is potentially a feedback loop that hasn't really been explored as much in the space.
显然不想说这些公司的坏话,但你知道,Mistral、Coherent 等公司的模型并不具有竞争力。甚至中国模型也才在前天开始变得有竞争力。
Obviously don't want to say bad things about some of these companies, but yeah, you know, the models from Mistral from Coherent and so on, they're not competitive. Even the Chinese models arguably they've only really started getting competitive the day before yesterday.
大概几个月吧。是的,GLM 5.2。所以感觉不错。虽然在 Arc 挑战中表现不太好,但也许那只是烟雾弹。我稍后可以谈谈这个。
Months or so. Yeah. GLM 5.2. So, the vibes are good. Although on the Arc challenge, it didn't do very well, but maybe that was just a red herring. I can talk about that in a minute, but yes.
很酷。但你知道,你几乎是在暗示,哦,那是因为他们没有努力做得更好。那么我们怎样才能制造出与那些前沿模型一样好的模型呢?
Very cool. But you know, like so is it because you were almost implying oh it's because they weren't trying to make it better. Like what how can we make models that are as good as those frontier models?
粗略来说,我认为有几个关键因素。嗯,也许是三个关键因素。第一是架构/原始模型大小。第二是活跃参数数量。第三是数据。我相信,如果我错了请纠正我,Mistral 迄今为止最大的模型是 675B 的 Mistral 3 large,对吧?稀疏的,与 DeepSeek 的架构非常相似,甚至相同。我认为,从根本上说,那个模型之所以存在,是因为它符合他们看到的用例。它可能符合他们在法国或欧洲试图销售的企业中看到的 GPU 部署配置。因此,务实地说,他们觉得,考虑到某些公司能获得的资源,这大概是他们能承受的最大规模。结果,你显然会限制模型性能的上限。所以有一篇博客文章做了一个非常有趣的分析,我记不起名字了,但可以在演讲后发给你,因为我发现它非常有趣,它分解了 Sonnet 和 Opus 等模型的可能大小。
Crudely, I think there are a couple of key things. Well, maybe three key things. I think one is architecture/raw model size. I think the second is active parameter count. And the third is data. I believe, and correct me if I'm wrong, I believe the largest model that Mistral have made to date is the 675B Mistral 3 large, right? Sparse, very very similar to DeepSeek's architecture if not the same. I think as a result, fundamentally that model exists in the way that it does because it fits a use case that they have seen. It probably fits a GPU deployment profile that they have seen in the enterprises they're trying to sell to in France or Europe. And as a result, pragmatically they're like, right, this is probably the biggest that we can get away with given what certain companies have access to. As a result, you're obviously going to cap out how far you can go in terms of model performance. So there was a very interesting analysis done in a blog post which I can't remember the name of, but I can send you post talk because I just found it very interesting, of a breakdown of the probable sizes of models like Sonnet and Opus.
是的,我看到了。
Yeah, I saw that.
你看到了吗?它通过 Vertex 完成的方式太酷了,根据我们对开放权重模型和延迟时间等的了解来推算。
Did you see that? It was so cool the way that it was done, right through Vertex and figuring out okay given what we know about open weight models and latency times and stuff like that.
我看到的那篇是他们提出了一堆问题,然后可以根据常识推断出来,是的。但那个有点粗略。
The one I saw was that they came up with a bunch of questions and they could infer based on the general knowledge that one as well, yeah. But that was a little bit sketchy.
我看到了一个更经验性的。我会发给你,因为那篇博客文章本身对我们为主权模型所做的架构决策起了很大作用。从中非常清楚的一点是,可能性是——再次强调,除了 Anthropic 之外没人真正知道——像 Sonnet 这样的模型大概有 1.3 到 1.5 万亿总参数,并且可能有 1000 亿以上的活跃参数,对吧?然后 Opus 可能处于 1.5 到 1.8 万亿范围,活跃参数大概在 1500 亿到 1800 亿之间,取决于我们讨论的数据类型,是 FP8 还是 FP4 等等。
I've seen a more empirical one. I'll send it to you because that blog post alone played a large part in our decisions architecturally we took for the sovereign model. One of the things that was very clear from that is that the likelihood is, and again no one really knows outside of Anthropic, that something like Sonnet is in the, I believe, like 1.3 to 1.5 trillion total params, and has probably 100 plus billion active, right? And then an Opus is probably in the 1.5 to 1.8 range and probably has 150 to 180 billion active, depending on the DT type we're talking about, whether it's FP8 or FP4, whatever.
那 Fable 呢?
What about Fable?
它上线时间不够长,我没来得及尝试。因为我本来想做的是把那篇博客文章给 Fable,然后指向一个端点,说,好,对自己做分析,告诉我你有多大。但没走到那一步,因为我认为它上线时间不够长。
That it wasn't up for long enough for me to try to, because I basically what I wanted to do is give Fable that blog post and then point it at an endpoint and be like, right, do the analysis on yourself and tell me how big you are. Never got that far though because I don't think it was up for long enough.
不过那是一个巨大的提升,不是吗?我看到——我相信你在 X 上也看到了——像 10 万亿这样的数字在流传。我不知道这是不是真的。
It was a ridiculous uplift though, isn't it? I saw, I'm sure you saw on X, like numbers like 10 trillion being knocked around. I don't know if that's true or not.
我不知道。我认为它显然比其他模型更大,活跃参数也更多。但我无法猜测。我真的不知道。但我认为那篇文章的要点是 Opus 大约在 1.5 到 1.8 万亿参数,活跃参数约 1500 亿。显然,其中涉及大量的算法和数据工作。但我认为,如果你至少不匹配那种架构,那么你已经很难达到那个天花板。我认为一个例子——这比我将要提出的基本论点要复杂得多——但如果你看看 DeepSeek V4 Pro,总参数 1.6 万亿,记不清确切数字,但活跃参数大约在 300 亿到 500 亿之间。
I have no idea. I think it's obviously bigger than the other ones and obviously has way more active parameters. But I couldn't speculate. I genuinely don't know. But I think the net of that article was that Opus was in the region of 1.5 to 1.8 with around 150 billion active. Obviously, there's a lot of algorithmic and data work that goes into it. But I think if you don't at least match that architecture, then you're already going to struggle to sort of reach that ceiling. And I think an example of this, and it's way more nuanced than this fairly basic argument I'm going to make, but if you look at like a DeepSeek V4 Pro, right, 1.6 trillion total, can't remember the exact number, but it's going to be in the region of 30 to 50 billion active.
对。
Right.
我的观点是,这些架构决策的原因主要是为了模型的推理。当然,你有那 1.6 万亿的顶级参数很好,但如果没人能运行它,因为他们需要两个节点的 B300 才能把它装进内存并以不错的 TPS 运行,那么实际上有多少人能利用它呢?
And my view is that the reason for these architectural decisions is largely for inference of the model. Like, sure, great that you have those top-line 1.6 trillion parameters, but if no one can run it because they need two nodes of B300s just to be able to fit it into memory and actually run it at decent TPS, then how many people can actually take advantage of that?
我认为,无论是中国的开源模型,还是欧洲或美国的开源模型,之所以未能达到闭源模型的性能,其中一个原因就是那个务实的问题:实验室拥有大量 GPU,并且有足够的内部需求来确保这些 GPU 的利用率达到一个他们不太担心闲置的水平;而开源社区则没有同样的理由。如果你在自己的硬件上运行这么大规模的模型,而你的组织并没有充分利用它,你就会担心那些 GPU 闲置会花掉多少钱。
I think that is one of the reasons that whether it be Chinese open source or European or American open source has not reached closed source performance is like there is that pragmatic question of okay well the labs have a huge number of GPUs and they have enough inbound demand to make sure those GPUs are utilized to a level where they're not that worried about having them up whereas the open-source community doesn't really have the same argument. If you're running a model of that scale on your own hardware and it's not really being utilized that much by your organization, you're going to worry about how much money you're spending on just those GPUs being idle.
关于第一点,从架构上讲,我认为总参数量显然很重要。我认为活跃参数量也非常重要。最后一点是数据。我认为实验室在数据方面显然有一些优势,既因为他们能够从数据经纪人那里采购大量数据,也因为他们拥有非常成熟的内部数据功能。我并不是说其他实验室没有这些,但他们在投入规模或成熟度上肯定不在同一水平。
So on that first point, architecturally, I think that overall parameter count is obviously important. I think active parameter count is also incredibly important. And then the last bit is data. I think that obviously the labs have some element of edge in data both because they are able to procure so much from the brokers who sell it and also because they have internal data functions which are very mature at this point. I'm not saying the other labs don't have that, but they're definitely not at the same scale both in terms of spend or maturity particularly.
我认为在某种程度上——当然这不完全正确——但某种程度上,预训练语料库,一旦你达到 30 万亿 token 的规模,还会有那么大的差异吗?基本上大家都拥有大致相同的内容。那时我们压缩了互联网的大部分内容。中期训练也类似。我认为后训练一直是并且仍然是这些实验室最有趣的领域之一。我曾与一些实验室合作处理后训练数据,因为我们显然非常擅长编码,所以看到他们如何采购数据,甚至包括他们用于原始数据的格式,以及他们在提出需求时感兴趣的不同用例,都很有趣。我认为这是他们差异化竞争的关键领域之一:拥有非常好的后训练数据,并且能够以惊人的规模运行强化学习。
I think to an extent, and this is definitely not true, but I think to an extent like the pre-training corpora, is there going to be that much difference once you're in like the 30 trillion token range? Okay, you kind of all have roughly the same stuff. We're compressing most of the internet at that point. Mid-training, similar story. I think post-training has been and remains one of the most interesting areas for these labs. Having worked with some of the labs on post-training data, because obviously we're very good at coding, it's been interesting to see how they've been procuring even the formats that they've been using for the raw data and the different use cases that they're interested about when they're putting out requests for we want this, we want that, we want the other things. I think that is one of the key areas they're differentiating: having really good post-training data and also just being able to run RL at ridiculous scale.
稍作停顿。智能体每天都在变得更聪明,但即使是最聪明的智能体,如果没有正确的上下文和工具,也会卡住。这就是 Notion 的用武之地。随着最近自定义智能体的推出,Notion 成为了团队和智能体并肩工作的协作式 AI 工作空间。而现在,他们的新开发平台正在将这个工作空间转变为开发者可以构建的基础设施。这正是我运营 MLST 的方式。整个节目都在 Notion 中。我的嘉宾、发布日历、商业方面,一切都在那里。但变化在于,它现在具有智能体能力。我只需与我的智能体对话。它可以是 Claude 或任何智能体框架。然后它通过 MCP 或 CLI 与 Notion 通信,事情就完成了。然后我可以在手机上访问它。这绝对是一个游戏规则改变者。
Quick pause. Agents are getting smarter every day, but even the smartest agents get stuck without the right context and the right tools. That is where Notion comes in. With the recent launch of custom agents, Notion became the collaborative AI workspace where teams and agents work side by side. And now their new development platform is turning that workspace into infrastructure developers can build on. Now, this is exactly how I run MLST. The whole show lives in Notion. My guests, the publishing calendar, the commercial side, everything is in there. But what's changed is that it's now agentic. I just talk to my agent. It can be Claude or any agentic harness. And then it then talks to Notion via the MCP or the CLI and it's just done. And then I can access it on my phone. It's an absolute game changer.
我们待会再谈强化学习,因为我知道你对强化学习训练非常重要有自己的看法。但关于你刚才说的,有几点。首先,这仅仅是推理速度的权衡,还是你认为在因式分解方面,混合专家模型实际上有一些有益的优势?
We'll get to the RL bit because I know you've got an opinion that RL training is really important. But there's a few things on what you just said. First of all, is it just a trade-off for inference speed or do you think there's actually some beneficial advantage in terms of factorization having an MOE?
我认为,至少我的观点是,如果我们能拥有万亿参数的密集模型,我们就会这么做。它会严格更好。我认为它只会更好。但同样,这是将其推向逻辑极端:你需要什么来推理它?你需要什么来实际运行这么大尺寸的模型?我认为一个有趣的例子是,虽然这不公平,但如果你拿小规模的模型比如 GPOSS 12B,我相信它有 50 亿活跃参数,我记不清了,大概是那样。然后你拿 DevStrol 2 123B,我相信它的架构类似于 Llama 70B,所以它是完全密集的。我不知道你是否之前交替使用过它们。DevStrol 感觉比 GPOSS 模型好得多,对吧?而且它们在架构上非常不同,尤其是在注意力机制上。但根本上,我认为很大一部分原因在于,一个模型每个 token 有 1200 亿活跃参数,另一个只有 50 亿。去年我们向客户部署了这两个模型。在感觉上以及他们在运行编码智能体时得到的结果上,差异是天壤之别。
I think at least my opinion is if we could have trillion parameter dense models, we would. It would just be strictly better. I think it would just be better. But again, that is taking it to the logical extreme of what would you need to inference that? What would you need to actually run a model of that size? I think an interesting example of this is, and it's not really a fair fight, but if you take at small scale something like a GPOSS 12B, right? That is, I believe it's 5 billion active, I can't remember, it was a while ago, but it was something like that. And then you take a DevStrol 2 123B, which is I believe like a similar architecture to Llama 70B, so it's fully dense. I don't know if you've used them back to back before. DevStrol feels so much better than the GPOSS model, right? And they are architecturally quite different, particularly in the attention mechanism. But fundamentally, I think that a huge part of that just comes from the fact that one you have 120 billion active parameters per token, the other you have five. And we've deployed both to customers last year. The difference was night and day in terms of how it felt and what they got out of the coding agent when they were running it.
哦,有意思。是的,因为关于其他方面,你谈到了数据、预训练、算法方面。我的意思是,也许算法方面已经趋同了,只是因为我们现在有了这个吸引子盆地,有内核优化器和整个生态系统,也许它已经趋同了,但数据方面很有趣,对吧?因为一定有大量的工程工作,而对于大语言模型来说,有点像“魔法词”是什么?如果你以正确的方式提出问题,它就有那种表征摩擦,并产生有趣的结果。所以我猜 Anthropic 做了大量的数据整理和修剪。但 Anthropic 做的另一件事是,他们有 Claude Code。他们有生态系统,并且每天都有轨迹数据源源不断地流入。
Oh, interesting. Yeah, because on the other stuff, you were talking about the data, the pre-training, the algorithmic stuff. I mean, maybe the algorithmic stuff is kind of converged only because we now have this basin of attraction where there's kernel optimizers and entire ecosystems around this and like maybe it's kind of converged, but the data thing is interesting, right? Because there must be an insane amount of engineering and with LLMs, it's a little bit like what's the magic word? Like if you frame the question in the right way, it has that representational friction and it does interesting things. So I'm guessing Anthropic do a whole bunch of data curation and pruning. But another thing that Anthropic do is they have Claude Code. They have the ecosystem and they have trajectories coming in all day every day.
正是如此。这些轨迹数据,因为你知道有一个重要的事情:重要的不是你最终到达哪里,而是你如何到达那里。我认为软件工程不是写代码。它实际上是创造过程。创造心理抽象、做实验、完善这些抽象、与团队分享。所以这个过程,迭代地运行代码、测试、完善抽象,而 Anthropic 可以访问大量这样的数据。这是多大的优势?显然,这是一个巨大的优势。我不知道他们的服务条款怎么说。我并不是说他们在所有这些数据上训练。我不知道他们是否这样做。
Exactly. And those trajectories, because you know you've got this big thing which is that it's not about where you end up. It's about how you got there. And I think software engineering is not about writing code. It's actually just creating process. Creating mental abstractions, doing experiments, refining those abstractions, sharing them with the team. So this process, iteratively I'm running code, I'm testing things, I'm refining my abstractions, and Anthropic have access to so much of that data. How much of an advantage is that? Obviously, it's a huge advantage. I don't know what their terms of service say. I'm not saying that they train on all that data. I don't know whether they do or not.
拥有轨迹数据的一个有趣之处在于——显然我们在内部会收集自己使用 Cosine 的轨迹,因为部署方式的原因我们拿不到别人的数据,但至少对于使用 Cosine 的 Cosine 员工来说,我们确实能获得相当数量的轨迹。虽然远不及 Anthropic 的量级,但轨迹对我们最有用的地方在于观察用户如何向模型提问。当你构建强化学习数据集时,可以设计出结构完美、表述清晰的问题陈述,明确告知预期输出。但现实是,用户根本不会那样提问。他们会说‘去你的,这不行,你他妈傻吗?为什么搞成这样?’而现实是,如果模型在真实世界中将面对这样的分布,那么训练数据也需要与之匹配。我猜测 Anthropic 从他们获得的轨迹中得到的很多优势显然是有帮助的。我不认为他们能规范地知道每条轨迹的好坏——这不是强化学习那种明确知道轨迹好坏的情况。当然,你可以做一些判断,观察用户的反应,比如‘嗯,此时用户可能对结果满意了’。但根本上,观察用户在对话中说了什么、怎么说的,这极其有用。即使在我们这个规模,关于这些数据对整个开发生态系统的代表性如何,那是另一个问题。但我们确实对工程师如何与这些产品互动、如何回应等有了大致了解,这有助于在训练时让数据更贴近现实。
Certainly one of the interesting things that having trajectories means, and something that obviously internally we collect our own trajectories of our own use of Cosine. Obviously we don't have anyone else's because of the way we deploy it, but for Cosine employees using Cosine at least, we do still get a fair number of trajectories. Nothing like the order of magnitude that Anthropic get, but the most useful thing for trajectories for us is seeing how users prompt models. So there's one thing when you're putting together RL datasets and you can come up with these beautifully formed problem statements that get fed into the model, that are very well structured and well formed, and they're very clear about what the expected outcome is. The reality is that users don't prompt models in that way at all. They're like, 'F you. It doesn't work. Why the hell are you dumb? Why have you done this, that, and the other thing?' And the reality is, if that's the distribution that the model's going to be exposed to in the real world, then you need training data that looks like that. And I think that I would speculate that a lot of the alpha that Anthropic get out of their trajectories they get back is obviously helpful. I don't think that they canonically know it's not like there is a grader attributed to that trajectory. It's not an RL thing where you know that this trajectory was good. Absolutely, I'm sure you can do some judging stuff and you can see how the users responded to be like, 'Okay, well at this point the user was probably satisfied that the job was done.' But it's also fundamentally seeing what is the user saying during the conversation, and how are they saying it. That is incredibly useful. And even at our scale, there's an argument to be made about how representative it is for the entire development ecosystem. That's a different question. But we do have an idea of, generally speaking, across months and months of usage, how do engineers interact with these products and how do they respond and all these things, and that's very useful to make things more realistic at training time.
没错,100%。因为你在网上谈到过‘垃圾代码’。这也是我特别喜欢的话题之一。所以,VI 编程在某种程度上是有效的——它通常能完成你要求的事情,测试也会通过,但实际上你在构建一个意大利面条式的怪物,你在好面条之后又扔进更多坏面条,本该一行代码修复的问题,它却给你加了 200 行。还有整个理解债务的问题。我们可以聊聊这个。但我们真正想要的是——神经网络在一定程度上能做到——它们学习代表某种抽象结构的统计不变性,这有助于泛化。显然我们希望它们抽象地学习问题,从而泛化到从未见过的新问题。所以思路是,如果我们做得足够好,就能捕捉思维过程,进而捕捉泛化。大致就是这样。你怎么看待垃圾代码?
Yeah, 100%. Because you've spoken about slop online. It's one of my pet topics as well. So, VI coding works in the sense that quite often it does the thing that you tell it to do, and the test will pass, but actually you're building a spaghetti monster and you're throwing more bad spaghetti after good spaghetti, and instead of doing a one-line fix that it's supposed to do, it'll give you 200 extra lines. There's the whole understanding debt thing. We can talk about that. But what we actually want to happen, and neural networks do this to a certain extent, they learn statistical invariances that represent some kind of abstract structure, and that helps them to generalize. Obviously what we want them to do is to learn problems in the abstract so that they can generalize to new novel problems that they've never seen before. So the idea is we capture the thought process and then we capture that generalization if we do it well enough. That's the rough idea. How do you think about slop?
垃圾代码也是我特别讨厌的东西。任何使用 Claude Code、Open Code 或其他智能体式编码工具的用户都会反复看到,这既是模型风格问题,也是垃圾代码问题。你的问题一针见血:你说测试通过了,对吧?技术上功能正确吗?当然。但代价是什么?这才是关键问题。我认为从根本上说,当你思考这些模型如何被训练来做软件工程时——这也是我们构建 Outpost 时的领悟,我相信你读过我们如何做的博客——在这个领域,通常(我们并不深入了解大实验室的做法),当你训练模型提升软件工程能力时,你会有一个软件工程问题、一个问题陈述、某种测试(通常是单元测试,但并非总是,取决于任务类型),该测试在之前状态失败,在之后状态通过。然后你把问题交给智能体,它执行轨迹、展开,最后你运行单元测试,如果通过,就认为它做对了,给予奖励并更新权重。显然问题有多个层面,但根本在于,它可能用最疯狂的方式完成任务——无论是运行了不安全的命令,还是写出了与应有代码相比完全垃圾的代码。当你仅基于正确性给予奖励时,所有这些都会被强化,无论你是否愿意。因此我们在强化学习过程中做了并持续在做一些事情来缓解这个问题。具体到垃圾代码,有几个关键点。一是正确性优先于一切:如果问题做错了,无论方式多优雅,都不给奖励。但除此之外,我们还有针对我们讨论的这些问题的其他奖励。另外,在某些情况下(并非全部),我们有参考实现。如果你从宽松许可的开源仓库中提取一个 PR 作为种子数据,你有人类编写的原始补丁,你可以进行一定程度的比较:比较智能体写的和人类写的,看看是否合理。如果长了 500 行,可能就不合理,也许不该给全额奖励。所以在纯奖励层面可以做些事情,但我们也在算法层面做工作,这个领域越来越关注轨迹中的信用分配。我认为当前强化学习的一大问题——我绝不是第一个这么说的人,Andrej Karpathy 一年多前就说过——从根本上说,你有一个可能长达 256,000 个 token 的展开(在某些极端情况下),最终根据模型的行为得到一个 1 或 0。
So slop is also one of my pet hates. When you are, obviously any of your viewers who use Claude Code or Open Code or whatever agentic coding harness they like, see ad nauseam, it's both a combination of model vibe and also slop problems. And you hit the nail on the head in your question, which is when you said that okay, the test passes, right? Is it technically functionally correct? Sure, right. But at what cost is the big question. I think fundamentally, when you think about how these models are trained to do software engineering, and this is the realization that we sort of had when we built Outpost, and I'm sure you've seen the blogs about how we did it. Within that realm, normally, and again we don't have that much insight over how this works in the big labs, but normally when you're training a model to be better at software engineering, you have some kind of software engineering problem, you have a problem statement, you have some kind of test, often a unit test but not always depending on the task type, that is either failing in the prior state and passing in the after state. And then you give the agent the problem. It goes and does its trajectory, its rollout, and then at the end you run the unit test, and if it passes, then okay, you got it right. Great. You get a reward and the weights update. Obviously the problem with that is a few fold, but fundamentally it could have come up with the most insane way of doing something, whether it be in terms of commands that it ran that could have been unsafe, or code that's absolute crap compared to what it should have actually done. And all of those things get reinforced whether you like it or not when you give that reward based on purely correctness. So there are a number of things that we have done and are continuing to do in the RL process to try to ameliorate this. In terms of slop specifically, there are a couple of key things. One is that correctness gates everything else. So if you get the problem wrong, regardless of whether you did it in an elegant way, you don't get rewarded. But beyond that, we do have other rewards that target the exact things that we've been talking about. And also in some cases but not all, we do have reference implementations for these things. If you are using a pull request from a permissively licensed open source repo as some seed data, you do have the original patch the human made, and you can actually do some level of comparison: let's compare what the agent wrote, let's compare what the human wrote, and does this seem reasonable? If it's 500 lines longer, probably not, maybe we shouldn't give the full reward for this. So there's stuff you can do on the pure reward level, but we're also doing stuff on an algorithmic level, which is being looked at more and more in the space, which is to do with credit attribution in trajectories. One of the big problems with RL as it stands, I think, and I'm by far not the first person to say this, I think Andrej Karpathy said this like over a year ago, but fundamentally this notion of you have a rollout of maybe 256,000 tokens, right? In some extreme cases. And that culminated in like a one or a zero depending on what the model did.
而我们现在在很多情况下说的是,好吧,所有这些词元在引导我们得到那个答案时都被同等加权。
And what we're saying at the moment in many cases is okay, all of those tokens are equally weighted in getting us to that answer.
仔细想想这很疯狂,因为显然不是这样。在很多情况下,轨迹中会有一些小决策或重要决策,它们是岔路口,可能导致坏结果,但模型选择了正确的方向并走到了终点。我们和这个领域的其他人现在试图做的是:如果你能找到那些高熵词元的范围或做出决策的地方,找到这个问题就解决了一半。然后,一旦你知道这是重要事件,判断它相对于最终结果是好是坏又是另一回事。但如果你能做到这一点,你的强化学习就会显著更高效,因为你不再依赖——我常用的类比是:假设你在准备英语 A 水平考试,你写了一篇 2500 词的论文给老师,老师只给了你一个 B。你不知道为什么是 B。然后你不得不写几百篇论文,有些得 A,有些得 B,有些得 C,最终你会发现当你这样做时更容易得 A,所以你认为这可能是值得强化的好做法。但如果老师圈出句子说‘这是垃圾,别这么说’,那就容易多了。这基本上就是我们试图引入强化学习的原理,因为这样你能从已有的算力中获得更多性能。同时,你也在教模型学习真正重要的东西,而不是填充内容。
Which when you think about it is insane because that's clearly not true. In so many cases, there will be small decisions or important decisions in a trajectory that were forks in the road that could have resulted in a bad outcome, where the model decided to go down the right one and ended there. What we and others in the space are trying to do right now is: if you can find out what those ranges of high entropy tokens or places where a decision was made, finding that is half the problem. Then figuring out once you know that this is an important thing that happened, determining whether it was good or bad relative to the final outcome is a different story. But if you can do that, your RL gets significantly more efficient because you're not relying—the analogy I always come up with is: say you were doing your English A level and you wrote a 2500-word essay for your teacher, and the teacher just gives you a B. You don't know what made it a B. Then you have to write hundreds of essays, and you'll get an A on some, a B on others, a C on others, and eventually you'll figure out that when you do this, you tend to get an A more, so you think this is probably a good thing to reinforce. But it would be far easier if the teacher circled the sentence and said, 'This is rubbish. Don't say this.' That is fundamentally the principle we're trying to bring into RL across the board, because you get so much more performance out of the flops that you have. Also, you're teaching the model to learn the things that are actually important and not just the filler.
我知道。机器学习的好处是它只是在抽象山的低处进行泛化。所以,从非常表面的统计泛化来看,机器学习的坏处是它会泛化。所以我完全同意你的看法。我们在基准测试和机器学习方面存在巨大问题。我们痴迷于一次通过、五次通过准确率,似乎不关心可靠性、一致性、安全性、抽象形成——这显然是最重要的。François 做了 ARC 挑战,不幸的是它们可以被暴力破解,现在他有了一个新版本,暴力破解极其困难,你必须形成抽象才能获得任何好的性能。所以你的意思是,有一种新形式的强化学习,可能不同于 DeepSeek 类型的强化学习,它不是仅仅因为得到正确答案而获得奖励,而是迫使它形成可重用的抽象,并爬得更高。
I know. The great thing about machine learning is that it just generalizes low down the abstraction mountain. So, from very superficial statistical generalizations, the bad thing about machine learning is it generalizes. So I completely agree with you. We have a huge problem with benchmarks and machine learning. We are obsessed with pass one, pass at five accuracy, and we don't seem to care about reliability, consistency, security, abstraction forming—which is clearly the most important thing. François did the ARC challenge, and unfortunately they were brute forceable, and now he's got this new version which is so difficult to brute force you have to form abstractions to get any kind of good performance on it. So you're saying there's a new form of RL, perhaps different from the DeepSeek type of RL, which is rather than just being rewarded for getting the right answer, you're forcing it to form reusable abstractions and go higher up the mountain.
是的,简而言之,你解释得比我好得多,但没错,这基本上就是我们想要达到的目标。好处之一是,我认为我们在这里讨论的本质上就是轨迹内的信用分配。这可以很好地移植到当前流行的一系列不同的强化学习算法中。你可以将其用于 GRPO,也可以用于 GSPO 以及该算法的所有不同变体。但根本上,拥有一种严格且重要的是不带偏见的方式来指出智能体完成的工作范围,并说‘这是好的,这是坏的’等等。对于你在那些不同轨迹中计算的优势,如果这些优势被不成比例地归因于这些范围,而不是轨迹中的其余词元,那么你的权重更新将更有针对性地使这些特征更频繁或更不频繁地出现。
Yes, in short, you've explained that far better than I did, but yes, that is essentially what we're trying to get to. One of the nice things is that I think essentially what we're talking about here is credit attribution within a trajectory. That ports quite nicely to a bunch of different RL algorithms that are in vogue at the moment. You can use it with GRPO, you can use it with GSPO and all the different flavors of that algorithm. But fundamentally, having a rigorous and importantly unopinionated way of pointing at ranges of work that an agent has done and being like 'this is good, this is bad' and so on. For the advantage that's being calculated across those different trajectories that you've made, for that advantage to be attributed to those ranges disproportionately to the rest of the tokens in the trajectory means that your weight updates will be more targeted at making those characteristics either appear more frequently or less frequently.
我们是否仍然存在认识论问题?因为机器学习的一个问题是它并没有真正拥有真与假的概念。所以我们可以从代码执行反馈中获得反馈。我们可以对工程师正在做的实际过程进行一大堆抽象视角的观察。但我们是否仍然存在这个差距,即我们并不真正知道它是否正确?
Do we still have an epistemic problem? Because the one problem with machine learning is it doesn't really have the notion of true and false. So we can do feedback from code execution feedback. We can do a whole bunch of abstract lenses on actual processes that engineers are doing. But don't we still have this gap that we don't really know whether it was correct or not?
这是最困难的事情之一。特别是,编码之所以成为一个热门用例,原因之一是在某些情况下你有可验证的奖励。我认为我刚才提到的一点是,我们正试图将一些更模糊的、与品味相关的东西变得可验证,但采用更浮动的尺度。对于法律等领域,对于其他不可验证的领域,这要困难得多。我认为这是我们在其他行业没有看到像编码、数学或物理学那样革命的关键原因之一,因为你不能静态地编译一些法律并看看是否得到 1 或 0。我不知道答案是什么。也许有新的方法来编码这些领域,或者你只是把所有东西都放在分布内,我认为这就是现在正在发生的事情。如果你只是让所有东西都在分布内,并针对每个用例,并且你有一定水平的人类评判者或经过人类训练或良好提示的 LLM 评判者,你就能接近目标。但我仍然认为,在许多情况下,你需要在训练时列举大量这些问题集,并确保模型擅长它们,因为否则全面泛化的梦想已经被证明并不那么存在。
It's one of the hardest things. Particularly, one of the reasons that coding has taken off as a use case is because you have verifiable rewards in some cases. I think one of the things I've just said is that we are trying to bring some of the more fluffy taste-related things and make them verifiable but on a more floating scale. For things like law, for other non-verifiable domains, that is way harder. I think that is one of the key reasons we haven't seen the same revolution in other industries as we've seen in coding, math, or physics, because you can't just statically compile some law and see whether you get a one or a zero. I don't know what the answer to that is. Maybe there are new ways of codifying those domains, or maybe you just bring everything in distribution, which I think is what's happening these days. If you just make everything in distribution and you target every use case, and you have some level of human judge or good LLM judge that has been either trained or well-prompted by a human, you get close. But I still think that in many cases you need to enumerate a lot of these problem sets at train time and make sure the model is good at them, because otherwise the dream of generalization across the board has already been shown to not really exist that much.
我知道,但我们正处于一个非常有趣的时代,因为每个新模型都变得更好,但我们希望系统能用更少的资源做更多的事,这就是你所说的。我们希望它们获得这些抽象。这真是一个奇怪的局面。你认为我们能否达到一个可以让人脱离循环的地步?因为现在我认为人工智能进步的基础是,你有非常有才华的人类,他们知道如何提问,他们有品味,他们朝着正确的方向前进,并且存在这种良性的共同创造循环。这非常好。
I know, but we're in such an interesting time because every new model comes out and it's so much better, but we want to have systems that do more with less, which is what you're saying. We want them to acquire these abstractions. It's a really weird situation. Do you think we'll ever get to a point where we can remove the human from the loop? Because right now I think the basis of AI progress is that you have very talented humans and they know how to ask the question and they have taste and they go in the right direction, and there's this virtuous co-creation cycle. It's very good.
所以你知道我们现在已经进入了不是“氛围编码”,而是智能体式工程的领域。这非常令人兴奋。但你认为它有可能在没有人类的情况下完成吗?
So you know we're now in the realm of it's not vibe coding, it's agentic engineering. It's very exciting. But do you think it could ever be done without humans?
是的,我认为可以在没有人类的情况下完成。但我觉得我们离那一步还很远。
Yes, I think it can be done without humans. I don't think we're anywhere near there yet.
那会是什么样子?你指的是在定义明确的问题上吗?我不是在说风格迁移。比如 Anthropic 构建了一个 C 编译器,那是一个定义明确的问题。但如果我给你一个新颖的应用,只能模糊地指定它呢?
What would that look like? Do you mean in well-specified problems? I'm not talking about style transfer. For example, Anthropic built a C compiler, which is a well-specified problem. But if I give you a novel application and can only vaguely specify it?
在很长一段时间内,我们都需要人类,不是吗?如果你希望它质量好、可维护,并且真正具有高级工程师的写作风格。目前来说,是的,你绝对需要人类参与。可能是在事后说:“好了,这里有一大堆你需要修改的东西”,以及所有那些在思维链中你认为很好但实际上因为真实原因而并不好的设计决策。但我确实认为我们会达到那一步。我认为这将通过模型改进、创建真正像你所说的那种问题的强化学习环境和问题,以及 harness 工程这三者的结合来实现。
We're going to need a human for a long time, aren't we? If you want it to be good, maintainable, and actually in the style of something a senior engineer would write. For now, yes, you definitely need a human there. Probably post-hoc to say, 'Okay, here's the mountain of stuff you need to change,' and all the design decisions that in your chain of thought you thought were good but actually weren't because of real reasons. But I do think we will get there. I think it's going to be through a combination of model improvements, creating RL environments and problems that really look like the kind of things you're talking about, and also harness engineering. A combination of all those three things.
是的。我们稍后会谈到 harness 工程。但与此同时,我们有了“意大利面条怪物缓解策略”。我认为一个大问题是代码审查,对吧?
Yes. And we'll get to harness engineering. But in the meantime, we've got the spaghetti monster mitigation strategy. And I think one of the big problems is code review, right?
哦,是的,完全同意。
Oh yes, totally.
因为 AI 的“精神错乱”已经表现为理解这一点。所以我越来越不清楚发生了什么,这实际上对我保持能力和未来演进软件非常不利。所以要问正确的问题。
Because the AI psychosis has manifested as understanding that. So increasingly I become less aware of what's going on, and that's actually really bad for me maintaining my competence and for me being able to evolve the software going forward. So ask the right questions.
是的。
Yeah.
绝对正确。那么我们应该怎么做?因为现在我们生成了海量的代码,对吧?这是用更多 AI 来做审查的情况,还是我们仍然需要人类参与审查?
Absolutely. Exactly. So how can we do this? Because now we are generating ridiculous amounts of code, right? Is this a case of using more AI to do the review, or do we still need humans in the review?
我认为我们需要更多对 AI 产出的运行时验证。这看起来像是代码审查会有所演变。我认为它会更像是一种证明,证明它声称在做的事情实际上正在发生。显然,AI 阅读 git diff 是没有用的。我认为它可以捕捉到一些问题,我过去也见过它捕捉到问题。但有一个类比,虽然不是针对代码审查,但我觉得很好:在我们的网络安全扫描产品中,它类似于代码审查,因为它基本上是用一个智能体群读取整个代码库。我们看到的其中一个关键点是,它会遍历整个过程,在大代码库中捡出很多东西,比如“这可能是个问题,这可能是个问题”。我认为我们在那里面和 AI 代码审查中看到的一个问题是,如果你孤立地看那段代码和那个函数定义,它可能看起来相当可疑。但实际上,代码路径从未被命中,或者有另一个函数先被调用,改变了这个变量,导致它不会做你认为它做的事,或者有一个运行时环境变量意味着这不会发生。所有这些都只是静态分析;你做不到。我们在那个产品中缓解这个问题的最佳方法之一是我们所谓的“漏洞验证”,具体来说就是:好的,智能体群提出它的列表,在它到达你之前,每一个——我们在虚拟机中启动应用程序,在尽可能接近生产环境的运行时中,然后我们告诉智能体:“你已经看到了源代码,所以如果它存在漏洞,你应该能够找出如何利用它;你应该能够构造你的可怕 zip 文件来利用你认为存在的东西。如果你不能,那么你就把它从列表中移除,因为它显然是误报。”我们一直在做的是将同样的逻辑应用到 PR 上。我们不一定在寻找网络安全漏洞,而是在寻找:好的,你声称已经构建了这个功能。你有这个新屏幕,里面有这个表格或这个表单,做这个。当你点击这个按钮时,它应该在数据库中产生一个新条目,并出现在你期望的所有地方。显然你可以看 diff,大部分情况下能获得很多信息,但也要向我展示它确实在这么做。在它到达我之前,以某种合理的方式证明它已经完成了,否则你就会陷入这种情况。我今天早些时候刚和一个客户聊过,他们说当他们刚开始采用智能体式编码工具时,他们要么陷入巨大的代码审查积压中——没人喜欢代码审查,说实话没人真正享受——要么有人会说:“一分钟后,看起来不错,合并吧。非常感谢。”然后你就 YOLO 合并到主分支。那也很糟糕。所以我认为,如果你在审查时能看到代码,但也能看到规范化的证明,至少在主路径上,它声称在做的事情确实在发生,那么认知负担会小得多。另一半,至少在目前,是对你构建的一切进行全面的端到端测试。构建这些测试很痛苦,但一旦你有了它,它能避免很多问题。因为我相信,如你所知,你见过,我在个人项目中也见过,你“氛围编码”一个下午,构建了 10 个新功能,然后突然之前工作的五个功能停止工作了。为什么会这样?哦,你之前的抽象,我把它搞乱了,现在它不工作了。所以是的,这是防御性措施(如端到端测试)和主动减轻人们心理负担的结合,比如“看,这里有一个屏幕录制或一些截图,向你展示我认为它是什么。”
I think we need more runtime validation of what AI is producing. And what that looks like is that code review will evolve somewhat. I think it will be more like proof that the thing it says it's doing is actually doing that thing. Obviously, AI reading git diffs is not useful. I think it can catch things, and I've seen it catch things in the past. But one analogy, and it's not actually for code review, but one analogy that I can tell you that's really good, when it comes to our cyber security scanning product, which I guess is analogous to code review because it's reading basically a whole codebase using a swarm. One of the key things we saw with that was that it would go through that process and pick up so many things across large codebases, like 'this could be a problem, this could be a problem.' And I think one of the things we see in that and in AI code review is that if you just look at that code in isolation and that function definition, it can look quite dodgy. But in reality, the code path is never hit, or there is another function that's called first that mutates this variable, meaning that doesn't do what you think it does, or there's an environment variable that happens at runtime that means this doesn't happen. All of these things are just static analysis; you can't do it. And one of the best ways we mitigated that problem in that product was that we had what we call exploit validation, which is literally: okay, the agents, the swarm, comes up with its list of things, and before it makes it to you, every single one—we spin up the application in a virtual machine in a runtime, as close to production as possible, and we tell the agent: 'Well, you've seen the source code, so if it's vulnerable, you should be able to figure out how to get through it; you should be able to craft your horrible zip file to exploit this thing that you think exists. And if you can't, then you just take it out of the list because it's clearly a false positive.' And what we have been working on is applying the same logic but to just PRs. And we're not necessarily looking for cyber vulnerabilities, but we're looking for: okay, you have allegedly built out this feature. You have this new screen that has this table in it or this form that does this. And when you click on this button, it should result in a new entry in the DB and show up in all the stuff you'd expect. Instead of obviously you can look at the diff and for the most part you can get a lot of mileage out of that, but also just show me it's doing that. Prove that it's done that in some reasonable way before it even makes it to me, because otherwise you end up in this situation. I was just talking to a customer earlier today where they were saying when they first started adopting agentic coding tools, they were in this spot where either they would end up with an enormous backlog of code review, and no one likes code review—let's face it, no one actually enjoys that—or you get people being like, 'After a minute, looks good to me, merge. Thank you so much.' And you just get this YOLO merging into your main branch. And that's bad as well. So I think there is going to be way less cognitive burden if, in however you're doing a review, you can see the code, but you can also see canonical proof that at least on the happy path, the thing it's saying it's doing is actually happening. And the other half, at least in the present day, is also really comprehensive end-to-end testing of everything you build. And that is something that really sucks to have to build out, but once you have it, it saves you from so many problems. Because I'm sure, as you know, you've seen, I've seen in personal projects and so on, you vibe code for an afternoon, you built 10 new features, then all of a sudden the other five you had before stopped working. Why has this happened? It's like, oh well, the abstraction you had, I've just messed with it and now it doesn't work for that thing. So yeah, it's a combination of defensive stuff like the end-to-end stuff and also just proactively lifting mental burden from people by being like, 'Look, here's either a screen recording or some screenshots or whatever of me showing you that this is what I think it is.'
我知道。这真的很有趣。
I know. It's really funny.
当 Claude 删除了你的生产数据库后,它会说:“哦,你说得对。我很抱歉这么做了。顺便说一句,你对此无能为力。” 但没错,我的意思是,这又是一个对齐问题,对吧?因为如果我们积累了理解债务,功能描述本身也会受到影响,因为我们不再理解功能描述了。而且不仅仅是功能描述,还有意图、行为,以及描述系统的所有不同层面,比如用户故事等等。不幸的是,这些都是盲人摸象的不同视角,对吧?它们不一定与现实产生摩擦。你看到问题所在了:我们正在失去与系统本该做的事情的联系。在很多情况下,Claude 和模型正在编写功能测试,然后它们又在某种程度上黑掉自己的测试。“哦,没通过。好吧,我改一下测试。现在通过了。太好了。搞定了。” 你经常看到这种情况。
After Claude deletes your production database, it'll say, "Oh, you're right to point that out. I'm so sorry for doing that. There's nothing you can do about it, by the way." But yeah, I mean, this is another alignment problem, though, right? Because if we are accumulating understanding debt, the functional descriptions themselves are going to suffer from that because we don't understand the functional description anymore. And it's not just functional descriptions; there are intents, there's behavior, there's all of these different levels of describing a system, user stories and stuff like that. And unfortunately, these are different views of the blind elephant, right? They don't necessarily have friction with reality. And you see the problem here: we're just losing touch with what it's supposed to be doing. In many cases, Claude and the models are writing the functional tests and then they're kind of hacking their own. "Oh, it didn't pass. Okay, I'll just change the test. Now it passes. Great. Here we go." You see it all the time.
我知道。所以,这是一个非常棘手的问题,但也是我们需要解决的问题。对我来说,我认为很大程度上与范围界定和约束有关,至少可以缩小问题范围。但我们还是继续吧。你对智能体式工程有什么看法?你们有一个智能体式框架。如果我没理解错的话,几年前,你们实际上强迫每个人都开始使用它,因为你们真的想把它优化到极致。你之前提到了强化学习部分。所以,也许智能体式框架和强化学习之间存在某种共同进化。这个智能体式框架中重要的是什么?
I know. So, it's a very difficult problem, but one that we need to fix. And for me, I think a lot of it is to do with scoping and constraints, to at least cut down the sides problem. But we should move on. What are your thoughts on agentic engineering? So, you guys have got an agentic harness. If I understand correctly, a couple of years ago, you actually forced everyone to start using that because you really want to optimize the hell out of it. And you were talking about the RL piece. So, maybe there's some co-evolution with the agentic harness and the RL. What's important in this agentic harness?
我认为从广义上讲,智能体式框架随着时间的推移变得越来越不重要了。
I think that broadly, agentic harnesses are getting less important over time.
哦,有意思。为什么?
Oh, interesting. Why?
模型变得越来越好了。我认为证据是,现在一个模型可能只用 bash 就能完成几乎任何任务,虽然更慢、消耗更多 token,但它可能仍然能做到。这并不是说智能体式框架不重要,但随着时间的推移,价值来自哪里?它来自模型,而不是框架。在我们构建自己的框架时,我们已经构建智能体式框架很长时间了。我们拥有的第一个智能体式模型是 2024 年 1 月训练的微调版 GPT-4 Turbo。那时编码智能体框架还不存在。Claude Code 也不存在。这些东西都不存在。我们不得不自己构建一个。我们与模型的设计共生地完成了这一点,今天我们仍然这样做,因为紧密耦合两者确实能获得更高的性能。从根本上说,框架工程仍然是我们非常关心的事情。如今,我们实际上更关心效率,而不是其他任何东西。在一个 token 成本和 tokconomics(这是我今天第一次听到的词,糟糕的词)对尤其是我们销售的企业越来越重要的世界里,我们希望我们的框架使用尽可能少的 token。句号。这就是我们正在努力做的。
The models are just getting so good. I think the proof point is that a model can probably do with bash only basically any task these days, more slowly and with more tokens, but it can probably still do it. That's not to say that agentic harnesses aren't important, but over time, where's the value coming from? It's coming from the model and not from the harness. In the way we built ours, we have been building agentic harnesses for a very long time. The first agentic model we had was a fine-tuned GPT-4 Turbo model that we trained in January 2024. Coding agent harnesses didn't exist at that point. Claude Code didn't exist. None of this stuff existed. And we had to build one out. And we did that symbiotically with the design of the model, which is something we still do today because you do get way more performance out of tightly coupling the two. Fundamentally, harness engineering is still something we care a lot about. These days, we actually care more about efficiency than anything else. In a world where token costs and tokconomics—which is a word I heard for the first time today, awful word—tokconomics is becoming increasingly important to particularly enterprises who we sell to. We want our harness to use as few tokens as possible. Full stop. That's what we're trying to do.
是的,我们一会儿再谈这个,因为如果模型能够进行认知量化,你原则上可以分配一个 token 预算来完成某些事情。但即使在我们谈到那之前,我想反驳一下框架工程,因为我发现的一件事是子智能体是游戏规则的改变者。
Yeah, we'll get to that in a sec because if the models can do epistemic quantification, you could in principle allocate a budget of tokens to get certain things done. But even before we get there, I want to push back on the harness engineering because one thing I have found is sub agents are a game changer.
是的,绝对如此。这是因为很多问题对于大语言模型来说太复杂了,无法一次性完成。所以我确信你有过类似的经历:随着问题变得更加具体,随着你减少模糊性,熵降低,模型变得更好,而且当模型的上下文中腐烂更少时,模型会更好。所以发生的情况是,工程师在做了一系列工程之后,发现他们基本上将问题分解为智能体式子任务,然后他们有一个全新的智能体,这个智能体有清晰的规范,并且只做这一件事。
Yes, absolutely. And that is because a lot of problems are too complicated for an LLM to do in a single pass. So I'm sure you've had a similar experience: as a problem becomes more specified, as you reduce the ambiguity, the entropy goes down, the models get better, and the models are better when they have less rot in their context. So what happens is the engineers, after doing a bunch of engineering, find that they kind of decompose problems into agentic subtasks and then they have a fresh agent, and the agent has a clear specification and only does this one thing.
正是如此。但作为工程师,你在逻辑上所做的是将一个问题分解成更小的子问题。你让智能体进行编排,并且你不会在主智能体的上下文中引入腐烂。你如何看待这一点随着时间的推移而演变?因为现在这是一个相当手动的过程,但你认为——好吧,你也许有这个 swarm 的东西。
Exactly. But what you're doing logically as an engineer is you're factorizing a problem into smaller sub problems. You're getting agents to orchestrate and you're not rotting the context in the main one. How do you see that evolving over time? Because now it's quite a manual process, but do you think—well, you've got this swarm thing maybe.
哦,谢谢你提到这一点,因为这正是我打算回答的。对我来说,Swarm 是子智能体编排的逻辑极致。我们很幸运能够做到这一点,因为我认为显然 Claude Code 有工作流,Codex 也有子智能体。Swarm 顾名思义:一个 swarm 就是大量子智能体以分层方式同时运行。而且因为 Cosign 并不试图每天服务数亿用户,我们能够提供类似 swarm 的功能。而我认为如果 Anthropic 为 Opus 提供一个 swarm,即使是 Colossus Swan 也会耗尽 token。Swarm 自动完成了你刚才概述的内容。我的 Twitter 和 LinkedIn 上有一个视频,我拿我们的 Lumen outpost 模型(这是后训练的 Kimi K2.6——绝对不是 Opus 或 Mythos 之类的)问它:“我想让你给我构建一个机械手表编译器。” 我出生在瑞士,所以我有理由这么做。我基本上问了,实际上有人用 Fable 做了完全相同的任务。但本质上,我想让你用 Python 给我构建一个 SDK,让我能够用代码指定机械手表,因为我不懂钟表学,但我还是想能够做到。我希望它在物理上是一致的。我希望你使用某种物理引擎。然后我还想要一个 3D 查看器,能够看到这个东西运行。而且所有这一切都需要在现实生活中可行。你不能让东西相互交叉,那是不可能的,等等。这是开箱即用的 Kimi 无法做到的。不可能,甚至差得远。事实上,即使是 Gemini 3.5、Opus 和 5.5 也无法真正做到。
Oh, thank you for bringing that up because that was exactly what I was going to answer. Swarm to me is sub agent orchestration taken to the logical extreme. We're kind of lucky to be able to do it because I think that obviously Claude Code has workflows, and Codex has sub agents as well. A swarm is what it sounds like on the tin: a swarm is genuinely a lot of sub agents running at the same time in a hierarchical way. And because Cosign isn't trying to serve hundreds of millions of people a day, we are able to serve a swarm-like feature. Whereas I think if Anthropic had a swarm for Opus, I think even Colossus Swan would run out of tokens. Swarm does exactly what you've just outlined automatically. There is a video on my Twitter and LinkedIn of me taking our Lumen outpost model, which is post-train Kimi K2.6—definitely not an Opus or a Mythos or anything like that—and I asked it, "I want you to build me a mechanical watch compiler." I am Swiss by birth, so I have a reason to do this. I basically asked, and actually someone did this exact same task with Fable. But essentially, I want you to build me an SDK in Python for me to be able to specify mechanical watches in code because I don't know horology, but I want to be able to do it anyway. I want it to be physically congruent. I want you to use some kind of physics engine. And then I also want a 3D viewer to be able to see the thing running. And it all needs to be possible in real life. You can't have things intersecting each other, which wouldn't be possible, and so on. That is something that out of the box Kimi cannot do. There's no way, not even close. In fact, even Gemini 3.5 and Opus and 5.5 can't really do it.
但一旦你把它们放进一个 swarm(智能体群),对于 Cosign 来说,swarm 的样子是:最顶层有一个编排器,它把问题分解成子问题,交给所谓的“产品经理”或随便你怎么称呼的角色。我们叫它们子规划器,它们各自负责一个垂直领域。所以在这个任务中,你会有一个子规划器负责 SDK,一个子规划器负责 3D 查看器,一个子规划器负责写文档,等等。然后这些子规划器可以再委派给工人,工人层是扁平的,数量不限。对于那个问题,我们使用了子智能体,我觉得这比你在 Claude Code 会话中看到的要多。我想你那样做很快就会达到使用上限。但当你这样做时,是有可能一次性完成整个项目的。我这是在严重自相矛盾,因为我刚才说 harness(工具框架)不重要,但在这方面它们显然重要。
But as soon as you put them in a swarm, and for what swarm looks like for Cosign, you have one orchestrator at the very top. It breaks down a problem into sub-problems for basically product managers or whatever you want to call them. We call them subplanners, but they own verticals of this. So within that task, you would have had a subplanner to do the SDK, a subplanner to do the 3D viewer, a subplanner to write the documentation, and so on. And then those subplanners could delegate to workers, and they have a flat layer of as many workers as they like. For that problem, we use subagents, which I think is more than you tend to see in a Claude Code session and so on. I think you'd probably hit your usage limit pretty quickly that way. But when you do that, it is possible to do that entire project in one shot. And I am contradicting myself quite badly because I've just said harnesses don't matter, but in that respect, they do obviously matter.
哦,确实,对此非常兴奋。但这引出了一个问题:当你开始拥有大量智能体时,你会面临更多的理解债务和更少的交互性,因为对我来说,缺乏交互性本身就是理解债务的一部分。有时你想插话,想说“哦,你那里有点不对,我想改变这个智能体的行为”。很多人在构建这些智能体系统时发现,智能体会互相干扰,覆盖彼此的工作,陷入死锁。你们是怎么处理这些问题的?
Oh indeed, very excited about that. But it raises the question: when you start to have loads and loads of agents, you have more understanding debt and less interactivity because for me, the lack of interactivity is part and parcel of the understanding debt. And sometimes you want to interject, you want to say, 'Oh, you've gone slightly wrong there. I want to change what this agent's doing.' And what many folks have found when they build these agent systems is that the agents interfere with each other, they kind of overwrite each other's work, they go into deadlock. How are you dealing with all that?
这是个难题,我们也经历过所有这些。我们做的一件关键事情是,让用户能够与最底层的智能体交互。比如,如果你有一个工人智能体,它比顶层低两级,你实际上可以直接和它对话,这很重要。至于其他问题,比如互相踩脚,你可以给文件加上写锁,这样一次只有一个智能体能编辑它。你还可以在智能体使用某些东西时提供上下文,比如一个智能体正在读它刚刚读过的文件。我们在 harness 中加入了这样的机制:“好的,另一个智能体刚刚编辑了这个文件,所以如果你看到它和上次不一样,别惊讶。”所有这些都有帮助。它们不是万能药,但确实有帮助。它们让智能体不那么惊讶,不会说“哦,这是哪来的?我上次编辑时没有这个”。但根本上,这又回到了我之前说的:是的,你会运行这个东西,它会很快给你带来巨大价值,但你之后还是得仔细检查。现实是,“我不喜欢你做的抽象,我不喜欢你做这件事的方式”。之后肯定需要做一些清理工作。
So it's a hard problem, and we experience all those things. One of the key things we did is we gave the ability to interject to an agent at the lowest level. So say you had a worker that was two levels down from the top one, you can actually talk to that one, which is an important thing. With regard to other problems like treading on each other's toes, you can put write locks on files so that only one agent can edit it at a time. You also can provide context to agents when they are using something, like an agent is reading a file that it just read. We have stuff in the harness which says, 'Okay, another agent has just edited this file, so don't be surprised if you see it slightly differently than how you saw it last time.' All these things help. They are not a panacea, but they certainly help. They make sure the agent is less surprised when they're like, 'Oh, where did that come from? That wasn't in my last edit.' But fundamentally, it comes back to the point I made earlier: yes, you will run this thing and it will provide you a huge amount of value very quickly, but you are still going to have to comb through it afterwards. The reality is, 'I don't like the abstraction you've done, I don't like the way you've done this.' There is going to have to be some sweeping afterwards, I think.
是啊,你对记忆有什么看法?
Yeah, what are your thoughts on memory?
很难做好。
Very hard to get right.
好的,多说一点。
Okay, tell me more.
非常难。实际上和强化学习类似,因为记忆不在于目的地,而在于你如何到达那里。是的,记忆很难做好。我们尝试了很多不同的方法。从根本上说,我认为目前存在的每一种记忆方法都有点像 hack(权宜之计),对吧?它就像一个工具。在很多情况下,它就像一个向量数据库或某个知识片段的嵌入版本,但智能体很难知道何时去查询。智能体也很难判断某个信息是否足够有用而值得写入记忆。而且保持这些信息更新也很困难。我们遇到过很多情况:智能体在执行一个轨迹,当它做了一件真正正确的事情时,它使用了记忆,但记忆已经过时了,然后智能体就想,“哦,记忆说应该这样做”,于是它改变了方向,而你作为工程师就会说,“不,请别那样做”。我们内部正在研究一些东西,比如持续学习等,试图避免记忆只是一个工具,而是让它成为模型潜在空间中的一部分。那也非常难。但我认为这是一个比单纯工具更直观、更优雅的解决方案。在强化学习期间,这个工具也很难做好,因为它为奖励黑客、信息泄露、以及能够查询未来本不该访问的信息提供了巨大的犯错空间,尽管你可以设置各种防护措施。就是很难。
Very hard. So, similar thing actually to the RL, because memory is not about the destination, it's about how you got there. Yes, memory is very hard to get right. We've tried a bunch of different approaches. Fundamentally, I think every approach to memory that exists right now is a bit of a hack, right? It's like a tool. In many cases, it's like a vector DB or an embedded version of some bit of knowledge, but they're very hard for the agents to know when to query. It's also fundamentally quite hard for the agent to know whether something was useful enough to write to memory. It's also difficult to keep these things up to date. We've had many situations where an agent's been doing a trajectory, and when it's been doing something that was genuinely the right thing to do, it's used its memory, and the memory's been old and out of date, and then the agent's like, 'Oh well, the memory says you should do it this way,' and it changes its tack, and there you're an engineer, you're like, 'No, please don't do that.' There are things that we're looking at internally with regard to continual learning and stuff like that to try to avoid memory being a tool and for it to just be something that's in the latent space with the model. That is also very hard. But I think it is a more intuitive and elegant solution to it just being a tool. It's also a tool that's very hard to get right during RL because it is a huge surface area for foot gunnery in terms of reward hacking, in terms of leakage, in terms of it being able to query something from the future that it shouldn't have access to yet, despite all the guardrails that you can put in place. It's just hard.
是啊,没错。从某种意义上说,这是 AI 精神病的另一个领域,因为我写了一个记忆 CLI,我几乎可以说现在你甚至不需要向量数据库之类的了。你只需要一个倒排索引,比如 SQLite,因为模型非常擅长从不同方向提问。所以,是的,这是一个大问题。它需要知道去检索什么。那是个大问题,但它确实有效,但只对我有效,因为它又制造了一团碎片化的意大利面条。所以记忆 CLI 里又有一团意大利面条,但它效果非常好,但在组织层面效果不佳,因为我的意大利面条怪物和约翰的意大利面条怪物不兼容。所以如果我们能解决那个问题,而且你也在描绘一幅有趣的图景,关于我们如何一起优化三明治的不同层次。
Yeah, exactly. And in a sense, this is another area for AI psychosis because I've written a memory CLI and I would almost argue that now you don't even need vector databases and so on. You can just have an inverted index, you know, just a SQLite because the models are so good at asking in different directions. So, yeah, there is a huge problem. It needs to know to retrieve that. That's a big one, but it actually works, but it only works for me because it creates this fractionated spaghetti mess again. So, there's another spaghetti mess in the memory CLI, but it works really really well, but it doesn't work very well at the organizational level because my spaghetti monster doesn't play with John's spaghetti monster. So if we could solve that problem and you're drawing an interesting picture as well of how we can actually optimize the different layers of the sandwich together.
是的。我认为一旦有人做对了,你使用它时立刻就能感觉到。我还没见过任何一个实现——不管是 include code,说实话,不管是我们的还是 ChatGPT 的,还是其他任何——能让我真正觉得,“哦,这不只是一个 hack。这不只是 RAG。这是传统意义上 RAG 残留的唯一痕迹。”我确信某处一定有更好的方法。
Yeah. I think that as soon as someone gets it right, you'll just know when you're using it immediately. I haven't seen a single implementation, whether it be include code to be honest, whether it be ours or chat GPTs or any of them, where I'm truly like, 'Oh no, this isn't just a hack. This isn't just RAG. This is the one remnant of RAG that still exists really in the more traditional sense.' And I am certain that there is a better way out there somewhere.
哦,当然。但我认为你提到的另一件事是专业化而非通用化。对我来说,很多智能体工程就像是涌现出的专业化。也就是说,我们用一个大的智能体来结晶出一个小的智能体,去做我们正在做的特定事情。
Oh definitely. But I think another thing you've said is specialization not generalization. I mean for me, a lot of agentic engineering is like emergent specialization. So it's like let's take a big intelligence to crystallize a small intelligence to do the particular thing we're doing.
但时间快到了,最后一个问题是关于合成数据生成的。你们在这方面做了什么?
But as we're nearly out of time, final question is synthetic data generation. So what are you guys doing about that?
很多。我不知道五分钟内能讲多少,但我会尽力。过去我们在多个领域做过合成数据生成,现在为了主权模型,我们做得更多了,因为主权模型不能只擅长软件工程,它必须全面有用。这意味着我们必须擅长合成数据生成,特别是在强化学习领域,不仅仅是编码。但核心还是编码。过去我们通过一个非常酷且复杂的流水线来做这件事,这个流水线我们已经构建了大约一年半。编码强化学习的一个大冷启动问题是,几乎所有的 PR 或提交都没有内置的评分器,无论你使用开源仓库作为种子数据还是闭源数据都一样。
Tons. I don't know how much of it I can get through in five minutes, but I'll do my best. There are a number of areas where we have done synthetic data generation in the past. We're doing a lot more of it in the more forward-looking sense for the sovereign model, because the sovereign model can't just be good at software engineering. It has to be useful across the board. That means we have to get good at synthetic data generation, particularly in the RL realm, in things that aren't just coding. The bread and butter though is coding. The way we have done this in the past is with a very cool and sophisticated pipeline that we have built out over the course of about a year and a half now. One of the big cold start problems in coding RL, particularly if you're using open-source repositories as seed data, and even if you're using closed source it doesn't matter, is that nearly all of the PRs or commits don't have a built-in grader.
对。
Right.
那些有评分器的通常是 bug 修复。经典例子,如果你足够极客,就是 SWE-bench 风格的问题:有一个 git issue,有一个修复它的 PR,因为是开源的,所以添加了回归测试,这就是你的种子数据。不幸的是,现实世界并非如此,广义的软件工程也不是这样。这意味着要做好强化学习,你仍然需要能够完成不是 bug 修复的任务,而这正是工程师所做的大部分工作。而且,你还需要能够判断智能体是否做对了。我们在 Cosine 的做法是,我们采用实际完成的工作。我们不会凭空捏造问题让模型解决,因为如果模型能想出问题,它很可能也能解决。但我们采用实际解决过的问题,无论是功能开发、重构还是其他。我们合成的是 ground truth 或评分器,或者衡量某件事是否完成的方法。这有点像雷区,因为特别是强化学习,显然不是监督学习。所以从根本上说,我们需要以一种与原始实现不太耦合的方式来测试这些东西。众所周知,条条大路通罗马。在实践中,这意味着你需要一种足够与实现无关的测试方法,同时又要足够严格以检查功能正确性。这是一条非常微妙的界线。我们在这方面做了很多工作,建立了一个很长的流水线。流水线中有我们定制的后训练模型,它们变得很好,因为我们在很多地方做了大量手动标注。但这使我们能够拥有一个自主的流水线,对企业尤其重要的是,我们可以指向几乎任何编程语言、任何类型的任务、任何技术栈等等,然后说“好的,我想要这个问题集的强化学习数据”,然后我们就能得到相当一部分。这就是为什么对于我们的 Outpost 模型,当我们确定要让模型擅长的语言时,比如 Java、Fortran、C++,还有很多,我们使用那个流水线来收集这些语言的强化学习数据。在很多情况下,有些编程语言甚至没有测试套件。这时就变得非常困难,你需要在衡量智能体是否做对方面有点创意。比如 Verilog 和 SystemVerilog,你实际上需要在环境中运行所谓的合成器,才能判断这个芯片是否工作。
The ones that do are often bug fixes. The classic example, if you're nerdy enough to be in the space, is a SWE-bench style problem. You have a git issue, you have a PR that fixed it, and because it's open source, you have some regression test that was added, and that's your seed data. The real world doesn't look like that unfortunately, and software engineering in the broad doesn't look like that. Meaning that to do RL well, you still need to fundamentally be able to do tasks that aren't bug fixes, which is the vast majority of what engineers do. And you still need to be able to tell whether the agent got it right or not, broadly speaking. What we do at Cosine is we take real work that was done. We're not magicking up made-up problems for the model to solve, because if the model can come up with a problem, it can probably solve it. But we take real problems that were solved, whether it be feature work, refactoring, whatever it is. What we're synthesizing is ground truths or graders or ways of measuring whether that thing has been done. It is a bit of a minefield because particularly with RL, obviously it's not supervised. So fundamentally we need to be able to test these things in a way that isn't too tightly coupled to the original implementation. There are many ways to skin a cat, as we know. What that looks like in practice is you need an implementation-agnostic enough way of testing that's still rigorous enough to check functional correctness. It is a very fine line to tread. We have done a lot of work around it. We have a long pipeline built on it. We have custom post-train models that live inside that pipeline that have gotten good because we had to do a lot of manual labeling in places. But what it has allowed us to do is have an autonomous pipeline where, and this is particularly important for enterprises, we can point at essentially any programming language, any type of task, any stack, anything like that, and say, "Okay, I want RL data for this problem set," and we can get a good chunk of it. That is why for our Outpost model, when we were coming up with the languages we wanted to get the model good at, things like Java, Fortran, C++, I could go on, we used that pipeline to gather the RL data for these things. In many cases, in some programming languages, there aren't even test suites. That's where it gets really hard. That's where you have to get a bit inventive in terms of measuring whether the agent has gotten something right or not. Things like Verilog and SystemVerilog, where you have to actually run what's called a synthesizer in your environment to be able to tell whether this chip actually works or not.
EDA。
An EDA.
对,没错。幸运的是,我不运行这个流水线,是一个叫 Ben 的同事在负责,他比我懂得多。但大致上,这就是我们在软件工程领域变得非常擅长的原因,而且我们也在将其推广到软件工程之外的其他用例。
Yeah, exactly. Fortunately I don't run this pipeline. A chap called Ben does, and he knows far more about that than I do. But yeah, that is broadly within software engineering how we've gotten very good at it, and it is something that we are generalizing out into other use cases that aren't just software engineering as well.
非常酷。那么最后,你可能会说美国政府给了你们一个商业优势,因为现在构建主权 AI 比以往任何时候都重要。你们今年年底会推出这个模型。那么你认为你们能做到吗?另外,从供应链的角度来看,你们是否仍然面临风险?因为很多硬件都受美国控制。这对你们来说会如何发展?
Very cool. So in closing, you might argue that the US government have handed you a commercial advantage here because now it's more important than ever to build sovereign AI. You guys have this model coming out towards the end of this year. So do you think you're going to be able to do it? But also, are you still at risk from a supply chain point of view? Because so much hardware is controlled by America. How's this going to pan out for you guys?
天真地说,我认为如果硬件已经在英国,我不确定他们能对此做多少。模型将要训练的所有基础设施已经存在并在英国运行,因为我们很快就要开始训练了。事实上,楼上的实验已经在进行中。至于最近发生的事情是否因为我们的定位而成为一份礼物,绝对是,毫无疑问。这可能是自公司成立以来我最忙碌的时期,人们来找我们说:“好吧,我们现在意识到你们在做的事情非常重要。我们如何参与?”你提到的公司联盟每天都在壮大。来自这些公司、政府甚至普通公民的参与和紧迫感已经爆棚。我感到非常幸运,在某种程度上也很幸运,我们早早地就占据了有利位置。当然,我们也确实把自己放在了那个位置,但也要感谢唐纳德·特朗普。
Naively, I think that if the hardware is already in the UK, I don't know how much there is they can do about that. All of the infrastructure the model is going to be trained on already exists and is up in the UK, because we're doing it very soon. In fact, experimentation already is happening upstairs. In terms of whether it has been a bit of a gift what's happened recently given our positioning, absolutely yes, without a doubt. It has probably been the busiest I have ever been since founding the company, in terms of people coming to us saying, "Okay, we now realize what you're doing is really important. How can we be involved?" The consortium of companies that you read out is growing by the day. The involvement and the urgency importantly from those companies, from government, from just like citizens as well, has gone through the roof. I feel very fortunate and to an extent lucky that we were well positioned to take advantage of this early. Also, we did put ourselves in that position, but also thank you Donald Trump.
但我们怎么没预见到呢?因为对很多人来说,这真是一个“哦”的时刻。
How did we not see this coming though? Because it was such an 'oh' moment for so many people.
我认为我们确实预见到了。我认为很多人确实预见到了。我的意思是,我们当然预见到了,但总是被搪塞说:“哦,当然。好吧,我想那可能会发生。”但我个人没想到它会这么快发生。我也对 5.6 的消息感到惊讶,我相信你也看到了,它即将推出。即使对我来说,也有很大一部分我在想:“哦,天哪,那真是太糟糕了。”
I think we did though. I think a lot of people did. I mean, we certainly did, but it was always fobbed off as, "Oh, sure. Okay, I suppose that could happen." But I personally didn't expect it to happen as soon as it did. I am also freshly surprised by the 5.6 news that I'm sure you've seen as well, where that's going to be rolled out. There is even for me a huge part of me being like, "Oh man, that really sucks."
我想试试那个模型,但不知道现在还能不能。这也许就是我们现在的处境,除非我们和其他人努力以其他方式达到那种性能水平。我想这就是我们现在的任务了。
I wanted to try that model and I don't know whether I'll be able to now. And that might just be the existence for us now unless we and others do work to get that level of performance out in some other way. And that's like our job now I guess.
是啊。我第一次感觉自己像个二等公民,因为那边美国那些人用的 AI 比我的好。
Yeah. I mean for the first time I feel like a second-class citizen because those folks over there in America they've got better AI than I do.
他们比我跑得快,我真的很讨厌这一点。
They are going faster than I am and I really hate that.
相信我,这比任何人都让我愤怒。嗯,我们会尽一切努力实现这个目标。你提到怎么做?我们就是要让它发生。嗯,我们别无选择,只能让它发生。
And believe me that boils my blood more than anyone else. Um, and we are going to do everything we can to pull this off. Um you mentioned like how are you going to do this? It's like we are just going to make it happen. Um we have no choice but to make it happen.
请务必做到。
Please do.
是的,我们会尽一切努力让它实现。
Yes, we are going to do everything we can to make it happen.
我代表英国所有人,请一定做到。
On behalf of everyone in the UK, please do.
我们会尽力。谢谢。
We'll do our best. Thank you.
很愉快。非常感谢。
It's been a pleasure. Thank you so much.
非常感谢你的邀请。
Thank you very much for having me.