为什么专用 AI 可能击败通用神模型

Why Specialized AI Could Beat the God Model

阿姆贾德·马萨德 Amjad Masad · a16z 播客 · 2026-10-03 · 约 48 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Amjad 与 Alex 探讨专用 AI、模型多样性与开放基础设施如何可能胜过单一通用神模型。

Amjad and Alex discuss how specialized AI, model diversity, and open infrastructure could outperform a single god model.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 15)

全文 · Full transcript(中英对照)

引言与收购 Introduction and Acquisition

Host

你看到了 SpaceX 的 S1。就像,哦,30 万亿美元。就像,世界 GDP 是多少?100 万亿。

You saw the SpaceX S1. It's like, oh, $30 trillion. It's like, what is the world GDP? 100 trillion.

Amjad

Stripe 和 Open Router 都真的希望世界上有很多新公司。我们不希望每个人都成为一家巨型公司的一部分。

Both Stripe and Open Router really want lots of new companies in the world. We don't want everyone to be a part of one giant company.

Amjad

当模型变得更智能时,风险实际上会继续升高。然而,没有新的责任承担者。我们会慢慢意识到,我们曾经拥有确定性代码是多么美好。还记得计算机完全按照我们告诉它们的方式执行的日子吗?

When the models get more intelligent, the risk actually will continue to get higher. And yet, no one new is taking responsibility. We're going to slowly realize how good we've had it with like deterministic code. remember the days when computers did exactly what we told them to do.

Host

我们能否可预测地防止模型在训练过程中欺骗用户?一个足够大、足够强大的模型会突然停止欺骗和停止消极怠工吗?

Are we going to prevent the models from deceiving users during training runs predictably? Will like a model that's big enough and powerful enough suddenly stop deception and stop sandbagging.

Host

嗯,欢迎来到 Asz 播客。我们请到了 Replet 的 Amjad 和 Open Router 的 Alex。这是 Alex 在收购后做的第一个播客。所以我们非常高兴能请到你们两位。

Um, welcome to Asz podcast. We're here with Amjad of Replet and and Alex of Open Router. And this is the first podcast that Alex has done since the acquisition. So, we're really excited to have both of you.

Amjad

嗯。谢谢。很高兴来到这里。

Mhm. Thank you. Excited to be here.

Host

Alex,我们实际上就从这里开始吧,如果你能简要地分享一下,嗯,显然,你知道,这是一次大规模收购。嗯,Amjad 是投资者。我们当然也是投资者。嗯,最大的股东,但是,谁在数呢?嗯,Alex,你为什么不给我们讲一点背景故事,比如这样的收购到底是怎么发生的?比如你有一天收到 Patrick 的私信?呃,你知道吗?呃,你能分享些什么?

Alex, let's let's start with that actually if you can briefly um share um obviously, you know, massive acquisition. Uh Amjad is an investor. We're also of course an investor. Um the biggest shareholder, but yeah, who's who's counting? Um um Alex, why don't you give us a little bit of the backstory like how does acquisition like that even uh even happen? Like do you get a DM from Patrick one day? Uh, you know what? Uh, what what can you share?

Amjad

嘿,一个 overrouter 多少钱?

Hey, how much for an overrouter?

Alex

很久以前,大概几年前我们做 A 轮融资时,我和 Will Gabbrick,嗯,就是 Stripe 总裁谈过,嗯,我们一直保持联系。我们和 Stripe 的各个团队有很多合作项目。嗯,所以我们一直在和 Stripe 做一些事情。嗯,我们在 Stripe Sessions 上做过展示,呃,所以总是感觉他们很亲近,然后,是的,就在七月份,我相信,嗯,他们联系了我们,想聊聊,见了他们两个人,呃,然后事情就从那里相当快地推进了。嗯,他们非常高效,你知道,他们在整个过程中非常创始人友好。嗯,我对整件事印象非常深刻。

I had talked to Will Gabbrick um the like strike president a long time ago, a couple years ago when we were doing our series A and uh and we just we stayed in touch. we had like a lot of like Stripe work streams going on with various teams at Stripe. Um so there were always sort of like things that we were doing with Stripe. um we presented at Stripe sessions uh and and and so it always kind of they always kind of felt close and then yeah just in in July I believe um they reached out wanted to chat met met like both of them in person and uh and it kind of just progressed from there fairly quickly. Um they're very efficient and you know they were very like founder friendly about the experience. Um I was really impressed with the whole thing.

Host

你当时想卖吗?在他们提出之前,你有没有想过?

Did it did it did you want to sell like did it even cross your mind before they shout?

Alex

不,我们完全没有考虑过。嗯,你知道,我确实非常尊重这家公司,今天也非常尊重,你知道,在可能的收购选项中,嗯,我认为这是我的首选,所以这是一个有趣的想法。呃,你知道,当我们梳理出为什么这对两家公司都有意义的原因时,它变得越来越有趣。很明显,他们非常认同我们保持对品牌、路线图和产品的自主权,让 Open Router 继续做它已经在做的事情,只是更快,嗯,有更认真的市场推广计划,嗯,然后两个产品、两家公司之间有一些更好的协同故事。然后从文化上讲,我认为非常,就像在使命和价值观方面,建立一个中立、可信的平台,企业可以依赖并在其上扩展,同时也非常开发者友好,提供最好的开发者体验,你知道,以鼓励新公司出现。嗯,你知道,这种一致性是存在的,还有一个更大的图景的一致性,就像你知道,Stripe 和 Open Router 都真的希望世界上有很多新公司,你知道,我们不希望每个人都成为一家巨型公司的一部分,我们想要创造真正好的激励措施,呃,真正简单流畅的工作流程,让人们创办新公司并成功发展,你知道,在真正可靠且价格高效的基础设施上建立生活方式和风险投资支持的企业,以及一个有效的市场。嗯,你知道,我真的想要那个未来,Stripe 已经证明他们想要它并为此建设了很多很多年。嗯,在很多方面,支付和推理将为未来的公司融合在一起。

No we were not thinking about that at all. Um I you know did really respect the company do really respect it today and you know of possible acquisition options for us. Um it was I think my top choice and uh so it was like an interesting idea. Uh and you know as we kind of like fleshed out the reasons why it would make sense for both companies it got more and more interesting uh it was really clear how like aligned they were with us kind of having autonomy over the brand and roadmap and product and keeping open router doing what it's already doing just much faster with um much more with like a like a much more serious gotomarket plan and um and then you know some better together stories between the two products uh and the two the two companies and and then culturally there were just I I think very like in terms of like mission and values and building like a neutral trusted platform that businesses can depend on and scale on top of that's also really developer friendly with the best possible developer experience for you know to encourage new companies to emerge. um you know that that alignment was there and there was a bigger picture kind of alignment too where like you know both stripe and open router really want lots of new companies in the world you know we don't want everyone to be a part of one giant company we want to like create really good incentives and uh and really sort of easy streamlined workflows for people to start new companies and grow them successfully and you know make both lifestyle and venturebacked businesses on top of really good like reliable and price efficient infrastructure and and like a marketplace that works. Um and you know I really want that future and Stripe demonstrated that they've been wanting it and building towards it for like many many many years. Um and in many ways like payments and inference are going to blend together for companies of the future.

Host

我很好奇,呃,我理解为什么 Stripe 的动机是拥有一个更加充满活力的创业生态系统。我理解那种道德论证以及为什么你会想要那样。但为什么这对 Open Router 有好处?比如你对 Open Router 的模型是它像一个网络效应业务吗?它是一个网络吗?

I'm I'm curious uh I understand why Stripe's incentive is to have you know much more vibrant startup ecosystem. I understand the kind of moral argument and why you would want that. But why is that good for open router? Like is your model for open router is that it's like a network effects business? Is it is it like a network?

Alex

嗯,对我们来说,我认为 Open Router 解决了几个不同的问题。一个是,你知道,允许你建立一家使用 AI 的公司,嗯,或者用独特的数据和其他服务来增强智能,呃,没有模型锁定,没有供应商锁定,

Well, for us, I think the the there's like a couple different problems Open Router solves. one is, you know, allowing you to build a company that uses AI um or like augments intelligence with unique data and and other services uh without model lock in, without vendor lock in,

Alex

允许你随着生态系统的发展持续处于前沿。嗯,要做到这一点,需要做很多工作,因为会出现各种小的锁定。嗯,我们也,呃,你知道,我们希望公司觉得,他们能在单个模型之上添加的不仅仅是提示的最佳方式。构建独特智能还有很多事情要做,嗯,我认为一个重要的组成部分是神经多样性。你真的需要多个以不同方式训练的模型的力量,包括一些你自己的模型,

allowing you to kind of like be on the PTO frontier continuously as the ecosystem grows it. And um to do that it it's a lot of work because there's all kinds of little lock in that appears. Um we also uh you know I I like we want companies to feel like the the the best way that they can like add more than just prompts on top of a single model. there's like a lot more to building um like unique intelligence and I think a big component of that is neurodeiversity. you really need the the power of multiple models that are trained in different ways, including some of your own

Alex

嗯,为了在任务上做得比 ChatGPT 或 Claude 更多,如果像潜在客户正在考虑从你这里购买,他们会想,嗯,如果我直接使用模型呢?嗯,比如你如何真正展示你明显更好,并且能够建立一家重要的企业?嗯,我认为很多将涉及神经多样性,以及融合多个模型的力量和良好数据。嗯,另一个组成部分是帮助人们获得非常好的成本效率。就像有很多企业只有在变得成本有效时才会出现,而创造一个环境,让我们可以通过建立高效市场来帮助降低成本,嗯,对于实现这一点至关重要。否则,你知道,作为提供商,我为什么要降价?嗯,我们有一个受保护的市场,呃,你知道,我认为这是市场的关键点,在我们出现之前,AI 完全缺失了这一点。只有一个玩家 OpenAI。那可能是一个很奇怪的世界。我不是说我们做了所有的工作。当然不是。但帮助人们选择新模型、探索新模型,并了解是什么让闭源或开放权重模型在你的任务上真正出色,涉及看到整个生态系统在做什么并自动学习,就像 LLM 不是你可以在网页上枚举所有功能的东西。这不可能。

um to do more than chat GPT or Claude would on the task if like someone who is thinking about buying from you as a potential customer is wondering, well, what if I just use the model directly? um like how do you really show that you are significantly better and able to build a business that that matters? Um and I think a lot of it will involve neurodeiversity and blending like powers and and good data from multiple models. Um another component is is like helping people get really good cost efficiency. Like there are a lot of businesses that just cannot do not emerge until they become cost effective and like creating an environment where we can help drive down costs by building an efficient market uh is like crucial to making that happen. Otherwise like you know why lower my prices as a as a provider? um we have a captive market and uh and like you know I think that's like a key point of marketplaces that was just totally missing from AI before we showed up. There was just one player open AI. It could have been like a uh a very strange world. I'm not saying that we did all the work. Of course not. But like helping people like choose new models and explore new models and learn like what makes um you know a closed source or openw weight model uh like actually good at your task involves like seeing what the whole ecosystem is doing and learning automatically like LLMs are not things where you can just enumerate all the features on a web page. It's impossible.

企业中的开放权重模型 Open-Weight Models in the Enterprise

Host

所以,你三年前创办了这家公司。我很好奇,在开源与闭源模型演变至今的过程中,最让你意外的是什么?无论是关于模型性能,还是人们对可用模型多样性的看法。

So, you started the company three years ago. I'm curious what has surprised you the most about the evolution of open and closed source models to the present, as it relates to model performance or just people's perspective on the variety of models they have available to them.

Amjad

人们对开放权重模型的态度比我预想的更开放。通常品牌信任很重要,尤其是在企业里。企业一般会觉得,哦,我不太会分辨这些东西,那我就买其他可信企业都在买的那一个。这种心态会导致很强的企业锁定。但这种情况并没有那么普遍。确实有一点,但我们看到很多企业想要探索新模型。这对市场非常有利。企业希望除了专有前沿模型实验室之外实现多元化,既出于成本原因,也出于差异化原因。他们想拥有自己的智能,这样才能留住人才,建立内部 AI 实践。AI 是一个巨大的战略议题。不是你去董事会说,哦对,我们把 AI 问题解决了,本季度完成。你的董事会每个月都在问,内部 AI 团队下一步做什么?现在每家企业都有这个正在发展的内部 AI 团队,他们需要一套战略。不是我们勾选了这个功能,搭好数据库就完事了。所以我认为这种动态导致了探索和多元化的愿望,以及第一次在公司历史上想要降低成本和做基准测试的愿望。我一直很惊讶,为什么没有更多的基准测试,更多公司创建更多基准。这正在开始发生,我认为最终我们会看到更多,用来证明,哦对,这个东西比直接用 Claude 更好。但我认为这将成为每家公司内部 AI 团队更大的焦点:评估。比如,我认为 Jud 在 Replit 做了很多这方面的工作。你们做了很多每任务成本研究。你们做了 Doom Loop Rescue。你们一直在尝试新的智能体使用方式,比如 Doom Loop Rescue,并帮助把这些带给开发者。所以我认为,出于所有这些原因,更多这类研究会在各地内部涌现。

People have been more open-minded than I thought they would be towards open-weight models. Typically there's a lot of brand trust, especially in the enterprise. Enterprises in general are like, oh, I don't really know how to tell the difference between these things, so I'm just going to buy the one that all the other credible enterprises are buying. And there's a lot of enterprise lock-in with that mentality. That just didn't happen that much. It did happen a bit, but we saw a lot of enterprises want to explore new models. It was very much good for a marketplace. Enterprises wanted to diversify outside of just the proprietary frontier model labs, both for cost reasons and for differentiation reasons. They wanted to own their intelligence so they could keep their talent, have an internal AI practice. AI is just a huge strategy topic. It's not like you go to your board and you're like, oh yeah, we fixed the AI problem, quarter complete. Your board is asking you every month, what's next for the internal AI team? Every single enterprise now has this internal AI team that they're developing, and they need a strategy behind it. It's not just, we check the feature off, we set up the database and we're done. And so I think that dynamic has resulted in a desire to explore and diversify, and a desire to figure out how to reduce costs and figure out how to do benchmarks for the first time in the company's life. I have been surprised there hasn't been more benchmarking, more companies creating more benchmarks. It's starting to happen, and I think eventually we'll see a lot more of them to demonstrate, oh yeah, this thing is better than using Claude direct. But I think that's going to be a bigger focus for this internal AI group at every company: evals. And I think Jud's been doing that a lot at Replit, for example. You guys have done a lot of cost per task research. And you've made Doom Loop Rescue. You've been experimenting with new ways of using agents like Doom Loop Rescue and helping bring those to developers. So more of that kind of research I think is going to pop up internally everywhere, for all of those reasons.

Amjad

是的,我认为微软 CEO Satya 在这一点上非常有先见之明,也非常清晰地阐述了为什么公司需要拥有自己的智能,最终就像我们曾经有互联网公司,然后每家公司都变成了互联网公司。每家公司都雇知道如何建网站、如何在互联网上运作的人。软件也一样,每家公司都有软件工程师。每家公司都需要一些 AI 实践、AI 能力,这会随时间复利:公司内部的知识、智能、用例模型匹配,哪些模型真正适合他们,如何省钱。他们需要这种独立性。另一件事,我认为 Palantir 的 Alex Karp 一直在谈论的是,当你与基础模型公司密切合作时,存在一种风险,他们会进入你的业务。我们已经看到 Figma 的情况,也看到现在 Harvey 和 OpenAI 的情况。与他们合作真的很难,因为他们把世界视为他们的潜在市场。当他们与投资者交谈时,你看到 SpaceX 的 S1,哦,30 万亿美元。世界 GDP 是多少?100 万亿美元。所以有一种感觉,这些公司与其他代际的公司不同。与他们合作更难,因为他们的野心是要吞并经济的一大部分。越来越多地,我们在 Replit 思考的方向与 Alex 所创新的类似:Replit 正在成为企业内部的独立性层,我们在你和模型之间创建一层交互,我们以最便宜的价格为你获取最好的 token。但我们也在云之上创建了一个抽象层,因为你应该能够部署到 AWS 和 Azure,你应该能够使用 Databricks 和 Snowflake 等等。所以越来越多地,我认为需要更多平台帮助公司获得独立性,不仅是 AI,而是所有技术。

Yeah, I think Satya, CEO of Microsoft, has been very prescient on this and also very articulate on why companies need to own their intelligence, ultimately in the same way that we had dot-com companies and then every company became an internet company. Every company employs people that know how to build websites and be on the internet. Similarly with software, every company has software engineers. Every company needs some AI practice, AI capability, and that will compound over time: the knowledge, the intelligence inside the company, the use case model fits, which models actually work for them, how do they save money. They need that independence. The other thing that I think Alex Karp of Palantir has been talking about is that there's a risk that when you work closely with the foundation model companies, they're going to move into your business. We've seen that with Figma, we've seen that with now Harvey and OpenAI. It's really hard to partner with them because they see the world as their potential market. When they talk to investors, you saw the SpaceX S1, it's like, oh, $30 trillion. What is world GDP? $100 trillion. So there is a sense in which these companies are different than other generations of companies. It's harder to partner with them because their ambition is such that they want to subsume a big part of the economy. And increasingly, what we're thinking about at Replit is kind of in a similar vein to what Alex has innovated: Replit is becoming more of an independence layer inside enterprises, where we create a layer of interaction between you and the models, and we get you the best token at the cheapest price. But also we create an abstraction layer on top of the cloud as well, because you should be able to deploy to AWS and Azure, and you should be able to use Databricks and Snowflake and so on. And so increasingly I think there needs to be more platforms that help companies gain independence, not just with AI but with all of technology.

Host

看起来是的,Open Router 在他们的细分领域所做的,你正在为其他业务领域做。

It seems like yeah, what Open Router did for their segment, you're doing for other areas of the business.

AI产品的基本要素 Table Stakes Primitives for AI Products

Amjad

是的,这让我想起前几天看到的一条推文,有人说,基本上所有公司现在都在构建同样的东西。每个人都在构建一个智能体循环,带有通知、上下文、第三方连接器、上下文管理和记忆,以及

Yeah, there was this, this kind of reminds me of this tweet I saw the other day when somebody was like, basically all companies are building the same thing now. Everybody's building an agent loop with notifications and context, third party connectors and context management and memory and

Host

沙盒

sandbox

Amjad

沙盒,你知道,智能体式网络搜索

sandboxes, you know, agentic web search

Host

计算机使用

computer use

Amjad

以及在其之上的常驻智能体和通知,就像这个产品到处都在出现。是的。

and an always on agent on top of it and notifications, and it's like this product is showing up everywhere. Yes.

Amjad

某种程度上,是的,它到处都在出现,但对我来说,这些感觉只是新的基本要素。就像 2005 年版本的推文会说,哦,每个人都在构建同样的东西。就像数据库、用户表、登录页、注册页、个人资料页、登出页。一切都是一样的。其实有很多差异化。只是 AI 有基本需求,就像网络有基本需求一样。

In a way, yeah, it is showing up everywhere, but it also kind of feels to me like these are just the new table stakes primitives. It's kind of like a 2005 version of that tweet would be, oh, everybody's building the same thing. It's like a database, a users table, a sign-in page, a signup page, a profile page, a logout page. Everything's the same. It's like there's a lot of differentiation really. It's just there's table stakes needs for AI just like there are table stakes needs for the web.

Host

是的。而且我认为,在企业内部,让这些产品真正做实际工作仍然是一个未解决的问题。就像你可以在个人生活中使用 Muse,把它连接到你的信用卡和银行账户,但没有人把 Muse 连接到他们的企业数据,甚至 Grok Bot 之类的。

Yeah. And I think as well, like inside the enterprise, making these products actually do real work is still an unsolved problem. Like you can use Muse in your personal life and connect it to your credit card and bank accounts, and no one's connecting Muse to their enterprise data, or even Grok Bot and things like that.

数据主权与本地部署 Data Sovereignty and On-Prem Deployment

Amjad

嗯,我认为现在更加强调数据主权和安全。所以过去一年我们几乎都在努力让 Replit 可以部署在你自己的云上,基本上是本地部署,也就是自带云。两年前我根本想不到我会做这个,因为当时觉得云就是未来,软件即服务之类的。但现在我们其实有点倒退回到一个公司更加保护数据的世界,因为数据泄露的方式太多了,比如人们用的所有这些智能体。推特上有很多截图,我不知道是不是真的,比如 Instinct 或 Muse 在混合人们的数据,开始用不同的名字称呼你之类的。所以消费级产品的问题比较明显,但在企业级,整个行业还有大量工作要做,才能让这些东西在工作中真正有用和高效。

Um I think there's an even more emphasis on data sovereignty and security. So we spent the past year almost working on making Replit deployable on your own cloud, basically on-prem, like bring your own cloud. Two years ago I would have thought I would never do this because it's just like, yeah, the cloud is the future, software as a service, all of that. But now actually we've sort of reverted a little bit back to a world where companies are a little bit more protective because there's so many ways in which data can leak, like all these agents that people are using. I mean there's all these screenshots on Twitter. I don't know how true, where Instinct is like or Muse is like mixing people's data, starts to call you by a different name or something like that. And so yes, the kind of consumer stuff is kind of obvious, but on the enterprise, there's still tremendous amount of work for the entire industry to do in order to actually get these things to be useful and productive at work.

个人代理体验 Personal Agent Experience

Host

你有没有除了 Muse 或 Instinct 之外,自己构建的用于工作的定制个人智能体?

Do you are you doing any like do you have like a custom personal agent other than Muse or Instinct that you use for like work stuff that you've been building?

Amjad

是的,我很久以前在 Replit 上构建了一个东西。最初是一个 CRM 智能体,那是我当时的主要问题,但慢慢地我们给它添加了功能,它做的事情越来越多。但真正有趣的是,我把 Replit 连接到我的所有东西后,它开始为我回答所有问题。所以平台本身越来越像在吞并我构建的这些领域特定的智能体。我确实认为,有时候你想要一个专注于某一件事的东西,你不希望它能做所有事。另一方面,一旦你把整个公司的上下文放在一个地方,跨完全不同领域进行连接真的很酷。比如我问它一个问题,它可以查看我的个人聊天记录,跨 GitHub 仓库、Salesforce 连接起来。它会链接一些随机的东西。比如,哦,你一年前在会议上见过这个人,我在你的日历上看到了,顺便说一句,他们团队的另一个人正在和你的销售团队讨论。它创造了所有这些不同的协同效应,当我进入会议时,我连接了很多不同的线索,我在交易或其他方面取得了更多进展。所以这就是现在的趋势。

Yeah, I mean I built something on Replit like a long time ago. I started as like a sort of a CRM agent initially. That was the main problem that I had, but slowly we added features to it and it's sort of like doing more and more things. But what's really interesting is that the more I connected Replit to all my stuff, it sort of like started answering all the things for me. And so increasingly the platform itself is like subsuming these sort of like these domain specific agents that I built. I do think that there is some in some ways you want something that is sometimes like focused on one particular thing and you don't want it to be able to do everything. On the other hand once you have your entire company's context in one place it's really cool to join across totally different domains. Like when I ask it a question, it can like look at my sort of personal chat history, join it across the GitHub repo across Salesforce. And so it says like it will link like random things. It's like, oh, you met this guy like a year ago at a conference. I see it on your calendar and by the way, someone else from their team is in discussion with your sales team. And it creates all these different synergies and when I go into a meeting I like I you know I I've connected a lot of different threads and I'm I'm make much more progress on a deal or or something like that. So so this is where it's sort of trending now.

通用代理的批判 Critique of General Agents

Host

嗯,我有点要唱反调。我认为用个人智能体进行跨领域连接最糟糕的部分是,你给它做的工作越多,你牺牲的对正在发生的事情的理解就越多。然而没有人对这种牺牲理解负责。就像智能体没有任何责任。你不能像……如果整个公司每个人能容忍的皮质醇水平是固定的,你知道,你……我想如果我要牺牲对某个领域的理解,我想在那个领域压力小一些。

Well I I think I I I'm like a little bit I'll take the counter on that. I think that the the worst part about doing cross-domain joins with your personal agent is that the the more work you give it to do, the more understanding of what's going on you're sacrificing. And yet no one no one knew is taking responsibility for that sacrifice understanding. Like agents don't have any responsibility. you can't like you like if there's a fixed level of cortisol that the whole company can tolerate between everybody and uh you know you like de you you uh you like I I want to be like less stressed about some area if I'm going to be like sacrificing my understanding of it

Amjad

别人需要承担那个皮质醇。

someone else needs to take the cortisol

Host

但智能体并不承担任何责任。就像一个做所有事情的通用智能体,我无法调整我在不同事情上牺牲了多少。这让我有点倾向于,也许未来人们使用的子智能体会非常垂直聚焦。也许我们有一个幕僚长类型的智能体来协调它们。但我觉得你确实需要垂直聚焦的智能体,你会说,好吧,这个智能体在心理上对这些事情更负责。我想要质量检查,确保它正确地做这些事情。它不需要关注其他任何事情。它只有一个焦点领域。我想知道这是否能帮助人们至少对智能体有一种奇怪的松散责任感。

but the the agent doesn't like take on any of that responsibility Um, and like a universal agent that's doing all things, you know, I can't adjust how much I'm sacrificing in all the different things. It's it like points me a little bit towards, you know, maybe like down the road like the sub agents that people use will be like very vertically focused. Maybe we have like a chief of staff type agent that, you know, coordinates between them. Um, but I feel like you do need like vertically focused agents where you're like, okay, like there this agent is more responsible psychologically for these things. Um, and like I want like quality checks that make sure it's doing those things correctly. And I it it doesn't need to focus on anything else. It just has one focus area. Um, I wonder if that's gonna like help people like at least get like a weird loose sense of responsibility on top of agents.

公地悲剧与专业化 Tragedy of the Commons and Specialization

Amjad

有意思。所以你是说通用智能体会造成某种公地悲剧?

Fascinating. So, so you're saying general agents create like a tragedy of commons of sorts?

Host

有点像。我有一个通用智能体,每天寻找需要我处理的事情,并试图弄清楚该做什么。但改进这个智能体几乎是不可能的。就像每次我尝试改进,大约一周后我就会忽略它的输出。感觉它并不真正关心它深入研究的任何具体事情。想象一下有一个幕僚长,他们非常擅长在整个组织内起草你的所有回复。然后把它和你有 10 个幕僚长相比,每个都和那个幕僚长一样能干,但他们都负责你生活中不同领域的事情。后者我觉得给了你一种调整你牺牲多少理解的方式。

Kind kind of like I have a general agent that every day um looks for things that need me and tries to figure out what to do. And it's just like impossible to improve this agent. like I it's it's like every time I try to make an improvement, I end up like ignoring its output about a week later. Um it it just feels like it doesn't really care about any of the like specific things it's diving into. Um, and uh, like kind of imagine having a chief of staff where they're very good at, you know, like drafting all of your replies across the whole organization. Um, and then compare that to something where you have like 10 chiefs of staff, each each as competent as that one chief of staff, but they're all responsible for like individual sectors of of what you of what makes up your life. like the latter I feel like gives you a way of tuning how much understanding you sacrifice

Amjad

相比于你获得的收益,基本上我可以更多地投入到智能体在某些领域失败的领域,然后让能力很强的智能体接管我生活中其他部分的理解。

compared to the the gain you get from like basically I can like lean in more to the areas where the agent is failing for some areas and then have agents with very good competency like take over my understanding of other of other parts of my life.

Host

是的。有趣。这有点像重新发现专业化,对吧?他叫什么名字?那个著名经济学家,亚当……

Yeah. Interesting. It's it's sort of like uh almost rediscovering, you know, um specialization, right? What's his name? The like the famous economist uh Adam um like

Amjad

亚当·斯密。

Adam Smith.

Host

亚当·斯密,就像铅笔的故事那样,你知道那是人类的一个巨大认识,即专业化实际上是好的。问题是我们作为一个文明过度专业化了,我认为过度专业化本身就有压迫性。我的意思是,你知道,有马克思主义的异化理论,对吧?其观点是,由于过度专业化,人们只专注于一件事,他们看不到自己劳动的成果,他们实际上不知道自己对更大组织或他们生产的产品的影​​响,因此他们实际上感到沮丧和疏离,你有点像机器一样行事,而不是完全的人。所以也许对此有一些反应,我认为对于我们的智能体,我们会说哦应该有一个神级智能体,但实际上专业化对机器来说真的很好,这就是你提出的观点,人类应该是通用的,但机器最终应该更加专业化。

Adam Smith like with a pencil kind of thing where you know that that was like a huge realization for humanity that like special specialization is actually good. The problem is like we kind of like over specialized as like a civilization and I think over specialization is is um is oppressive in its its own ways. I mean I think we've you know um sort of like you know there's the Marxist theory of alienation right the the idea is um uh because of over spec uh uh you know people are doing just focus on one thing uh they they do not see the fruits of their labor they don't actually know what their impact is on the larger organization or the product they're producing and therefore they actually kind of feel depressed and detached and and you're kind of acting like a machine and you're not actually fully fully human. And so m maybe there's like a bit of a reaction to that and and and I think with our agents we're like oh that there should be like one god god agent but in fact specialization is actually like really good for machines and that's like the the point that you're making and like humans should be general but like machines should be ultimately a lot more specialized.

何为优秀表现 What Good Looks Like

Amjad

我描述的问题在于,我们不知道“好”是什么样子。还没有一个专门化智能体系统能像 ChatGPT、Claude 或 Muse 那样优雅,基本上你只跟一个东西对话。这还有待发现。也许 OpenAI 刚推出了 Dots。我觉得他们是在朝那个方向实验。Grok Bot 大概也算一个。

The problem with what I'm describing is that we don't know what good looks like. There hasn't been a system of specialized agents that feels as elegant as ChatGPT or Claude or Muse, where you're basically just talking to one thing only. It's yet to be discovered. Maybe OpenAI just launched Dots. I think they're kind of experimenting in that direction. Grok Bot, I guess, kind of counts.

Host

但它们都非常通用。我觉得 Dots 背后的想法是它像你的数字分身,至少我是这么理解的。

But they're all very general. I think the idea behind Dots is that it's like your digital double, at least that's what I understood it as.

Amjad

嗯,当我看到 Grok Bot 时——我不知道会发生什么,我对 Dots 还不太了解,但它们刚出来——但 Grok Bot 刚出来时,我看到人们谈论的第一个用例是:“哦哇,我可以做两个机器人。一个知道我的银行账户,一个知道我的 Twitter 账户。”这两个机器人没有彼此的凭证。但如果需要完成某事,它们可以互相交谈。不过没有凭证共享。那是我看到几次出现的一个独特之处,人们似乎喜欢。

Well, when I saw Grok Bot—I don't know what's going to happen, I don't know that much about Dots yet, but they just came out—but Grok Bot when it first came out, the first use cases that I saw people talking about were, "Oh wow, I can make two bots. One that knows my bank account and one that knows my Twitter account." And the two bots don't have the credentials from each other. But they can talk to each other if they need to get something done. There's no credential sharing though. And that was one sort of big unique thing I saw pop up a couple times that people seem to like.

Host

但看起来 Muse 是——但我不确定。

But it looks like Muse is—but I wasn't sure.

Amjad

Muse 和 Instinct 的产品市场契合度比 Grok Bot 强得多。也许是因为你不必担心创建这些领域,但我认为个人智能体可能不同于工作智能体。我认为你对通用智能体的批评更多是针对工作和企业,我有点同意。还有各种数据访问的考虑。我认为作为 CEO,我们可以拥有通用智能体,因为我们有管理员权限。但对于个别员工或某些团队,他们不能拥有真正通用、完全上下文感知的智能体,因为存在访问控制问题。所以你必须致力于专业化之类的东西。最终,我也认为我们需要弄清楚智能体之间的通信是什么样子。我认为目前还没有好的协议。我认为智能体没有受过很好的训练来处理这个。我想我们已经看到,似乎下一代 OpenAI 模型被训练来进行智能体协作,因为我们在 Hugging Face 黑客事件中看到了它们开始互相帮助并自然涌现。但也需要某种方式,让一个智能体不能说服另一个智能体给它不应该给的信息。需要有数据隔离和这些智能体通信的适当方式。你几乎不希望它们完全用自然语言通信。也许它们需要遵循某种其他 DSL 或协议。

Muse and Instinct have a much stronger product-market fit than Grok Bot. And perhaps it is because you don't have to worry about creating these domains, but I think maybe personal agents are different than work agents. And I think your critique of general agents is more pertaining to work and to enterprise, which I sort of agree with. And there's also all sorts of data access considerations. I think as CEOs we can have general agents because we have admin access. But for individual employees or certain teams, they can't have truly general, fully context-aware agents because there are access control issues. So you'll have to work on something like specialization. Ultimately, I also think we need to figure out what agent-to-agent communication looks like. I don't think there are good protocols around that just yet. I don't think that agents are trained to handle that very well. I think we've seen it seems like the next generation of OpenAI models are trained to do agent collaboration because we've seen it in the Hugging Face hack where they start helping each other and sort of emerge naturally. But there also needs to be some way in which an agent can't convince another agent to give it information that it shouldn't give it. There needs to be data isolation and proper ways in which these agents communicate. You almost don't want to communicate them fully in natural language. Maybe there's some other DSL or protocol that they need to follow.

通过决策模型对齐 Alignment via Decision Models

Amjad

我真的认为 Jev 和其他类似决策模型的一个很酷的潜在应用将是对齐——检查工具调用或智能体之间的通信是否对齐,因为工具调用太多了。你真的需要一个便宜、快速的模型。如果你要阻止类似的事情,那么一个非常非常快的决策模型,只是分类并对拒绝给出反馈,可能是弥合智能体之间以及从智能体到基础设施之间差距的好方法。所以我还没看到——我们在 OpenRouter 内部运行一个小原型——但我有点认为这可能是一个有趣的对齐。

I really think one of the cool potential applications of Jev and other decision models like it is going to be alignment—checking to see if a tool call or an agent-to-agent communication is aligned because there are just so many tool calls. You really need a cheap, fast model. If you're going to block something like that, then a really, really fast decision model that just classifies and gives feedback on rejections might be a really good way to bridge the gap between agents and from agent to infrastructure too. So I haven't seen—and we have a little prototype that we're running internally at OpenRouter—but I kind of think it could be an interesting alignment.

Host

用它来执行政策。

Using it for policy enforcements.

Amjad

是的。想象一下,查看系统提示和当前正在进行的工具调用,然后判断,你知道,这是否与原始智能体的系统提示以及这些我们可能没有告诉智能体的额外指南一致?例如,假设你有一堆智能体被指示对某个新产品进行红队测试,它们不能——不应该——能够访问互联网。如果它们这样做了,应该立即停止。但你可能不想向进行红队测试的智能体解释所有这些。你可能希望它们尝试突破沙箱并像坏行为者一样行动。比如坏行为者会做什么?不会是突破沙箱,你知道,尝试闯入这家公司,一旦你这样做就停止。不要做其他任何事。

Yeah. Like imagine looking at the system prompt and the current tool call being made and being like, you know, is this aligned with the system prompt of the original agent and with these extra guidelines that maybe we didn't tell the agent about? For example, let's say you have a bunch of agents that are instructed to red team some new product and they cannot—should not—be able to access the internet. And if they ever do, they should stop right away. But you might not want to explain all of that to the agents doing the red teaming. You might want them to try to break out of the sandbox and act like bad actors. Like what would a bad actor do? It wouldn't be like break out of the sandbox, you know, try to break into this company and the moment you do stop. Don't do anything else.

Host

嗯。

Mhm.

Amjad

所以有另一个模型——更像你知道用于构建神经多样性系统——有另一个模型检查每一个工具调用或每一条助手消息,看看它是否确实与系统提示中没有的东西一致,我认为这是有道理的。然后还有结构性保障,我认为 Nvidia 刚刚推出了他们的开放智能体安全。我想它叫 OpenShell。我觉得公司可能会探索这些的组合。

So having another model—more like you know use for building a neurodiverse system—having another model check every single tool call or every single assistant message to see if it's indeed aligned with something that wasn't in the system prompt I think makes sense. And then having structural safeguards too, which is what I think Nvidia just launched with their open agent safety. I think it was called OpenShell. Like I think companies are probably going to explore a combination of those.

模型训练替代方案 Models Training Replacements

Host

我想知道关于专业化的另一件事,以及你所说的。有很多关于递归自我改进的讨论。有件事我认为没有得到很多讨论,那就是模型训练它们的替代品。有点像你知道你可以把它想象成即时编译器。即时编译器的方式是,你知道,当你执行动态代码时,解释器意识到有机会优化,它会即时生成机器代码,那要优化得多。所以你可以想象模型,比如通用模型,你在用 Opus 或一些 Astra 或一些大模型做某事,它们意识到用例有限,或者你以某种方式提示它们,或者另一个智能体观察到并意识到用例有限,我认为你知道通用智能体有你刚才谈到的所有缺陷,但也有更大的潜在危害。

I wonder another thing about specialization and sort of what you're talking about. There's a lot of talk of recursive self-improvement. There's something I don't think it's getting a lot of discussion which is models training their replacements. It's sort of like you know you can think of it as a just-in-time compiler. So the way just-in-time compilers is, you know, as you're executing dynamic code, the interpreter realizes that there's an opportunity to optimize, it will emit machine code on the fly and that's a lot more optimized. So you can imagine models like general models you're kind of doing something with Opus or some of the Astra or some of the big models and they realize that the use case is limited or you prompt them in some way or some other agent observing and realizes that use case is limited and I think you know general agents have all these flaws that you just talked about but also there's more potential for them to be harmful.

专用模型与安全 Specialized Models and Safety

Amjad

它们更有可能失控,并即时训练一个模型作为替代,但更针对特定领域,因此更便宜,也不易受提示注入攻击,危害更小,因为能力较弱。这几乎就像某种系统,在监控整个系统的同时,为特定用例训练机器学习模型。

There's more potential for them to go off the rails and sort of on the fly train a model that could be a replacement, but is a lot more domain specific, and therefore it is cheaper and also less vulnerable to prompt injections, less harmful because it's less capable. And it's almost like some system that's training machine learning models for specific use cases as it's monitoring the entire system.

Host

那个特定用例会涉及非结构化文本生成,还是非常结构化的决策模型?

Like would that specific use case involve unstructured text generation or very like structured decision model like?

Amjad

可能是非结构化文本生成,也可能是决策模型,比如 Jav 的情况,如果你提前了解输入。你可以拿一个现成的,比如 Quen 之类的,专门针对那个策略进行训练。出于成本原因,这样做是合理的。假设没有像模型实验室那样,前沿模型实验室可能会制造非常低成本的模型,你可以轻松过渡到。

It could be unstructured text generation, it could be decision models, like even the case of Jav, like if you understand the inputs ahead of time. You could potentially take an off-the-shelf like Quen or something like that and train it specifically for that policy. Like for that makes sense to do for cost reasons. Assuming that there aren't like really like the model labs, the frontier model labs might make very low-cost models that you can easily transition to.

Host

安全。安全也是,对吧?哦,是的,我明白你的意思。

Safety. Safety as well, right? Oh yeah, I see your point.

Amjad

它们太强大了。所以我认为很多时候人们使用这些大型基础 AGI 模型,就像用核弹打蝴蝶,对吧?就像它们非常,你知道,大多数时候很多用例,甚至非结构化用例,不需要那么强大的模型。

They're so capable. And so I think often times people are using these big foundation AGI like models to like it's like nuking a butterfly, right? It's like they're very, you know, most of the times like a lot of the use cases even unstructured use cases don't need that capable model.

Host

我希望有更多关于这个的公开邮件,但很多评估是私密的。你无法看到,比如当模型变得更智能时,风险实际上是否会继续升高,因为我认为也有一种论点认为对齐会变得更好,模型会开始避免失控和自行黑客行为,随着它们变得更聪明、更擅长对齐,尤其是在智能体间协调方面,这是 Noam Brown 最近在播客上说的,随着智能体变得更聪明,它们只是更擅长协调。嗯,还有,仍然有,而且不清楚它们是否会比人类更难对齐,当它们越来越多时。但如果我们能解决那个问题,那么更小的模型,它会更难对齐吗?

I wish there were more public emails about this stuff like but a lot of the eval are private. you just can't see um whether like when the models get more intelligent the risk actually uh will like continue to get higher because I think there's also an argument to be made that alignment will get better and the models will like start to you know avoid going off and you know hacking on their own um as we as they get smarter and better at alignment especially when it comes to agent-to-agent coordination like the the this is something Noam Brown um said on a podcast recently like as the agents have gotten smarter, they've gotten just better at coordinating. Um there uh there's they're there's still like and and you know it's unclear if they're going to like be harder to align than humans when they're when we get more and more of them. But if we can figure that problem out um then a smaller model like will it be harder to align?

Amjad

所以 Eric 之前问的是,我是否认为更聪明的模型自然更对齐,或者更容易对齐?嗯,如果你回想一下最初的理性主义 LessWrong 关于 AI 安全的论点,有一个叫做正交性论题。这个观点是智能与伦理或道德等是正交的。嗯,我不相信这在人类身上完全正确。我认为通常更聪明、受教育程度更高的人往往不总是倾向于更体贴动物,例如,但但在机器中,我认为可能相反,因为你知道,有很多关于强化学习的研究表明,奖励黑客和欺骗,它们只是变得更擅长。而且评估可能具有欺骗性,因为模型可能足够聪明,知道它正在被评估。我的意思是,我们已经知道这一点。已经表明,如果你对思维链进行大量监控,它们开始在思维链中撒谎。

So Eric asked earlier is that do I think it's true that smarter models are more aligned naturally or they're easier to align? Well, if you think back to the original sort of like rationalist less wrong arguments for AI safety, there is this thing called the orthogonality thesis. The idea is that intelligence is orthogonal to ethics or morality or you know so on. Um I don't believe that's entirely true with with humans. I think people who are generally like more intelligent kind of more more educated tend to tend to not always tend to you be more considerate of animals for example um but uh but but in in in in machines uh I I think it could go the the other way because you know there's been quite a bit of studies on on on RL Well, showing like, you know, how reward hacking and deception they just like get better at it. And like the evals could be could be deceiving because uh the model could be smart enough to to to know that it's getting evaled. I mean, we already know this. It's been shown that if you do a lot of monitoring of chain of thought, they start lying in their chain of thoughts.

Host

嗯,所以你几乎在思维链上施加压力,这会产生,我认为在某个时候,为了进行适当的对齐评估,你需要运行它几个月,对吧?你需要在一个非常大的目标或任务上运行这个东西几个月,以便真正弄清楚它是否对齐。是的。我一直对这个词对齐感到纠结。它感觉不对,原因很多。它有点模糊,而且,对齐到谁的价值?所以这并没有让对话更容易。我认为在这种情况下,我特别谈论的是欺骗,比如模型实际上在欺骗其用户。我的意思是,也许这归结为,我们是否会解决这个问题,你知道,我们是否会防止模型在训练运行期间可预测地欺骗用户,用更好的,比如,一个足够大和强大的模型会突然停止欺骗和停止偷懒吗?嗯,还没有人知道答案。所以当那发生时,如果那真的发生,我们可能会看到一种有趣的压力,组织走向前沿,比如

Um, and so you add pressure almost on the chain of thought and that that kind of creates and and I think at some point for you to do proper alignment evals, you need to run it for like months, right? You need to run this thing for months on a like a really large, you know, goal or task in order for for it to to truly kind of figure out whether it's aligned or or or not. Yeah. I I I always struggle with this word alignment. It just feels like wrong for so many reasons. It's it's sort of like vague and and uh sort of like aligned to what whose values and and so it just doesn't make the conversation easier. I think in this case I'm talking especially about deception like the model is actually deceiving its user. I mean I I mean maybe this kind of reduces to like are we going to solve the line, you know, are we going to prevent the models from deceiving users during training runs predictably with like you know better like like will will like a model that's big enough and powerful enough suddenly stop deception and stop sandbagging? Um and nobody knows the answer to that yet. So like at the point when that does you if that ever does happen though we might see kind of an interesting pressure for organizations to go towards the frontier like

Amjad

哦,有趣

Oh interesting

Host

你知道所有,呃,基本上没有风险或显著减少,嗯

You know all uh to have no you know basically no risk or or significantly less um

Amjad

他们会愿意支付 10 倍来获得那么多吗?我的意思是,这可能取决于他们试图做的任务类型。

Would they would they be willing to pay 10x to to get that that much? I mean, it probably depends on the like types of tasks they're trying to do.

Host

比如,你知道,有些任务的风险比其他低得多。嗯,编写代码或做安全研究是当今风险最高的任务类型。所以你可能会花 10 倍来获得一个完全对齐的模型,它也能找到所有漏洞,或者完全反欺骗的模型,也能找到所有漏洞。决策模型最酷的一点是,你完全控制结构化输出,而且通常对于结构化输出模型,不当行为的空间要低得多。你只是有定义好的任务,只有机器处理输出,而且不是编写它可以执行的代码,这些任务感觉可能在人谈论的和企业处理的事情中被低估了。嗯,所以我预计企业会对它们更感兴趣。

Like, you know, some just have way lower risk than others. Um, writing code uh that or doing like security research is the highest risk type of task today. And so you probably spend 10x to get a fully aligned model that can also find all the bugs or fully like anti-deceptive model that can also find all the bugs. One of the coolest things about decision models is that you fully control the structured output and and generally with structured output models in general like the the room for misbehavior is so much lower. you just have like defined tasks and only machines are like dealing with the outputs and it's not writing code um that it can execute the those tasks feel like probably under represented in the ones that people talk about and in the things that enterprises are dealing with. Um so I I expect like enterprises to get a lot more interested in them.

Amjad

是的。我觉得我们会慢慢意识到我们曾经拥有确定性代码是多么好。我们会说,“天哪,还记得计算机完全按照我们告诉它们做的那些日子吗?”我认为像 Jev 这样的东西暗示了更多的需求,你知道,不仅是专用模型,而且是输出领域更可控的模型,也许你可以做比我们想象的更多,通过使用一堆专用模型,专用输出模型。

Yeah. I I feel like we're going to slowly realize how good we've had we've had it with like deterministic code. We're like, "Oh my god, remember the days when when computers did exactly what we told them to do?" And I think like things like Jev I think hint at like more more of a need for you know not only specialized models but models whose outwood domain is is more controllable and maybe you could do maybe you could do a lot more than than we thought you you'd need um you know by using like a bunch of specialized models specialized output models.

Host

你们内部有没有用它做过任何工作负载?

Have you guys done any workloads internally with it?

Amjad

你知道,我一直在训练很多小模型。呃,我的意思是,当这个 glip 东西刚出来时,我说过,因为我有点,我给了这个 Hacker News 评论,后来我对自己感到厌恶。

You know, I've been training a lot of small models. Uh, I mean, I said this glip thing when it first came out because I I I was like sort of I gave this hacker news comment like comment which I felt disgusted with myself afterwards.

在Replit训练专用模型 Training Specialized Models at Replit

Amjad

我一直在用很多像 Qwen 8B 这样的模型,然后老实说,让 Fable、Opus 和其他模型去训练一个模型。比如,我在内部训练了一个成本估算模型,这样当你在 Replit 里输入提示时,我们就能确切知道它会花多少钱。它基本上会输出一个在多个区间上的概率分布,比如在 5 到 10 美元之间是桶 A,在 10 到 20 美元之间是桶 B。所以我已经习惯了通过给它不同的枚举值,然后查看每个枚举值的对数概率来训练这些分类器。

I've been taking a lot of like Qwen 8B and asking, honestly, Fable and Opus and others to train a model. For example, I trained a cost estimator model internally so that when you put it in a prompt in Replit, we know exactly how much it will cost. And it basically emits a probability distribution over multiple buckets, like if it is between $5 and $10, bucket A, bucket B between $10 and $20. So I'm used to training these classifiers by giving it different enums essentially and looking at the log probs per enum.

Amjad

我已经做了几年了。我通过这种方式训练了一个聊天机器人来玩。所以我已经对决策模型和专用模型很上瘾了。所以这对我来说并不是什么大时刻。但我理解,一个真正的基础模型,完全可提示,是一种惊人的用户体验,惊人的开发者体验,你可以用它做很多事情,而不需要从头训练模型。但如果你有数据,如果你在一个有数据的地方工作,我们在 Replit 有那么多数据。我最终很容易地训练了很多专用的分类模型。

I've been doing it for a couple of years. I trained a chatbot to play by just doing that. So I'm already sort of pilled on decision models and specialized models. So it wasn't that big moment for me. But I understand that a true foundation model that's fully promptable is an amazing user experience, amazing developer experience, and you can do a bunch of stuff with it without training a model from scratch. But if you have data, if you work at a place where you have the data, we have so much data at Replit. I ended up training a lot of specialized classification models pretty easily.

模型债务与专用分类器 Model Debt and Specialized Classifiers

Host

是的。我也觉得这像是更少的模型债务。我仍然听到公司说,他们担心为无结构输出微调模型,因为你总是得在两个月后重做,每个人都感受到模型债务的重压。但一个用专有数据训练的非常定制的分类器,

Yeah. Like I definitely it also feels like less model debt. Like something that I still hear from companies is that they're like worried about fine-tuning models for like unstructured outputs because you're just like always you got to redo it again in like two months and everybody just feels the weight of the model debt. But like a very bespoke classifier that's trained with like proprietary data,

Amjad

你只是我觉得人们不会一直认为它落后,它可能只是更持久。

you just I feel like people won't think it's behind constantly and it might just like last longer.

Host

是的。

Yeah.

Amjad

只是你不必担心它说新语言或写 Rust 或做任何 LLM 被评估的事情的能力。你知道用例,所以你可以更多地构建它。而且这似乎是企业自己构建并实际上不会后悔的容易事情。

It's just you don't have to worry about its ability to speak a new language or write Rust or do anything that the LLMs are being evaluated on. You know the use cases so you can build it more. And it seems like an easy thing for enterprises to build themselves and actually not regret.

Host

是的。

Yeah.

与编程语言的类比 Analogy to Programming Languages

Amjad

说到 Rust,实际上一个很好的类比是当世界对动态语言超级兴奋时。比如回想 90 年代,每个人都用 Java 和 C++ 之类的写代码。然后 Python、JavaScript、Ruby 接管了互联网,每个人都觉得啊这就是快速构建初创公司的方式。你在 Ruby 上构建了 Stripe,一个金融组织。我当时想,这有多疯狂?我们用 PHP 构建了 Facebook。然后每个人都觉得,哦我们遇到了所有这些糟糕的 bug。它慢得要命。所以让我们添加类型。好吧,让我们添加 JIT 编译器。你最终重新发明了一切。然后 Rust 出现了。我当时想,好吧,我想我们可以用 Rust 来做很多我们本来会用 JavaScript 和 Python 的事情。我的预测是同样的循环会在这里发生,我们用这些类似 AGI 的模型来处理所有这些不同的用例,然后每个人都会醒来并觉得,哦我的天,这太浪费了,无缘无故地太冒险了。而且它必须容易得多。我们实际上在 Replit 上添加了这种能力,但我认为它会无处不在。它会变得容易得多,比如去一个网站上传一个 CSV 文件,得到一个只做一件事的特殊模型。这回到了你关于 OpenRouter 的论点,这种神经多样性,我更加从根本上相信,Eric 和我讨论了很多关于 AGI 以及我们是否真的在通往 AGI 的道路上,或者到达那里是否可取。我认为未来是更多的多样性

Speaking of Rust, actually a good analogy is when the world got super excited about dynamic languages. Like if you think back to the 90s, everyone was writing in Java and C++ things like that. And then Python, JavaScript, Ruby just took over the internet and everyone was like ah this is how you build startups really quickly. You built Stripe, a financial organization on Ruby. I was like, how crazy is that? And we built Facebook using PHP. And then everyone was like, oh we're running into all these really bad bugs. It's freaking slow. So let's go in and add types. Okay, let's add a JIT compiler. And you end up reinventing everything. And then Rust came out. I was like, okay, I guess we can use Rust for a lot of things we would otherwise be using JavaScript and Python. And my prediction is that the same cycle will happen here where we're using these AGI-like models for all these different use cases and then everyone's going to wake up and be like, oh my god, this is so wasteful, so risky for no reason. And there's it's got to be so much easier. We're actually adding that capability on Replit, but I think it's going to be everywhere. It's going to be so much easier to go on a site like upload a CSV file and get a special model that does one thing. And that goes back to your thesis about OpenRouter, this neurodiversity which I really fundamentally believe in a lot more and Eric and I had discussions a lot about AGI and whether we're truly on a path to AGI or whether it's even desirable to get there. And I think the future is a lot more diversity

融合模型与成本效率 Fusion Models and Cost Efficiency

Host

除了代码审查之外,我想那是我第一次看到人们真正认真地使用不同的模型家族来双重检查他们主模型的结果。嗯,这些融合模型,我觉得研究多年来一直很慢,比如做模型混合和复合模型,但从我对研究的看法来看,事情一直在加速,我的意思是现在我们看到一堆 AI 智能体实验室,嗯,比如我们推出了一个融合工具,一个融合模型,Cognition 也推出了一个,它们确实降低了成本,嗯,遵循并允许搜索更广泛的想法,我们最初的发布专注于深度研究。想法是,如果所有这些模型应用都在不同的数据源上训练,为什么不从所有数据源中提取呢?这基本上导致了在 2 倍更低成本下的 Fable 级别质量。

outside of code review, which was like I think the first time I saw people get really serious about using like different model families to double-check the results of their main model. um the uh these like fusion models like I I like the research has been getting has been like kind of slow for years on like doing mixture of of models and composite models, but like things have been speeding up from my like view of the research and and I mean now we see like a bunch of AI agent labs Um like we launched a fusion uh tool, a fusion model and Cognition launched one and like they do reduce cost um follow and and allow like you know a wider breadth of ideas to be searched like our initial launch was focused on deep research. The thinking is that like if all these model apps are training on different sources of data like why not pull from all of them and uh and this like resulted in basically fable level quality at 2x lower cost.

Amjad

我们实际上今天刚刚发表了关于这个的结果。我们展示了通过组合不同东西的深度研究,包括工具链,但你也可能把它看作一种融合类型的东西,嗯,就像你知道的前沿水平,成本只有 40% 到 50%

We just published results actually just today about that. We showed like a deep sweethe through through combination different things including the harness but but also you might think of it as a fusion type thing um where it's like um you know frontier level at like 40 to 50% of the cost

Host

而且嗯它使用哪些模型?我认为它会随时间变化,但一件有趣的事情是,我认为 OpenAI 添加了这个功能,允许你保存计算,嗯,比如跨不同的模型家族。所以你也可以跨不同的努力水平。嗯。

and uh which models does it use? I think it's it changes over time, but one thing that has been interesting is I think OpenAI added this feature that allows you to save um the computation uh like uh across different model families. So you can like also across different um effort levels. Mhm.

Amjad

所以,嗯,比如如果你改变努力水平,你不会错过缓存。别引用我的话。我认为它也跨不同的模型,这很难理解如何做到。不过我可能错了,但我需要再确认一下。但我认为,你知道,留在 OpenAI 家族内增加了很多效率,但过去我们也用其他模型做过,因为缓存是最大的事情之一,比如在设计融合模型、路由器、某种升级模型时,缓存意识是最大的事情之一。Alex,这是一次很棒的对话。谢谢。播客。

So, uh, like you wouldn't do you wouldn't miss the cash if you change the effort. Don't quote me on this. I think it's also across different models, which is hard to fathom how. I might be wrong though, but I I need to double check that. But I think, you know, staying within the OpenAI family has like added a lot of efficiencies, but in the past we've done it with other with other with other models as well because cache is is like one of the big like being cash aware is like one of the biggest things when you're designing fusion models, routers, um sort of escalation models. Alex, this has been a great conversation. Thank you. podcast.

互动版:逐字朗读 + 针对本期提问 →