OpenRouter 创始人 Alex Atallah:以 70 亿美元出售给 Stripe,以及正在展开的 AI 经济

OpenRouter's Alex Atallah on Selling to Stripe for $7B+ and the Unfolding AI Economy

亚历克斯·阿塔拉 Alex Atallah · Sabrina Halper Show · 2026-09-29 · 约 57 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

OpenRouter 创始人 Alex Atallah 畅谈以 70 多亿美元出售给 Stripe、服务 1000 万开发者和每日处理 10 万亿 token,以及为何 AI 经济必须保持多元异构。

OpenRouter founder Alex Atallah discusses selling to Stripe for $7B+, serving 10 million developers and 10 trillion tokens a day, and why the AI economy must stay heterogeneous.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 17)

全文 · Full transcript(中英对照)

介绍与OpenRouter概览 Introduction and OpenRouter Overview

Host

Alex,欢迎来到我的节目。你是 OpenRouter 的创始人,你们刚刚达成协议被 Stripe 收购。你们服务 1000 万开发者,每天处理 10 万亿个 token。在此之前,你创立了 OpenSea,那是全球最大的 NFT 市场。所以过去几周你过得非常疯狂,今天能请到你我很兴奋。感谢你抽时间过来。对于不了解或一直与世隔绝的人来说,OpenRouter 是什么?你们在做什么?

Alex, welcome to my show. You are the founder of OpenRouter and you guys just entered an agreement to be acquired by Stripe. You serve 10 million developers and you process 10 trillion tokens a day. Before this, you founded OpenSea, which was the largest NFT marketplace in the world. So you've had a crazy couple of weeks and I'm so excited to have you here today. Thank you for taking the time to be here. For anyone that does not know or has been living under a rock, what is OpenRouter? What are you guys doing?

Alex

OpenRouter 是首个也是最大的 LLM 市场和网关,我们帮助人们路由到适合任务的正确模型。但我们做的很大一部分是帮助人们发现新模型,并弄清楚它们好在哪里。语言模型的问题是,你不能只在网页上罗列一个语言模型的所有功能、能力、优势和劣势。你必须使用它,必须看看别人是怎么用的,才能理解它能做什么,以及你如何从中获得价值。很长一段时间里,人们比较语言模型的唯一方式是静态偏好,比如基准测试,这些就像 SAT 一样的测试,你拿来跑模型,它们很容易被钻空子。它们可以被纳入训练数据。所以,虽然你可以在 OpenRouter 上探索基准测试,你仍然可以让你的智能体从中受益。我们喜欢基准测试,尤其是那些真正创新的社区基准测试。我们相信,开发者、企业和公司对语言模型和其他 AI 模型所显示出的偏好非常有价值,把它们集中放在互联网上的一个地方非常重要。我们刚开始时,模型在互联网上没有这样一个地方,让你能看到所有模型,并看到人们在现实世界中如何使用它们。这就是 OpenRouter 的起点。随着我们成长,我们在 8 月有一天处理了 19 万亿个 token,换个角度看,这大约相当于每 30 秒处理一个英文维基百科,而且过去一个月我们的 token 量环比增长了约 50%。随着我们成长,数据变得更有用,我们可以提供纯基准测试,帮助人们发现最适合他们用例的模型。我们还帮助企业监控、观察、了解其边界内的一切运行情况并进行治理。你必须限制人们能花多少钱,以及不同团队可以访问哪些模型。然后,我们帮助每个人为他们关心的每个模型获得最好的价格和最佳的性能。

So, OpenRouter is the first and largest LLM marketplace and gateway and we help people route to the right model for the job. But a big part about what we do is help people discover new models and figure out what's good about them. Like, the thing with language models is that you can't just enumerate all the features and capabilities and strengths and weaknesses of a language model on a web page. You have to use it and you have to see how other people are using it to understand what it's capable of and how you can get value out of it. For a long time, the only way people had to compare language models was static preferences like benchmarks, which are these just SAT-like tests that you run on the model and they're very gameable. They can be incorporated in training data. And so while you can explore benchmarks on OpenRouter, you can still let your agents benefit from them. And we love benchmarks, especially really innovative community ones. Our belief is that the revealed preferences that developers and businesses and enterprises have for language models and other AI models are really valuable and really important to have in one place on the internet. Models when we got started didn't have a place on the internet where you could see them all and see how people were using them in the real world. And so that's how OpenRouter got started. And as we've grown and we had a 19 trillion token day in August, which to put in perspective is about one English Wikipedia every 30 seconds, and we grew about 50% month over month in terms of tokens this past month. As we've grown, the data becomes more useful and we can provide pure benchmarks and help people discover the best models for their use cases. We also help enterprises monitor, observe, see how everything is going on within their borders and govern it. You have to put limits on how much people can spend and different teams and which models different teams can access. And then we help everybody get access to the best prices and best performance for each model that they end up caring about.

初衷与早期岁月 Motivation and Early Days

Host

你在开始构建 OpenRouter 时是什么心态?你觉得自己有什么需要证明的吗,还是只是感到兴奋?你刚刚已经打造了一个了不起的东西。你当时脑子里在想什么?

What mindset were you in when you were going into building OpenRouter? Did you feel like you had something to prove or did you just feel like excited? You had already just built something amazing. Where was your head at?

Alex

我真的很想开始一些新东西。我有点渴望在一个新领域里带一个小团队,老实说,这基本上就是我当时的想法。你知道,OpenSea,那是在 FTX 之前,我离开时还留在董事会。是的。我想休息几个月什么的,但我觉得 11 月 ChatGPT 发布时,大约是我离开 OpenSea 两个月后,我就想,哦,休假大概结束了。这是一个疯狂的新事物。我最大的问题是,这会不会成为一个赢家通吃的市场?这可能像谷歌式的垄断,所有价值都流向一家公司,就像荷兰东印度公司,但规模是它的千万亿倍。如果所有知识工作最终都依赖一家公司,那看起来非常危险,而且,你知道,我也很好奇,谁在使用其他模型实验室,比如谁在使用 OpenAI 以外的模型,他们用这些模型做什么,似乎没有人有关于这些的任何经验数据。所以我做了一个叫 Window AI 的实验,它有点受加密货币启发,有点像 MetaMask,但你可以自带用户,用户可以选择他们想与 Web 应用一起使用的模型,并直接将他们的模型选择带到那个应用,它允许那些想要模型无关或想让用户使用开放权重模型的应用出现。Alpaca 刚刚问世,那是我见过的第一个可用的开放权重模型。它只花了 600 美元就做出来了。是斯坦福的一个小团队搞出来的。它只是用一堆合成数据微调了 Llama。我就想,哇,如果这可行,那会有大量这样的东西,也许成千上万或数万个模型。这是一种新的数据销售方式。这是一种新的知识变现方式,也为经济创造了一个新的原始构建块。有没有一个主页?有没有什么方法可以发现这些东西,或者互联网上有没有一个地方可以找到它们?有没有什么方法可以获得高质量访问?有没有办法比较随时间出现的各种提供商?有没有办法监控或管理、观察和治理所有这些?没有。所以 OpenRouter 就是这样开始的。

I really wanted to start something new. I was kind of feeling the itch to have a small team in a new domain and that was mostly where my head was at honestly. Like, you know, OpenSea, this was before FTX when I left and I stayed on the board. Yeah. I wanted to take a few months off or something but I think in November ChatGPT launched about two months after I left OpenSea and I was like oh sabbatical is probably over. This is a crazy new thing. And my biggest question was like is this going to be a winner take all market? This could be like a Google style monopoly where all value flows to one company and it would be like the Dutch East India trading company but like times a quadrillion in magnitude. If all knowledge work ends up depending on one company that seems very dangerous and very, you know, I was also curious like who's using the other model labs like who's using anything other than OpenAI, what are they using them for like nobody seems to have any empirical data on any of this stuff. And so I made an experiment called Window AI that was kind of crypto-inspired, kind of like MetaMask but where you'd bring your own users could choose the model they want to use with a web app and bring their model choice directly to that app and it allowed apps to kind of emerge that wanted to be model agnostic or wanted to let users use open-weight models with them. Alpaca had just come out which was the first usable open-weight model I'd ever seen. It was made and it only took $600 to make it. It was a small team at Stanford that set it up. And it just fine-tuned Llama using a bunch of synthetic data. And I was like, wow, if this is possible, then there's going to be tons of these things, maybe like thousands or tens of thousands of models. And it's a new way of selling data. It's a new way of monetizing knowledge and also creating a new primitive building block for the economy. And is there like a homepage? Is there any sort of way to discover these things or any sort of place on the internet for them? Is there any sort of a way to get like high-quality access to them? Is there a way to compare the various providers that emerge over time? Is there a way to monitor or manage, observe and govern all of these things? No. So that's how OpenRouter got started.

竞争与差异化 Competition and Differentiation

Host

Sam Lessin 发推说:“OpenRouter 背后没有真正的技术。每家公司都会构建自己的。” Ramp,一家专注于完全不同领域的公司,也推出了自己的路由器。

Sam Lessin tweeted something like, "There's no real technology behind OpenRouter. Every company is going to build their own." Ramp, which is a company focused on an entirely different set of things, launched their own router.

Alex

我们最大的根本区别在于,这对我们来说不是支线任务。这是我们的主线任务。我 100% 专注的一切就是构建最好的市场网关和路由器。这是我们与 Stripe 最清晰的使命一致之一。我们都相信,增长全球 GDP 的最佳方式是构建异构经济,让价值从世界各个角落涌现,而不是只来自一家巨型公司、一家银行或一个模型实验室。

Big fundamental difference with us is that it's not a side quest for us. It is our main quest. Everything I'm 100% focused on is building the best marketplace gateway and router. This is one of the clearest mission alignments that we have with Stripe. We both believe that the best way of growing global GDP is by building heterogeneous economies that allow value to emerge from all corners of the world instead of just from one mega company, one single bank or one single model lab.

Stripe收购的合理性 Why the Stripe acquisition makes sense

Host

我不知道这笔交易是怎么促成的,但当你和 Patrick、John 坐下来一起梳理这件事时,为什么你觉得推进这次收购是合理的?然后你觉得为什么对他们来说也合理?

I don't know how the deal came together, but when you sat down with Patrick and John and thought through this, why did it make sense for you to go ahead with the acquisition? And then why do you think it made sense for them?

Alex

除了 Stripe,我想不出还有多少公司更擅长鼓励全球新企业的成长——真正帮助人们创建自己的企业、接受支付、把生意做起来。它一直是一家以自下而上的方式推动 GDP 增长的榜样公司。而在很大程度上,我们正试图为 AI 社区、为 AI 生态做同样的事——帮助人们获取模型,帮助人们为自己的用例发现合适的模型。这是我们与 Stripe 最清晰的使命契合之一。我们都相信,增长全球 GDP 最好的方式是构建异质的经济体,让价值从世界的各个角落涌现,而不是只来自一家巨型公司、一家银行或一个模型实验室——然后把所有这些异质的来源组织成一种极其易用、安全、可管理、有用的开发者体验和商业体验。这正是 Stripe 在过去十多年里为支付所做的事。而这也是我们为智能所做的事。所以,构建企业很大一部分不只是接受支付,现在还要获取智能。它既是新企业销货成本中最大的一项,也正迅速成为任何企业运营支出中最大的一项。AI 是今年美国 GDP 增长的最大贡献者,我认为它将成为经济中至关重要的一部分。除了接受支付以及资金如何在公司之间流动之外,智能如何被获取也将变得重要,两者会变得非常紧密。这就是为什么我们和 Stripe 在同一支团队里做这件事是合理的。幸运的是,我们保留了品牌、保留了愿景、保留了产品和路线图,以及所有模型——当然还有我们的模型中立性。这是两家公司都深信的东西:构建真正优秀、可信、中立的基础设施,覆盖这个我们想要保持异质的极其狂野的世界。我们希望它是多元的,我们希望新进入者能轻松地进来、创新、推向市场。你需要中立、可靠的基础设施才能让那个世界成为现实——并避免一个只有一家公司的世界。

I don't know very many companies other than Stripe that are better at encouraging the growth of new businesses around the world — really helping people create their own businesses, accept payments, and get them off the ground. It's always just been a role model company for growing GDP in a bottoms-up way. And in a big way, we're trying to help do that for the AI community, for the AI ecosystem — helping people access models and helping people discover the right models for their use cases. This is one of the clearest mission alignments that we have with Stripe. We both believe that the best way of growing global GDP is by building heterogeneous economies that allow value to emerge from all corners of the world, instead of just from one mega company, one single bank, or one single model lab — and then organizing all of those heterogeneous sources into an extremely easy-to-use, safe, manageable, and useful developer experience and business experience. That's a key thing that Stripe has been doing for payments for over the last decade. And it's what we do for intelligence. So a massive part of building businesses is not just going to be accepting payments. It's now sourcing intelligence. It's the biggest line item both for new businesses' cost of goods sold and is rapidly growing to be the biggest line item for any business on their operational expenses. So AI was the biggest contributor to GDP growth in the US this year, and I think it's going to be a critical part of the economy. In addition to accepting payments and how money flows between companies, it will also be how intelligence is sourced, and the two will become very tied. And that's a big reason it makes sense for us to do this on the same team with Stripe. And fortunately, we're keeping the brand, keeping the vision, keeping the product, the road map, and all the models — and of course our model neutrality. That's something both companies deeply believe in: building really good, trusted, neutral infrastructure across this extremely wild world that we want to be heterogeneous. We want it to be diverse, and we want it to be easy for new entrants to come in, innovate, and get to market. You need neutral, reliable infrastructure to make that world happen — and to avoid a world where there's just one company.

奇点与一月拐点 The singularity and the January inflection

Host

有一封泄露的投资人信说,他们相信奇点始于 2026 年 1 月 1 日。如果你认同这个说法,你觉得今年我们身处的这个起飞阶段,根本上发生了什么变化?

There was a leaked investor letter that said they believe the singularity started on January 1st, 2026. What do you think has fundamentally changed this year about this takeoff period we're in, if you agree with that?

Alex

嗯,我不对那封信发表评论——

Well, without commenting on the letter —

Host

我不是要你评论那封信。

I'm not asking you to comment on the letter.

Alex

我觉得有件非常——我同意。

I think something very — I agree.

Host

你以一种非常独特的方式看到这些数据。你实际上能看到人们怎么使用这些模型。他们在用它做什么?用哪些模型?

You see the data in a really unique way. You're actually seeing how people are using these models. What are they using it for? Which models?

Alex

如果你去我们的排行榜页面,有一个小按钮写着线性还是对数,针对我们的 token。如果你点对数切换到对数模式,你会看到它一直在增长,直到 1 月 1 日,然后斜率急剧上升,并且一直比之前更高,大概到 1 月 7 日左右。确实发生了一次增长的变化,一次加速,就在 1 月 1 日前后。所以就时机而言——

So, if you go to our rankings page, there's a little button that says linear versus log on our tokens. And if you click on log to switch it to log mode, you'll see that it was growing until January 1st, and then the slope dramatically picks up and has been consistently higher than it was up until January 7th or something like that. There really was a change of growth, an acceleration that happened around January 1st. So in terms of timing —

Host

新年决心,大家就像在说“多用点 AI”。

New Year's resolutions, people were like, "Use more AI."

Alex

是的。我记得新年之后或新年前后,我就想,“天哪,我得生成一些论文,因为我现在真的能和模型一起头脑风暴了。”而且它还挺有效的。Opus 能产出一些不错的想法。编码也好太多了。我觉得我们可以重构代码,而且大多数时候真的能指望它跑通。这以前一直是我对模型的个人评测:嘿,我能不能以结构上合理的方式重构代码库的这一部分,并且它还能正常工作?突然之间,这就能做到了。所以这相当不可思议。而且我认为它表明我们必须改变工作方式。所以我告诉整个公司,嘿,发生了疯狂的事。我们必须改变工作方式。如果你大多数时候没用智能体,如果你还在用自动补全,是时候切换了。别再自动补全了。我们必须利用云端智能体,不过本地智能体也行。我们必须更好地理解我们的速度,我们必须给自己更多杠杆。外面就是那样,而工具还很原始。所以我们开始——我们发布了智能体工具。我们发布了 Ory harness,它帮助人们非常轻松地把任何智能体与 OpenRouter 配合使用,任何智能体框架。我们开始在工具上深入,帮助开发者获取服务器工具和推理无关的基础设施,或者推理相邻的基础设施,这些确实能让你的最终结果好很多,有更好的 grounding 或与你的组织更相关。在很多很多方面,它改变了这家公司。

Yeah. Well, I remember after New Year's or around New Year's, I was just like, "Oh my god, I need to generate some papers because I can actually brainstorm with the models now." And it's kind of effective. There are some good ideas coming out of Opus. And the coding was so much better. I felt like we could refactor code and actually count on it to work most of the time. And that was always my personal eval with models before: hey, can I refactor this part of our codebase in a way that makes sense structurally and have it still work? And suddenly that was working. So it was pretty incredible. And I think it showed that we have to change the way we work. So I told the whole company, hey, something insane has happened. We have to change the way we work. And if you're not using agents most of the time, if you're using autocomplete, it's time to switch. No more autocomplete. We have to leverage cloud agents, but local agents are fine, too. And we have to better understand our velocity, we have to give ourselves more leverage here. It's out there and the tooling is just primitive. So we started — we've launched agent tools. We've launched Ory harness, which helps people use any agent with OpenRouter really easily, any agent harness. We started going very deep in tools and helping developers get access to server tools and inference-agnostic infra, or inference-adjacent infrastructure, that does make your end outcome a lot better with better grounding or better relevance to your org. And in many, many ways it changed the company.

回应“无真正技术”批评 Responding to the "no real technology" critique

Host

当这笔收购在 X 上宣布时,有很多噪音、很多评论、很多帖子。有一种观点,像 Sam Lessin 发推说,OpenRouter 背后没有真正的技术。每家公司都会自己造。Ramp,一家专注于完全不同领域的公司,也发布了自己的路由器。你怎么回应这个?

When the acquisition was announced on X, there was a lot of noise, a lot of comments, a lot of posts. And there was one take from people like Sam Lessin tweeted something like, there's no real technology behind OpenRouter. Every company is going to build their own. Ramp, which is a company focused on an entirely different set of things, launched their own router. What do you respond to that?

Alex

是的,现在有很多公司在造跟风的路由器。我认为我们一个根本性的不同在于,这对我们不是支线任务,而是主线任务。这是我 100% 专注的全部:打造最好的市场、网关和路由器。而拥有这种程度的专注,我认为会让事情发生巨大变化。这意味着我们在路由器的质量、市场的广度、我们制定的策略,以及我们招揽的人才上都领先。是专注在这个问题上的人才。

Yeah, there are a lot of companies building me-too routers right now. And I think a big fundamental difference with us is that it's not a side quest for us. It is our main quest. This is everything I'm 100% focused on: building the best marketplace, gateway, and router. And having that level of focus, I think, changes things dramatically. It means that we are ahead both in the quality of the router, the breadth of the marketplace, the strategy that we set, and the talent that we acquire. It's the talent that's focused on this problem.

OpenRouter生态使命 OpenRouter's Ecosystem Mission

Alex

我们和生态做了很多工作,因为 OpenRouter 是一家生态驱动的公司。我们希望帮助模型和提供商找到他们的受众。如果有人来自新实验室或新的推理提供商,有很棒的创新,我们希望他们尽快进入市场,我们希望全世界都知道他们。我们的使命是改善神经多样性和 AI。这让我们能够提供天然的市场优势:最大的市场有这种天然的护城河,所有有创新的实验室和提供商都希望加入。反过来,开发者和企业可以在 OpenRouter 上找到最好的价格、最多的种类和最高质量的推理。确保这种市场魔力能够规模化运作,是业务的核心焦点,而且像所有市场一样,它很难被复制。

We do a lot of work with the ecosystem because OpenRouter is an ecosystem-driven company. We want to help models and providers find their audience. Someone with a great innovation from a new lab or a new inference provider—we want them to get to market as quickly as possible, and we want the whole world to know about them. Our mission is to improve neurodiversity and AI. That allows us to provide natural marketplace benefits: the biggest marketplace has this natural moat where all the labs and providers on the supply side that have innovations want to be on it. In turn, developers and businesses find the best prices, the most variety, and the highest quality inference all on OpenRouter. Just making sure this marketplace magic works at scale is the core focus of the business, and like all marketplaces, it's pretty hard to replicate.

Host

你觉得会不会有一个世界,它真正多样化到所有类型的模型?一旦有了很棒的生物学模型和机器人模型,人们可以去 OpenRouter 访问一切?

Do you see a world where it really diversifies to all types of models? Once there are amazing biology-focused models and robotics models, people can go to OpenRouter to access everything?

Alex

是的。我们有种类繁多的非语言模型。有视频模型、图像模型、语音转文本、文本转语音、嵌入模型、重排序模型,现在还有非常令人兴奋的方式来比较它们。我们刚刚推出了这个基准测试浏览器,你可以按基准测试探索模型,并理解这些基准测试到底是什么意思。很多人引用 MMLU,却根本不知道一个 MMLU 问题长什么样。现在你可以看到了。我们还创建了自己的基准测试——我们有图像和视频基准测试,你可以看到所有模型尝试渲染一个半满的酒杯,你会看到模型就是不太能搞定,每个杯子里的酒量都不一样。但很容易就能看到并比较它们。我们只是找到这些很酷、非常通用的评估,帮助你选择你想选择的东西,也帮助你的智能体选择。所以你会发现非常酷的模型多样性,未来还会更加多样化。

Yes. We have an enormous variety of non-language models. There are video models, image models, speech-to-text, text-to-speech, embedding models, reranking models, and now there are really exciting ways of comparing them all. We just launched this benchmarks explorer where you can explore models by benchmark and understand what the hell these benchmarks mean. Tons of people reference MMLU without having any idea what a single MMLU question looks like. Now you can see. We also created our own benchmarks too—we have image and video benchmarks where you can just see all models try to render a wine glass that's half full, and you can see the models just not quite get it, with different levels of wine in each glass. But it's really easy to just see and compare them all. We just find these cool, very general evals that help you choose what you want to choose and help your agent choose too. So you'll find a very cool diversity of models that will get more diverse in the future.

代理经济未来 Future of Agentic Economy

Host

如果我能请你戴上一副未来护目镜,想象这种未来的智能体经济,每个人都有大量的智能体。我们刚在来的路上遇到一些人,每个人都在谈论 Instinct,以及现在人们如何拥有他们的智能体,可以真正进行购买并在世界上运作。再加上 Stripe 提供的支付基础设施。当你把这些了不起的公司拼凑在一起协同工作时,你认为你或 Patrick 和 John 想象的这个未来经济是什么样的?甚至他们收购 Bridge,也有一组非常清晰、有趣的事情。

If I could ask you to put on a pair of future goggles and imagine this sort of future agentic economy where everyone has a ton of agents. We just ran into people on our way in here and everyone was talking about Instinct and how now people can have their agents that can really make purchases and operate in the world. Plus the payments infrastructure that Stripe offers. What do you think is this future of economy that you or Patrick and John kind of imagine when you're piecing all these amazing companies to work alongside each other? Even the acquisition of Bridge that they made, there's sort of this very clear, interesting set of things.

Alex

是的。我的意思是,我有时和人们谈论智能体支付,人们会说,这永远不会成功。我喜欢购物。我想去亚马逊。我想在买之前看到我要买的东西。我永远不会用这个。而我个人不是那样。我知道如果智能体支付运作得很好,我会买很多东西。我知道我终于可以按计划收到牙膏,而不是每两个月走过街去药店买。如果我能用智能体做到,并且只用文本就能取消或修改计划。我知道我甚至会对衬衫这样做。

Yeah. I mean, I sometimes talk to people about agent payments and people are like, this is never going to work. I enjoy shopping. I want to go to Amazon. I want to see the things that I'm buying before I buy them. I would never use this. And I personally am not like that. I know I would buy a lot of stuff if agent payments worked really well. I know I would finally get my toothpaste delivered on a schedule instead of walking across the street to the pharmacy to buy it every two months. If I could do it with an agent and cancel the schedule or modify the schedule with just text. And I know I would even do that for shirts.

Host

我们刚才还在说,你几乎穿了六年前在我们播客上穿的那件衬衫。他可以多买几件衬衫。

We were just saying you almost wore the same shirt that you wore six years ago on our podcast. He could get a few more shirts.

Alex

我没有那么多衣服。如果我的智能体像你一样知道我的尺码,我肯定会买更多衣服,因为我过去订购过东西,它知道一些关于我的事情,也知道一些关于世界上重要品牌的事情。

I don't have that many clothes. I would definitely get more clothes if my agent knew my size like you, because I ordered something within the past, and knew some things about me and knew some things about brands that matter around the world.

Host

有研究证明,人们享受计划旅行的过程,胜过实际旅行。有些人从这些体验中获得很多感激。我不是那种人。仍然有一些体验我可能想手动做,但在 AI 的许多事情中,有你做并且喜欢的任务范围,然后有你不喜欢的任务范围,你甚至不知道那个范围有多宽,因为你不做那些任务。你不理解那部分有多大,因为你自然地避免它。所以如果一个智能体要为你接管那整套任务,我认为它会比所有人想象的都要大得多,因为每个人都在低估它。你不知道你会为你不喜欢购买的东西做多少购买,因为你不做,所以你无法知道它有多大。

And there are studies that prove that people enjoy the process of planning a trip more than the actual trip. Like some people find a lot of gratitude in those experiences. I am not one of those people. There are still some experiences where I might want to do it manually, but with a lot of things in AI, there's the span of tasks that you do and love, and then there's the span of tasks that you don't love and you don't even know how wide that span is because you don't do those tasks. You don't understand the magnitude of how big that section is because you avoid it naturally. So if an agent is going to take over that whole set of tasks for you, I think it's going to be way larger than everybody thinks because everybody's underestimating it. You don't know the magnitude of purchases you're going to make for things you don't like to purchase because you don't do it, so you can't know how big it is.

Host

多买点。

Buy more.

Alex

是的。是的。我认为人们会买更多,人们会买他们通常不买的东西,人们会以程序化方式购买他们通常不买的东西,或者创建团购实体,为整个群体购买一切。人们会送更多礼物,人们会使用更多可支配收入,人们会捐赠更多,人们可能会退回更多商品,如果退货变得更容易。所以有各种 1% 问题的规模,一旦你构建了一个非常好的产品,针对人们不做的 1% 用例,那 1% 就会爆炸成一个更大的百分比。这在 AI 中不断发生。就像,哦,人们自己无法完成的智能任务的 top 1% 突然爆炸成为他们的大部分支出,因为现在做研发和发现并在那里花钱效率高多了。

Yeah. Yeah. I think people will buy a lot more and people will buy things they don't normally buy and people will programmatically buy things that they don't normally buy or create group buying entities that buy everything for a whole group. People will give more gifts, people will use more of their disposable income, people will donate more, people will possibly return more goods if it gets easier. And so there's all kinds of magnitude of the 1% issues where once you build a really good product that targets the 1% of use cases that people aren't doing, that 1% blows up into a much bigger percent. This is constantly happening with AI. Like, oh, the top 1% of intelligence tasks that people aren't able to do themselves suddenly blows up to become most of their spend because it's now just so much more efficient to do R&D and discovery and spend there.

模型偏好与品牌 Model Preferences and Branding

Host

上周,OAX Alpha 有点病毒式传播,基本上是一个匿名模型。没人知道是谁做的。现在,我们弄清楚了它来自 Z.AI。观察那个经历,它向你证明了什么关于模型偏好,以及人们有多在乎品牌和这些事情?

This past week, OAX Alpha went a bit viral, which was basically an anonymous model. No one knew who made it. Now, we figured out that it's from Z.AI. What did watching that experience prove to you about model preferences and how much people care about brand and some of these things?

Alex

是的。所以,当它处于隐身模式时,那是我们 2024 年启动的一个项目。

Yeah. So, when it was in stealth mode, which is a program that we have that we started back in 2024.

隐身模型计划 Stealth Model Program

Host

那你为什么要启动这个项目?

And why did you start that?

Alex

我们和 OpenAI 合作,为即将发布的 GPT-4.1 做了这个隐身模型项目。我们想做的,基本上是让用户在不带品牌偏见的情况下去评估一个模型,同时给模型实验室提供一种更好的、能拿到有意思反馈的方式。另外也是想创造一种途径,让人们能自己构建基准测试,看看这个隐身模型在社区基准上的表现——而不是模型实验室内部通常评估的那些东西。所以这是一种很好的方式,能在不带人们固有偏见的情况下观察显示性偏好。后来我们把这个项目延续了下来,和很多模型实验室都做过。上个月我们还和 Z.AI 合作,为他们的 GLM 5.3 Flash 发布做了这个。那个模型叫 Ox Alpha。

So we worked with OpenAI to do the stealth model program for the upcoming launch of GPT-4.1. And we wanted to basically let users evaluate a model without showing bias around the brand that made it, and get a better way of providing interesting feedback to the model lab. But also to create a way for people to build their own benchmarks and see how the stealth model performs on community benchmarks — benchmarks other than the things the model labs normally evaluate internally. So it was a good way of doing revealed preferences without the implicit bias that people have. And we've continued the program and done it with many model labs since then. And this past month we did it with Z.AI for their launch of GLM 5.3 Flash. The model was called Ox Alpha.

Host

哦,Ox Alpha。

Oh, Ox Alpha.

Alex

对,Ox Alpha。很多人试用了它。试用它的独立用户数超过了任何前沿模型。它非常非常受欢迎。按 token 量算,这是 OpenRouter 历史上规模最大的一次模型发布,而且它确实非常强。我们还把隐身模型免费提供出来,很多开发者发现它在长周期任务上非常强,文笔也很好。

Yeah, Ox Alpha. And a lot of people tried it. More unique users tried it than any frontier model. It was very, very popular. It was the largest model launch in OpenRouter history in terms of tokens, and it was just very capable. We also offer the stealth models for free, and a lot of developers just found it very capable at long-horizon work and very good at prose as well.

Host

那它退出隐身模式之后,相比其他模型还是相当便宜,对吧?

And when it exited stealth mode, it's still pretty cheap compared to other models, right?

Alex

对。GLM 5.3 Flash 的每任务成本表现非常非常好。我记得他们在头四天里跑了大约 5 万亿 token,这是非常非常高的。模型公开之后,业内很多人非常关注每 token 成本,也就是每 token 价格。我们正在引导行业去思考每任务成本和每会话成本,因为不同模型在 token 效率上的表现差异太大了。真正重要的是完成任务要花多少钱,而不是每 token 的价格。所以如果你去看我们的排行榜页面或模型页面,现在就能看到估算数据:用每个模型在 Claude Code 或 Codex 或任何智能体里跑一次合理会话,实际要花多少钱。如果你今天去我们排行榜页面的那个位置看,就能看到 GLM 5.3 Flash 排名非常高。

Yeah. GLM 5.3 Flash is very, very good cost-per-task performance. And I believe they did about five trillion tokens in their first four days, which is very, very high. After the model was revealed, a lot of the industry is very focused on cost per token, like price per token. We're guiding the industry towards thinking about cost per task and cost per session, because the models have such different profiles around token efficiency. And it matters how much it costs to complete the task, not price per token. So if you look on our rankings page or on our model pages, you can now see estimates of how much it actually costs to have a reasonable session in Claude Code or Codex or any agent using each model. And if you go to that spot in our rankings page today, you can see GLM 5.3 Flash ranked really high.

Host

哇。看到围绕它的那种文化讨论也很酷。大家都很兴奋、很好奇。那非常有意思。如果你把眼光放得很远,你觉得人们会不会就只为结果付费?就像电一样。你不在乎它是哪里发的、怎么发的,你只是说:“哦,我愿意为得到这个结果花这么多钱。”然后你们来处理这一切。于是人们就不在乎品牌、前沿模型、中国开源权重模型,或者现在这些很重要的各种概念了。

Wow. It was also just cool to see that cultural conversation around it. Everyone was so excited and curious. And that was very interesting. If you were to look very far out, do you think people will just pay for an outcome? Like this will all be sort of like electricity. You don't care where it's made or how it's made, but you just are like, "Oh, I'm willing to spend this much on getting this outcome." And then you guys handle all of that. And so people don't care about brand or frontier model or Chinese open-weight model or these different ideas that are a big deal right now.

Alex

我觉得两种市场都会存在——一种是在乎模型、在乎背后用的是哪种智能的人,另一种是不在乎的人。比如很多消费级应用,人们其实不太在乎底层用的是哪个模型。他们只在乎事情有没有办成。B2B 就有点不一样。企业确实倾向于在乎模型,而且实际上想自己去理解并构建路由器。我相信大多数公司都会变成一种新型的模型路由器。我们从 2024 年左右就一直这么说——所有初创公司、所有软件初创公司都会变成模型路由器。他们想知道自己在用哪些智能输入,也想理解这些选择如何影响他们的销货成本(COGS),以及他们为产品努力最大化的关键绩效指标(KPI)。而选择模型、理解模型如何被提示,就是这些等式里的输入。所以作为企业,你想要控制权;但作为终端用户,你有时候就没那么在乎。

I think there will be markets for both — people who care about the model and the kind of intelligence being used behind the scenes, and people who don't. Like a lot of consumer apps, for example, people don't really care about which model they're interacting with under the hood. They just care about the job getting done. B2B is a little different. Businesses do tend to care about the model and actually want to understand and build routers themselves. Like I believe most companies are going to become model routers of a new type. And we've been saying this since like 2024 — like all startups, all software startups will become model routers. They want to know which intelligence inputs they're using, and they want to understand how those choices affect their cost of goods sold, their COGS, and their key performance indicators, their KPIs that they're trying to maximize for the product. And choosing the models and understanding how the models are prompted are the inputs to those equations. So you want control as a business, but you don't really care as much as an end user sometimes.

网络的代理驱动未来 The Web's Agent-Driven Future

Host

你觉得未来 10 年网络会怎么变化?它会为人类优化,还是为智能体优化?

And how do you think the web changes over the next 10 years? Like is it going to be optimized for humans or for agents?

Alex

嗯,从传输字节数来看,智能体对网络的使用量已经非常巨大了。我觉得很多公司会意识到,为智能体优化会成为他们增长支付量的最快方式。嗯,你知道,当人们可以直接从智能体那里购买东西的体验越多,为智能体优化就越合理。而为智能体优化的人越多,就会有越多公司涌现出来,让人们用智能体买东西。所以这两个部分就像一个向上的良性循环,带来越来越多的智能体 GDP。嗯,我认为这个驱动力会让智能体成为……你知道,可能成为 GDP 增长的最大构建者。我我觉得,人类这边,人类经济仍然会有一个很大的市场,但你知道,这非常合理。如果智能体在不久的将来与大部分 GDP 增长相关,我一点也不会感到惊讶。

Well, we're seeing pretty massive usage of the web by agents in terms of bytes transferred and you know like I think a lot of companies are going to realize that like optimizing for agents becomes their fastest growing way of op of of of growing their payment volume. Um, you know, the more like experiences that exist where people can buy things directly from an agent, the more it makes sense to optimize for agents. And the more people optimize for agents, the more companies will start up that let people buy things with an agent. And so those two components are like a virtuous cycle upwards of more and more agent GDP. Um, that's the the driver that I think will like cause agents to become, you know, potentially the biggest uh architect of GDP growth. And I I think like humans there will still be like a big market for like the human economy, but you know, it would make a lot of sense. I would not be at all surprised if agents, you know, are tied to most GDP growth in the near future.

稳定币与代理支付 Stablecoins and Agentic Payments

Host

你觉得这是不是那种加密货币最能发挥作用的场景?比如人们会使用稳定币,而不仅仅是你知道的跨境支付和它今天被使用的那些地方?

Do you think this is sort of the the situation where like crypto makes the most sense like it will find like people will be using stable coins more than just you know across country payments and sort of where it's being used today?

Alex

嗯,我认为现有的稳定币可能更适合智能体支付。部分原因是它们比美元更容易管理,也更具确定性。嗯,我觉得我们会看到很多新兴用例,其中兑换成稳定币的便捷性变得非常高,以至于消费者甚至不知道是稳定币在支撑这个产品。你知道,他们可以在任何地方使用自己的资金,但如果他们直接使用稳定币,可能会获得更好的经济性和更低的手续费,企业也会利用这一点来激励人们转向稳定币。所以,我认为我们会看到更多的采用。

Well, I think the the stable coins that exist are probably better for agentic payments. Um, in part because they are so much easier and so much more deterministic to manage than dollars. and uh and you know I think we'll have like a lot of emerging use cases where the the ease of converting to a stable coin become so so high that like consumers don't even know that stable coins are powering um the product and you know they can use their funds wherever they are but if they if they use stable coins directly, they get potentially better economics and lower fees and businesses use that to kind of like go down and incentivize people to move to stable coins as well. So, I I think we'll see more adoption.

GDP增长与代理信任 GDP Growth and Agent Trust

Host

你对 GDP 增长有什么预期?比如过去几年我们看到的,对比未来三年,因为有一种观点是:我们是否已经实现了 AGI?它会很快发生吗?会在一年内发生吗?有很多关于递归自我改进的讨论,以及这些潜在的巨大进步,然后这如何转化为我们使用智能体的方式,以及我们对它们的信任程度。

What do you expect for GDP growth? Like what we've seen the past couple of years versus the next three years because there's this idea of like have we already achieved AGI? Is it happening soon? Is it happening in a year? There's a lot of uh discussion about recursive self-improvement and like these sort of potential massive steps and then how does that translate to the way we can use agents and how much we trust in them.

Alex

我认为这些模型进步会让智能体更值得信赖、更高效,并且总体上能力更强。更值得信赖意味着,对于我知道它们能完成的任务,我会觉得,哇,这将会非常可靠。比如,我知道它以前做过很多次。我的朋友告诉我它做过很多次。嗯,我知道它会立刻完成。我知道它会以低成本完成。我会把它交给一个智能体。嗯,更高效意味着当智能体完成时,我最终得到的比之前版本的智能体或之前的 RSI 前模型要多得多。嗯,然后更通用智能。我可以给它更多测试,我的智能体也可以给它任务,它可以递归。它可以用于流水线,可以作为另一个业务的组件,而那个业务又是另一个业务的组件。你可以更可靠地堆叠这些组件,并降低价值链中每一层的价格。

I think these model advancements will make agents more trustworthy um more effective and more generally capable basically. So more trustworthy means that, you know, for the task that I know they can achieve, I feel like, wow, this is just going to be so reliable. Like, I know it's done this many times before. My friend tells me it's done it many times before. Um, I know it's going to do it right away. I know it's going to do it at a low cost. I'm going to give it to an agent. um uh more effective means that when the agent does it, I end up getting a lot more out of it uh than I was in the previous version of the agent or previous like pre RSI models. Um and then more, you know, more generally intelligent. I can just give it more tests and my agents can give it tasks too and it can be like recurs. It can be used in pipelines that it can be used as a component to another business that's used as a component to another business. And you can stack these components much more reliably um and drive the price down of each layer in the value chain.

人类优势与代理卸载 Human Edge and Agent Offloading

Host

你认为我们最后才会委托给智能体的事情是什么?我的意思是,哪些任务仍然以人类为核心,并且在你工作中人类仍然做得好得多?过去两年里,你有没有能够卸载掉的事情?当我说卸载,我指的是交给智能体去实际执行,而不是 Alex 必须亲自作为核心的事情?

What do you think are the last things that we will entrust agents with? Like what I mean is like what what tasks are humans still key to and so much better at like in your job? Are there things you've been able to offload over the last two years? And when I mean offload, I just mean like hand to an agent to actually carry it out versus like what does Alex have to be at the core of?

Alex

人类喜欢彼此一起做的事情,以及人类喜欢自己拥有并愿意花钱自己拥有的东西,可能仍然会有很多价值。你知道,他们实际上想成为批准者,而不是让智能体成为批准者。嗯,他们想拥有品牌。他们想说是他们创造了品牌。他们想说是他们设计或构想出了某个概念。他们想,你知道,他们想获得名声或功劳。嗯,我觉得,你知道,很多……我认为人类可能会被这些吸引,因为智能体正在承担如此多的实际知识工作,以及软件工作,最终还有机器人硬件工作。嗯,而且这并不全是反乌托邦的,人类基本上会不断向上移动到价值创造的更高层,他们总是试图思考:智能体不做什么?模型不思考什么?我的意思是,它们不会思考我喜欢做的事情,不会思考我大脑中正在发生的事情,以及我如何帮助那些与我有相似经历的人。

There's probably going to be a lot of value in the things that humans like to do with each other and um and uh and that humans like to own themselves and will pay to own themselves. you know, they they actually want to be the approver rather than having an agent be the approver. Um, they want to own the brand. They want to say they created the brand. They want to say they like, you know, designed or envisioned some concept. They want to, you know, they want to claim the the the fame or claim the credit. Um, and I like I think you know a lot it'll be like humans will I think we'll probably gravitate to those things because agents are taking on so much of the actual knowledge work and and um and software work and eventually hardware work with robotics. Um and uh and and and like it's not all dystopian like humans will basically move up the stack of value creation I think continuously where they're just always trying to think of like well what are the agents not doing like what are the models like not thinking of I mean they're not thinking of like the things I like to do um the things that are going on in my brain and uh how can I like help people who who have similar experiences to me.

工程师即科学家与上下文 Engineers as Scientists and Context

Host

我读了你发表的一篇旧博客,在 2024 年你写道,工程师将成为科学家,人类通过从非数字场所获取上下文来保持优势。嗯,你仍然同意这一点吗?你从哪里获取你的上下文?

So, I read an old blog that you published and in 2024 you wrote that um engineers will become scientists and that humans keep an edge by acquiring context from non-digital places. Uh do you still agree with that and where do you acquire your context?

Alex

是的,我仍然同意这一点。我认为,比如我们的很多工程师基本上一直在我们的博客上写论文,关于对模型和网络搜索提供商进行基准测试,以及不同的开放权重模型提供商在不同推理质量维度上的表现。嗯,成为一名严谨的科学思考者现在非常受欢迎。嗯,这是 OpenRouter 非常重要的一部分,也是公司的一大支柱,即极高质量的基准测试,开发者可以在此基础上构建。这样他们就可以依靠 OpenRouter 进行基准测试,并为其用例找到最高质量的推理和最高质量的基础、最高质量的模型。然后,你知道,我们鼓励开发者构建自己的基准测试和评估。嗯,我认为这就是价值如何堆叠起来的方式,你不必每次都从大楼的一楼开始。你可以从 10 楼开始,更快地向上构建。所以,嗯,我我觉得,成为一名优秀的科学思考者,知道如何,你知道,对你试图提出的论点进行 steelman。嗯,知道如何理解方差和置信水平,以及你知道最终结论中的潜在弱点。这些都是你需要训练自己变得更好的肌肉。嗯,我们现在有工程师在做这样的工作,嗯,我们也在招聘传统的研究人员。

Yeah, I do still agree with that. I think uh a lot of our engineers for example have been writing papers basically in our blog about benchmarking uh models and web search providers and like like different openweight model providers along different like inference quality dimensions. Um and like being a rigorous scientific thinker is now like very in demand. Um like it's a very important part of open router and very like big pillar of the company is extremely high quality benchmarking that developers can build on top of. So they can count on open router to benchmark and find the highest quality inference and the highest quality you know grounding the highest quality models for their use cases. And then you know we encourage developers to build their own benchmarks and eval um and I think that's like how value stacks up where like you don't have to start on floor one of the building every time. You can start on floor 10 and and build upwards more quickly. So, um, I I I think like being a good scientific thinker, like knowing how to, you know, steelman the other side of, you know, an argument that you're trying to make. Um knowing like how like understanding variance and you know confidence levels and um you know potential weaknesses in you know in the you know the end conclusion. Like these are all things that are muscles that you need to like train yourself to get better at. and um and like we have engineers now like doing that work um and we're also hiring you know traditional researchers as well.

预测下一个大事件 Predicting the Next Big Thing

Host

你联合创办了 OpenC,现在又在做 Open Router。我觉得你很擅长深入挖掘,真正看清开发者如何工作的趋势,以及市场上缺什么。如果你向前看——这对正在听、想创业的人来说可能是个启发——有没有其他领域让你觉得“下一个开放的就是这个”,在哪个行业?

You co-founded OpenC and you're building Open Router today. I think you're someone that is very good at diving deep and really seeing these trends of how developers work and what was missing. If you were to look forward — and this might be an idea for someone else listening who wants to start a company — are there any other things where you're like, the next open is what, and in what sector?

Alex

嗯,我觉得未来三到四年会非常疯狂。很多事情都会改变,现在真的很难预测,因为我们每四个月就能看到显著的智能进步。寻找创业公司的一个方法是,找出最佳可能答案与你认为 12 个月后模型给出的答案之间的瓶颈——不只是答案,还有行动、服务,以及任何人类或智能体可能付费的可变现的东西。所以机器人领域还有很多事要做。我们还没有一个人人都有的消费级机器人。这看起来很疯狂。为什么会这样?你可以问很多问题,也有很多深入的方式。也有一些有前景的公司。还有上下文——我们如何让模型获得更好、更连续的上下文来帮助回答问题。当模型在记忆方面变得更好时,对于一个对过去发生的一切都有非常非常好、甚至完美记忆的模型,需要什么样的基础设施?需要什么样的连接?人类和智能体之间需要什么样的协作?我认为你也可以花时间思考,很多任务将被解决并变得确定,我们会有擅长做这些特定事情的模型,但许多任务将是非确定性的、研究性的,全都是关于弄清楚一些极其有价值、极其困难的问题,这些问题本来需要 2 万美元或 2000 万美元的推理才能完成。但让我们试着在一天内完成。思考这些问题,以及思考当一切都能在一天或一周内完成时会发生什么,真的很难。很多人不知道他们可以直接让 ChatGPT 或大多数智能体阅读他们关心的所有新闻来源,并每两天发给他们一份摘要。有时候你只需要问出这个问题。

Well, I think the next three to four years are going to be really insane. A lot of things are going to change, and it's really hard to make predictions right now because we're still seeing significant intelligence advancements every four months. One way of looking for companies is to figure out what the bottleneck is between the best possible answer and where you think answers from models are going to wind up in 12 months — not just answers but actions and services and any kind of monetizable thing that a human might pay for or an agent might pay for. So there's still a lot to be done in robotics. We don't have a consumer robot that everybody has. Seems crazy. Why is that? There are a lot of questions you can ask there and a lot of ways to go deep. There are also some promising companies too. There's context — how can we get models better and more continuous context to help answer their questions. When models get better at memory, what will be needed for a model that has really, really good, even perfect memory about everything that has happened in the past? What kind of infrastructure will be needed? What kind of connections will be needed? What kind of collaboration will be needed between humans and agents? I think you can also spend time thinking about how a lot of tasks will be solved and deterministic, and we'll have models that are good at doing these specific things, but many tasks will be non-deterministic and researchy and all about figuring out some incredibly valuable, incredibly difficult problem that would have taken $20,000 or $20 million in inference to do. But let's try to do it in a day. Thinking about those problems and thinking about what happens when it all becomes doable within a day or within a week is really hard. A lot of people don't know they can just ask ChatGPT or most agents to read all the news sources they care about and send them a summary every two days. You just have to ask the question sometimes.

Host

是的。所以我认为保持非常好奇是一个很好的通用策略,同时思考在价值曲线末端被解锁的那一小部分任务也是另一个有用的策略。

Yeah. So I think being very curious is just a good general tactic, and also thinking about the little portion of tasks that get unlocked at the end of the value curve is another useful tactic.

Alex

我采访过一位教授,他实际上与一些实验室合作,分析人们实际如何使用他们的模型。他的一个结论是,很多模型改进并不在于模型本身的改进,而在于用户真正学习模型以及如何提示它。我仍然认为,在大多数人知道如何与模型互动、如何提示它以及如何最好地使用它之间,还存在很大差距。这对大多数人来说并不是超级明显的。它并不是完美地为人类创造的。你必须学习它。而且随着模型如此快速地改进和变化,这也很有趣。你必须不断重新学习这个过程吗?当人们对某些模型产生强烈依恋时,我认为这会很有趣——如果它们改进得非常快,人们会感觉到变化吗?比如从 9 岁起,他们就用这些模型来谈论自己的问题,那么模型如何在与那个人的互动方式上保持很多一致性?这些想法很有趣。

I interviewed a professor who actually worked with some of the labs on analyzing how people are actually using their models, and he had a takeaway where so much of the model improvement isn't about the model improvement. It's about that user really learning the model and how they prompt it. And I still think there's a big gap there between the way most people know how to interact with the model and how to prompt it and how to use it best. It's not actually super obvious to most people. It's not perfectly created for humans. You have to learn it. And it also is interesting as the models improve and change so quickly. Do you have to keep redoing that process? And as people get super attached to certain models, I do think it's going to be interesting — if they are improving really fast, do they feel a change when it's like, since they're 9 years old, they use these models to talk about their problems, and how does that model maintain a lot of consistency in the way it interacts with that person? And these ideas are interesting.

Host

是的。人们与很多这样的模型进行非常私人的对话,这——你怎么看?

Yeah. People have very personal conversations with lots of these models, and it is — what do you think about that?

Alex

当我和朋友在一起,问他们“你的 ChatGPT 或 Claude 或聊天记录是什么样的?”时,看到别人的会话是极其酷的。这就像窥视某人的思想。太疯狂了。顺便说一句,这真的很强烈。我有一个朋友,她的男朋友读了她的 ChatGPT 历史记录,而不是读她的手机。我希望他们不会看到这个。

Extremely cool to see somebody else's ChatGPT session when I'm with a friend and just asking them, what is your chat or Claude or chat look like? And it's like peering into somebody's mind. It's crazy. By the way, it's really intense. I had a friend whose boyfriend read their ChatGPT history instead of reading their phone. And I hope they don't watch this.

Host

天哪。

Oh my god.

Alex

实际上,这比你读到的任何短信都更具侵入性,因为大多数人并不是——

Actually, that is so much more invasive than any text you could ever read because most people are not —

Host

是的。

Yeah.

Alex

你知道,这以一种疯狂的方式窥视某人关于一段关系的最深层问题和疑虑以及真相寻求。我在实验室的朋友们对我向 ChatGPT 敞开心扉的亲密程度感到震惊,我需要收敛。但我只是觉得这是一种——我不知道。这是一种免费的来源,让人能够吸收大量信息,并以某种结构吐出来,我认为这对人们来说感觉很好。很多时候,当你与某人交谈时,你实际上在寻找的就是这个——被倾听,或者有人让你的问题感觉有结构、可行,并有前进的方向。所以这会很有趣。在这一点上,因为 Open Router——不是每个人都用它来编程;人们也用它来做其他活动。

You know, it's peering into someone's actually deepest questions and qualms and truth-seeking about a relationship in kind of an insane way. And I have friends at the labs that are shocked with the level of intimacy in which I open up to my ChatGPT and I need to not. But I just find it to be this — I don't know. It's this free source of someone being able to take in a lot of information and spit it out in some type of structure, which I think feels really good for people. And so often when you're talking to someone, that's actually what you're looking for — being heard or for someone to make your problems feel structured, doable, and have something to move forward with. So it'll be interesting. And on this point, because Open Router — not everyone is using it for coding; people are using it for other activities as well.

Host

我敢肯定,会有一系列诉讼,甚至围绕 Character AI 发生的事情,比如当人们深深依恋模型并自残或做危险的事情,人们发现他们与 AI 有过关于这些事情的对话。在这些情况下,你如何看待模型问责制?因为我敢肯定这将是一个即将出现的问题。

There's going to be a set of lawsuits, I'm sure, around even what happened with Character AI, like when people become deeply attached to models and self-harm or do dangerous things, and people find that they've had conversations with their AI about these things. How do you think about model accountability in these situations? Because I'm sure this will be an upcoming issue.

Alex

显然,对于某些类型的对话,模型要正确处理,风险非常高。我认为人们——当你有很多自我怀疑时,当你不信任自己的直觉时,你要么陷入螺旋、毫无进展,要么寻求帮助,向人类寻求建议。除了提供很多好的结构外,模型还提供了一种更容易缓解自我怀疑的方式,让你感觉自己对观点和直觉进行了事实核查、直觉核查和氛围核查。它取代了很多人与人之间的互动和人与人之间的建议。

Obviously, the stakes were really high for some types of conversations to get it right for the model. I think people — when you have a lot of self-doubt, when you don't trust your instinct, you either spiral and go nowhere or you ask for help, you ask a human for advice. And in addition to providing lots of good structure, the models also provide a way easier way of relieving your self-doubt, feeling like you fact-checked and gut-checked and vibe-checked your opinions and your instinct. And it replaces a lot of human-to-human interaction and human-to-human advice.

人类错误与问责 Human Mistakes and Accountability

Host

人类在给别人建议时也会犯大量错误。

Humans make tons of mistakes too when they give advice to people.

Alex

当然,这也很大程度上取决于刚刚发生的事。就像我们忍不住会把自身的偏见带进给别人的建议里。

Of course, it's so based off of the thing that just happened too. Like, we can't help but put our own biases into our advice for others.

Host

是啊。

Yeah.

Alex

我是说,人类也会犯大量错误。平均而言,模型犯的错可能比人类还少。但模型缺少人类拥有的一样东西,就是问责。人类要为自己给一个真正陷入困境的人的建议负责,能承担责任,能亲自去见对方。还有很多面对面的互动,模型也做不到。

I mean, humans make tons of mistakes, too. Probably on average, the models will make fewer mistakes than humans. But what they don't have that humans do have is accountability. Like a human is on the hook for the advice they give someone who's really struggling, and can take responsibility and can go in person to meet with them. And there's a lot of in-person interaction the models can't do too.

Host

而且我有时发现,我会提示它来推我一把。我会说,我想让它反驳我,因为之前——

And sometimes I found, like, I'll prompt it to push me. So I'll say, I want it to push back on me because before—

Alex

我最近在这些事上用得没那么多了。但在经历不同的分手时,我确实发了很多消息。

I haven't been using it as much recently for these things. But throughout different breakups, I definitely text a lot.

Host

然后就像,你知道,反驳这一点,我这样看问题。我是不是漏了什么?然后它就会改变回答。我就想,好吧,难道我还得告诉你去反驳吗?你懂我意思吧?之前它挺肯定的,所以有点——

And it was like, you know, push back on this, this way I'm looking at it. Is there something I'm missing? And then it will change its response. I'm like, well, did I have to tell you to? You know what I mean? Before it was pretty affirmative, so it's kind of—

Alex

而且,显然我有偏见,但我觉得最酷的事情之一,就是看不同的模型如何以不同方式回应你,就像给你组一个专家小组——不同的模型。

And I mean, obviously I'm biased, but I think one of the coolest things to do is to see how different models respond to you differently and to get like a panel of experts on your—different models.

Host

如果你必须——我敢说你没试过——我甚至不知道你有没有对模型敞开心扉过。我是说,完全抛开编程,把编程和做东西放一边。

If you had to—I'm sure you haven't gone and—like I don't even know if you ever open up to the models. I mean, totally outside of coding, putting coding and building things aside.

Alex

我个人觉得 Fable 最擅长头脑风暴,感觉它更能汲取一些想法。给模型做基准测试真的很难。所以很难判断我是不是对它有些偏见。它是最贵的模型,这让一些人觉得它也是最大的那个。更大的模型我会认为有优势,能更有创造力,凭空拉出更疯狂的想法。所以如果我对 Fable 有偏见,觉得它更擅长汲取创意,那也说得通。GPT 5.6 Soul 在语言上更容易控制,感觉更可操控一些。Kimmy K3 我觉得语调和语气也很好,同样可操控。有一次我妈妈想弄清楚我的狗出了什么问题,所有医生都不知道。它有神经方面的问题,走不了路。我就把所有症状输入那三个模型,然后把结果都发给我妈,我说,嘿,哪个听起来最好?结果她选了 Kimmy K3,让我很惊讶。其实是 Kimmy K 2.5。但这说明语气很重要,即使所有想法都相当不确定。

I found Fable to be the best at brainstorming personally, where it just feels like it's drawing on ideas a little. It's really hard to benchmark the models. So it's hard to know if I just don't have some bias towards it. It is the most expensive model, which makes some people think that it's the biggest one as well. Like bigger models I would think would have an advantage of being a little more creative and pulling crazier ideas out of thin air. And so it would make sense if I have a bias towards Fable at just being better at drawing creative ideas. GPT 5.6 Soul is much easier to control the language. It feels a little bit more steerable. And Kimmy K3 I think also has very good tone and voice and is similarly steerable. One time my mom was trying to figure out what was wrong with my dog and none of the doctors knew. She had a neurological issue and wasn't able to walk. And I just put all the symptoms into those three models and sent them all to my mom and I was like, hey, which of these sounds best? And I was surprised that she picked Kimmy K3. Or it was Kimmy K 2.5 actually. But it shows that the voice matters a lot, even when all the ideas are pretty uncertain.

Host

我的最后一个问题是这个。在我们的生活中,尤其是在旧金山以外,当我在洛杉矶甚至纽约时,我和朋友们出去玩,他们真实的感受就像,我讨厌 AI,因为我没有要求这种反乌托邦式的转变发生在世界上。我喜欢现在的世界。我真的很喜欢我的生活。而大多数人听到的关于它的一切都是失业、全民基本收入,就像你陷在反乌托邦的 AI 媒体循环里,你看到的全是这些。你实际上相信未来会是什么样子?什么让你兴奋?

And my last question is this. In our life, especially outside of San Francisco, when I'm in LA or even New York, I hang out with my friends and their feeling genuinely is like, I hate AI because I didn't ask for this dystopian shift to take part in in the world. Like, I like the world as it is. I really like my life. And everything that most people hear about it is like job loss, UBI, like you're in your dystopian AI media loop and that's all you see. What do you actually believe the future is going to look like? Like what excites you?

Alex

是的,我乐观的看法是,当我们能识别出人类创造力时,我们低估了自己对它的欣赏。比如,《沙丘 3》预告片出来时,有人通过 Slack 消息发给我。所以我在 Slack 里看了预告片,那个情境下它可能是粉丝制作的预告片,我也不会知道。我不知道这预告片是制片厂出的。我点了按钮,心想,好吧,这大概是个假的,你知道,粉丝用 AI 做的预告片。然后到了配乐,那种声音我从来没听过类似的东西。它像一种疯狂的新黑暗氛围,和前两部《沙丘》电影非常不同。我就想,天哪,这不可能再是假预告片了。这非常有创意。这就像我从未听过这样的东西。那是个不可思议的时刻,因为我想,天哪,我能看出这部分像是人类做的。也许他们用了——我怀疑他们用了 AI,但未来也许他们会用 AI 来创作其中的组件,但整体架构——如果我们能看到,天哪,这是我从未听过、见过或感受过的东西,那么追逐新奇将是一场疯狂的竞赛,对好莱坞这样的创意产业来说会是件大事。

Yeah, my optimistic take is that we underestimate our appreciation for human creativity when we can identify it. For example, when the Dune 3 trailer came out, it was sent to me in a Slack message. So I watched the trailer from Slack and it was in a context where it could have been a fan-made trailer and I wouldn't have known. I didn't know the trailer was from the studio. And I clicked the button and I was like, okay, this is probably a fake, you know, fan-made AI trailer. And then it got to the soundtrack, like this sound that I've just never heard anything like this sound before. It was like a crazy new dark vibe that was very different from the first two Dune movies. And I was like, oh my god, there's no way this is a fake trailer anymore. Like, this is very creative. This is like I've never heard anything like this before. That was an incredible moment because I was like, oh my god, I can tell that this part is like some human. Maybe they used—I doubt they used AI, but in the future maybe they'll use AI to create components of it, but the overall architecture—if we can see, oh my god, this is I've never heard or seen or felt this way before, that novelty chasing is going to be a bananas race and a really big thing and big deal for creative industries like Hollywood.

Host

这真的很有趣。我其实从没听过这样的观点,但我觉得真的很酷。

That's really interesting. I've actually never heard like that take, but I think that's really cool.

Alex

AI 很难想出可验证的东西——并且真正擅长创造力。我不知道怎么——我觉得很多研究者相信创造力可以通过强化学习达到,就像更好地执行任务一样。也许我们能达到。我不知道。但是的,模型会折叠到最明显、最让人安心的回答,这和最富创造力的艺术作品来自怪异、独特且不显眼的东西这一理念有些冲突。所以是的。

It's very hard for AI to come up with something verifiably—and like get really good at creativity. I don't know how—like I think a lot of researchers believe that creativity is something you can get to with reinforcement learning just like executing tasks better. Maybe we get there. I don't know. But yeah, models fold into the most obvious, like assuring response that is sort of in competition with the idea that the most creative pieces of art came from something weird and unique and like nonobvious. So yeah.

Host

是啊,我们拭目以待。目前人类还是那些怪胎。总之,Alex,非常感谢你来上节目,感谢你抽出时间。

Yeah, we'll see. Humans are still the weird freaks for now. So anyway, Alex, thank you so much for coming on the show and thank you for taking the time.

Alex

谢谢邀请我。很高兴能回来。

Thanks for inviting me. It was great to be back.

Host

如果你喜欢这期节目,请花一分钟订阅。这真的能帮我发展节目,并继续每周推出新内容。

If you enjoyed this episode, please take a minute to subscribe. It really helps me grow the show and continue to put out new episodes every week.

互动版:逐字朗读 + 针对本期提问 →