LangChain CEO Harrison Chase 谈他押错的基础设施赌注

LangChain CEO Harrison Chase on the Infrastructure Bet He Got Wrong

哈里森·蔡斯 Harrison Chase · Browserbase · 2026-09-15 · 约 52 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

LangChain 联合创始人兼 CEO Harrison Chase 回顾框架的起源、构建它的社区,以及他希望更早下注的基础设施。

LangChain co-founder and CEO Harrison Chase reflects on the framework's origins, the community that built it, and the infrastructure bet he wishes he had made sooner.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 26)

全文 · Full transcript(中英对照)

介绍与嘉宾欢迎 Intro and guest welcome

Host

大家好,我是 Paul Klein,Browserbase 的创始人,我们正在录制 Navigators,这是我们与 AI 领袖对话的系列节目,探讨 AI 的未来走向。今天请到的是 Harrison Chase,LangChain 的联合创始人兼 CEO。LangChain 是我最初用 AI 构建时使用的公司之一。在智能体还被称为智能体之前,它对于构建智能体的方式至关重要,但自那以后发生了很多变化。所以我很高兴能和 Harrison 聊聊这些。Harrison,感谢你加入我们。

Hey everybody, I'm Paul Klein, founder of Browserbase, and we're here on Navigators, our series where we meet with AI leaders navigating the future of what's happening in AI. Today I have on the show Harrison Chase, co-founder and CEO of LangChain. Now LangChain is one of the companies I first used when I was building with AI. It was instrumental to the way that you could build agents before they were called agents, but a lot has changed since then. So I'm excited to talk to Harrison about all this stuff. Harrison, thank you for joining us.

Harrison

谢谢邀请,Paul。很高兴来到这里。

Thanks for having me, Paul. Excited to be here.

LangChain的起源故事 LangChain's origin story

Host

请多讲讲 LangChain,你刚开始时它是什么,现在又是什么。

Tell me more about LangChain, what it was when you started and what it is now.

Harrison

好的,LangChain 最初是一个副业项目。当时我在上一份工作,知道自己要离开,但不知道要做什么。那是 2022 年秋天,我参加了很多聚会,和人聊天。在聊天过程中,我发现那正是生成式 AI 的时期,特别是图像方面。Stable Diffusion 刚刚出来,大家都在做那个,但也有一些古怪的人在玩 LLM。我看到他们做的一些事情,那是 RAG 的早期阶段。我记得有一些非常基础的数学或计算机使用的东西,所以我看到了一些常见模式,就把它们放进了 LangChain,那是一个开源的 Python 包。我发了推文。我还记得 Karpathy 点赞了我的推文,那是个巨大的胜利。然后我就继续做,没有抱什么宏大的幻想,只是出于乐趣,然后继续构建并更多地发推文。一个月后 ChatGPT 出来了,很明显每个人都会想要构建一些 AI 体验。而我们关注的那些 AI 体验很多被称为智能体。我清楚地记得那年感恩节假期,我们试图给 LangChain 中的一个类命名,我们叫它 agent executor,agent 部分是对的,executor 部分是错的。那部分很糟糕。但早期我们就在思考这些像 LLM 在循环中运行的东西。所以,是的,把它作为开源发布,然后几个月后围绕它创办了一家公司。所以我最终离开了我的工作,然后

Yeah, so LangChain started as a side project. So I was at my previous job, knew I was going to leave, didn't know what I was going to do. And so this was fall of 2022, was going to a bunch of meetups and talking to folks. And as I was talking to folks, saw that this was actually the time around generative AI, specifically the image stuff. So Stable Diffusion had just come out and so everyone was doing stuff with that but there were a few wacky people doing things with LLMs. And so I saw some of the things they were doing and it was like early days of like RAG. I think there was a few very basic like math or computer use things and so saw some common patterns, put that in LangChain which was an open source Python package. Tweeted it out. I still remember Karpathy liked my tweet and that was like a huge victory. Yeah. And then just kind of kept on, didn't have any grand illusions with it or anything, just was doing it for fun and then kept on building on it and tweeting about it more and a month later ChatGPT comes out and it becomes pretty clear that everyone is going to want to build some AI experience. And a lot of the AI experiences that we were focused on were called agents. And I actually distinctly remember over Thanksgiving break that year we were trying to figure out how to name a class in LangChain and we called it agent executor and the agent part was right, the executor part was wrong. That was a terrible part. But we were thinking of these things like LLMs running in a loop for a while early on. And so yeah, launched that as an open source and then started a company around it a few months after that. So ended up leaving my job and

社区与早期时光 Community and early days

Host

我们实际上是在你刚创办那家公司时认识的。我在一次黑客松上碰到你,我走向你是因为我在 LangChain 里发现了一个 bug。我记得你有一个 PR。嘿,你能合并这个吗?所以从一开始 LangChain 周围的社区支持就非常病毒式传播,现在这变得有点常见了,但当时围绕你构建的框架有前所未有的社区热爱,这是一个巨大的成就。

We actually met when you had just started that company. I caught you at a hackathon and I walked up to you because I found a bug in LangChain. I think you had a PR. Hey, can you merge this thing? So the community support around LangChain from the very beginning it was super viral and it's becoming a little more commonplace but that was like unprecedented amounts of community love around a framework that you built which is a huge accomplishment.

Harrison

是的,我觉得我明显做的一件事是,我并不是有意围绕它创办公司。我只是用它来探索这个领域,所以我也用它来联系老同事和其他人,并和其他人一起构建。比如 Sam Whitmore,我想我们都认识,她创办了一家初创公司,现在在 Cursor,她是早期 LangChain 的主要贡献者之一,我和她在 Kencho 一起工作过,那是一次有趣的重聚。所以,是的,无论是已经认识的人还是整个社区,绝对是早期的重要组成部分。我现在看,我认为开源实际上因为一些事情而发生了很大变化,好坏参半。第一,我认为这些项目的人气会出现巨大的峰值。AutoGPT 是一个更好的例子,它在 LangChain 几个月后出现,是有史以来最受欢迎的东西,然后你看到 OpenClaw 也有类似的情况。所以我确实认为它们,我怀疑是不是因为编码变得更容易了,所以有更多人在做。我认为你看到这些项目出现巨大的峰值和转折。但如果你现在看很多这些项目,我实际上不认为有同样的社区感,因为你有所有这些 AI 生成的拉取请求或问题,它们可能好也可能不好,很难解析,我知道包括我们在内的项目都在试图弄清楚如何处理这种新涌入的 AI 生成的 PR 或问题。所以我回顾早期,我确实认为事情有点不同。我确实认为今天的项目少了一点那种社区感。

Yeah I think one of the things that I distinctly did was again like I wasn't intending to start a company around it. I was just using this to explore the space and so I was also using it as a way to just like connect with old co-workers and other people and build with other people. So Sam Whitmore who I think we both know who started a startup and now is at Cursor she was one of the main contributors to early kind of like LangChain and I had worked with her at Kencho and it was a fun kind of like reconnection there and so yeah like I think whether it was people that I already knew or the whole community absolutely that was a big part of early things. I look now and I think open source has actually changed a bunch for better or worse because of a few things. One I think yeah there are these just like massive spikes in popularity of these projects. So AutoGPT is an even better example that came out a few months after LangChain and was like the most popular thing ever and then you see a similar thing with like OpenClaw. And so I do think they're they're I wonder if it's because like coding's become easier and so you have more people doing it. I think you see these just like massive spikes in inflections in these projects. But also if you look at a lot of these projects now I actually don't think the same sense of community is there because you have all these AI generated kind of pull requests or issues that may or may not be good and it's really hard to parse through and I know projects including ours are trying to figure out how to deal with this influx of new AI generated again PRs or issues and so I do you know I look back on the early days and I do think things are a little bit different. I do think in projects today there's a little bit less of that sense of community.

LangChain作为公司的演进 Evolution of LangChain as a company

Host

是的。所以你拿走了那个初始框架,为 AI 构建这些框架和技术非常困难,因为 AI 变化太快了。所以我猜你是出于热爱有机地开始做这个东西,而模型在你下面不断变化。你是如何把它变成一家公司的,LangChain 这个框架随时间如何变化?你怎么看 LangChain 这家公司也随时间成长?

Yeah. And so you took that initial framework and it's so hard to build this framework these technologies for AI because it's AI was changing so much. So I imagine you started with this thing organically out of the lot to it and the models kept changing underneath you. How did you turn this into a company and how is LangChain the framework changed over time? What you think about LangChain the company also growing over time?

Harrison

是的,这是个好问题,它们肯定都进化了很多。所以也许按时间线来看,你知道,我们在 2022 年 10 月发布了 LangChain。我们在 2023 年 1 月或 2 月创办了公司。从那时起,我们就在做我们的第一个商业解决方案,那就是 LangSmith,用于智能体的可观测性和评估。

Yeah, it's a good question and they've definitely both evolved a lot. So maybe like going kind of like in a timeline view, you know, we launched LangChain October 2022. We started the company in like January or February of 2023. And from that point we were working on our first kind of like commercial solution which is LangSmith, observability and eval for agents.

构建可靠的智能体 Building Reliable Agents

Harrison

我们意识到、并且从那以后基本上指导了我们所有产品的一点是:构建智能体的难点在于让它们可靠地做你想让它们做的事。这涉及上下文工程、框架工程、提示词。有时它们能做到,有时不能。它们需要什么上下文?需要什么工具?而正确的做法一直在变。但我们从框架到评估能力所构建的一切,确实帮助工程师们让这些智能体更可靠。

And the thing that we realized that actually has guided basically all of our products since then is that the hard part of building agents is getting them to do what you want reliably. It's the context engineering, the harness engineering, the prompting. Sometimes they do it, sometimes they don't. What context do they need to have? What tools do they need to have? And the right way to do that has changed over time. But everything that we've built from the frameworks to evalability has really helped with letting engineers try to make these agents more reliable.

Harrison

所以我们开始构建评估可观测性。它不是开源的。这算是我们的商业平台。我们在 2023 年夏天正式发布。

So we started building eval observability. So it's not open source. This was kind of like our commercial platform. We launched that GA in summer of 2023.

Host

然后发生的一件事是,正如你指出的,事情变化得非常快。

And then one of the things that happened is as you noted like things were just changing really rapidly.

Harrison

早期的时候,构建智能体的许多基础组件就像我们脚下的流沙,东西一直在变。所以在 2024 年初,我们发布了第二个框架叫 LangGraph。我们看到的趋势是,人们想要对自己构建的智能体有更多控制。LangChain 对人们来说太高层了,所以他们想要更底层。LangGraph 就强调底层控制,没有隐藏的提示词,没有隐藏的认知架构,你对一切都有完全控制。所以我们在 2024 年做了 LangGraph,继续做可观测性和评估。然后到了 2025 年左右,我觉得在如何构建智能体方面,事情实际上稳定了一些。从技术角度看,智能体的概念一直就是让 LLM 在循环中运行并调用工具,这是 AutoGPT 的核心思想,也是 LangChain 中一些智能体的核心思想。问题在于当时模型还不够好,但后来它们变得足够好了。这真正开始发生在 2025 年初。

Like early days it felt like a lot of the primitives for building agents were kind of like quicksand underneath our feet like stuff was just changing and so in 2024 early 2024 we launched a second framework called langraph and so this I think the trend that we saw was that again to this to this theme of people wanting more control over the agents that they that they built people LinkedIn was too high level for people and so they wanted to go lower level and so langraph really emphasized like low-level control like there's no hidden prompts there's no hidden like cognitive architectures you have full control over everything and so and so we did langraph 2024 kept on doing observability and eval um and then in in kind of like 2025 I think things have kind of like actually stabilized a little bit in terms of how we build agents specifically I think this idea of an agent from the technical point of view has always really been like just run an LLM in a loop and let it call tools that was the core idea of of autog um that was the core idea of some of the agents in lang chain the issue is the models just weren't good enough um at that point in time, but then they got good enough. And this really started to happen in in early 2025.

Harrison

Claude Code 就是一个例子,我认为这是它第一次真正开始并行工作。我们也开始发现基本上所有会进入框架的技巧。比如让它访问文件系统,用它来写计划,卸载大型工具调用等等。我认为编码智能体带来的许多基础组件实际上可以推广到其他更广泛的智能体类型。所以在 2025 年,地面不再像流沙,我们能够开始层层叠加地构建。我认为那时我们发布了——那时我们对开源进行了大规模重构。我们把 LangChain 重构为基于 LangGraph。然后我们构建了 Deep Agents,这是一个在 LangChain 之上更固执己见的智能体框架。我们继续构建 LangSmith 的可观测性和评估。然后最近六个月左右,随着这些框架变得更稳定,许多框架的基础设施也变得更稳定。所以我们开始在 LangSmith 这个商业平台中构建部署、沙箱、网关和无代码智能体构建器,真正把框架和运行所需的基础设施结合起来。这就是随时间推移的进展。基本上最初 100% 开源,然后加上可观测性和评估。然后最近六个月左右,更多是运行时部署方面。

And so Claude Code is an example of when I think this first started really working in parallel. We also started to discover basically all the tricks that would go into harnesses. So giving it access to a file system and using that to like write plans and like offload like large tool calls and things like that. Like I think a lot of the primitives that come with coding agents can actually be generalized to other more broad types of agents as well. And so I think in 2025 the the the the kind of like ground stopped feeling like quicksand and and we were able to start building more on top of each other. And I think that that's when we released uh that's when we did a big rearchitecture of the open source. So we rearchitected lang chain to be on top of langraph. We then built deep agents which was an even more opinionated agent harness on top of lang chain. We kept on building lang with observability and eval. And then the last like six months or so uh again as these harnesses have become more stable the infrastructure for a lot of these harnesses has also become more stable. And so we've started to build into Lang Smith our commercial platform things like deployments and sandboxes and a gateway and a no code agent builder which really kind of like um bring together the the harness and the infrastructure needed to to run it. And so that's kind of been the progression over time. Basically initially 100% open source then add on observability eval. And then very recently in the past six months or so more of this like runtime deployment aspect.

产品顺序的反思 Reflections on Product Order

Host

是的,这一切看起来都很自然,这也是我们在 Browserbase 采取的方法:从浏览器的基础设施开始,然后是一个控制它的框架,然后上面加更多层。我经常问自己这个问题。我很好奇,如果你现在回顾开始做 LangChain 以及你所构建的产品和层次的进展,你会怎么想?你会做任何不同的事情吗,还是你认为顺序是对的?

Yeah, it all seems so organic and it's the same approach that we've taken at browser base starting with you know infrastructure for the browsers and like a framework to control it and then like more layers on top. I often ask myself this question. I'm curious what you would think if you look back now at starting laying chain and the progression of products and layers that you've built. Would you do anything different or you think it was the right order?

Harrison

嗯,我认为我们肯定错过了一些东西。我觉得我们的网关建得太晚了。我记得一年或一年半前作为团队讨论过这个。我们当时真的不知道在 LiteLLM 或 OpenRouter 之上能提供什么不同的东西。虽然我认为 LiteLLM 当时确实是主要的东西。嗯,我认为有一件事我一直低估了,那就是对推理的需求有多大。它是巨大的。另外,编码智能体会变得多么庞大和通用,那个市场也是巨大的。但是的,我认为我们应该更早构建网关。我认为那是核心构建部分。这可能是我们做得有点太晚、应该更早做的一件事。

Um I think there's some things we missed for sure. I think we built a gateway way too late. Um, I remember talking about this as a team a year year year and a half ago. Um, and and we just didn't really know like what we could provide that was different on top of like a light LLM or an open router. Although I think light LLM was really like the main thing at the time. Um, and uh, I think there I I think clearly there's one of one of the things that I've just consistently underestimated is just the demand for just inference that is out there. Like it's massive. Uh the other thing by the way is just how big and general purpose coding agents would be that market also massive. Um but yeah I I I think we should have built kind of like a gateway earlier on. I think that is a core building part. Um I I think that's probably one thing that we did a little too late and we should have done earlier.

Harrison

我其实也很好奇你们的历程,因为我知道你们也在为浏览器智能体构建框架和构建工具之间平衡。比如我觉得你们刚做了 Stagehand v4,它非常专注于工具,让人们——如果我说错了请纠正——更专注于工具,让人们构建自己的框架,而之前你们更多有自己的框架。我很好奇你们是怎么考虑的。

I'm actually curious for your guys' kind of like journey through things as well because I know cuz you guys have also kind of like uh uh balanced the line of like building a harness for browser agents versus building like the tools, right? Like I think you just did stage hand v4 which was really focused on the tools and letting people and correct me if this is off but like more focused on the tools and letting people build their own harness and I think previously you had your own harness more. I'm curious how you've thought about that.

Host

是的,你知道这很有趣。我认为这个领域、AI 基础设施领域的每家公司,我们可能都在自己的类别中打磨好了基础组件,然后你尝试在基础组件的边缘试探,找出什么对什么错,或者我们在哪里构建了太高的抽象或太低。所以 Browserbase 的核心浏览器产品 Browserbase Cloud 基本上打磨好了,那是必要的,有很多挑战,所以那是核心。然后我们在上面有一个框架 Stagehand,它基本上打开了一个 SDK 来控制那个浏览器。我们早期构建 Stagehand 的原因其实是,我们可以像 LangChain 捆绑提示词那样捆绑提示词,让用户的操作——比如点击这个按钮、点击页面上的立即购买按钮——更容易翻译成实际命令。但随着模型变得越来越好,你对实际上下文的提示词抽象——在我们的情况下就是网页——已经变成了一个不必要的层。所以我们仍然为那些真正喜欢它、在代码库中添加了半确定性的遗留用户保留它。

Yeah, you know it's it's funny. I think like every company in this space in the infra space for AI we probably nail the primitive like in our categories and then like you try and poke at the edges of the primitive and find out what's right or what's wrong or where we've built too high of an abstraction or too low. So browserbased like the core browser offering browser and cloud pretty much nailed that like that is necessary there's a lot of challenges with it so that's the core and then we had a framework on top of it stage hand which opened up basically a a SDK to control that browser and the early reason we built stage hand was actually like oh we could just bundle in the prompts similar to how lang chain had bundled prompts to make it easier to translate a action from a user like click this button click the buy now button on the page to the actual command but as models have gotten better and better. Your prompt abstractions over the actual context which in this case for us is like the web page has become like an unnecessary layer. So we still have that for legacy users who actually really like it who've like added semi-determinism into their code bases.

子智能体与工具层优化 Sub-agents vs. tool-layer optimization

Harrison

我们那些越来越进阶的客户会说:不,直接给我这个工具最好的一版,能把鼠标移到这个点上——真正去优化工具层,而不是提供分层的子智能体。我觉得子智能体作为一个想法,在实际实现方式上已经证明是有点更难的。子智能体需要拥有一个任务的完整理念,而不是执行任务的某一部分。那太小了——就像我们会有个子智能体去点那个按钮来节省上下文。所以,构建优秀智能体、或者说能调用系统的智能体的工具和技术,随着时间已经演进了太多。

Our more and more advanced customers are like, no, just give me the best possible version of the tool to move the mouse to this point — really optimize the tool layer, as opposed to offering layered sub-agents. I think sub-agents as an idea has proved to be a little more challenging in how they're actually implemented. Sub-agents need to own complete ideologies of a task, not execute a single part of a task. It's very small — like we'd have a sub-agent to click that button to save context. So it's just that the tools and techniques of how to build great agents, or agents that can call systems, have evolved so much over time.

Harrison

而且我觉得作为创始人——我相信你也有同感——很难摆脱你过去那些想法。你必须不断重置自己对什么才合理的先验。这一直是我们在构建过程中始终面临的挑战。

And I think as a founder — I'm sure you feel the same way — it's hard to pull away from the ideas you had in the past. You have to constantly reset your priors on what makes sense. And that's been something that's always challenged me as we've been building.

Harrison

而当我展望模型能力时,我现在下的赌注一直是这样的:我相信模型能力会向人类能力收敛。因此,转向文件系统——智能体将拥有自己的文件系统——其实非常合理,因为人就是这么做的。如果模型正在向人的能力收敛,那你大概会想去掌握人的工具,因为它们就是在那上面训练的。所以市场似乎兜了一整圈——从“我们为 AI 智能体构建一个完全定制的东西”,到“哦,干脆给它们我们已有的工具,因为它们其实就是在所有这些上面学习的”。

And as I look towards model capabilities, the bets that I've been making right now have always been like: I believe model capabilities will converge on human capabilities. Therefore, this move to the file system, where agents are going to have their own file system, just actually makes a lot of sense, because that's what people will do. And if models are converging on people's capabilities, you probably want to master the people tools, because that's what they're trained on. So it seems like the market has gone full circle — from being 'let's build a completely bespoke thing for AI agents' to 'oh, let's just give them the tools that we already have, because they're actually learning on all of those things.'

拟人化LLM的风险 The risk of anthropomorphizing LLMs

Host

这其实是我很纠结的一件事,因为一方面我同意你说的,而且我觉得很容易、也很诱人,把这些 LLM 拟人化,说“哦,如果我是这个 LLM,这就是我会想要的”。但它们其实是不同的东西。所以我不喜欢——到底在什么层面上这样做才是对的?我真的不知道。而且就像你说的,也许你在很多由人类标注的任务上做强化学习,它就开始向那收敛。但这些 LLM 擅长的是人类不擅长的事情,我们应该保留这些,也应该去发挥这些。

This is something I actually struggle with a lot, because on one hand I agree with what you're saying, and I think it's very tempting and very easy to kind of anthropomorphize these LLMs and be like, 'Oh, if I was this LLM, this is what I would want.' But they are just different things. And so I don't like — what is the right level to do that at? I genuinely don't know. And to your point, maybe if you RL on a lot of these tasks that are labeled by humans, it starts to converge to that. But these LLMs are great at things that humans are not great at, and we should keep those and we should lean into those as well.

编码框架与浏览框架 Coding harnesses vs. browsing harnesses

Host

所以这其实是我——接着你说的文件系统那点。我们之前稍微聊过这个,你觉得编码智能体的 harness 和一个好的网页浏览 harness 有多相似?你从那里学到东西了吗?有没有一些你觉得就是不同的、需要以某种形式加进去的东西?

So that's something I'm actually — so, going off of the file system thing that you said. We were talking about this a little bit earlier, but how similar do you think coding agent harnesses are to a good harness for browsing the web? Did you learn things from there? Are there things that you think are just different and need to be added in in some form?

Harrison

对。代码模式,或者说写代码来执行工具的智能体,被称为 API。我觉得这是对“智能体会像人一样使用工具”这个说法的一个很好的反驳,对吧?因为,它们只是去执行代码。它们更像是编译器或操作系统。而这个叙事已经被证明非常非常好。我觉得智能体写代码带来了很大的灵活性。但确实存在那个自主性滑块:我们信不信它去写所有这些代码?它能可重复地写代码吗?它应该复用那些代码吗?我发现代码会创造记忆,或者说创造技能,可以在未来的智能体中被引用,而对于高度可重复的浏览任务来说,这真的很有益。

Right. Code mode, or agents that write code to execute tools, are called APIs. I think that's a great counterargument to the 'agents will use tools like humans' idea, right? Because, well, they're just going to execute code. They're more like compilers or operating systems. And that narrative has proved to be really, really good. I think that agents writing code gives a lot of flexibility. But there is that autonomy slider of: do we trust it to write all this code? Can it write code repeatably? Should it reuse that code? What I have found is that code creates memory, or creates skills that can be referenced in future agents, and for browsing tasks, which are highly repeatable, that's really beneficial.

Harrison

所以写代码来控制浏览器的智能体,其实正是我们在最近这个框架阶段 MV4 里采用的方法,因为我们想:智能体调用我们的框架,几乎更像是写一些小脚本来做这些任务,然后随着时间你会编译出越来越大的脚本。这就是为什么我们最近的框架完全没有集成任何模型。它全是工具层。所以代码模式对于做重复性任务的智能体来说似乎非常相关。

So agents that write code to control a browser is actually the approach we took in our most recent framework stage, MV4, because we were like, well, agents will call our framework almost more like writing little scripts to do these tasks, and you'll compile bigger and bigger scripts over time. And that's why our recent framework doesn't have any models integrated at all. It's all just the tool layer. So code mode seems very relevant for agents doing repetitive tasks.

Harrison

我觉得很多智能体式的工作都是重复性的。比如我跟企业交流时,他们其实有一份文档写着“这是我想做的所有步骤”,其中有些模糊的部分。如今我很少见到即时的智能体式用例,也就是智能体在做完全全新的事情。我觉得试图把所有智能体都归到某一种特定类型里是非常困难的。我一直思考的是领域特定的智能体——比如浏览网页的智能体、做软件工程的智能体、做深度研究的智能体。我很好奇,你是怎么思考智能体的不同领域,然后反推它们实际应该拥有什么架构的?我认为做不同任务的不同智能体需要有不同的架构。

I think a lot of agentic work is repetitive. Like when I talk with enterprises, they actually have a document that says, 'Here's all the steps I wanted to do,' with some fuzzy parts. It's very rare that I see just-in-time agentic use cases these days, where the agent is doing something completely brand new. I think trying to bucket all the agents down into one certain type of agent is very challenging. And I've always thought about domain-specific agents — like agents that browse the web, agents that do software engineering, agents that do deep research. I'm curious, how have you thought about the different domains of agents and then working backwards to the architecture they actually should have? I think different agents that do different tasks need to have different architectures.

从定制认知架构到智能体框架 From custom cognitive architectures to agent harnesses

Host

是的,我的意思是,我觉得这就是我们想着力实现的核心用例——基本上就是如何帮助人们为他们的产品、他们的领域、或他们的用例构建智能体。我觉得答案随着时间变了很多,而且还在变。

Yeah, I mean, I think this is the core use case that we think about enabling — basically how to help people build agents for their product, or for their domain, or for their use case. I think the answers have changed a bunch over time and are still changing.

Host

所以早期我们构建 LangGraph 时,想法是你为你做的每件事都创建一个完全定制的认知架构。你会像你说的那样把它铺开——比如有这些流程,有这些不同的步骤,或者你想做的不同的事情序列——于是你把它铺成一张图,然后有些模糊的部分,你就把它放进一次 LLM 调用里,而那也许可以像是这些图里的一个条件边。

So early on, when we built LangGraph, the idea was that you'd create a completely custom cognitive architecture for each thing that you do. You'd lay out kind of how you were saying — like there were these processes, and there are these different steps, or different sequences of things you want to do — and so you'd lay that out in a graph, and then there was some fuzzy part, and so you'd put that in an LLM call, and that could maybe be like a conditional edge in some of these graphs.

Host

我觉得对于那些你有更多这类众所周知的流程、并且在某种程度上在意确定性的场景,这真的很好。我们仍然看到人们为此使用 LangGraph,也正因为这个原因,它在金融服务和金融行业非常流行。

And I think that's really good for places where you have more of these well-known processes and you care about determinism to some extent. And we still see people using LangGraph for that, and it's really popular in financial services and financial industries for that reason.

Host

与此同时,我觉得人们用那些东西做的很多事情,老实说只是更偏智能体式的。我们早期实际做过的一个例子是,我们写了一个深度研究的示例,它是一个 LangGraph 应用,它大概是:首先你要做规划,然后你会得到这三个要点,然后你会扇出,然后你会扇回来,然后你可能会审查你的工作。

At the same time, I think a lot of what people were doing with those things were honestly things that were just more agentic. So an example of this that we actually did early on was we wrote a deep research example, and it was a LangGraph application, and it had like: first you're going to plan, and then you're going to get these three bullet points, and then you're going to fan out, and then you're going to fan back in, and then you might review your work.

Host

而深度研究其实就是一个智能体式的任务。我真的这么认为——你可能会深入太多不同的方向,很难把它放进一个图式的工作流里。所以基本上我们看到这样一个趋势,就是更多地转向这些用于做深度研究之类事情的智能体 harness,然后你定制它的方式开始变成技能、工具、提示词。我们有一个概念叫中间件,但那类似于 Claude Code 里的钩子之类的东西。

And deep research is really just an agentic task. Like I really think it is — there's so many different things you might go down that it's kind of tough to put that into a graph-like workflow. And so basically we've seen this trend towards more of just these agent harnesses for doing things like deep research, and then the way that you customize it starts to become skills, tools, prompts. We have a concept that we call middleware, but that's similar to like hooks in Claude Code or something like that.

为智能体框架添加确定性 Adding Determinism to Agent Harnesses

Harrison

你也可以加进去——我觉得很多这种定制化的做法,本质上就是加入一些确定性的东西。比如加一个类似目标模式的检查。嘿,当你觉得你做完了,再跑一次检查,看看你有没有满足目标。那是一个显式的确定性检查,你去做,然后如果它失败了,就回到起点。所以这就是在里面加入一些确定性。很好,你可以用中间件来做,也可以用钩子之类的东西来做。这就是你开始定制这些 harness 的方式。

And you can also add in — I think a lot of customizing it this way is about adding in deterministic things. So like adding in a check like goal mode. Hey, when you think you're done, run another check to check whether you satisfied the goal. That's an explicit deterministic check that you do, and then if it fails, it goes back to the start. So this is adding some determinism in there. And so great, you can do that with middleware, you can do that with hooks and things like that. And so that's how you start to customize some of these harnesses.

Harrison

我确实在想,谈到不同领域之间的差异——比如编程,很多编程任务你需要一个完整的沙箱,因为你想安装任意依赖,你想启动一个服务器之类的。在考虑文档处理时,我们看到过一些反复。你需要多少?你只需要访问一个文件吗?所以我们在 deep agents 里有一个虚拟文件系统的概念,你可以直接和一个文件系统交互。你可以用编程智能体一样的方式读写文件,但你不必启动一个沙箱,而沙箱是有点重的基础设施。所以就容易多了。

And I do wonder, talking about the difference between different domains — like coding, for a lot of coding tasks you need a full-on sandbox because you want to install arbitrary dependencies and you want to spin up a server or something like that. We've seen some back and forth when thinking about document processing. How much of that do you need? Do you need just access to a file? So we have a concept in deep agents of a virtual file system where you can just interact with a file system. You can read and write files using the same way that a coding agent would, but you don't have to spin up a sandbox, which is a heavyish piece of infrastructure. So it's a lot easier.

Harrison

我们有一个客户,早期他们做文档分析那部分时就只是那样做。我前几周和他们聊了聊,他们说,你知道吗,我们其实把所有东西都搬进了一个完整的沙箱,因为人们想用 PowerPoint 的 CLI 之类的——是啊,你知道,那个你没法真正模拟。你确实需要一个完整的沙箱来做那个。Vercel 有他们那个 just bash 的东西,也是这一类里的另一个,它试图模拟出 bash 文件系统。Pyantic 有 Monty,有点像是一个假的 Python 解释器。所以我觉得——一方面,是的,也许对某些其他任务,你不需要完整的编程那一套。也许你需要,你就是想要所有这些不同的 CLI 都装好,那就是做法。

And so we have one customer that early on they were just doing that for their document analysis bit. And I caught up with them the other week and they're like, you know, we actually moved everything into a full-on sandbox because people want to use like PowerPoint CLIs and like — yeah, you know, you can't really mock that. You kind of need a full-on sandbox for that. Vercel has their just bash thing, which is another thing in this vein where it tries to mock out the bash file system. Pyantic has Monty, which is kind of like a fake Python interpreter. And yeah, so I think — so on one hand, yeah, maybe for some of these other tasks, you don't need the full-blown coding stuff. Maybe you do and you just do want all these different CLIs installed and that's the way to do it.

Harrison

所以我现在想得很多的一件事是,是的,比如好的技能、MCP、指令、中间件——那些绝对是你定制它的方式。但除此之外还有没有更深层的差异?我不知道正确答案。我觉得我们还在琢磨这个。

And so one of the things I'm thinking a lot about right now is like, yeah, what like great skills, MCPs, instructions, middleware — absolutely those are how you customize it. Are there like deeper differences than that? And I don't know the right answer. I think we're still thinking through that.

Host

是啊,太对了。而且就算在沙箱里装一个浏览器,那也不太够,对吧?你得做很多额外的步骤才能让浏览器在网络上正常工作。有这些 web bot off 协议让智能体能去某些地方,或者用可观测性来记录浏览器会话。这些东西你有多少是搭进沙箱本身的?有多少是依赖第三方提供商的?我确实觉得存在某种清晰的栈,而且某种程度上浏览器就是一个沙箱,对吧?因为它在内存里跑自己的 JavaScript,web assembly 也可以在那里运行。

Yeah, it's so true. And like even installing a browser on a sandbox, that's not very sufficient, right? There's so many extra steps you'll have to do to get the browser working on the web. There's these web bot off protocols to allow agents to go places, or observability to record the browser session. Like how much of that are you scaffolding into the sandbox itself? How much are you relying on third party providers for? I do think that there is some clear stack, and in a way a browser is a sandbox, right? Because it's running its own JavaScript in the memory, web assembly can be working there too.

Host

那好,关于这个有个问题,因为我觉得在这个沙箱用法里冒出来的一件事是,智能体和沙箱的正确架构是什么?智能体是跑在沙箱里,还是——是的,把大脑和手分开,跑在外面再连接?所以我很好奇,比如浏览器这类东西——你提到某种程度上浏览器就像一个沙箱。你看到人们大多是把这个浏览器之类的当成一个独立的东西来跑,让智能体在另一边作为独立的东西,只是把它当作工具来对话,还是你其实必须把它们放在一起?

Well so okay, question about that, because I think one thing that's popped up in this sandbox usage is like, how — what's the right architecture for the agent and the sandbox? Does the agent run in the sandbox, or does the agent — yeah, to separate the brains and hands and run outside and connect? And so I'm curious for like browser stuff — you mentioned in some ways a browser is like a sandbox. Do you see people mostly running the browser, whatever, as a separate thing and having the agent over here as a separate thing and just talking to it as a tool, or do you actually have to colocate them?

Harrison

是的,我越来越多地看到人们把智能体的大脑和工具调用分开,尤其是因为你想能查看中间的数据,确保它没有把任何坏数据发给工具。比如说——我们先撇开沙箱和浏览器,保持在高层面。比如说你的智能体有一个发 API 请求之类的工具。它只是有一个网络工具。你大概会想要在智能体发出这个网络请求之间有一层,去查找 PII 或社会安全号码,并在它发到网上之前把它剥掉。所以把那些分开,那个控制层很有意义。另外,为了避免提示注入从工具回到智能体,把它们分开也很重要。你多少想要在智能体逻辑周围有一个安全区。

Yeah, I've seen more and more that people are separating the agent brain from the tool calls, especially because you want to be able to look at the data in between and make sure it's not sending any bad data towards the tool. Like let's say — let's remove sandboxes and browser for a second just to keep it high level. Like let's say your agent has a tool to make an API request or something. It just has a network tool. Like you probably want some layer between the agent saying issue this network request that looks for like PII or social security numbers and strips it out before it goes out to the web. So separating those and that control layer makes a lot of sense. Also to avoid like prompt injection coming back from the tool into the agent, having those separate is important. You kind of want to have a safety zone around your agent logic.

Harrison

但然后,是的,对浏览器,我觉得这适用于所有工具——所有工具大概都应该跑在智能体循环之外,或者说智能体 harness 之外,在某种安全的、可以被销毁的环境里,比如一次性使用。你大概不想让工具在同一个沙箱或同一个虚拟化层里被复用,因为那样的话,如果上一个被调用的工具造成了某种污染怎么办,或者你怎么管理这些工具之间的文件系统?所有这些事情对我们来说都超级难搞清楚。到最后我总是看向我们的客户在做什么。这很棘手,因为你听到客户说的东西非常不一样。

But then yeah, for the browser I think that maps to all tools — like all tools probably should run outside the agent loop, or sorry, the agent harness, in some sort of safe environment that could be destructible, like one-time use. You probably don't want to have tools reused across the same sandbox or same virtualization layer because then what if some pollution happens from the last tool that's called, or how do you manage the file systems across those tools? Like all these things are just super challenging for us to figure out. And in the end I'm always looking towards what are our customers doing. It's tricky because you hear customers speaking very different things.

Harrison

而且我觉得你的客户大概比我们的还要多样,因为比如我的一些客户,我形容为那种很猛的工程团队,真正在推动事情往前走,然后另一些客户才刚开始他们的 AI 之旅,想同时为两者构建是非常非常难的,你几乎必须把抽象层从那些很猛的团队往上抬,去适配更新的团队或更企业化的团队。你怎么看这个?比如你在可能有的这两类不同客户之间,是怎么移动抽象层的?

And I think your customers are probably even more diverse than ours, because like some of my customers I describe as the cracked engineering teams that really push things forward, and then other customers are really getting started on their AI journey and trying to build for both is very very challenging, and you almost have to move up the abstraction layer from the cracked teams to the more like newer teams or enterprisy teams. How do you think about that? Like how are you shifting the abstraction layer across the two different types of customers you might have?

Host

我觉得我们看那些 AI 构建工程师,是为了看智能体的未来走向,百分之百。那里发生的事并不是每件都会达到企业级,甚至不一定是正确的做法,但人们在实验。我觉得看到这个真的很酷。而且我觉得如果你看到人们用三种略微不同的方式在实验同一件事,你就会想,好吧,那里有点意思。所以当人们做定制化的东西时,如果这是我们第一次看到,他们说,嘿,我想做这个定制化的东西,在 deep agents 里没有简单的方式来做那个。

I think we look at those AI build engineers for where the future of agents is going, 100%. And not everything going on there is going to be enterprise grade or even the right way to do things, but people are experimenting. And I think that's really cool to see. And I think if you see people experimenting with the same thing in three slightly different ways, you're like, okay, there's something kind of interesting there. And so when people do bespoke things, if it's the first time we see it and they're like, hey, I'm trying to do this bespoke thing, there's nothing easy in deep agents to do that.

开源策略与抽象层 Open Source Strategy and Abstraction Layers

Harrison

我认为这就是我们现在对开源故事感到满意的地方,有不同的抽象层级,比如 LangGraph 在最底层,所以如果你想做任何事情,你都可以在 LangGraph 中完成,就像你可以用 LangGraph 做任意的事情,比如……

I think that's where now we feel good about our open source story where there are different levels of abstraction and there's like LangGraph at the lowest level and so if you want to do anything you can do it in LangGraph like you can use LangGraph for just like arbitrary things like...

Host

>> 你可以分叉它,总是有逃逸舱口进入更低层级

>> you fork it to like there's always an escape hatch into the lower levels of the

Harrison

>> 确切地说,就像 LangGraph、LangChain、深度智能体,根据你想做什么,你可以沿着这些不同的层级往下走。但一旦我们看到人们做同样的事情三次,我们就会说,好吧,你知道,举个例子,就是工具调用的卸载。工具返回非常大的响应。我认为现在这通常是标准做法,但一年前还不是,基本上就是把它转储到文件系统,然后说:“嘿,这是前一千个字符,但如果你想读剩下的,你可以直接浏览这个文件。”所以现在这成了深度智能体的一部分。我们尝试这样做,因为绝对有很多团队想要采用……所以我实际上认为这里有两件事。我认为,是的,也许他们不像其他人那样 AI 药丸,但我也认为,我们从编码智能体中学到的所有这些经验实际上比你在本地运行的编码智能体更广泛。所以当很多人想要把这些最佳实践用于他们的第一方智能体时,这些智能体可能不是编码智能体,也可能不在沙箱中运行之类的。但是的,我们想要把这些常见的东西拿出来,让每个人,无论他们是否在构建编码智能体,无论他们是否 AI 构建,都能采用它们并用它们来构建自己的智能体。所以这就是我们思考平衡的方式。

>> exact so it's like LangGraph, LangChain, deep agents and depending on what you want to do you can kind of go down those different layers and so but then like once we see people doing like the same thing like yeah three times we're like okay you know um like like an example of this would be um offloading of tool calls. So tools return really large responses. I think it's generally standard practice now, but wasn't like a year ago to basically dump that to a file system and say, "Hey, here's the first like thousand characters of that, but if you want to read the rest, you can like go just navigate this file." And so now that's part of deep agents. And so we try to do that because absolutely there are a bunch of teams that want to take um I so I actually think there's two things here. I I I think like um it's yes like maybe they're not as AI pill as others, but I also think it's like there's all these lessons that we're learning from coding agents that are actually just more broad than just like you know a coding agent that you run locally. And so um when a lot of these people want to take these uh uh best practices and use them for their first party agent which may not be a coding agent and may not be um may not even run in a sandbox or things like that. But yeah, we we want to take these common things and bring them so that everyone regardless of whether they're building coding agents or not, regardless of whether they're AI build or not, can kind of like uh take them and use them to build agents of their own. So that's kind of how we think about balancing it.

评估前沿创新 Evaluating Bleeding Edge Innovations

Host

>> 是的。这很合理。我开始更多地思考模型将走向何方,因为我认为如果我们回顾两年前,事情来了又去。比如记得有一个实验室推出了这个 DAG 构建器工作流,然后我想一切都结束了,然后我再也没听说过它,对吧?计算机使用推出时,有人说这东西很烂,永远不会成功,操作员也不太好,这实际上阻碍了我的行业一点,就像计算机使用很糟糕,不,它已经改变了,对吧?我们经历了这些高潮和低谷,以及技术变革的假象,这使得预测什么会持续、什么不会变得非常困难。你如何看待这些前沿创新?比如你怎么判断某件事是长期的创新还是只是暂时的?

>> Yeah. And it makes sense. I mean, I'm starting to think a lot more about where the models are going to go because I think like if we look back two years, things like things have come and gone quite a bit. Like uh remember one of the labs launched like this DAG builder this workflow thing and then like I was like everything's over and then I've never heard about it again right like computer use launched there was like this thing sucks it's never going to work like operator was not that good that actually held back my industry a little bit like it's computer use is bad like no it's changed right like we've had these highs and lows and these head fakes around technology changes that can make it really hard to predict what's going to stick and what not what's not. How do you think about that for these like bleeding edge innovations? Like how do you tell if something is like a long-term, you know, innovation or just a temporary thing?

Harrison

我的意思是,我想是的,我肯定会尝试很多东西,所以每当我想真正理解一个产品或技术时,我就会直接用它构建或使用它,或者尝试以那种方式构建一些东西,然后如果我认为它合理且令人兴奋,无论是因为它让事情更容易,还是它实现了以前没有的功能,这就是我获得最大信念的方式。我认为这是一种奇怪的混合,既要对现状现实,又要对未来乐观。比如,是的,我们在 2022 年 11 月左右在 LangChain 中有一个叫做 agent executor 的东西,但我们并没有全力投入,因为我们意识到当时它根本不起作用。所以,你知道,反事实地,如果我们当时说,嘿,我们将在 2022 年开始构建一个智能体框架,

I mean, I think so I um I definitely hack around with a lot of stuff and so like whenever I'm trying to really understand like a product or technology, like I'll just go build with it or use it or try to build something in that way or something like that and then uh if I think it's like reasonable and and exciting for whatever reason, either it makes it easier or it enables something that wasn't before, that's how I get like the most conviction of things. I think it is like a weird mix of being realistic about where things are but also optimistic about the future. Like yeah, we had something in LangChain called agent executor in like November 2022 but like we did not go all in on that because we realized that you know it just didn't work at the time. And so like you know counterfactual what if we had what if we were just like hey we're just we're going to start building an agent harness in 2022.

Host

>> 我不知道会发生什么,你知道,人们会使用我们吗?可能不会,因为你真的不能用它构建任何东西。

>> I like I don't I don't know what would have happened like you know would like would people have used us? Probably not because you couldn't really build anything with that.

Harrison

所以我认为这是一种奇怪的混合,既要考虑现在能构建什么、现在能实际应用什么,然后是的,你能为未来架构什么。我认为我的总体哲学是尽量不要做出任何疯狂狂野的……或者有更好的说法吗?我认为我的总体哲学是建立一种团队、流程或精神,即快速迭代,而不是超级教条地坚持“这是必须这样做的方式”。显然,如果你能做出这样的赌注,嘿,这非常正确,我要坚持三年,那真是令人难以置信,可能有很多价值可以捕获。

So I think it's a weird mix of like what can you build now and what can you be practical with now and then yeah what can you architect for the future. My I think my general philosophy has been try to not like make any crazy wild kind like or what's a better way to say this. I think my general philosophy has been like um build build up kind of like a a a team or a process or an ethos of just iterating super rapidly rather than like being super dogmatic and sticking to like yeah this is kind of like the way it has to be done. And obviously if you can make that bet on like, hey, this is like uh uh outlandishly true um and and and correct bet and I'm going to stick to it for like 3 years that is like incredibly impressive and like you'll there's probably a lot of value to be captured there.

浏览器自动化与创始人信念 Browser Automation and Founder Conviction

Host

>> 那是基于浏览器的,对吧?我的意思是,我同意。我认为我在 2023 年对浏览器做出了这种教条式的赌注,比如可以使用浏览器,但当时没有太多证据,而且在大约一年里,大型语言模型完成的浏览器自动化非常粗糙,比如我们过去在东西上画边界框,然后发送截图,我们会为 BLM 标注截图,变化太大了,但我确实认为,拥有那种核心创始人信念,在某些原则和想法上,这就是让我们的公司走得很远的原因,同时你必须有一定程度的怀疑,比如什么才是适合世界的正确方式,就像世界观和现实

>> That's browser based, right? I mean like I I I do agree though. I think I made this dogmatic back on the browser in 2023 like a can use a browser and like there wasn't much evidence of that and like for the first I'd say almost a year browser automation done by a large language model was super hacky like we used to draw bounding boxes on stuff and like send it to screenshot and we'd annotate the screenshots for the BLM and so much has changed uh but I do think like those like having that core founder conviction in certain principles and ideas that's what makes our companies, you know, go really far and then at the same time you have to have some degree of skepticism about like what is the right way that fits onto the world like it's a worldview and then there's reality

Harrison

>> 你也必须对此进行迭代,就像你之前说的,最初人们想要更高层级的东西,但现在他们想要低层级的,你知道,就像 LangChain 的核心思想一直是人们想要为他们的特定领域和特定任务构建自己的智能体,但具体形式已经彻底改变,所以我认为,是的,拥有一种使命和专注。

>> you have to like iterate on that as well like I mean when you were talking earlier about like yeah you know initially people wanted the higher level thing but now they want the lowle like that you know like the core idea of lang chain has always been that people want to build their own agents for their particular domains and their particular tasks what that has looked like has absolutely changed and so I think like there's like yeah having kind of like a mission and and a focus.

二阶效应与智能体身份 Second-Order Effects and Agent Identity

Harrison

然后还有更多那种小的微观的事情,实际上非常重要,因为如果你搞错了,你可能会说:“是啊,我在 2022 年构建了一个智能体框架,但没人用。为什么?”方向上有趣,人们想构建智能体,但微观上,当时那不是正确的做法。所以我认为存在这种平衡,这很难。

And then there's more kind of small micro things which actually matter a ton because if you get those wrong, you can be like, "Yeah, I built an agent harness in 2022 and no one used it. Why not?" It's like, well, directionally interesting, people want to build agents, but the micro was like, that's not the right way to do it at the time. So I think there is this balance which is tough.

Host

是的,我的路线图越来越侧重于二阶效应。对于浏览器相关的东西,我现在花很多时间在智能体身份上,以及我们浏览器如何确保这些是可信的。未来,一家基于浏览器的公司可能更像是一家身份公司,而不是纯粹的基础设施公司,因为这是一个非常重要的问题,需要与许多对手方一起解决。

Yeah, and my roadmap has become more and more dominant about second-order effects. For the browser stuff, a lot of time I'm spending right now is on agent identity and how we, the browsers, make sure these are trusted. In the future, a browser-based company might be more of an identity company than a pure infrastructure company because that's a really important problem to solve with a lot of counterparties.

Host

对你来说,我认为你在可观测性方面做的一些事情,比如智能体的二阶效应,你必须观察它们,然后你可以利用这些数据来实际改进它们,让它们自我改进并反思。你从 LangChain 或基于 LangChain 构建的智能体中看到了哪些二阶效应,你如何围绕这些调整你的产品?我想数据和轨迹方面的一些东西对你的客户非常有用。

For you, I think some of the stuff you did around observability, like the second-order effects around agents, you have to observe them and then you can use that data to actually make them better and self-improve and reflect back. What are some of the second-order effects that you've seen from LangChain or agents built on LangChain, and how have you oriented your product around that? I imagine some stuff on the data and the trajectory side has been very useful for your customers.

Harrison

是的,我的意思是,我认为开源和可观测性与评估的核心一直都是:我如何构建一个智能体来完成特定任务?你必须以某种方式定制它,所以你必须要有旋钮和框架来定制它,并且你必须要有可观测性和评估的仪器来测量你定制的效果。所以这就是将它们结合在一起的原因。定制的方式绝对随着时间的推移而改变,我实际上对它们的发展方向非常兴奋,特别是很多后训练和强化学习相关的事情。所以到目前为止,人们主要通过上下文工程或框架工程或提示和规模来定制东西。

Yeah, I mean, I think the core of both the open source and the observability and eval has always been basically: how do I build an agent to do a particular task? You have to customize it in some way, so you have to have the knobs and the harness to customize it, and you have to have the instrumentation with observability and eval to measure how well you're customizing it. So that's what brings them all together. The ways of customizing it have absolutely changed over time, and I'm actually very excited about where they're going, particularly with a lot of post-training and RL things. So primarily to date, people have largely customized things by doing context engineering or harness engineering or prompts and scales.

Host

查看评估,来回运行东西之类的。

Looking at the evals, running the thing back and forth type of stuff.

Harrison

没错。今天他们还没有真正改变模型的权重。是的,他们会 100% 更换模型,但不会改变底层模型的权重。随着后训练技术变得越来越普遍,开源模型在基础能力上赶上,以及人们开发这些使用更多的智能体,实际上拥有这些数据,我们绝对发现自己越来越被拉向这个方向。所以有很多不同的技术,主要要么涉及使用轨迹,LangSmith 帮助整理这些,要么使用评估,同样 LangSmith 帮助整理这些。所以这是一个我们绝对发现自己越来越被拉向的方向。

Exactly. And today they've stayed away from actually changing the weights of the model. Yes, they'll swap models 100%, but they won't change the weights of the underlying model. And that is a direction we absolutely find ourselves getting pulled in more and more as post-training techniques become more common, as open source models catch up in just the base strength, and as people develop these agents that are getting used more where they actually have this data. So there's a bunch of different techniques that largely either involve using the traces, which LangSmith helps curate, or using the eval, again LangSmith helps curate those. So that's been a direction that we've absolutely found ourselves getting pulled in more and more.

Host

是的,这也是我没有预料到的。我们整个播客都在谈论我们没有预料到或已经改变的事情,这很好。因为我认为你和我几年前经历过,比如我没有预料到开源模型会赶上这么多,后训练会变得如此真实。我的意思是,微调在早期是一个东西,它有点帮助,但和现在完全不一样。我真的震惊,不仅是基础设施如何适应需求的变化,还有模型能力以及我们客户中不同类型模型的涌入。我想知道,你已经很好地帮助客户训练模型或后训练模型。你认为他们现在如何使用?你看到更多客户使用自己的模型吗?你看到更多客户说前沿吗?如何平衡?

Yeah, that's something I also didn't predict. We're spending this whole podcast talking about stuff that we didn't predict or has changed, which is good. Because I think you and I lived through a couple years ago now, like I didn't expect how much open-source models would catch up and how much post-training would be real. I mean, fine-tuning was a thing in the very early days and it kind of helped a little bit, but it's nothing like it does now. I've been really shocked not just about how the infrastructure has changed to the needs, but also the model capabilities and the influx of different types of models for our customers. I'm wondering, you've been set up really well to help customers train models or post-train models. How do you think that they're using that right now? Do you see a lot more of your customers using their own models? Do you see a lot more of them saying frontier? How does it balance?

Harrison

对于高流量的用户,我认为有很多兴趣。我认为实际上可能数量惊人地多的人在使用后训练模型,通常不是用来驱动核心智能体循环,而是用于更简单的事情,比如分类或提取之类的。但总的来说,实际这样做的客户仍然是少数。所以,我认为在 LangChain,我们一直试图做的一件事是,将一些高级主题,无论是使用 LLM 构建、构建智能体框架,让它们对人们来说变得平易近人和容易。所以,我们花了很多时间与更 AI 化的公司一起,试图弄清楚他们后训练的常见做法是什么,以及在哪里,比如什么类型的问题实际上是智能体循环还是分类。我认为现在我们看到的一般是更简单的事情,尽管他们正在尝试更核心的智能体循环,然后我们希望让这对每个人来说都尽可能容易。

A lot of interest, I think, for people who have high volume. I think actually probably a shockingly large number of those are using a post-trained model generally not to drive the core agentic loop but maybe for simpler things like classification or extraction or something like that. But that's still, in the grand scheme of things, a small number of customers that are actually doing it. And so, one of the things that I think at LangChain we've always tried to do is take some of the advanced topics, whether it's building with LLMs, building agent harnesses, and make them approachable and easy for folks. And so, we're spending a lot of time with the more AI-pilled companies trying to figure out what exactly are common things they're doing to post-train and where, like what types of problems is it actually the agent loop or is it the classification. And I think right now we're seeing it's generally simpler things, although they're experimenting with the more core agent loops, and then we want to make that as easy as possible for everyone.

Harrison

我记得你提到过微调曾经是个东西,然后消失了,我记得大约两年前和某人聊天,当时我对微调非常怀疑,我认为事后看来那可能是对的。我认为 OpenAI 弃用了他们的微调服务之类的。所以这次又有一波,我不知道像什么……

I remember you mentioned kind of like fine-tuning was like a thing and then went away, and I remember chatting with someone two-ish years ago and I was very skeptical on fine-tuning then, and I think in hindsight that was probably right. I think OpenAI deprecated their fine-tuning service or something like that. And so there's another wave of it this time, and I don't know like what...

Host

你认为这次有什么不同?我的意思是,由于某种原因,我现在对它稍微更看好了。我不确定你是否也有同感。你认为这个后训练时代现在有什么不同?

What do you think is different this time? I mean, I feel a little bit more bullish about it for some reason now. I'm not sure if you feel the same way. Like what do you think is different about this post-training era now?

Harrison

是的,我认为核心的开放权重模型更好了。你能够后训练的东西更好了。我认为我们对需要什么类型的数据有了更好的想法。总体上知识更多了。而且推理更加可用。我不知道 2023 年我们有多少 GPU 在流通。我认为有一段时间很难获得。但现在似乎你可以以更低的成本进行实验,并实际看到一些类似的结果。我早期做的少数微调事情,你只需要大量数据才能做出真正好的东西。数据必须非常干净。我认为现在你可以做的事情的误差范围已经增加了。这引出了我想和你谈论的一个话题,那就是感觉每个人优化的目标已经改变了。

Yeah, I think the core open-weight model is just better. The stuff that you're able to post-train on is better. I think we have better ideas about what types of data we need. There's more knowledge going on overall. And inference is much more available. I don't know how many GPUs we had floating around in 2023. I think there was a moment there where it was very hard to get. But it seems like nowadays you can experiment for a lot lower costs and actually see some semblance of results. The few fine-tuning things I did early on, you just needed a lot of data to make something really good. It had to be really clean data. I think now the error bars have just increased on what you can do. And that kind of leads into a topic I'm curious to talk to you about, which is it feels like what everyone is optimized for has changed.

准确性、成本与速度权衡 Accuracy, Cost, and Speed Trade-offs

Host

当然,每个人都想要准确性,但准确性几乎成了基本门槛。我不会说它完全就是,但确实越来越接近了。现在我看到在成本和速度之间有一些非常有趣的权衡。如果非要给这三个排序,我认为人们关心的是准确性、成本和速度。但越来越多的客户,尤其是那些构建通用智能体的,非常关心推理延迟,因为对于一些前沿的超大模型,推理延迟可能非常慢。你看到你的客户是如何优化这些的?

Of course, everyone wants accuracy, but accuracy is almost becoming table stakes. I'm not going to say it completely is, but it's really getting there. And now I'm seeing these really interesting trade-offs on cost and speed. If I had to rank the three, I think people care about accuracy, cost, and speed. But more and more of our customers, especially those building general agents, care a lot about latency of inference, because latency of inference can be really slow for some of these frontier, really large models. How have you seen your customers optimizing these?

Harrison

我们过去两年做了一项调查,我想,我们问的是:你最大的挑战或优先事项是什么?我忘了确切的措辞,但就是你认为构建智能体难在哪里?在过去两年里,质量或准确性基本上是速度和成本的两倍。而速度和成本差不多。所以我很期待今年再做一次,我会告诉你结果。我仍然认为对于通用智能体,我会押注成本。对于那些你真正想深入特定用例的场景,我确实认为质量仍然很重要。我仍然认为那是一个障碍。我还认为,随着模型变得更好,也许现有的用例现在已经饱和了,但我们是一个有创造力的物种,我们会想出新的东西来扔给这些模型,我认为那个标准会不断上升。所以我确实认为准确性和质量可能仍然是很多人的最大瓶颈。但我很期待做这个调查。如果结果比以前更接近,我不会感到惊讶。我确实听说也有不同类型的智能体,对吧?对于编码智能体,我认为现在 100% 的人都在意识到成本是主要问题之一。我知道在内部我们对此非常关注,我们认为对于很多任务,质量基本上已经足够好了——你选择你的框架,选择你的模型——对于 50% 的任务可能已经足够好了。我不知道正确的百分比是多少,但对于一部分任务来说是这样。所以现在我们大量思考成本,我可以告诉你,在内部我们花更多时间思考这些模型的成本和这些框架的成本,而不是试图衡量,哦,它是 95% 还是 96%?

So we ran a survey for the past two years, I think, where we asked, what are your biggest challenges or priorities? I forget the exact phrasing, but what do you think is hard about building agents? And for the last two years, quality or accuracy was basically double that of speed and cost. And those were about the same. So I'm excited to run it again this year and I'll let you know what the results are. I still think for general purpose agents, I would bet cost. For things where you're really trying to drill into a particular use case, I do think quality still matters a lot. I still think that's a blocker. I also think that as models have gotten better, maybe the existing use cases are now saturated, but we are a creative species and we come up with new things to throw these models at, and I think that bar keeps on rising. So I actually do think accuracy and quality is still probably the biggest bottleneck for a lot of folks. But I'm excited to run the survey. I wouldn't be shocked if it was closer than before. I definitely hear there's also different types of agents, right? For coding agents, I think 100% people right now are waking up to cost as one of the main things now. I know internally we are looking at this a lot and we think for a lot of tasks quality is basically good enough based on—you choose your harness, choose your model—it's probably good enough for 50% of tasks. I don't know what the right percentage is, but for a chunk of tasks. And so now we're thinking a lot about costs, and I can tell you internally we spend a lot more time thinking about the cost of these models and the cost of these harnesses as opposed to trying to measure, oh, is it 95% or 96%?

Host

是的,特定领域的质量要好得多。编码确实已经解决了。在成本和延迟方面,我一直在做一个思想实验:好吧,假设我们解决了质量,假设成本变得更容易接受,或者我们可以适当地路由来解决这两个问题。推理速度,假设它变得非常非常快。我认为这给基础设施本身,比如智能体基础设施,带来了很多有趣的压力。我认为现在很多基础设施公司可以容忍几百毫秒的差异,比如在沙盒大战中。我见过关于启动时间 P99 的争论,就像我们快了一毫秒,好吧,谁在乎呢?但随着模型变得非常快、便宜和准确,这些事情将变得非常重要。我很好奇,作为另一位基础设施创始人——我们这样的人不多——你如何看待模型的需求,即为智能体提供服务的基础设施在未来会怎样?

Yeah, domain-specific quality is very much better. Coding is really, really solved. And on the cost and latency side, there's a thought experiment I've been running about: okay, let's say we solve quality, let's say cost becomes much more palatable or we can route appropriately to solve those two. Inference speed, let's say that gets really, really fast. I think that puts a lot of interesting pressure on the infrastructure itself, like the agentic infrastructure. I think right now a lot of infra companies can get away with a few hundred milliseconds difference, like in the sandbox wars. I've seen battles over the P99 of startup time and it's like, we're one millisecond faster, like, okay, who cares? But as models get really, really fast and cheap and accurate, those things are really going to matter a lot. And I'm curious, as another fellow infra founder—there's not many of us—how do you think about the needs of models, the infrastructure serving the agents in the future?

工具与智能体的基础设施 Infrastructure for Tools and Agents

Harrison

是的。嗯,我实际上认为更大的负担可能是为智能体访问的工具提供服务的基础设施。所以就像沙盒一样,如果你把它作为工具运行,我认为你需要能够查询它,无论这意味着启动它还是从热池中提供服务,你需要能够快速查询它,快速得到结果。但每个工具都一样,访问你希望这些智能体使用的任何东西。一个例子是,我们实际上并不是因为这个原因才做的,但我认为会受益匪浅。我们有 LangSmith。它存储了一堆追踪记录,这些追踪记录通常对编码智能体或任何智能体都很有用,原因有很多,包括反思其他智能体做了什么、运行实验等等。我们最近做的一件事是重新架构了 LangSmith,使其构建在我们内部构建的数据库 SmithDB 之上。我的联合创始人 Ankush 做了所有的工作。所以如果你问我更多信息问题,我会让给他来回答。但它在查询追踪、查询反馈和附加反馈等方面快得多。所以现在当我们把它交给智能体时,它们使用这些东西要快得多。我认为这种加速真的很重要。再次回到我们之前谈到的,为人类构建,为智能体构建。我们为人类做了那件事,是为了更快地填充 UI。但事实证明,对于编码智能体来说,这重要得多,因为随着它们运行时间越来越长,它们会不断敲打这些 API。而这些 API 只是一个数据库,对吧?那不是智能体特定的基础设施。这就是我所说的,我认为一般的基础设施都会变得有趣。如果你把一个数据库暴露给一个智能体,它会被大量查询敲打。所以你必须考虑如何为这种情况扩展。

Yeah. Well, I actually think probably the bigger burden is the infrastructure serving the tools that the agents access. So like sandboxes again, if you're running it as a tool, I think yeah, you need to be able to query it and whether that means spinning it up or serving it from a warm pool or whatever, you need to be able to query it fast, get a result back fast. But same with every tool, access to anything that you want these agents for. So an example of this that we actually didn't do for this reason, but I think will benefit a lot. So we have LangSmith. It stores a bunch of traces and those traces can often be useful to give to coding agents or any agents for a variety of reasons including reflecting on what other agents have done, running experiments, things like that. One of the things that we did recently is we rearchitected LangSmith to be built on top of SmithDB, which is a database that we built in house. My co-founder Ankush did all of it. So if you ask me any more info questions I'm going to yield and ask him. But it's a lot faster at querying traces and querying feedback and attaching feedback and things like that. So now when we give it to agents, it's a lot faster for them to use these things. And I think that speed up really matters. And again, going back to something we talked about earlier, actually, like building for humans, building for agents. We did that for humans. We did that so we could populate the UI faster. But it actually turns out to matter I think a lot more for coding agents because again as they get longer and longer running, they're just going to be hammering these APIs. And so these APIs which are just a database, right? That's not agent specific infra. And this is what I mean by I think just infrastructure in general is going to be interesting. Like if you're exposing a database to an agent, it's going to get hammered by a ton of queries. And so you got to think about how to basically scale for that.

Host

是的。这几乎是一种乘法效应,对于每个人来说可能有数百个智能体,而这些数百个智能体可能调用数百个工具,这些工具也可能在反思自己。这个智能体对智能体的世界将会非常有趣,看看这一切如何发展。我认为协议也允许人们调用像 MCP 这样的工具,而围绕智能体的治理是我思考很多的二阶效应之一。似乎我们至少有一种使用工具、调用任意服务的方式,但配置它们、允许它们代表这个人行事,这是我看到很多活动的地方。你如何将其视为你更大平台或战略的一部分,客户在要求什么?

Yeah. It's almost like a multiplicative effect where for every human there might be hundreds of agents and those hundreds of agents might be calling hundreds of tools that also might be reflecting on themselves. And this agent-to-agent world is going to be very interesting to see how it all plays out. I think that the protocols too allow people to call tools like MCP and the governance around agents is one of the second order effects I think a lot about. It seems like sure we may have at least a way to use tools, call arbitrary services, but provisioning them, allowing them to act on behalf of this person, that's been something that I've seeing a lot of activity around. How have you thought about that as part of your greater platform or strategy and what are customers asking for?

Harrison

嗯,你之前提到你们正在深入身份领域。

Well, you had mentioned earlier that you guys are kind of getting deep into identity.

多人模式下的智能体凭证 Agent Credentials in Multiplayer Mode

Host

我其实很想听听你们那边在做什么,因为我们现在也遇到这个问题了。人们在构建智能体,想让它接触各种工具,而 MCP 现在是一个更好的协议,大家想用。于是就有了这个问题:智能体如何与这些工具交互?它们有自己的凭证,还是代表用户行事?然后当你进入多人模式时——想象一个共享的 Slack 频道,你在和一个智能体聊天。它用谁的凭证?用我的凭证吗?还是固定的凭证?实际上有不同的项目用不同的方式处理这个问题。据我了解,Claude Tag 的处理方式基本上是:有一组凭证,当你在频道里和那个智能体交互时,无论你是谁,它都会用那组凭证。还有很多其他项目采取了不同的方法。所以我其实不知道最好的方法是什么。我们开始思考这个问题。我很好奇——你是在浏览器的语境下提到这个的。我很好奇你在那里看到了什么,以及那是怎么运作的。

I would actually love to hear what you guys are doing there, because that's coming up for us now. People are building agents and they want to give them exposure to tools, and MCP is a protocol that is better now and people want to use. And yeah, there's this question of how do the agents interact with these tools? Do they have their own credentials or do they act on behalf of users? And then when you start getting into multiplayer mode — imagine a shared Slack channel and you're chatting with an agent. Whose credentials does it use? Does it use my credentials? Is it fixed credentials? There are actually different projects that tackle this in different ways. So Claude Tag, the way they handle it as far as I understand, is basically there's a set of credentials and when you interact with that agent in a channel, no matter who you are, it will use those credentials. And a bunch of other projects take different approaches there. So I actually don't know what the best approach is. And so we're starting to think about that. I'd be curious — you mentioned it in the context of browsers. I'd be curious what you're seeing there and how that works.

登录网络中的身份 Identity on the Logged-In Web

Harrison

对,你们走的是容易的路。至少这类东西有协议和 API,对吧?在我们这边,我们面对的是并非 AI 优先的 Web。登录态的 Web 才是智能体身份需要被解决的地方,但并没有一个我们可以直接叠加的协议。有 SSO、SAML 和 OAuth 能帮上忙,但看起来智能体会试图代表人们去访问服务。也许这又回到你和我在这里的不同观点或不同信念。我可能认为我的智能体应该是我的代表,因为我大概只有有限的访问权限。你也可以为智能体创建一个独立账户,如果你想让它以自己的身份登录、拥有更少的访问权限。但我认为,我们在 Web 上会看到的大多数智能体,都会以所有者的权限范围行事。在企业场景中限制这一点的最好方式,就是让它们用 Okta 登录,对吧?和智能体对话的人在 Okta 或他们的 SSO 平台上有一定程度的访问权限,而智能体大概需要镜像这一点。所以我认为,与现有的 Web、现有的凭证系统、现有的认证后端和 SSO 协作,能让我们快速推进很多事情,因为我们所触及的那部分 Web、那部分世界,并没有 API 和 MCP。所以我正在把我们的很多原语映射到人类已有的身份原语上,这就是我们能够真正成为现有 Web 之上的兼容层、同时为互联网网站添加 AI 原生连接器的方式。

Yeah, you have the easy path. At least there are protocols and APIs for this stuff, right? In our world, we're working with the web that's not AI-first. The logged-in web is where agentic identities need to be solved, but there isn't a protocol that we can bolt on top of. There's SSO and SAML and OAuth that can help here, but it seems like agents are going to try and access services on behalf of people. And maybe once again this goes back to how you and I have different views here or different beliefs. I might believe that my agent should be a delegate of me because I probably have a limited scope of access. And you can create a separate account for an agent if you want it to log in as its own self and have even less access. But I think for the majority of agents that we're going to see out on the web, they're going to act to the extent of which the owner has permissions. And the best way to limit this in an enterprise context is to have them sign in with Okta, right? The human talking to the agent has some degree of access in Okta or their SSO platform, and the agent probably needs to mirror that. So I think that working with the existing web and the existing credentialing systems and the existing auth backends and SSO allows us to speedrun a lot of these things, because our part of the web, the part of the world we touch, doesn't have APIs and MCPs. So I'm mapping a lot of our primitives onto existing identity primitives that humans have, and that's been the way that we can really be this compatibility layer over the existing web but add an AI-native connector for internet websites.

超越Okta的企业访问 Enterprise Access Beyond Okta

Host

那你在企业里有没有发现,你想让智能体访问的一切都在 Okta 后面,或者在这种 SSO 登录后面?

And have you found that in enterprises, everything that you'd want to give agents access to is behind Okta or behind this SSO sign-up?

Harrison

哦,当然不是。有时候它藏在你见过的最破烂的网站登录后面,是 2006 年某个人做的,对吧?需要有密码共享,或者我们怎么让你以一种类似 Plaid 的体验让智能体登录一次,然后保存连接,以便将来你能代表它们登录。所有这些东西,我想你看着它们就会觉得,天哪,一定有更好的办法,而随着时间推移会有更好的办法。但看起来对智能体的需求如此之高,而在我看来模型能力也正在达到能解决大多数基于 Web 的任务的程度,但存在一种模型能力的悬置:模型可以去解决任务,却实际上无法访问解决任务所需的系统。而模型能力的这种扩散,正是我们认为身份可以发挥重要作用的地方。我们正在和 Okta 或 1Password 这样的伙伴大量合作,尝试找到安全地与智能体一次性共享凭证的方法,让它实际上不会传到 LLM,或者我们如何配置 SSO 和 OAuth,让你的智能体请求访问,而你收到一条小小的推送通知问能不能登录。随着计算机变得更强,这些模式越来越多地出现,而且在计算机使用中已经有过几个根本性的时刻,事情发生了,人们意识到,哦,AI 会使用计算机,这就引发了这些讨论。所以我很兴奋,随着它们变成越来越有成效的对话,人们把智能体使用浏览器和 Web 视为必然,而不只是一种取巧。

Oh, certainly not. Sometimes it's behind the jankiest website login you've ever seen that was built by somebody in 2006, right? There needs to be password sharing, or how do we allow you to let your agent log in one time in a Plaid-like experience, then save the connection so that you can log in on behalf of them in the future. All of these things, I think you look at them and you're like, man, there has to be a better way, and there will be a better way over time. But it seems the demand for agents is so high, and model capability in my opinion is getting there to solve most web-based tasks, but there's this overhang of model capability where the models can go solve the task but they can't actually access the systems they need to solve the task. And that diffusion of model capability is where we think identity can really play a big part. We're working a lot with partners like Okta or 1Password to try and find ways to securely share credentials with an agent one time so it doesn't actually go to the LLM, or how do we configure SSO and OAuth to allow your agent to request access and you get a little push notification saying can I log in. More and more of these patterns are popping up as computers get good, and there have been these several fundamental moments in computer use where things have happened and people realized, oh, AI will use the computer, and it's brought up these discussions. So I'm excited as they become more and more productive conversations and people view agents using a browser in the web as an inevitability and not just a hack.

浏览器使用身份验证的成熟度 Maturity of Identity for Browser Use

Host

你觉得这里的一些身份相关的东西,尤其是对浏览器使用而言,有多成熟?比如我们现在是在现编,并没有真正的标准?你觉得你们现在这一刻有产品吗?

How mature do you think some of the identity things are here for browser use in particular? Like are we making it up right now and there's no real standards? Do you feel like you guys have a product here right now at this moment?

Harrison

我们做身份这块已经大约一年了,一路上也在宣布各种合作。去年我们和 Cloudflare 做了第一次合作,谈论一个叫 Web Bot Auth 的东西,一种让智能体通过表明自己身份来获准浏览某些网站的方式。几个月后,我们宣布了 Browserbase 加 1Password,一种让我们允许智能体实际上把它们的 1Password 保险库连接到我们浏览器并安全共享凭证的方式。所以这几乎很相似——我们拥有这些原语已经很久了,但我认为人们现在正到达一个点:模型实际上会跑到这些部分,而不是在登录步骤之前或之后把所有东西硬编码,然后智能体接管。所以我仍然认为有大量工作要做,而其中很多归结为对这些原语是什么、应该是什么的社会化。类似于智能体式工程,我觉得当我们看到 2025 年冬天的 Claude Code 热潮时——我得补充一句,那对我来说是个去度假的疯狂时机——我们开始在我们行业的各个部分越来越多地看到这一点。而且看起来,如果你能以正确的方式进入讨论,成为人们谈论这个的地方的一部分,我猜是 Twitter,很多这些原语就会非常非常快地成形。但我是一个非常务实的人。试图让智能体以尽可能简单的方式完成它们想做的任务,正是我在 Browserbase 的智能体身份产品中寻求做到的。还有很长的路要走。互联网是一个非常大的地方。而且大概没有一刀切的解决方案。你只能一块一块地攻克 Web 的各个部分,找出连接它的最佳方式。

We've been working on identity for about a year and have been announcing partnerships along the way. We did our first partnership with Cloudflare last year to talk about this thing called Web Bot Auth, a way to say agents are allowed to browse to certain websites by identifying themselves. A few months later, we announced Browserbase plus 1Password, a way for us to allow agents to actually connect their 1Password vault to our browsers and securely share credentials. So it's almost similar — we've had these primitives for a long time, but I think people are now getting to the point where the models are actually running to these parts as opposed to hard coding everything up until or after the login step and then the agent takes over. So I still think there's a ton of work to do, and a lot of this comes down to the socialization of what these primitives are and should be. Similar to agentic engineering, I felt like when we saw the Claude Code rush of winter 2025 — a crazy time for me to take a vacation, I have to add — we're starting to see that more and more in different parts of our industry. And it does seem like if you can get in the discourse in the right way and be a part of, you know, Twitter, I guess, where people are talking about this, a lot of these primitives take shape really, really quickly. But I'm a very pragmatic guy. Trying to enable agents to do the task they want to do in the most simple way possible is really what I seek to do with our agentic identity products at Browserbase. And there's a long road to go. The internet is a very large place. And there's probably not a one-size-fits-all solution. You're just going to take down parts of the web piece by piece to figure out the best way to connect to it.

网页浏览智能体的突破时刻 The breakout moment for web browsing agents

Host

我觉得你刚才其实已经触及这一点了,就是某些项目的热度会突然出现尖峰,比如 Claude Code 或者 open claw。你觉得什么样的项目会让网页浏览这件事也出现这种爆发?

I think you were kind of getting to this with like there's just these spikes that happen in kind of like popularity of projects where there's Claude Code or open claw. What do you think like the project would be that would make this happen for kind of like web browsing?

Harrison

我觉得这已经发生了。我感觉 Codex 的 computer use 真的火起来了,我总看到有人在推特上讨论。大家在求别人去用 Codex 的 computer use。有意思的是,所有人都给家里买了 Mac mini,说“对,我们就在上面跑个小智能体”。那本质上就是一个设备,让你把流量通过家里的 IP 地址隧道出去,来应对你用网页时遇到的一些封锁。所以当我看到大家买 Mac mini 时,我就想,天哪,这是一个需要被解决的基础设施问题,因为显然模型已经足够好了,而那个方案是昂贵的基础设施解法。所以我觉得这是一系列不同的发布、不同的上线、不同的事情在发生,而且感觉智能体的消费者用例常常会带动 B2B 用例。我经常看 AI 的消费者用例,就像我看 AI 构建团队一样,通过大家广泛使用的消费者用例来判断企业侧在往哪走、B2B 用例在哪,因为很多时候基础设施是可以类似地迁移过去的,只是可能需要更多的治理和政策。

I think it's already happened. I feel like Codex computer use has really taken off and I see people tweeting about all the time. People are begging other people to use Codex computer use. And it's funny, everyone bought Mac minis for their house and they're like, "Yeah, we'll just run a little agent on it." And like that essentially is a device that allows you to tunnel traffic through your home IP address to handle some of the blockers you face when using the web. So when I saw people buying Mac minis, I was like, man, this is an infrastructure problem that needs to be solved because clearly the models have gotten good enough and that is the like expensive infrastructure solution. So I think it's been a series of like different announcements, different launches, different things that are happening and it feels like the consumer use case of agents often drives the B2B use cases. It's like I often look just like I look at like the AI build teams. I look at the consumer use cases of AI that a lot of people are using to look where enterprise is moving and where the B2B use cases are because often times the infrastructure will carry over similarly but with probably a more governance and policy that's required.

Host

我知道我对 Codex 的 computer use 了解不多。它是在你本地电脑上运行,还是像 browser base 那样把它发到某个远程环境?

I know I don't know that much about Codex computer use. Is that running like on your local computer or do they do something like browser base and send it off to some remote kind?

Harrison

大部分情况下是在你本地电脑上运行。

It runs on your local computer for the most part.

用智能体运行LangChain Running LangChain with agents

Host

我想多听听你自己作为 AI 构建公司创始人的经历,因为 LangChain 现在挺大了,对吧?

I want to hear a little bit more about your journey as an AI build founder yourself cuz LangChain's pretty big now, right?

Harrison

对,我们大概有 360 人。

Yeah, we're like 360 people.

Host

你怎么做到的?我是说,那是什么感觉?你有智能体在管理 LangChain 吗?我想听听你们用 AI 来运营 LangChain 本身的情况。

How do you do that? I mean, like what is it like? Do you have agents managing LangChain? Like I want to hear about your AI use to run LangChain itself.

Harrison

我们在一些地方用了智能体。其实我觉得,当我跟候选人聊,他们能聊的全是用智能体来运行和完成工作,我反而觉得这有点负面信号。我认为在如何应用智能体、在哪里用智能体上,还是需要很多品味。但我们确实在一些地方用了智能体。就使用场景而言,我觉得我们用得最多的是在市场推广组织。我们有一个相当不错的市场推广智能体,负责对入站线索做资格筛选,也帮助外呼。那是内部构建的。我们实际上有两到三名应用 AI 工程师全职在做这个,还有一位市场负责人基本上就专门负责这件事。现在它已经扩展到了——一开始主要是在销售,现在我们有一个,我忘了名字,好像叫 content 什么的,但我们用它来生成所有的 GIF 和图片之类的东西,我们叫它 content studio。所以那是一个内部智能体,全部只在 Slack 里暴露。所以 Slack 是——我想我们技术上有个网页应用,但 Slack 是主要界面。然后我们还大量使用编码智能体。所以我们有自己版本的、基本上是基于云的智能体。我们叫它 open suite。它其实是开源的,如果大家想去看看的话。但我们同样在 Slack 里暴露它,然后用得很多。所以这其实是我现在提交大部分 PR 的方式,因为我没有太多时间,但我看到有问题,就在 Slack 里 @ 一下。然后我基本上还有一个执行助理类的智能体,基本上在过我的邮件。所以做分诊,有需要我注意的事情就 ping 我,尝试起草回复,找日历时间,诸如此类。所以这是我个人做的,也是我们在组织里做的。

We have agents in some places. I actually think when I'm talking to candidates and all they can talk about is using agents to run and do their job, I actually think that's a bit of a negative signal. I think there still needs to be a bunch of taste in terms of how you apply agents and where you use agents. But we do use agents in some places. In terms of where we use it, I actually think we use it most in kind of like the go-to-market org. So we have a pretty good go-to-market agent that does qualifying of inbound leads and helps with outbound as well. And that's built in-house. And we actually have two to three applied AI engineers that are building that out full-time, along with a marketing leader who works basically just on this. And that's now extended into — it started off mostly in sales and now we have kind of like, I forget the name of it, I think we call it content something, but we use that to generate all of our GIFs and images and things like content studio is what we call it. And so that's like an in-house agent all exposed just in Slack. So Slack is — I think we technically have a web app but Slack is the primary interface. And then we use coding agents a bunch. So we have our own version of basically a cloud-based agent. We call it open suite. It's actually open source if people want to go check it out. But we again expose that in Slack and then use that pretty heavily. So that's actually how I put up most of my PRs these days because I don't have a ton of time but I see something wrong and then Slack at something. And then I have basically an executive assistant kind of agent that's basically going over my email. So triaging it, pinging me when there's things that need my attention, trying to draft responses, finding calendar times, things like that. So that's what I've done personally and that's what we've done in the org.

Host

对,我自己也有同感。我有自己的智能体用来处理邮件等等。但维护自己的智能体可能很费功夫。你花多少时间来构建这套帮助你作为 LangChain 创始人执行工作的机器,还是你一次就搞定了,所有工具一直都运行得挺好?所以,你是说当有人来 LangChain 面试,说他们什么都用智能体,这是个危险信号?

Yeah, I found that myself. I have my own agents I use for email, etc. But maintaining your own agent can be a lot of work. How much time do you spend building the machine that helps you execute as founder of LangChain or did you just one-shot it and all of your tools are working pretty well all the time? So, you're telling me that when someone interviews for LangChain and they say they use agents for everything, it's a red flag?

Harrison

有一点。

A little bit.

Host

为什么?再多说说。

Why is that? Tell me a little bit more.

Harrison

我觉得有些事是智能体擅长的,有些是它们不擅长的。而且我认为你产出的工作,不管是你手动做的还是让智能体做的,都是你的责任。所以,如果做出来很烂,那你就得负责。我不在乎你是不是用智能体做出了烂东西。你以不正确的方式使用了工具。我为什么要关心那个?所以我觉得,关注工具多于关注结果不是好事。我觉得当人们先谈他们想要的结果,然后再进入具体哪些地方用智能体可能合理,我喜欢那种回答。我觉得当我问一个问题时,如果对方一上来就说“对,用智能体做那个”,或者硬要把智能体塞进某件事里,我觉得那——我们最近不得不应对大量的 AI 垃圾内容。我不知道你有没有看到

I think there's things that agents are good at and there's things that they're bad at. And I think the work that you produce, whether you've done it manually or you've had agents do it, is your responsibility. And so, if it comes out looking like crap, then yeah, you're on the hook. And I don't care if you've used the agent to create something that's bad. Like you've used a tool in an incorrect manner. Like why should I care about that? And so I think the focus on tools more than outcomes is not a good thing. I think when people start talking about the outcome they want and then get into specific places where it could make sense, I think I like that answer. I think when I ask a question it's like yeah use an agent for that right off the bat or like trying to force an agent into something. I think that's — we've had to deal with a ton of AI slop recently. I don't know if you've seen the

Host

你们怎么对抗垃圾内容?我见过。我很好奇你们内部在做什么。

How are you fighting slop? I have. I'm curious what you're doing internally.

Harrison

这是个好——我不认为我们有完美的答案。我觉得我们做过的事情是,我们让工程 VP 基本上谈了 AI 垃圾内容,谈了——Clay 写了一篇很好的博客,讲他们关于 AI 写作的四点。

It's a good — I don't think we have a perfect answer for this. I think things that we've done are we've had our VP of engineering basically talk about AI slop, talk about — Clay wrote a really good blog on their four points for kind of like AI writing.

Host

写作。对。

Writing. Yeah.

Harrison

对。所以基本上就是谈那个。我想我们加了一个 Slack 表情,叫 AI slop。我看到东西时越来越多地用 Pangram。再说一次,如果你写得好,我不会在意——我会把它发出来或转出去之类的。如果你写得烂,而且我能看出是 AI 写的,我就会去 Pangram,把它贴进去,它会说这是 100% AI,然后我就会说,你为什么要这么做?这很糟糕。所以我觉得你多少得点名批评。我觉得——对。我觉得我们看到的,而且我不认为我们在这方面是独特的。我觉得很多很多人都在应对,到处都是 AI 垃圾内容。对。而且你知道,作为一个帮助别人构建智能体的人,我不知道这么说是不是奇怪或虚伪,但它让我烦。看到所有这些 AI 的东西让我很恼火。我觉得它真的没有经过深思熟虑。

Yeah. And so basically talk about that. I think we added a Slack emoji that's AI slop. I've been using Pangram more when I see things. Again, if you write something good, I'm not going to care about — I'm going to post it or send it or something like that. If you write something bad and I can tell that it's AI, I'm going to go to Pangram, post it in, it's going to say it's 100% AI and then be like, why are you doing this? This is bad. And so I think you kind of have to call people out. Like I think — Yeah. I think we're seeing and I don't think we're unique in this. I think a ton of people are dealing with like there's AI slop all over. Yeah. And you know, I don't know if this is weird or hypocritical to say as someone who helps people build agents, but it annoys me. It pisses me off a lot to see all this AI stuff. I think it's really not well thought through.

在变化的世界中构建 Building in a Changing World

Harrison

是啊,挺有意思的。我们是在一个不断变化的世界里构建基础设施,不同的模型、要解决的问题都在变。但与此同时,我们也在一个不断变化的世界里经营公司,公司怎么运作、发生了什么、你用不用 slop、它怎么变,这些都在变。所以感觉就像身处这两个世界,不断面对变化的环境,还得真正去挑战它们。挺有意思的。你知道,我们正活在一个很棒的时代。

Yeah, it's funny. We're building the infrastructure on an ever-changing world of different models and problems to solve. But we're also simultaneously building companies on an ever-changing world of how companies operate and what's going on and if you use slop or not or how it changes. So it feels like these two worlds where we're constantly seeing changing environments and having to really challenge them. It's pretty fun. You know, we get to live in this great time.

Host

你觉得 LangChain 未来几年会怎么发展?从这里出发,它要往哪里去?你接下来想攻克哪些尚未解决的问题,好让我下次做这档播客时能跟进一下?

How do you imagine LangChain develops over the next few years? Like where is it going from here? What are the unsolved problems that you want to take on next that I can check in on next time I do this podcast?

智能体的受控运行时 The Governed Runtime for Agents

Harrison

我觉得我们现在考虑的一件大事是,对于很多这类智能体来说,合适的运行时是什么。所以,接着你之前提到的关于质量的观点——我仍然认为质量是个问题,但还有其他因素在起作用。比如,我觉得我们正在超越可观测性和评估,我们在部署、沙箱、网关方面做了很多工作,真正在探索“受治理的运行时”这个理念。当你想要运行一个不在你笔记本电脑上的智能体,想要为大量终端用户大规模运行它,还要有恰当的权限管理、恰当的护栏、恰当的成本控制,并且能扩展而不崩溃——你需要哪些东西?我觉得其中一些组件我们已经知道了,另一些还不知道。但我认为沙箱、持久化执行、模型网关、某种 off 网关——而且这可能会因你是在 MCP 里用,还是像我今天了解到的在浏览器里用而有所不同。但 off 非常重要。所以就是思考我们如何能构建或集成那里所有最好的东西,并提供一种……因为我觉得正在发生的一件事是,现在有所有这些组件开始发挥作用,所以大规模部署实际上变得越来越棘手。你需要把这些小乐高积木一块块拼起来。所以我们发布的一个东西,我个人花了很多时间在上面,就是托管式深度智能体。托管式深度智能体基本上把执行框架和基础设施绑定在一起。我对此如此兴奋的原因是,我确实认为有这么多基础设施组件,不得不一个个组装起来真的很耗时,我们不希望人们这么做。我们希望他们能轻松地在生产环境中把智能体跑起来。

I think the big thing we're thinking about now is what is the appropriate runtime for a lot of these agents. So I guess kind of to your point earlier around quality being a little bit — again I still think it's a concern but other factors coming into play. Like I think we're going beyond observability and evals and we're doing a bunch around deployments and sandboxes and gateway and really this idea of a governed runtime. When you want to run an agent that's not on your laptop and you want to run it at scale for a bunch of end users with proper permissioning and proper guardrails and proper cost controls and have it scale up and not fall over — like what are all the things you need there? And I think we know some of those components. I don't think we know other components. But I think sandboxes, durable execution, model gateway, some form of like off gateway — and that might look different for whether you're using it in an MCP or, as I learned today, in browsers. But off is super important. And so just thinking about how we can either build or integrate with all the best things there and provide kind of like — because I think one thing that's happening is there's all these pieces that are now coming into play and so deploying this at scale is actually becoming trickier and trickier. You need to assemble all these little Lego pieces. And so one of the things we launched that I'm spending a lot of time on personally is managed deep agents. So managed deep agents basically binds the harness with the infrastructure. And the reason I'm so excited about that is I do just think there are all these pieces of infrastructure and having to assemble them one by one is really time-consuming and we don't want people to do that. We want them to get up and running with agents easily in production.

智能体工程永无止境 Agent Engineering Is Never Solved

Host

所以一旦我们把所有这些都建好了,就大功告成、可以收工了,对吧?智能体工程就解决了。

So once we build all of that, it's mission accomplished and we're done, right? Agent engineering is solved.

Harrison

不,还会有更多事情。我不知道是什么,但还会有更多事情。

No, there'll be more things. I don't know what, but there'll be more things.

Host

好吧 Harrison,我相信你会继续把它们搞清楚的。今天能请你来聊这一场真的很有意思,谢谢你抽时间,兄弟。

Well Harrison, I'm confident that you'll keep figuring them out. It's been really fun having you on this conversation today and thank you for taking the time, man.

Harrison

我非常享受。谢谢。

I've enjoyed it a lot. Thank you.

互动版:逐字朗读 + 针对本期提问 →