Harnesses and Evals: Owning Your Intelligence
打开互动全文版(中英对照 + 朗读 + 问答)→LangChain 的 Harrison 探讨了掌控智能体框架、模型和上下文的重要性,并解释了如何通过中间件定制核心智能体循环。
Harrison from LangChain discusses the importance of owning your agent's harness, model, and context, and explains how to customize the core agent loop using middleware.
Harnesses。我认为这是一个非常重要的话题。你们很多人现在都在思考如何构建自己的 harness。嗯,我非常高兴地介绍 Harrison。我第一次注意到 Harrison 是在 2022 年的 Twitter 上,当时还是 GPT-3 时代。Harrison 是最早思考“好吧,我们有这些模型,我们如何在它们周围构建一个完整的 harness,让它们不仅仅是自动补全任务,而是开始充当虚拟协作者或智能体?”的人之一。嗯,Harrison,自从 2022 年以来,生态系统已经发展了很多,我也看到你在思考如何构建智能体、构建 harness、如何评估它们等方面也成长了很多。嗯,所以,我非常高兴今天能请你来谈谈。我认为这次演讲将同时涉及 harness 和评估。然后形式再次是大约 15 分钟的演示内容,15 分钟的问答。感谢你加入我们,Harrison。
Harnesses. I think this is a very important topic. A lot of you are thinking through building your own harnesses right now. Um, I'm very excited to introduce Harrison. I first noticed Harrison on Twitter in 2022, back in the GPT-3 era. And Harrison was one of the first people thinking about, "Okay, we have these models. How can we build an entire harness around them so that they're not just um auto-complete tasks, but that they start acting as virtual collaborators or agents?" Um and Harrison, like the ecosystem has grown so much since 2022, and I've seen you also grow a lot in terms of how you think about building agents, building harnesses, how to eval them, etc. Um so, I'm very excited to have you talk today. I think the talk's going to be both about harnesses and evals. And then format again will be 15 minutes or so of presentation content, 15 minutes of Q&A. Thanks for joining us, Harrison.
酷。嗯,我叫 Harrison,LangChain 的联合创始人兼 CEO。我想在“拥有你自己的智能”的背景下谈谈评估和 harness。所以,当我们谈论智能时,我们通常指的是智能体。智能体到底由什么组成?在 LangChain,我们认为大致有三个主要部分。有一个 harness 来编排模型和一些上下文。如果你谈论的是拥有你自己的智能,你可能想要拥有这三部分。所以,拥有模型,我不打算多谈。Lin 来自 Fireworks,谈到了开放权重模型和拥有它。嗯,这部分的一个重要方面是切换模型的能力。嗯,过去有一种概念,就像云无关,能够切换云,那是以前的事。同样的事情也存在,但针对模型。你想要能够切换以避免锁定,但也要在最好的模型可用时使用它。上下文,你想要拥有你的智能体使用的所有上下文。无论是记忆,嗯,无论是语义知识,嗯,无论是之前的对话。这些可以帮助个性化和引导智能体前进。然后最后一部分是 harness,这是我真正想关注的。那么,你如何真正拥有你的 harness?这到底意味着什么?harness 的主要工作是什么?harness 的主要工作是在正确的时间点将上下文带给模型。所以,它围绕固定上下文、动态上下文做所有的编排。它把它带入模型的上下文窗口,展示一些东西,得到一些响应,然后对此做些什么。所以,智能体需要做所有这些不同的事情来完成它们的工作。它们还需要做大量的领域特定的事情,但它们需要与外部系统交互。这些外部系统,当你与它们交互时,它们会发出更多的上下文,这些上下文可以被反馈到智能体循环中。所以,harness 是真正将所有这些编排在一起的东西。
Cool. Um my name's Harrison, co-founder CEO of LangChain. I want to talk about evals and harnesses in the context of kind of owning your own intelligence. So, when we talk about intelligence, we're normally talking about agents. What exactly makes up an agent? At LangChain, we think there's kind of like three main parts. There's a harness that orchestrates a model and some context. And if you're talking about owning your intelligence in general, you probably want to own all three parts of these. And so, owning the model, I'm not going to talk too much about. Lin was here from Fireworks and talking about open weight models and owning that. Uh big part of this is also the ability to switch models. Uh there used to be this concept of kind of being like cloud agnostic and being able to switch clouds uh back in the day. Same thing exists, but for models. You want to be able to switch to avoid lock-in, but also to just use the best model when it's available. Context, you want to own all the context that your agent uses. Whether that is memory, uh whether that is semantic knowledge, um whether that is previous conversations. You can these can help personalize and guide the agent as it goes along. And then the last bit is the harness, and that's what I really want to focus on. So, how do you really own your harness? What does that even mean? What's the main job of a harness? The main job of a harness is to bring context to the model at the right point in time. And so, it does all the orchestration around the fixed context, the dynamic context. It brings it into the context window of the model, shows it something, gets some response, and then does something with that. And so, agents need to do all these different things in order to accomplish their jobs. There's a ton of domain-specific stuff that they need to do as well, but they need to interact with external systems. These external systems, when you interact with them, they emit more context that can get fed back into the agent into the loop. And so, the harness is the thing that really orchestrates all of this together.
智能体在最简单的情况下,当每个人谈论智能体时,他们真正谈论的只是一个 LLM 在循环中运行并调用工具。嗯,这是一个非常简单但非常通用的架构。一些请求进来,LLM 做出一些生成。该生成可能包含要调用的工具。如果有,你调用这些工具,并将观察结果传回给 LLM。这是当今几乎所有智能体背后的核心架构。但它们都以稍微不同的方式有所不同。所以,在左边这里,这是那种基础核心循环。但是在你的特定 harness 中,你可以在不同阶段做很多不同的事情。所以,这是呃,在这里,这是所以,我们构建了 LangChain,这是一个非常非常基础的最小 harness,那就是这里的 LangChain。然后这是 Deep Agents。Deep Agents 有点像我们的模型无关且更通用的 Quad Code 版本。所以,它做更多的事情。它连接到文件系统。呃,它有技能。它有子智能体。它构建在这个非常简单的 harness 之上,但我们通过使用这里的这些杠杆来定制它。所以,你可以在智能体被调用之前、每次模型调用之前运行特定的代码片段。你可以包装这些模型调用,你可以包装工具调用,你可以通过使用我们称之为中间件结构的小东西,以许多非常强大的方式定制这个核心简单循环。还有其他定制 harness 的方法,但这已经出现了,因为在许多编码智能体中也有钩子和插件的概念,这基本上就是它们所做的。它们采用这个正在运行的基础循环,并在各个点添加小钩子或插件,让你定制它。所以,当你定制时,你可以做很多事情,这一切都是通过中间件的概念,通过修改核心循环来完成的。所以,智能体仍然在循环中运行,它仍然使用相同的简单架构,但通过这种方式,你可以让它访问沙箱,你可以让它访问文件系统,你可以让它访问子智能体,你可以让它访问记忆,你可以进行摘要。所以,摘要,如果我们回到这件事,摘要会在模型之前出现。在模型被调用之前,你检查上下文是否太长,然后你总结它。所以,你可以通过中间件的概念将其添加到这个核心循环中,上下文卸载也是如此,嗯,这是一种基本上取大型工具调用并转储它们的方式。这有点包装了工具调用。所以,关键是有一个非常简单的通用智能体架构。所有这些更高级的智能体 harness 基本上都在做那个循环,但在运行核心循环的同时添加了一堆东西。所以,当你考虑构建或定制你自己的 harness 时,这些是你可以插入东西的不同位置。你可以添加你自己的摘要步骤,你可以添加你自己对特定工具调用的处理,这是你可以将智能体定制到你的特定领域和将 harness 定制到你的特定领域的一种方式。另一种定制智能体运行的 harness 的方式是拥有更明确的认知架构。所以,这曾经是很多人在 2023 年、2024 年构建智能体的方式,因为模型不够好,无法在循环中运行。所以,为了让它做特定的事情,你会有这些非常定制的认知架构。
Agents at their kind of like simplest, when everyone talks about agents, what they really talk about is just an LLM running in a loop calling tools. Um and uh this is a really simple, but really general architecture. Some request comes in, the LLM makes some generation. That generation may include a tool to call. If it does, you invoke those tools, and you pass that observation back to the LLM. And this is the core architecture behind pretty much every agent out there today. But they're all different in like slightly different ways. And so, on the left here, this is kind of like the base core kind of like loop. But there's a bunch of different things that you can do in your particular harness at different stages. And so, this is uh over here, this is So, we built LangChain, which is a really really base minimal harness, and that's LangChain over here. And then this is Deep Agents. Deep Agents is kind of like our model-agnostic and more general-purpose version of Quad Code. And so, it does more things. It connects to file systems. Uh it has skills. It has sub-agents. It's built on top of this really simple harness, but we customize it by using these uh these levers over here. So, you can run particular code snippets before the agent's invoked, before each model call. You can kind of like wrap these model calls, you can wrap the tool calls, and you can customize this core simple loop in a lot of really powerful ways just by using kind of like small what we call kind of like middleware constructs. There's other ways to customize the harness as well, but this is kind of emerged as uh there's a concept of hooks and plugins in a lot of the coding agents as well, and that's essentially what they're doing. They're taking this base loop that's running, and they're adding little hooks or plugins at various points to let you customize it. And so, a lot of the stuff that you can do when you can customize, this is all done by that concept of middleware, by just modifying that core loop. So, the agent's still running in a loop, it's still doing that same simple architecture, but through that you can give it access to a sandbox, you can give it access to a file system, you can give it access to sub-agents, you can give it access to memory, you can have summarizations. So, summarization, if we go back to this thing, summarization would come in before the model. Before the model's invoked, you check if the context is too long, and then you summarize it. And so, that you can add into this core loop through this concept of middleware, same with context offloading, um which is a way of basically taking large tool calls and dumping them. That kind of wraps the tool call. And so, the point is there's this really simple kind of like general architecture of an agent. All of these more advanced agent harnesses are basically doing that loop, but adding in a bunch of stuff while still running the core loop. And so, as you think about kind of like building or customizing your own harness, these are the different places that you can insert things into. You can add your own summarization step, you can add your own handling of particular tool calls, and that's one way that you can customize kind of like the agent to your particular domain and the harness to your particular domain. The other way that you can customize the harnesses that the agent runs in is by having a more explicit kind of like cognitive architecture. So, this used to be the way that a lot of people would build agents in kind of 2023, 2024 because the models weren't good enough to run in a loop. And so, in order to get it to do particular things, you would have these very bespoke cognitive architectures.
所以,这边这个是一个深度研究的例子,它会生成一些子问题,分派出去,然后去执行。然后这边这个是代码审查机器人。你可以看到这些非常定制化的步骤。现在很多这样的东西已经融入了 harness。我说的 harness,指的是它仍然是这个核心循环。这些可能会作为对核心循环的特定修改被添加进去。但对于很多非常特定的流程,我们确实看到人们仍然使用这样的认知架构来以特定方式引导它。
And so, this one over here is for a deep research example where it would generate some sub-questions, fan them out, and then go and execute them. And then this one over here is for a code review bot. And you can see that there's these very kind of bespoke steps. A lot of this has gone into the harness now. And by the harness, I mean it's still this core loop. These might be added as particular kind of modifications to that core loop. But for a lot of really particular kind of flows, we do see people still using cognitive architectures like these to really guide it in particular ways.
我们建议大家的一件事是从一个通用的 harness 开始。那是最容易上手的,也是最快能产生价值的。然后当你逐渐聚焦到你想擅长的用例上,你就可以开始添加更多这样的门控和检查,来引导它以特定方式运行。
One thing that we recommend to people is to start with a general harness. That's the easiest to get started. It's going to be quickest to time to value. And then as you kind of narrow in on the use case that you want to be excellent at, you can start to add more of these kind of gates and checks around it to guide it into particular ways.
我们经常收到的一个问题是,什么时候该考虑构建自己的 harness,而不是使用现成的 harness。很多现成的 harness 都是针对特定模型工作的。所以,现成的 harness 包括像 Claude Code 或 Claude Agent SDK,它们适用于 Anthropic 的模型;还有 Codex,适用于 OpenAI 的模型。我认为这是行业中的一个重大开放问题。我的答案通常是,你越接近模型训练的数据分布,现成的 harness 就越好用。一旦你开始越来越偏离分布,你可能就需要以某种方式调整你的 harness。
One question that we get a lot is when to think about building your own harness versus using an off-the-shelf harness. A lot of the off-the-shelf harnesses work with particular models. So, the off-the-shelf harnesses include things like Claude Code or Claude Agent SDK, which works with Anthropic models, Codex, which works with OpenAI models. I think this is a big open question in the industry. My answer generally is the more in distribution you are of what the models are trained on, then the better the off-the-shelf harness will be. As soon as you start to move further and further out of distribution, then you'll probably want to tune your harness in some way.
调整 harness 也有不同的方式。所以,模型可能在你正在做的某个特定任务上是分布内的,即使整个任务是分布外的。我的意思是,如果你考虑像法律 AI 这样的东西,Gabe 刚才谈到了,我想他提到他们有自己的 harness,在法律 AI 中有一些东西仍然在主模型的分布内。例如,编辑文件是主模型都经过强化学习(RL)训练过的任务。而且它们实际上是以非常特定的方式被 RL 的。所以,OpenAI 和 Claude 模型在他们的 harness 中以不同的方式编辑文件,结果,他们的模型实际上在不同方式上最擅长编辑文件。现在,模型本身在法律 AI 这个更大的任务上是分布外的,但在编辑文件这个任务上是分布内的。所以,如果你考虑构建一个在那里工作的 harness,你可能需要一个自定义的 harness,但你会希望它使用对你正在使用的模型来说是分布内的编辑文件工具。
There's different ways to tune the harness as well. So, the models may be in distribution on particular things that you are doing on an out of distribution task. So, what I mean by that is if you think about something like legal AI, which Gabe just talked about, and I think he mentioned how they have their own harness, there are things that are in legal AI that are still in distribution of the main models. So, for example, editing files is something that the main models have all been RL'd on. And they've actually all been RL'd in very particular ways. So, OpenAI and Claude models edit files in different ways in their harnesses, and as a result, their models are actually best at editing files in different ways. Now, the models themselves are out of distribution on this larger task of legal AI, but they're in distribution on this task of editing files. So, if you think about building a harness that works there, you'll probably want a custom harness, but you'll want it to use the edit file tool that is in distribution for the model that you're using.
所以,例如我们在 Deep Agents 中做的一件事,Deep Agents 是我们的可定制 harness,我们实际上有模型配置文件的概念,对于像编辑文件这样在模型分布内的事情,我们基本上会根据正在使用的模型切换不同的编辑文件实现。所以,我认为这是一个例子,当整个 harness 对某个任务来说是分布外的,但保持较小的分布内部分尽可能接近模型层。
So, one of the things we do in Deep Agents, for example, so Deep Agents is our customizable harness, we actually have this concept of model profiles, where for things that are in distribution of models like editing files, we basically switch between different edit file implementations depending on which model's being used. And so, I think that's an example of customizing the overall harness when it's out of distribution for a task, but keeping smaller in distribution parts as close to the model layer as possible.
我想讲的第二个重要部分是评估和可观测性。所以,我认为当你在实验智能体的各个部分时,无论是模型、harness 还是上下文,你都会想知道系统内部发生了什么,并且希望能够评估它。所以,这些是你可以使用的有用工具,再次强调,不仅适用于自定义 harness,也适用于自定义模型。
The second big part of what I want to talk about is evals and observability. And so, I think as you're experimenting with all parts of an agent, whether it's the model or the harness or the context, you're going to want to know what's going on inside of this system, and you're going to want to be able to evaluate it. And so, these are useful tools that you can use, again, not just for custom harnesses, but also for custom models.
所以,Satya 两周前在 Twitter 上发表了一篇很棒的文章,他谈到了很多这些概念。其中有三句话特别让我印象深刻。第一,创建你的私有评估,因为评估定义了组织内部什么是好的。第二,保留对你组织记忆、轨迹、反馈(虽然那个加粗是我加的)、决策和制度背景的所有权。第三,你创建自己的持续学习循环爬山机器,这将使你的 AI 投资能够复利增加你公司的价值。所以,我认为这些说明了评估和可观测性以及它们驱动的学习循环的重要性,真正拥有你的智能并使其复利增长。
So, there was a great Twitter article that Satya wrote 2 weeks ago, where he talked about a lot of these concepts. And there's three quotes in particular that kind of stood out for me. One, create your private evals because eval defines what good looks like inside the organization. Two, retain ownership of your organization's memory, traces, feedback, though that bold is mine, decisions, and institutional context. And then three, you create your own continuous learning loop hill climbing machine that will allow your AI investments to compound the value of your firm. And so I think these speak to the importance of evals and observability and the learning loop that they power in really owning your intelligence and compounding it.
那么他们具体是怎么做的呢?所以是评估。Gabe 在这里谈到了他们如何为法律领域构建基准。我认为每个公司在构建关键任务智能体时,都会为该智能体构建基准。你可以用它来定义和捕获回归,或者你可以在这个基准上进行爬山优化。同样,要么通过调整 harness,要么通过调整模型。
So how exactly do they do that? So evals. Gabe was here talking about how they built benchmarks for the legal domain. I think every company when they're building a mission-critical agent, they will build benchmarks for that agent. You can use it to define and catch regressions or you can hill climb on that benchmark. Again, either by adjusting the harness or adjusting the model.
我们看到成为定义这些基准的行业标准的是 Harbor。Harbor 是一个开源的评估运行器。它由 Terminal Bench 2 的创造者创建,Terminal Bench 2 是用于基准测试编码智能体的行业标准基准之一,并且它在各种领域已经变得相当流行。它让你能够做的是,这是 Frontier Bench,另一个编码基准。你得到这个漂亮的基准,你可以比较不同的智能体 harness、不同的模型、不同的推理努力,你可以得到这个漂亮的基准,你可以看到所有这些不同的 harness 和不同的模型在你的任务上表现如何。所以,为你的任务拥有一个基准,当你试图定义它时,将变得非常非常重要。
Things that we see becoming the industry standard for defining these benchmarks is Harbor. Harbor is an open-source eval runner. It's created by the makers of Terminal Bench 2, which is one of the industry standard benchmarks for benchmarking coding agents, and it's become pretty popular for a variety of domains. What it lets you do, so this is Frontier Bench, which is another coding benchmark. You get this nice benchmark and you can compare different agent harnesses, different models, different reasoning efforts, and you can get this nice benchmark and you can see how all these different harnesses and all these different models do on your task. And so having a benchmark for your task will become really really important when you're trying to define it.
Harbor 到底是什么?它非常简单。从高层次来看,它包含一个智能体,你让它针对一个数据集运行。数据集有一堆不同的任务。通常,它们在沙箱中运行,因为有很多不同的任务,你可能想要并行化它们。而且,正如我稍后会谈到的,每个任务都有自己的环境。
What exactly is Harbor? It's pretty simple. At a high level, it consists of an agent that you run against a data set. A data set has a bunch of different tasks. Generally, they're run in sandboxes because there are a lot of these different tasks and you might want to parallelize them. And as I talk about in a little bit, each task has its own kind of environment.
所以,这就是 Harbor 任务的样子。在右边,你可以看到它有一个环境。这是你定义智能体运行环境的地方。很多这些运行时间更长、更有状态的智能体需要与环境交互。所以,你基本上启动一个沙箱,给它一个在 Dockerfile 中定义的环境,并在那里运行它。然后有一个解决方案,基本上是一个黄金解决方案,你用它来做健全性检查,所以它不太有趣。测试更有趣。这基本上是智能体运行的验证器。测试脚本可以做任何事情。它们可以运行代码,可以运行单元测试,可以运行另一个 LLM 作为评判者,可以运行一个智能体作为评判者。你基本上定义了智能体在这个测试中如何被评分。然后 instruction.md 是给智能体的提示。这有点像 Harbor 的核心。
So this is what a Harbor task looks like. So, on the right, you can see that it has an environment. This is where you define the environment that the agent runs in. A lot of these longer running, more stateful agents need to interact with their environment. And so, you basically spin up a sandbox, give it its own environment that's defined in a Dockerfile, and run it there. There's then a solution, which is basically a kind of golden solution that you use to sanity check it, so it's not that interesting. Test is more interesting. This is basically the verifier for the agent run. The test scripts can do anything. They can run code, they can run unit tests, they can run another LLM as a judge, they can run an agent as a judge. You basically define how the agent is scored in this test. And then instruction.md is the prompt that the agent is given. And that's kind of like the core of Harbor.
你定义这些任务,它们是打包好的、可以在沙箱中运行的东西,然后你针对智能体运行一批任务。智能体又由模型和 harness 组成,你给它们打分。当你完成所有这些,你会得到什么?你会得到一些可以比较的好结果。所以,这就是 LangSmith,我们为评估和可观测性构建的平台。你可以在这里看到一系列不同的实验。我们与 Harbor 有很好的集成。你可以看到反馈分数。在这种情况下,它是一个单一的奖励函数。你还可以跟踪延迟和 token。当你对智能体进行基准测试时,你可能不只是关心准确性。你可能还关心延迟和成本。所以你会想跟踪所有这些。对于你运行的特定实验,这些就是 Harbor 数据集中的不同任务。
You define these tasks, which are bundled up things that can be run in a sandbox, and then you run a bunch of them against agents. Agents again consist of models and harnesses, and you score how they do. When you do all of that, what do you get? You get some nice results that you can compare. So, this is LangSmith, the platform that we build for evals and observability. You can see here a bunch of different experiments. We have a great integration with Harbor. You can see the feedback scores. In this case, it's a single reward function. You can also track latency and tokens. When you're benchmarking agents, you probably don't just care about accuracy. You also probably care about latency and cost. So you'll want to track all of those. For a particular experiment that you run, these would be the different tasks that are in a Harbor data set.
稍微谈谈可观测性。可观测性听起来很基础,但我认为它对于智能体来说非常重要,而且实际上被低估了。当智能体出错时,它们出错是因为 LLM 调用出了问题。为什么会出错?可能有两个原因。一是模型不够好。二是 LLM 接收的上下文不够好。我实际上认为第二个原因更常导致问题。对于进入模型上下文窗口的内容,以及上下文是如何累积的、执行了哪些步骤、运行了哪些工具、上下文是如何到达那里的,拥有非常好的可观测性,所有这些对于在智能体出错时进行调试都非常重要。所以这是我们拥有的一个可观测性视图。这旨在成为一个更用户友好的视图,我们实际上在其中表示它。这类似于你在 Claude Code 中可能看到的。我们隐藏了一些工具调用,所以你可以看到上面有七个工具调用。我们试图让它非常容易浏览。如今大多数智能体路径都以轨迹的形式出现。轨迹基本上就是你在 Claude Code 运行时看到的消息列表。当你运行 Claude Code 或其他智能体时,你输入一条人类消息,然后它进行一系列工具调用。这些都是底层消息,然后它响应,然后你再输入另一条人类消息。这有点像这种消息轨迹,它正成为这些智能体越来越核心的部分。但这还不足以完全调试它,所以我们还有这个完整的 trace,你可以点击特定内容,确切地看到模型内部发生了什么。这种可观测性对于了解正在发生的事情非常重要。
Talking a little bit about observability. Observability sounds basic, but I think it's really important and really underrated for agents, actually. When agents mess up, they mess up because an LLM call goes wrong. Why might it go wrong? It might go wrong for one of two reasons. One, the model is not good enough. Two, the context that the LLM received isn't good enough. I actually think it's the second one that more often than not causes issues. Having really good observability into what is going into the context window of the model and then how that context is accumulated, what steps were run, what tools were run, how does that context get there? All of that is really important for debugging your agent when it goes wrong. So this is one view of observability that we have. This is intended to be a more user-friendly view where we actually represent it. This is similar to what you might see in Claude Code. We hide some of the tool calls so you can see seven tool calls up there. We try to make it really easy to skim through this. Most agent paths these days come in the form of trajectories. Trajectories are basically the list of messages that you see like Claude Code running. When you run Claude Code or another agent, you type in a human message, it then makes a bunch of tool calls. Those are all messages under the hood and then it responds and then you type in another human message. That's kind of like this message trajectory that is becoming more and more of a central part of these agents. But that's not enough to fully debug it and so we also have this full trace and you can click into particular things and see exactly what goes on inside the model. This type of observability is pretty important for knowing what's going on.
评估和可观测性确实让你能够建立这个数据飞轮,并随着你开始使用智能体、你的用户开始使用智能体并开始获得反馈,从而复合智能。这是我们团队成员在 Swyx 的 AI 工程展上展示的一张幻灯片,实际上是关于持续改进智能体的配方。在高层次上,这非常简单。你构建一个智能体,开始运行它,收集大量 trace,然后整理 trace 数据,然后在你创建的数据上运行实验。这非常简单,但当然底层有很多复杂性。对此非常重要的一件事是反馈。从环境或从合成来源获取反馈。从环境来看,我认为在智能体设计中被严重低估的一件事实际上是 UX 设计,即你如何向用户展示智能体。如果你以非常智能的方式展示它,你实际上可以从用户那里获得大量反馈。他们可能不会明确点击赞或踩。没有人真正这样做。但如果你巧妙地设计 UX,你可以获得一些这样的反馈。你可以做的另一件事是开始获得合成反馈。你可以运行我们所说的在线评估器来评判这些 trace。Gabe 谈到了我们与 Harvey 做的一个实验,我们显著降低了某些 LLM 作为评判者的成本。如果你想象对进入系统的每一个 trace 运行 Opus,那将产生巨额账单。你想要一种非常便宜和快速的方法来做这件事。所以我们微调了一些 SLM 来实际做这件事,但你当然也可以使用现成的模型和自定义提示来做。或者,如果你要测试的一些事情足够简单,你也可以直接使用代码。
Evals and observability really let you set up this data flywheel and compound the intelligence as you start to use the agent, as your users start to use the agent and you start to get feedback. This is a slide that one of our team members presented at Swyx's AI engineering fair, actually, around a recipe for continuously improving agents. At a high level it's really simple. You build an agent, you start running it, you collect lots of traces, you then curate the trace data, and then you run experiments on that data that you create. It's really simple, but of course there's a lot of complexity under the hood. One thing that's really important for this is feedback. Getting feedback either from the environment or from synthetic source. From the environment, one thing that I think is really underestimated in agent design is actually UX design of how you present the agent to your users. If you present it in a really intelligent way, you can actually end up getting a lot of feedback from them. They may not click thumbs up or thumbs down explicitly. No one really doing that. But if you design the UX in a clever way, you can get some of that feedback. The other thing you can do is you can start to get synthetic feedback. You can run what we call online evaluators over these traces to judge things. Gabe was talking about an experiment that we did with Harvey where we significantly reduced the cost of some of these LLM-as-a-judge type things. If you imagine running Opus over every single trace that comes into your system, that's going to rack up a big bill. You want a really cheap and fast way of doing this. So we've fine-tuned some SLMs for actually doing this, but you can of course use off-the-shelf models with custom prompting to do it. Or you can just use code if some of the things that you want to test are simple enough.
所以这是它的完整部分。整理 trace 数据,反馈是其中很大一部分。然后另一件事是,当你使用这些数据来更新所发生的事情时,你可以用这种系统更新智能体的任何部分。你可以通过 harness 工程来更新 harness。你可以通过微调来更新模型。你可以通过记忆来更新上下文。我们在 LangChain 最关注的部分是 harness 工程部分。我想快速演示一下我们添加的其中一个功能,以帮助实现这一点。但我想 Trajectory 接下来会讲一些可以做的微调。这是一个非常相似的过程,你运行智能体,获得一些 trace,以某种方式使用这些数据来改进系统。系统是什么?就是这三个部分。其中任何一个都可以以某种方式更新。所以,是的,这是你可能想要做的完整端到端流程。我们思考的一件事是如何尽可能自动化这个过程,因为这很棘手且耗时。这是我们在过去几个月里一直在思考的事情之一。我想快速演示一下我们称之为 LangSmith Engine 的东西,它基本上是一个位于你的 trace 之上的智能体,负责完成所有这些工作。如果我们看看那项工作是什么,你有这些 trace,从那里开始的工作是整理 trace 并运行一些实验,建议对三个部分之一进行修复。正如我提到的,我们主要关注 harness 工程部分。在演示中,我想展示这是什么样子,以及它如何表示这一点。希望这会成功。如果不成功,也没关系。完美。好的。所以这是 LangSmith。这是我们收到的一堆 trace。我们这边有一个叫 engine 的标签。这是一个智能体。它在后台运行。它创建我们所谓的问题板。这是整理数据的一部分。它会寻找……它基本上在底层是一个编码智能体,可以访问我们的 LangSmith CLI。
So this is the full part of it. Curating the trace data, feedback is a big part there. And then the other thing is when you use that data to update what happens, you can update any part of the agent with this kind of system. You can update the harness by doing harness engineering. You can update the model by doing fine-tuning on that. You can update the context by doing memory. The part that we think most about at LangChain is the harness engineering part of that. I want to show a really quick demo of one of the things that we added to help with that. But I think Trajectory is talking next on some fine-tuning that can be done. It's a very similar process where you run the agent, get some traces, use that data in some way to improve the system. What is the system? It's these three pieces. Any of these can be updated in some way. So yeah, this is the full end-to-end flow that you might want to do. One of the things that we think about is how can you automate this as much as possible because this is tricky and takes a lot of time. That's one of the things that we've been thinking about for the past few months. I want to do a quick demo of what we call LangSmith Engine, which is basically an agent that sits on top of your traces and does all this work. If we look at what that work was, you've got these traces, the work from there is curating the traces and running some experiments, suggesting fixes to one of the three things. As I mentioned, we mostly focus on the harness engineering bit. In the demo I want to show what this looks like and how it represents that. Hopefully this will work. If not, it's not that big of a deal. Perfect. Okay. So this is LangSmith. This is a bunch of traces we have coming in. We have this tab called engine over here. This is an agent. It runs in the background. It creates what we call issue boards. This is the part of curating data. It will look for... It will basically under the hood is a coding agent that has access to our LangSmith CLI.
通过 LangSmith CLI,你可以按反馈等条件筛选追踪记录。所以我们给它一个很大的提示词和一些子智能体,帮助它去探索这些数据、识别问题、看看有哪些常见情况。然后它就在这里创建这些问题。这里它创建了一个问题,给出了描述,还链接到追踪记录,这样我就能去看一些佐证。然后下面,我猜这是对提示词的一些简单修改。但这里它更新了部分上下文,这里还更新了一些指令。我们还能看到它添加了一些代码进入 harness。这是我们过去几个月推出的东西,我认为它体现了这个数据飞轮,再说一次,这是一个非常简单的事情:运行智能体,获取追踪,发现模式,修复。而这是我们自动化这一过程的尝试。我就讲这些,欢迎就 harness 或评估提问。
With the LangSmith CLI, you can filter traces for feedback and other things like that. So we give it a nice big prompt and some sub-agents that help it basically go out and explore this data and identify issues and see what common things are. And then it will create these issues right here. And so here it's created an issue. It gives a description of it. It links to the traces so I can go see some supporting evidence. And then down here, I guess this is very simple changes to the prompts. But here it's updating part of the context in this case. Here it's also updating some instructions. And we can see here that it's adding some code to go into the harness as well. And so this is something we launched in the past few months and I think speaks to this data flywheel, which again is a very simple thing. Run agent, get traces, see patterns, fix. And this is our attempt at automating it. That's all I've got. Happy to take any questions on harnesses or evals.
顺便说一句,这很棒。非常感谢你的整个演示。Engine 本身就是一个智能体,对吧?给它一个提示词,然后它去搜索各种东西。你让 Engine 跑过 Engine 吗?
This is great, by the way. Really appreciate the whole presentation. Engine itself is an agent, right? That's given a prompt and go search over things. Have you run Engine on Engine?
我们确实在运行它,是的。所以 Engine 也接入了 Slack,会发送一些关于它自己的报告。是的,这就是我们吃自己狗粮的方式。我们还为 Engine 创建了所谓的 issue bench,这基本上又是一个 Harbor 格式的基准测试,我们不断在上面测试不同的模型和不同的 harness。我想 Gabe 之前也提到过一点,但拥有基准测试的一个好处是,你可以在各种不同的 harness 上运行它,看看它们各自的优劣。所以几周前,我们在这上面运行了我们自己的 deep agents,然后是 Codex,再然后是 Claude Code。我们发现 Codex 做了一个非常有趣的事情:它会给自己写一堆小脚本去跑这些追踪记录,而且它非常激进地这么做,这实际上让它表现得非常好。所以我们做了一个冲刺,就是我们所谓的 Engine 的 codex 化,基本上就是把那个经验带进核心的 Engine harness。所以我认为拥有基准测试的另一个好处是,你可以在上面运行各种不同的东西,看看它们实际表现如何,然后把那些东西带回你的核心智能体 harness。
We have it running, yeah. So, Engine also hooks up to Slack and sends it kind of like reports about itself. And yeah, that's how we dogfood it. We also created what we call kind of like issue bench for Engine, which again is like a Harbor-formatted benchmark, basically, that we're constantly benchmarking different models and different harnesses on. And so, I think Gabe talked about this a little bit, but one of the benefits of having a benchmark is you can benchmark it on a bunch of different harnesses and see what they're good and bad at. So, we ran our own kind of like deep agents and then Codex and then Claude Code on this a few weeks ago. And we saw that Codex was doing a really interesting thing, where it would write itself a bunch of small scripts to run against these traces, and it was doing that really aggressively and actually allowing it to perform really well. So, we did a sprint to do what we call the codis- codexification of Engine and basically take that learning and bring it into the core Engine harness. And so, I think that's another benefit of having a benchmark is you can just run a bunch of different things on it and see how they actually perform and then bring those things back into your core agent harness.
非常酷的演讲。你认为 harness 会在多大程度上收敛为一个统一的东西,用户会被教育去那样做,模型也会为此做到最好,而不是走向多样化——每家公司都有自己的做事方式,并针对自身进行优化?
Super cool talk. To which extent do you think that harnesses will converge into one thing and users will be educated to do that and the models will be best for that versus diversifying here, every company has their own way of doing things optimized for them?
是的,非常好的问题,我们经常思考这个问题。我和 Factory 的 Eno 聊过,他也经常思考这个。我想我们讨论过一些相关的东西。我觉得模型,我不知道,这是诚实的回答。我看到的一些情况是,通用型 harness 已经足够好,至少在你刚开始时,能处理很多基础任务。所以我建议从现成的 harness 开始,无论是 Deep Agents、Codex 还是 Claude Code 之类的。因为我认为模型现在已经足够好,而且我们学到的关于什么让这些模型表现出色的东西——文件系统访问、子智能体等等——这些已经足够好了。我认为我们经常看到,你越偏离常规分布,你就越会想要定制 harness。这是一个谱系,对吧?所以在谱系的极端,你可能会想构建一个完整的认知架构,真正专注于某些事情。顺便说一句,你可能想这样做的另一个原因是可预测性和控制。所以我们有很多金融服务行业的客户,他们需要可预测性。所以我们给他们看 Deep Agents 这样的东西,他们会说:“哇哇哇,这智能体对我们来说太吓人了。我们想要更多这种自定义的认知架构,这样我们才能真正控制事情。”但另一方面,你可以直接使用现成的 harness,中间还有一些东西,比如钩子或中间件,你可以使用。所以这也是一个谱系。你越偏离常规分布,你就越会想要拥有更定制的 harness。然后还有其他一些奇怪的事情,比如,我认为 OpenAI 和 Anthropic 都在编程上变得非常出色,但他们在编辑文件的方式上却采用了实际上相当不同的方法。我认为他们有一些基准测试,而且他认为一种方式在严格意义上优于另一种。所以我认为模型实验室会有所收敛,它们似乎都非常擅长编程,如果它们继续沿着这条路走下去,harness 也会收敛到非常擅长编程。但与此同时,这些非常小的差异仍然存在,我也不知道如何解释它们。而且目前我认为这些差异最具体地体现在小事情上,但你可以想象,如果某个实验室真的深入生物领域,那些 harness 会变得非常擅长生物智能体的事情,那么 harness 本身就会开始分化。所以我不知道,这是对这个快速变化领域的回答。这就是为什么评估和可观测性很重要,我认为我们需要衡量所有这些。
Yeah, really good question, one that we think a lot about. And I chatted with Eno from Factory, who also thinks a lot about this. And I think there's some stuff we talked about this. I think the model, I don't know, is the honest answer. I think some things that I've seen is that the general purpose harnesses have gotten good enough to work for a lot of basic tasks at least when you're getting started. So, I would recommend getting started with an off-the-shelf harness, whether it's Deep Agents or Codex or Claude Code or something like that. Because I think the models are now good enough and the things that we've learned about what makes these models good, access to file systems, sub-agents, things like that, those are kind of good enough. I think we often see that the more out of distribution you get, the more you're going to want to customize the harness. And it's a scale, right? So, at the extreme end of a scale, you might want to build a complete cognitive architecture that is really focused on things. Another reason you might want to do that, by the way, is for predictability and control. And so, we have a lot of customers in financial services where they need predictability. And so, we show them something like Deep Agents and they're like, "Woah, woah, woah. That's way too scary an agent for us. We want more of this kind of custom cognitive architecture where we can really control things." But then on the other end, you could just use an off-the-shelf harness and there's things in the middle like hooks or middleware that you can use. So it's a spectrum as well. The more out of distribution you get, the more custom harness you're going to want to have. And then there's other weird things where, again, I think both OpenAI and Anthropic are getting really good at coding, but they've landed on different ways to edit files that are actually pretty different. I think they have some benchmark and I think he thought that one way was just strictly superior to the other. And so I think the model labs will kind of converge and that they all seem to be really good at coding, the harnesses will converge to being really good at coding if they keep on going down that path. But at the same time, there are these really small differences and I don't really know how to explain those either. And right now I think those show up most concretely in small things, but you could imagine what if one lab really goes down bio and those harnesses become really good at bio agent things, then the harnesses themselves start to diverge. So I don't know is the answer to the fast-moving space. That's why evals and observability are important, and I think we need to measure all of that.