网页交互的未来:AI 代理与浏览器自动化

The Future of Web Interaction: AI Agents and Browser Automation

黛维·帕里克 Devi Parikh · TWIML AI 播客 · 2025-11-18 · 约 55 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Devi Parikh 探讨 AI 代理如何将网页交互从手动点击转变为高级任务描述,并分享他从计算机视觉到创立 Yutori 的历程。

Devi Parikh discusses how AI agents will transform web interaction from manual clicking to high-level task description, and shares his journey from computer vision to founding Yutori.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 26)

全文 · Full transcript(中英对照)

AI代理的愿景与介绍 Introduction and vision for AI agents

Host

我们将不再像现在这样与网络互动。我们不会再点击按钮、摆弄网站和浏览器上的表单。我们将在一个更高的抽象层次上与网络互动,描述需要完成的事情。也许我们的助手会主动注意到需要做什么,然后后台的智能体开始代表你在网络上执行这些工作流程。好了,各位,欢迎收听另一期 TwiML AI 播客。我是主持人 Sam Sharington。今天,我们邀请到了 Devi Parikh。Devi 是 Yutori 的联合创始人兼联合 CEO。在开始之前,请花点时间点击订阅按钮,无论你在哪里收听今天的节目。Devi,我们上次聊天已经有一段时间了,欢迎回到播客。

We will no longer be interacting with the web in the same way that we do right now. We won't be clicking buttons, fiddling with forms on websites and browsers. We'll be interacting with the web one level higher in the abstraction where we're describing what needs to be done. Maybe our assistant is proactively noticing what needs to be done and sort of agents in the background are starting to execute these workflows on the web on your behalf. All right, everyone. Welcome to another episode of the TwiML AI podcast. I am your host, Sam Sharington. Today, I'm joined by Devi Parikh. Devi is co-founder and co-CEO of Yutori. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Devi, it has been a while since we caught up last. Welcome back to the podcast.

Devi Parikh

谢谢。谢谢你再次邀请我。

Thank you. Thank you for having me again.

Host

是啊,五年过去了。其实也没发生什么大事,对吧。

Yeah, five years later. Not much has happened at all. Right.

Devi Parikh

从某些方面来说,发生了很多事,但从某些方面来说,我又觉得,“哇,已经五年了。”所以,是的。

In some ways a lot has happened, but in some ways I'm like, "Wow, it's been five years." So, yeah.

Host

我知道。我知道。所以,我们打算聊一聊 AI 浏览器和浏览器使用智能体,以及你在 Yutori 正在构建的东西,但我希望你能花几分钟时间,跟我们讲讲你最近在忙什么。

I know. I know. I know. So, we're going to be talking a bit about AI browsers and browser use agents, and what you're building at Yutori, but I'd love to have you take a few minutes and catch us all up on what you've been up to recently.

Devi Parikh

是的。我可以再往前追溯一点,不只是五年,简单聊聊我的背景。我在 AI 领域已经工作了大约 20 年。最初我的博士论文是关于计算机视觉的,后来我逐渐对探索如何让人们更自然地与这些系统互动产生了兴趣,于是转向了视觉与语言交叉的多模态问题。比如,给定一张图片,你能用一句话描述它吗?你能回答关于它的问题吗?你能就图片内容进行来回对话吗?那是在 2014 年左右。那是在深度学习模型最初的热潮之后,但已经开始让人觉得,等等,这些模型确实在做一些事情。东西真的开始奏效了。但那远在如今围绕生成式 AI 等热潮之前。所以这些模型远没有今天这么好。嗯,在可能性的边界上摸索还是挺有趣的。然后我开始对探索如何将 AI 用作创意表达的工具产生了兴趣。这就是我涉足图像、视频、音乐等模态的生成模型的原因。我在学术界待了一段时间,在弗吉尼亚理工大学和佐治亚理工学院任教,然后在 Meta 工作了大约八年,先是在 FAIR,后来在 GenAI,担任高级总监,领导那里的许多多模态研究工作。像 Emu、Emu Video、Emu Edit 这样的图像和视频生成与编辑模型,都在 Meta 的各个平台上发布。我的团队参与了这些工作,Llama 3 中的多模态能力也来自我的团队。这一直持续到去年初,我和我的联合创始人离开了 Meta,创办了 Yutori。

Yeah. And I can go a little bit further back than five years, just to talk about my background a little bit. So, I've been working in AI for about 20 years now. Originally my PhD thesis was in computer vision and then over time I got interested in seeing if we can find ways in which people can interact with these systems more naturally and that's how I moved towards multimodal problems at the intersection of vision and language. So things like given an image can you describe it in a sentence? Can you answer questions about it? Can you have a conversation going back and forth about the content of an image? And this was back in 2014 or so. So it was after that initial excitement of deep learning models, but it was starting to feel like wait, these models are doing something. Stuff is actually starting to work. But it was well before all of the current excitement around GenAI and others and so on. So these models weren't really as good as they are today. And yeah, so it was kind of fun to tinker on the boundaries of what's possible. And then I started getting interested in seeing if we can find ways in which we can use AI as a tool for creative expression. And that's how I got involved with generative models for images and videos and music and other modalities like that. I was in academia for a while, faculty at Virginia Tech and then Georgia Tech, and then I was at Meta for about eight years, first in FAIR then in GenAI, where I was a senior director leading a lot of the multimodal research efforts there. So models like Emu, Emu Video, Emu Edit for image and video generation and editing, they were shipped across Meta surfaces. My teams were involved in that and the multimodal capabilities in Llama 3 were coming from my teams as well. And this was up until early last year, where my co-founders and I left Meta to start Yutori.

Host

如果我记错了请纠正我,但你对时尚感兴趣过。我记对了吗?你是不是做过一些关于时尚数据集之类的论文?

Correct me if I'm misremembering this but at some point you were interested in fashion. Am I remembering that correctly? Did you do some papers on like fashion datasets or something?

Devi Parikh

是的,我做过。嗯,我和其他合作者在这个领域做过几个项目。我想上次我们聊天时,我正好处在寻找下一个方向的边缘,开始探索能否将 AI 用作创意表达的工具。所以当时我做了一堆有点奇怪的小项目。其中一些出于热情的项目更正经一些。我不会说那些奇怪,因为我是和其他合作者一起做的,但确实如此。

I did. I did. Yeah, I had done a couple of projects in that space with other collaborators. I think the last time we talked I was right at that edge of looking for the next thing and I was starting to tinker in the space of like can we use AI as a tool for creative expression. So there were a whole bunch of kind of weird little projects that I had done at the time. Some of the passion ones were more legit. I won't say those were weird, like I was doing them with other collaborators but yeah.

Host

跟我们讲讲你想用 Yutori 解决什么问题。

Tell us a little bit about the problem that you're aiming to tackle with Yutori.

Devi Parikh

我们在 Yutori 的工作是朝着这样一个愿景:我们将不再像现在这样与网络互动。我们不会再点击按钮、摆弄网站和浏览器上的表单。我们将在一个更高的抽象层次上与网络互动,描述需要完成的事情,或者我们的助手会主动注意到需要做什么。然后后台的智能体开始代表你在网络上执行这些工作流程。这些智能体将始终在线,主动且个性化。因此,我们在 Yutori 的工作是同时构建底层技术和产品体验,以逐步引领这一变革。

So what we're working on at Yutori is towards this vision that we will no longer be interacting with the web in the same way that we do right now. We won't be clicking buttons, fiddling with forms on websites and browsers. We'll be interacting with the web one level higher in the abstraction where we're describing what needs to be done or maybe our assistant is proactively noticing what needs to be done. And sort of agents in the background are starting to execute these workflows on the web on your behalf. These agents will be always on. They'll be proactive. They'll be personalized. And so what we are working on at Yutori is both building the underlying tech and the product experiences to sort of usher in this change over time.

Host

你是如何确定这个问题的?也跟我们讲讲创始人吧。你的联合创始人之一是你丈夫 Duv,对吗?

And how did you settle on that problem? And tell us a little bit about the founders also. One of your co-founders is your husband Duv, is that right?

Devi Parikh

是的。而且我记得他好像也上过你的播客,如果我没记错的话。可能我记错了。但确实如此。嗯,我可以谈谈我们是如何确定这个问题的,然后也可以聊聊我们三个人。总的来说,效率、生产力,或者更广泛地说,设计一个对你更有意义的生活,这个领域是我个人的动力来源。Yutori 这个名字其实是一个日语词,指的是因精神上的宽敞感而体验到的幸福感。也就是说,你不需要每隔几分钟就切换上下文。不会有铺天盖地的通知向你涌来。你不想做的事情在后台被处理好了,你就有空间和时间专注于任何对你有意义的事情。所以这个领域是我个人的动力。然后在技术方面,感觉这些网络智能体、数字助手还没有完全准备好。比如,当时你不能直接把 GPT 放在一个循环里,就指望这些智能体可靠地自主执行网络工作流程。同时,它似乎也不需要像机器人技术那样等上十年才能获得可靠性。所以感觉这是一个最佳点,凭借我们的研究背景和我们能组建的团队,我们能够在底层能力上取得实质性进展,并为人们带来他们日常使用的实际产品用例。所以这感觉是一个最佳点。嗯,这就是我们得出这个结论的方式。

Yeah. Yeah. And I think actually he has also been a guest on your podcast if I remember correctly. I could be wrong there. But yeah. That also. But yeah. So I can talk about how we settled on this problem and then I can also talk about the three of us. So the general space of sort of efficiency, productivity or I think more generally just sort of designing a life that is more meaningful to you is something that's personally motivating. And Yutori actually the name is a Japanese word for the sense of well-being that you experience as a consequence of mental spaciousness. So sort of you're not trying to context switch every few minutes. You don't have a bazillion pings coming your way. Things that you would rather not be doing are being taken care of for you in the background and you have the space and time to focus on whatever it is that is meaningful for you. So that general space was personally motivating. And then on the technical front, it was feeling like these web agents, these digital assistants, is something that wasn't quite ready yet. Like at the time you couldn't just take GPT, put it in a for loop and sort of expect these agents to do workflows on the web reliably autonomously. At the same time, it didn't seem like it was going to take 10 years before we can start getting that reliability, unlike something like robotics for instance. And so it felt like that sweet spot where with our research backgrounds, with the kind of team we can put together, we'll be able to make a solid dent on the underlying capability and bring to people actual product use cases that they're using on a day-to-day basis. And so that felt like a sweet spot. And so yeah, that's how we arrived at that.

机器人背景与创始团队 Robotics background and founding team

Host

说到机器人技术,我发现那期与 Droo 的节目,他谈到在盲 AI 智能体中构建地图和空间感知的工作,那只是两年前的事。

And speaking of robotics, I found that episode with Droo that was talking about his work building maps and spatial awareness in blind AI agents and that was only two years ago.

Devi Parikh

我明白。是的。没错。所以他当时在做机器人研究,在 FAIR 领导了很多具身 AI 的工作。其中涉及序列决策、智能体采取行动、行动有后果等,很多东西都延续了下来。但在构建网络智能体时,这些实体在物理世界中围绕你的许多挑战都被消除了。所以这影响了我们的视角。简单介绍一下三位创始人。三位创始人是我本人、Abhishek Das(大家叫他 Das,那是他的姓),以及我们刚才提到的 Tuvatra。Duv 和我是夫妻。我们从认识以来就一直一起工作。相识的 20 年里,有 19 年我们受雇于同一家公司。我们的办公室相邻,办公桌相邻,有相同的经理,等等。Das 是 DH 在佐治亚理工的博士生。因为 Duv 和我一起管理实验室,我也与 Das 密切合作过。现在,我们三人是非常好的朋友。六七年前我们就讨论过一起创业。我们甚至有一个每周晚餐,最初叫“头脑风暴”,因为我们在 brainstorm 如果做这件事会做什么。后来就变成了社交聚会。我们连续 brainstorm 了六七年。但就是这样。

I see. Yeah. Yeah. So exactly. So he was working in robotics like he was leading a lot of the embodied AI efforts at FAIR. And so there are sort of certain sequential decision-making, these agents taking actions, there being consequences to the actions, a lot of those things carry over. But a lot of the challenges of these entities just being physically around you in the physical world are taken out when you're building web agents. And so yeah, that has influenced our perspective. And so a little bit about the three founders. The three founders are myself, Abhishek Das, who goes by Das, which is his last name, and Tuvatra, who we were just talking about. Duv and I are married. We've worked together the whole time we've known each other. We've had the same employer 19 of the 20 years we've known each other. Our offices have been next to each other. Our desks have been next to each other. We've had the same managers. The whole thing. Das was DH's PhD student at Georgia Tech. And because Duv and I ran our labs together, I have collaborated very closely with us as well. And at this point the three of us are just really good friends. We had talked about starting something together going back six, seven years. We even had this weekly dinner that was originally called brainstorming because like yeah we were just brainstorming on like if we did this what would we do it on? Over time we were just socially hanging out. We were brainstorming for six, seven years straight. But yeah.

赌注与愿景 Stakes and vision

Host

那么,这是否意味着这个想法的风险很高,不允许有任何转向?

So does it feel like the stakes are really high for this to be the idea and no pivots are allowed?

Devi Parikh

我不这么认为。我们确实相信这个愿景——我们与网络互动的方式将发生巨大变化,我想很多人也会认同这一点。我认为关键在于我们如何执行。随着发展,我们推向市场的产品用例是什么,这些都是实验,对吧?我们推出一些东西,从中学习,调整,然后继续前进。是的。

I don't think so. We do genuinely believe in this vision that how we're interacting with the web is going to change drastically and I think a lot of others would also buy that. I think the devil is in the details of how we execute on it. What are the product use cases that we bring to the market as we go along and those are all experiments, right? We put something out there, learn from it, tweak it, and go from there. Yeah.

空间概览与差异化 Space overview and differentiation

Host

那么,我们不妨深入探讨这个广阔领域,尝试涵盖为什么人们对使用浏览器智能体自动化网络如此兴奋,以及你如何看待自己的工作与市场上其他产品的区别。我想到 OpenAI 的 Atlas、Perplexity 的 Comet 等浏览器,但可能还有一二十个甚至几百个。请为我们阐述你对这个领域的看法。

So let's maybe dig into that broad space and try to cover why folks are so excited about automating the web with browser use agents, how you see what you're working on versus some of the other things that are out there. OpenAI's Atlas comes to mind, Perplexity's Comet come to mind in terms of browsers, but there are probably a dozen or two more, if not hundreds. Lay out the way you think about the space for us.

Devi Parikh

我认为有几个维度值得评论。一是你选择关注技术栈的哪一部分。有很多工作专注于底层模型和底层技术,使其可靠。还有一些工作更专注于产品体验本身。但我认为在这个领域,至少在这个时间点,鉴于技术的现状,你需要跨整个技术栈进行创新。你需要推动底层模型能力和架构,同时思考基于你对技术现状的了解,你能可靠地交付哪些产品体验。这两者需要齐头并进。如果你只做其中一项,那是不够的。如果你只关注技术,你会有看起来很酷的技术演示,但没人日常使用。如果你孤立地思考产品体验,而不以模型能力为基础,那么你要么承诺过多,用户第一次尝试时无法兑现。所以你需要跨技术栈推进。第二点,更相关的是你提到的 AI 浏览器,无论是 Atlas、Comet 还是 DIA,这有点像是把我们今天与网络互动的方式(通过浏览器)用 AI 功能增强,这很有用。但我认为我们的思考方式更倾向于它甚至不应该是这样的。我们不应该以今天的方式待在浏览器里,以今天的方式查看网页。如果我想在网站上完成 X,而你想完成 Y,既然我们做的是不同的事情,网站就应该以两种不同的方式呈现给我们,因为我们的意图不同,目的不同。没有理由你我都看到同一个网站。第二点,我们坚信这些智能体应该在后台运行,不占用你的空间。这涉及到几个方面。一是,假设你合上笔记本电脑走开,如果智能体在你的设备上工作,它们会怎样?这是其一。二是,这些系统天生具有多智能体价值,大量智能体并行为你执行工作流。如果它们都在你的浏览器里,那将无法扩展。它会接管你的设备。所以出于几个不同的原因,我们设想的产品体验与现有浏览器中的 AI 功能非常不同。

I think there are a couple of dimensions worth commenting on. One is what part of the stack you choose to focus on. There are a good number of efforts focused on the underlying models and the underlying tech, getting those to be reliable. And then there are efforts focused more on the product experiences themselves. But I do think that in this space, at least at this point in time, given where the tech is, you need to be innovating across the stack. You need to be pushing on the underlying model capabilities and architectures, and in tandem thinking through what product experiences you can deliver on reliably based on what you know the status of the tech is. Those need to go hand in hand. If you do one or the other, that's not going to be sufficient. If you focus just on the tech, you have tech demonstrations that are awesome to look at, but no one uses them on a day-to-day basis. And if you think through product experiences in an isolated way, not grounded in the modeling capabilities, then you either end up promising too much that doesn't live up to it the first time the user tries it. So you need to be pushing across the stack. The second, more relevant to what you were talking about with AI browsers, whether it's Atlas, Comet, or DIA, that has a little bit of the flavor of taking how we are interacting with the web today, which is through these browsers, and enhancing that with AI features, which is useful to do. But I think the way we think about it is more that it shouldn't even look like this. We shouldn't even be in browsers the way they are today, looking at web pages the way we do today. There's no reason if I am trying to get X done on a website and you are trying to get Y done on the website, given that we're trying to do two different things, the website should just show up to us in two different ways because our intent is different. Our purpose is different. There's no reason you and I are both looking at the same website. And the second bit is that we are strong believers of these agents being in the background, out of your space. The way that's relevant is a couple of different things. One is, let's say you shut the lid of your laptop down and walk away, and if the agents are working on your device, what happens to them? That's one. The second is there is a lot of value to these systems inherently being multi-agent, where there are a whole bunch of agents in parallel executing on these workflows for you. And if they are all in your browser, that's not going to scale well. It's just going to take over your device. So for a few different reasons, the kinds of product experiences we envision are very different from having AI features in existing browsers.

产品路径:用例vs平台 Product approach: use cases vs platform

Host

那么,这是否意味着你必须逐个用例地处理这些体验,并且最终会发布一系列服务于不同需求的分散产品,还是这些实验或步骤会引导你走向一个更广泛的平台方法,可以跨用例应用?

So is the implication of that that you have to tackle these experiences use case by use case and does that mean that you end up releasing a bunch of disparate products that serve different needs or are these experiments or steps that lead you to some broader platform approach that can be applied across use cases?

Devi Parikh

是的,我认为更像是后者。我们的方式——我可以谈谈我们推出的第一个产品,你会明白我的意思,然后我会进一步讨论后续可能的产品。我们推出的第一个产品叫做 Scouts。Scouts 监控网络上任何你关心的事情。所以如果你想随时了解任何新的 AI 公告,你可以为此设置一个 scout。

Yeah, I think it's more the latter. The way we've—I can talk about the first product that we put out there and you'll see what I mean, and I'll talk more about what could follow. So the first product that we put out there is called Scouts. Scouts monitor the web for anything that you care about. So if you want to stay updated on any new AI announcements, you can set up a scout for that.

Scout产品概览 Scout product overview

Host

如果你想买某个产品,但它缺货或价格太高,你想在补货或价格低于阈值时收到通知,你可以为此设置一个 Scout。如果你在找实习或工作,想在任何符合特定条件的新职位发布时收到通知。如果你在找公寓,也是一样。如果你在做商业和竞争情报分析,想在任何竞争对手的产品收到差评时收到通知,因为那时你可以联系那个人去推销。所以,任何你想在网络上保持更新的东西,你都可以设置一个 Scout,它会通知你。这里的关键是,它并没有承诺整个世界,对吧?它不是说“我来帮你预订”或“我来帮你购买产品”或“我来处理你所有的数字杂务”。它非常具体。它监控网络上任何你可能关心的信息。所以它的能力很窄,但领域非常通用,对吧?你可以监控网络上任何你感兴趣的东西。因此,我们的想法是,随着时间的推移,我们会为它增加更多能力。目前它只是监控。未来它可以告诉你“这个有货了,你想让我帮你买吗?”如果你说“是”,它就可以直接去做。所以随着时间的推移,它可以完成你试图完成的工作流程中越来越大的部分,同时在你应用它的领域上保持相当通用。这就是我们采取的方法。

If there's a certain product that you're looking to buy and it's out of stock or the price is too high and you want to be notified whenever it's in stock or the price is below threshold, you can set up a scout for that. If you're looking for internships or jobs and you want to be notified anytime there's a new listing of a certain characteristic. If you're looking for apartments, same thing. If you're doing business and competitive intelligence, you want to be notified anytime your competitor's product gets a bad review because now you can reach out to that person to sell. So anything that you are interested in wanting to stay up to date on on the web, you can set up a scout and it will notify you. So what's relevant here is that it is not promising the world, right? It's not saying that I'm going to make the reservation or I'm going to purchase the product or I'm going to take care of all your digital chores. It's very specific. It monitors the web for any information that you might care about. So it's narrow in the capability, but it's very general in the domain, right? You can be monitoring anything on the web that is of interest to you. And so the way we think of it is over time we will add more capabilities to this. Right now it just monitors. In the future it can let you know that this is available. Do you want me to buy it for you? And if you say yes then it can go ahead and do that. And so over time it can do a larger and larger chunk of the workflow that you're trying to get done, while all along being fairly general in the domains on which you are applying this. So that's the approach that we've taken.

技术挑战与架构 Technical challenges and architecture

Host

我希望你能更深入地谈谈你们构建 Scout 的方式,它的技术方法和架构,以及你们遇到的挑战。我想到的是,当你描述这个用例时,它听起来很容易——我们一直都有像 Google 快讯这样的通知——但我也尝试过用 AI 构建这类东西,真的很难。即使是简单的事情,比如我做过一个爬取某个汽车论坛帖子的小工具,光是 cron 作业之类的东西就让人头疼。所有事情都很麻烦。而且那是一个论坛,并不是有人故意阻止我这么做。但即便如此,事情还是会出错,LLM 会误解指令,不能很好地遵循指示。请谈谈其中的细微之处,以及为什么这对你来说是一个有趣的实验。

I'd love for you to dig a little bit more deeply into the way you've approached Scout, its technical approach and architecture, but also the challenges that you run into. What is coming to mind for me is that as you describe this use case, it both sounds easy — we've had notifications like Google Alerts forever — but also having tried to build that kind of thing with AI, it's really difficult. Even simple things like I did something to look at a particular forum for car postings, and just the cron jobs and stuff was a pain in the butt. Everything was a pain in the butt. And this was a forum, not something where someone was actively trying to prevent me from doing it. But even then things would break, the LLM would misinterpret things, it wouldn't follow instructions very well. Talk a little bit about the nuances and why this is an interesting experiment for you.

网络监控与浏览器使用细节 Nuances of web monitoring and browser use

Devi Parikh

是的。你描述的情况我们经常从用户那里听到。他们中的一些人曾尝试自己拼凑一些工具来实现这些功能,但要么是因为基础设施挑战,要么就像你说的,LLM 不会按你期望的方式工作。甚至像记住它之前告诉过你的内容,并确保它接下来告诉你的信息相对于之前是新的,这些都需要思考和编排。另一个很多人有时会忽略的点是,网络上的信息,即使是公开信息,也常常隐藏在轻量级表单后面,对吧?所以如果你在想当地网球场的预订,你会说“周一早上 7 点有空位时通知我”。但并不是所有日期和时间都有一个列表。你必须去选择日期,选择时间,然后才能看到是否有空位。所以即使对于这种监控能力,你通常也需要一个浏览器使用模型,一个能自动在网站上导航以找到你需要的信息的导航器。我们编排 Scout 的方式是,它可以访问大量的 API 和 MCP 服务器。所以只要信息以非常智能体友好的方式可用,它就会优先使用这些。但在那些信息需要你在网站上执行某些操作才能获取的“长尾”网站情况下,我们会启动一个远程浏览器,打开网页,并使用我们内部训练的导航模型,它会点击按钮、填写轻量级表单来获取你需要的信息。这就是信息获取方面的一个复杂性。另一个是覆盖范围很重要,对吧?当你说“有关于某某主题的新闻时通知我”,这些新闻可能出现在任何地方。可能是社交媒体、传统新闻文章、Reddit 上的讨论,网络上的任何地方。所以系统需要构建成能持久运行,首先优化覆盖范围,但当你生成面向用户的报告时,你又希望它高精度、总结得好、易于阅读和消化,并且能结合之前告诉过你的上下文来呈现。这些都是我们构建的架构的不同部分。这是一个非常通用的架构。当你谈到 API 访问、MCP 服务器、这个浏览器使用模型时,你可以看到它如何为 Scout 提供动力,但也可以想象随着时间的推移,这个架构还能做其他很多事情。所以对我们来说,重要的是以相当横向的方式构建底层技术,同时为用户带来特定场景的产品体验,在这些场景中我们可以设定预期,知道我们能可靠地交付,并随着时间的推移建立信任。

Yeah. And what you're describing is something that we've heard from a bunch of our users frequently. Some of them have tried to put together some tools on their own for these capabilities, and either because of infra challenges or like you said, the LLM won't do what you expect it to do. Even things like having the context of what it has already told you in the past and making sure that whatever it tells you next is actually new relative to it. All of this needs thought and orchestration. The other bit here that I think a lot of people sometimes miss is that information on the web, even publicly available information, is often behind lightweight forms, right? So if you're thinking about your local tennis court reservation and you're like, 'Let me know whenever there's an opening for Monday at 7:00 a.m.' It's not like there's a list somewhere for all dates and all times. You have to go pick the date, you have to go pick the time, and only then you can see if it's available or not. So even for this monitoring capability, you often need to have a browser use model, a navigator that's automatically navigating on the website to find the information that you need. The way we've orchestrated scouts is that it has access to a whole bunch of APIs, a whole bunch of MCP servers. So whenever information is available in a very agent-friendly way, that is what it would choose to use. But in situations which are the large heavy tail on the web of websites where information is sitting behind some actions that you need to take on the website, for that we will spin up a remote browser, we will open up the web page, and we use a navigator model that we've trained in-house that will click on buttons, fill out lightweight forms to get you the information you need. So that's one complexity just around access to information. The other is coverage is important, right? When you're saying 'let me know whenever there is news about blah topic,' that news could be anywhere. It could be on social media, traditional news articles, conversations on Reddit, anywhere on the web. So having the system built in a way that it's optimized for a lot of persistence, optimizing for coverage to start with, but then by the time you make it a user-facing report, you do want it high precision, well summarized, easy to read, easy to digest, presented in a way that has the context of what it has told you before. All of these are various pieces of the architecture that we built. And it's a very general architecture. When you talk about access to APIs, MCP servers, this browser use model, all of that you can see how it powers scouts, but then you can also imagine all the other things this architecture can do over time. So that's been important to us: to build the underlying tech in a fairly horizontal way, but bring product experiences to users for specific things where we can set expectations, we know we can deliver reliably, and build that trust over time.

MCP服务器示例 MCP servers examples

Host

你们依赖的 MCP 服务器有哪些?能举一些例子吗?

And what are some of the MCP servers or what are examples of the MCP servers that you're relying on?

Devi Parikh

这些工具涵盖各个领域。比如 Airbnb 的可用性查询,或者各种社交媒体如 LinkedIn 或 Twitter 上的公开信息——这些 Scout 不能绕过付费墙。还有用于常规网页搜索和天气查询的搜索 API。我们的技术栈可以访问大约 80 到 90 个这样的独立工具。而对于长尾的其余部分,我们就启动这些浏览器使用模型。

So it's tools that go across the board. Things like Airbnb availability, or various social media like LinkedIn or Twitter that is publicly available — these scouts can't go behind paywalls. Search APIs for just regular web search, for the weather. Yeah, there are 80 or 90 of these individual tools that our stack has access to. And then for the heavy tail, for the rest, is where we spin up these browser use models.

浏览器使用模型方法 Browser Use Model Approach

Host

你创建的浏览器使用模型,它们会脱离常规操作吗?

And do the browser use models that you've created, do they navigate off walls?

Devi Parikh

目前还没有,我们当前发布的产品中还没有这个功能。模型本身具备这种能力,随着时间的推移,我们会推出这个功能。所以你可以授权 Scout 监控你有权限访问的各种信息流,并授权它代表你登录。但这还不是产品当前的功能。

Not currently, not what we've shipped in the product right now. The model itself has the ability to do that. Over time that is something that we'll put out there. So you can authorize the scout to monitor various feeds that you have access to and authorize it to login on your behalf. But that's not something that's in the product right now.

Host

谈谈你们为这个功能设计模型的方法。我和 Dan Jeff 等人聊过,他提到一些看似简单的事情其实非常复杂,比如让模型在 Google Flights 表单中从弹出窗口选择正确的日期就非常困难。也许你们通过避开整个旅行预订用例来避免了这个问题,但也许没有。我猜如果你们做的是网球相关的事,可能仍然需要这个功能。不过,还是谈谈你们遇到的那些问题吧。

Talk a little bit about the way you've approached the model for this. I've had some conversations with folks like Dan Jeff, where he talks about the complexity of really simple things that you think would be easy, like going to a Google Flights form and trying to get the model to choose the right date from the popup being really difficult. Maybe you have avoided that by sidestepping the whole travel booking use case. But maybe not. I guess if you're doing tennis, you probably still need that perhaps. But yeah, talk about that kind of thing that you've run into.

Devi Parikh

是的,我认为任何尝试训练这些导航器(即浏览器使用模型)的人,都会因为日期选择器的难度而产生共鸣。这是一个反复出现的例子。在我们的案例中,我们可以使用 Scout 来监控航班的可用性或价格等。我们选择训练模型的方式是依赖视觉信息。我们截取你正在查看的网页截图,然后训练模型预测它接下来可以执行的操作,无论是点击、输入、滚动、点击单选按钮,还是点击日期选择器中的适当位置等。所以它以一种非常视觉化的方式来处理,这让我们能够应对日期选择器和其他问题,而如果使用底层 DOM 信息来构建这些智能体,这些问题会困难得多。很多人确实采用那种方法,但我们认为它无法很好地泛化。我们刚开始训练这些模型时,这有点反直觉。我们的假设是,网页是为人类消费而渲染的;机器没有理由需要通过视觉模态来感知网页。网页只是底层 DOM 信息的渲染结果,所以我们应该直接使用 DOM 信息。为什么要增加这一层呢?但随着时间的推移,我们意识到这样做很难可靠地实现。所有不同的网页构建方式差异很大。两个网页在视觉上可能非常相似,但底层 DOM 信息却截然不同。所以随着时间的推移,我们意识到像人类一样通过查看视觉截图来消费这些网页,要可靠得多,也通用得多,我们可以通过积累越来越多的这类数据来扩展,让这些模型越来越好。

Yeah, so I think anyone who's tried to train these navigators, these browser use models, all sort of bonded over how hard date pickers are. So that is a recurring example that comes up. In our case, yeah, we can use a scout to monitor the availability or price of a flight and things like that. The way we've chosen to train the model is by relying on visual information. So we take a screenshot of the web page that you're looking at and the model is trained to predict the next action that it could take, whether it's clicking, typing, scrolling, clicking a radio button, clicking on the appropriate place in the date picker, that kind of thing. So it approaches it in a very visually grounded way, which is what lets us deal with date pickers and other things that would be much harder to do if you used the underlying DOM information to try to build these agents. A lot of people do approach it that way, and we don't think that's going to generalize well. This was a little counterintuitive to us when we first started working on training these models. Our hypothesis was that the web page is rendered for human consumption; there's no reason a machine needs to be perceiving the web page through that visual modality. The web page is just a rendering of the underlying DOM information, so we should just be using that directly. Why go through this added layer? But over time we realized that's very, very hard to do reliably. All different web pages are built very differently. Two web pages can visually look very similar but have very different underlying DOM information. So over time we realized that consuming these web pages in the same way that humans do, by just looking at visual screenshots, is way more reliable and way more general, and we can just scale by having more and more data of that kind and make these models better and better.

后训练策略 Post-training Strategy

Host

你们是在后训练一个通用的 VLM 用于浏览器使用,还是针对你们的特定用例后训练一个浏览器使用模型?

And are you post-training a generic VLM for browser use or are you post-training a browser use model for your specific use case?

Devi Parikh

我们是在后训练一个通用的 VLM 用于浏览器使用。

We are post-training a generic VLM for browser use.

Host

有没有现成的开源通用浏览器使用模型?我们是否知道如何交付一个后训练的、普遍有用的浏览器使用智能体,然后再进行后训练?问题的一部分是:如果它现在不存在,似乎很快就会存在,就像你们正在做的事情很可能在某个时候被商品化。你怎么看?

Are there off-the-shelf open-source generic browser use models? Do we know how to deliver a post-trained browser use agent that is generally useful and then post-train it? Part of the question is: it seems like if it doesn't exist, it would exist sometime soon, like it would be a thing that you're doing that is likely to be commoditized at some point. How do you think about that?

Devi Parikh

我确实预料到了这一点,而且我相信随着时间的推移,这些 VLM 也在各种数据上进行了训练,包括计算机使用截图和屏幕截图,以及其他视觉信息。所以我认为在计算机使用和浏览器使用方面的接地能力一直在提升,我预计这种情况会持续下去。我认为我们的后训练仍有持续的价值,原因有三。第一,我们可以针对产品中出现的用例进行定制,从而提升可靠性。第二,随着底层基础模型的开箱能力越来越强,它们能够处理越来越长的工作流,但如果我们进行后训练,就能在任何给定时间点推高我们所能达到的上限。我们可以追求比现有能力更复杂的工作流。想想我们在网络上做的事情,工作流的复杂性和长度是无限的。第三是成本原因。我们能够用自己的浏览器使用模型(即我们自己的导航器)来服务产品。这比将所有请求路由到现有的 API 提供商更能控制成本。

I do anticipate that, and I'm sure over time these VLMs have also been trained on various kinds of data including computer use screenshots and screenshots in addition to other visual information. So I do think that grounding capabilities in the context of computer use and browser use has been getting better over time, and I anticipate that will keep happening. I do think there is continued value to our post-training for, I would say, three reasons. One is we can cater them for the kinds of use cases that are showing up for our product, so we can push on reliability there. The second is as the capabilities of the underlying foundation models get better out of the box, they'll be able to do longer and longer workflows, but then if we post-train, we'll be able to push the ceiling of what we can do at any given point in time. We can just go after even more complex workflows than what's possible. And if you think about the kinds of things we do on the web, the sky is the limit of how complex and how long these workflows can get. The third is cost reasons. We are able to serve our product on our own browser use model, our own navigator. That keeps our costs under check way more than if we were routing everything to one of these existing API providers.

基础模型选择 Base Model Choice

Host

你们是不是用了像 Qwen 这样的基础模型?

And are you using something like a Qwen as the base model?

Devi Parikh

我们用的是 Qwen 模型。目前我们会尝试任何新出现的模型。随着时间的推移,我们可能会换掉它,但现在我们用的是 Qwen。

We are using a Qwen model. Right now we experiment with any new models that come out. Over time we might swap that out, but right now we're using Qwen.

系统演进方法 System Evolution Approach

Host

当我思考如何改进这样的系统时,除了与用户交流、了解他们想做什么以及试图解决的问题之外,似乎有一个相当简单的循环可以遵循:按他们希望这个工具使用的领域排序,然后从列表顶部到底部,让产品在这些领域表现更好。这是你们思考系统演进的一部分吗?

When I think about how to evolve a system like this and make it better, beyond talking to your users and understanding what they're trying to do and the problems they're trying to solve, it seems like there's a pretty easy loop to follow: sort by domain that they're trying to have this thing use and just go from the top of the list to the bottom and make your product work better in those domains. Is that part of the way you're thinking about evolving the system?

Devi Parikh

我们训练的模型相当通用。它并没有根据我们看到的用法过度拟合到任何特定领域。它是在这些监控和信息搜索任务上训练的。所以任务的性质影响了它在后训练期间看到的数据类型。使用情况确实会产生一些影响,但我们不会采取“这是我们关心的前 10 个网站;确保模型在这些网站上表现良好,然后不断扩展”的方法。我们确实在整个栈上有一整套评估,从单个步骤级别开始,根据导航器当前的位置,评估它决定下一步采取的行动。

The model that we've trained is fairly general. It hasn't been overfit to any specific domains based on where we've seen usage. It is trained on these monitoring and information seeking tasks. So the nature of the task is playing a role in the kind of data that it has seen during post-training. There is a consequence of the usage that shows up there, but we don't approach it as 'here are the top 10 websites that we care about; let's make sure the model works well there and then keep expanding.' We do have a whole bunch of evals across the stack, all the way from the individual step level that given where the navigator is right now, this is the action that it decided to take next.

评估与用户反馈 Evaluation and User Feedback

Devi Parikh

这样做是否合理?所以从单步预测准确度到轨迹层面——也就是它沿途做的所有事情——它是否完成了正确的任务?一直到我们发送给用户的报告层面,这不仅仅关乎这个导航器。它还使用了很多其他工具。所以考虑的因素包括:信息是否与用户寻找的内容相关?我们提供的报告中是否有足够的引用,让用户可以点击了解更多?这些引用是否足够具体?是否有链接失效(这会是糟糕的结果)?它是否重复了之前已经说过的话?所以我们考虑了所有这些更面向用户的因素。我们有评估——其中很多是自动评估,但我们也对很多项目进行人工评估。这就是我们用来跟踪进展、确保质量达标的方法。此外,我们还从用户那里获得大量反馈。目前产品通过候补名单提供。我们非常谨慎地分批邀请用户,以便随时间收集反馈。根据用户反馈,我们在产品中构建了许多功能。顺便说一句,如果有听众想要使用该产品,可以在 utori.com 上注册候补名单。如果他们注明是从“tool”听说的,我们很乐意优先处理他们的访问请求。

Was that a reasonable thing to do or not? So literally one-step prediction accuracy up to sort of trajectory level — these are all the things that it did along the way. Did it get the right thing done? All the way to the level of the report that we sent to the user, which isn't just about this navigator. It also uses a whole bunch of other tools. So things like: was the information relevant to what they were looking for? Were there enough citations in the report that we provided so that users can click on those to find out more? Were those citations specific enough? Were any of the links broken, which would be a bad outcome? Is it repeating anything that it had already said before? So all of these more user-facing factors that we considered. We have eval — a whole bunch of these are automatic eval, but we also have human eval for a lot of these things. So that is what we track to make progress and ensure quality is up to the mark. Then there's a bunch of feedback that we get from our users. Right now the product is available behind a waitlist. We've been quite careful with letting in cohorts of users so we can get feedback over time. There's a bunch of stuff that we've built out in the product in reaction to that user feedback. And by the way, if any of the listeners do want access to the product, they can sign up on the waitlist at utori.com. If they mention 'tool' for where they heard about it, happy to prioritize their access.

Scout作为起点,而非报告生成器 Scouts as Starting Point, Not Just Report Generator

Host

你提到这个过程的最终结果是一份报告。这感觉是一种非常静态且死胡同的方式来结束这次交互。而我可能想把这个东西作为我启动的其他过程的起点,或者进一步处理这些信息。例如,如果我在总结或获取并总结竞争费用,我可能想把它从你的平台拿出来,放到我创建的包含其他内容的更广泛报告中,或者放到 Slack 上,或者做点什么。你怎么看待 scouts 作为其他过程的起点,而不是一个报告生成器?

One thing that you mentioned was that the end result of this process is a report. That feels like a very static and kind of dead-end way to terminate this interaction. Whereas I might want to use this thing as the beginning of some other process that I kick off, or to further work with this information. For example, if I'm summarizing or fetching and summarizing competitive fees, I might want to take that off your platform and put it into some broader report that I create that includes other things, or put it on Slack, or do something. How do you think about the scouts as the starting place for other processes, as opposed to a report generator?

Devi Parikh

非常同意。我们确实把 scouts 视为起点。它是一个只读操作——目前它为你监控信息。它不会改变世界的状态。但意图是随着时间的推移,它会为你做越来越多的事情。例如,产品现在可用,它会通知你。但现在需要你自己去购买,点击链接并完成操作。它试图让事情变得简单——它会直接给你一个你所要求产品的链接,这样你就可以轻松点击购买。但随着时间的推移,它应该变成这样:'我找到了。你想让我帮你买吗?'你说好,它就去做。第二点是我们为感兴趣的人提供了 scouting API 和 webhooks。所以有些用户会从 scouts 获取报告,并在其之上构建其他工作流——自定义工作流。这也是一个选项。第三点是,对于某些事情,它有点一次性。比如你监控产品一段时间,一旦找到并购买,就不再需要 scout 了。这时可以引导用户进行他们可能感兴趣的下一步。但还有其他事情,比如'让我知道各种技术主题或 AI 新闻的更新'——这是一个持续监控的事情,没有自然的终点。另一个点是,我们最近推出了一个功能,你可以回复 scout 报告并给出反馈,告诉它将来应该如何以不同方式执行——你希望报告如何改变。这使得你和 scout 之间的互动更多,并且随着时间的推移,它会为你提供越来越高的信号。但很多用户问我们:'我能和它聊天吗?它为我找到了所有这些酷信息,我只想了解更多。'所以随着时间的推移,我们可能会优先考虑这一点。

Very much so. We very much think of scouts as that starting point. It's a read-only action — it's monitoring information for you right now. It doesn't change the state of the world. But the intent very much is that over time it will do more and more for you. For example, the product is now available and it lets you know of that. But now it's on you to go purchase it, click on the link, and do that. It tries to make it easy — it will give you a link very directly to the product that you had asked for so that you can click and buy it easily. But over time it should just be like: 'I found it. Do you want me to buy it for you?' You say yes and it goes ahead and does it. The second bit is that we have scouting APIs and webhooks available for people who are interested. So there are some users who take these reports from scouts and build other workflows on top of it — custom workflows. That's an option as well. The third thing is that for some of these things, it is a bit of a one-off. Like you were monitoring the product for a while and then once you found it and bought it, you no longer have use for the scout. There it could be relevant to walk the user through the next thing they may be interested in doing. But there are other things where you're like 'let me know of any updates on various technical topics or AI news' — that's just an ongoing monitoring thing over time. There is no natural endpoint to that. Another bit is that we recently shipped this feature where you can respond to the scout report and give it feedback in terms of how it should do this differently going forward — in what way you want the report to be different. That makes it a bit more of an interaction between you and your scout, and it's getting higher and higher signal for you over time. But a lot of users have asked us: 'Can I just chat with it? It found me all this cool information and I just want to learn more about it.' So that's something that over time we might prioritize.

反馈作为代理循环的一部分 Feedback as Part of Agent Loop

Host

所以他们给出的反馈会进入模型的上下文,并影响下一次的输出——有点像系统指令或个性化指令。

And so the feedback that they're giving goes into the context of the model and shapes the output the next time — kind of like a system instruction or a personalization instruction.

Devi Parikh

完全正确。他们给出的反馈会被共享——成为下次运行时智能体循环指令的一部分。所以它更有可能为你提供更高的信号。

Exactly. The feedback they give is shared — it's part of the agent loop instruction the next time it runs. So there's a higher chance of it being higher signal for you.

浏览器使用的领域特异性vs泛化 Domain Specificity vs. Generalization in Browser Use

Host

我想回到领域特定性与苦涩教训的问题——只是收集大量数据并训练这个通用化的东西。我经常想,我有时不知道普通人如何使用网络。显然他们会用,但例如,有很多奇怪的事情,比如如果你有多个 Google 账户,你无法真正使用它们,除非你知道可以在 URL 中把 U1 改成 U2 或 Uzero 改成 U1。这种情况比两三年前少了,但仍然不罕见——我试图在网站上解决问题或完成一些基本操作时,会打开开发者工具并摆弄他们的 CSS。你们是否在训练浏览器使用智能体普遍地做那些奇怪的 hacky 事情?你们是否在训练它们做这些事情?你们是否有针对特定重要网站的启发式方法?还是你们目前认为这无法实现?

I want to return a little bit back to the domain-specificness versus the bitter lesson — just collect a bunch of data and train this generalized thing. I often think about how I sometimes don't know how normal people use the web. Obviously they do, but for example, there are all these weird things like if you have multiple Google accounts, you can't really use them unless you know that you can go into the URL and change the U1 to U2 or Uzero to U1. It's less frequent than it was two or three years ago, but it's still not infrequent that I'm trying to solve a problem or get something basic done on a website and I'm opening up dev tools and fiddling with their CSS. Are you training the browser use agent to do those kinds of weird hacky things generally? Are you training them to do those things? Do you have heuristics that you use for specific important sites where you need that kind of thing? Or are you just carving that out as not attainable at this point in time?

Devi Parikh

是的。我们训练这些模型的方式是从监督微调开始的——我们让人去访问网页、点击、完成各种任务,然后用这些作为训练数据来训练我们的模型。

Yeah. So the way we've trained these models is we started with supervised fine-tuning, where we got people to go to web pages, click around, get various tasks done, and then use that as training data to train our model.

数据扩展与平台期 Data Scaling and Plateauing

Devi Parikh

在一段时间内,数据越多,尤其是确保质量高的情况下,准确率会持续上升,但最终会开始趋于平稳,对吧?然后我们就开始使用一种叫做拒绝采样的方法,让智能体去尝试完成一个轨迹。我们有自动化的方法来判断这个任务是否完成得好、是否正确。如果正确,就把这些轨迹加回训练数据,继续推进。这最终也会开始趋于平稳,然后自然过渡到强化学习。这就是我们目前模型所处的阶段。是的,我们做了监督微调,做了拒绝采样,现在正在进行强化学习。但正是在这个阶段,我认为模型有可能发现一些取巧的做法,无论人们是否那样做。嗯,所以这可能会涌现出来。我们目前还没有看到明显的迹象。但这是可能发生的地方。不过,在我们还在做监督微调的时候,除非我们的标注员知道这些网站上各种取巧的做法,否则模型是不会学到这些的。

And there for a little while, the more data you have, especially if you're making sure quality is high, sort of accuracies keep going up, but then you eventually start plateauing, right? And then that's where we started using something called rejection sampling where the agent goes out and sort of tries to complete a trajectory. We have automatic ways of telling whether this was done well or not, whether this was done correctly or not. And then if it was done correctly, use those trajectories to sort of add it back to the training data and sort of keep going from there. That also eventually starts to plateau and sort of becomes a natural transition into reinforcement learning. And that is where we are at like our models currently. Yeah, we've done supervised tuning, we've done rejection sampling and we're doing reinforcement learning right now as we speak. But that's where I think there is potential for the model discovering certain hacky things whether or not people do it that way. And yeah, so that might emerge. We haven't seen obvious signs of that so far. But this is where there is potential for that to happen. But while we were still doing supervised finetuning, unless our annotators knew of these hacky ways of doing various things on these websites, the model wouldn't have picked up on that.

Host

想必,你们还没有发现有必要针对特定网站做很多定制化的工作来取得进展。但模型能够仅仅依靠这种通用知识来完成你们希望它们做的大部分事情。

Presumably, you haven't found it necessary to do very site-specific things to make progress. Like you're, but the models are able to rely on solely this general knowledge to do most of the things you want them to do.

Devi Parikh

是的。到目前为止,我们从执行层面看到的情况就是这样。当我们做这些探索性实验,比如强化学习,刚开始时为了控制范围,我们可能会从一个网站开始,确保它确实能在这个网站上工作,然后再扩展到其他网站。所以我们也曾在更窄的领域做过实验。但生产环境中的模型是一个在多个不同网站上以相当通用的方式训练出来的模型。我确实认为,我之前提到过,这一点非常重要:我们是在网页截图上进行训练。我们选择使用视觉模态而不是底层的 DOM,这使我们能够达到现在的水平。早期我们尝试使用 DOM 并编写解析器来提取元素信息时,那让我们走上了非常特定于网站的道路,因为你需要为不同的网站以不同的方式解析信息。那时我们意识到这种方法并不可靠。即使在一个网站上,也经常会出现各种边缘情况,而当你尝试转向另一个网站时,几乎得从头开始。所以我认为,我们转向使用视觉的转变非常关键,它让我们能够规模化,并拥有一个跨网站的通用模型。

Yeah. So far that is what we've been seeing just from a sort of execution and like when we're doing these experiments on these exploratory things like reinforcement learning and when we were first getting started there just to keep it scoped we might start with one website and make sure that this is actually working on one website before we expand to others. So for that kind of a thing we've done experiments in narrower domains as well. But the model that's in production is one model that's been trained across many different websites in a fairly general way. And I do think I mentioned this earlier but I do think it's quite important like the fact that we are training on screenshots of web pages. The fact that we've chosen to go with the visual modality and not the underlying DOM is what has enabled us to get to this point. Earlier in our journey when we were experimenting with DOM and trying to write parsers for it to extract element information that is what was taking us down a very specific website route because you'll need to parse that information in different ways for different websites and that's when we realize that one is just not reliable. There are constantly edge cases that keep coming up even from one website and then when you try and go to a different website you kind of have to just start from scratch. So I think that shift that we made to using vision was quite significant in terms of just scale and have one general model across websites.

结构化RL调优 Structuring RL Tuning

Host

你能谈谈你们是如何为强化学习调优构建问题的吗?

Can you talk a little bit about how you structure the problem for RL tuning?

Devi Parikh

这和我之前说的类似,我们至少可以为某些任务设置自动奖励。例如,某件商品是否有货、是否可以预订等等。如果你执行了正确的操作,最终进入了一个显示可用性信息的页面,那么你只需查看那个页面就能知道模型是否正确。也就是说,如果你能把查询描述为可用性搜索,你就能从结果屏幕判断模型是否进入了选项展示页面,诸如此类。

So it's similar to what I was saying that we have these automatic rewards that we can set up at least for certain tasks. So for example like whether something is available for sale or for reservation or whatever the case may be. If you've done the appropriate actions and you find yourself on the page that is now showing that availability information you can look at just that page and know whether the model got it right or not. Meaning if you can characterize the query as an availability search, you can tell from the resulting screen whether they got to a presentation of options, that kind of thing.

Host

是否达到了。是的。

Whether they got. Yeah.

Devi Parikh

所以有些任务,你可以只看最终状态,就能自动判断模型的响应是否正确。你不需要查看整个轨迹来评估。因此,这类任务很适合设置这些奖励,然后用来训练模型。

So there are some tasks where you can look at the final state of it and automatically decide whether the response that the model gave you is correct or not. You don't need to look at the whole trajectory along the way to be able to assess that. And so those kinds of things, for example, set themselves up well for having these rewards that you can then use to train the model for those.

超越Scout:未来方向 Beyond Scout: Future Directions

Host

你对 Scout 之后的发展有什么想法吗?你们认为这些是串行实验还是并行实验?

Do you have a sense for what's beyond Scout? Like you think of these as serial experiments or parallel experiments?

Devi Parikh

我们目前是一个相对较小的团队,这限制了我们可以并行做的事情。我们确实认为 Scout 是有价值的。最初推出时,我们并不确定。我们都觉得这很有道理,也很有价值。但推出时,不清楚人们是否会理解——这不是 ChatGPT,对吧?这不是你可以实时对话的东西。也不是 Google 搜索。它本质上是监控未来将要发生的信息。所以现在当我描述它时,我觉得大家都能理解。但当时我们不确定这是否能吸引用户,是否在他们的生活中足够频繁地出现,是否足够有价值。不过,从我们看到的用户、留存率和付费转化来看,我们确实觉得找对了方向。但我们不想陷入局部最优,觉得这个很好就只在它附近优化。在越来越接近帮人们处理数字杂务方面,还有很多事情要做。所以,未来会有一系列新功能陆续推出,包括能够监控需要实际完成网页任务才能获取的信息,以及与其他工具和你的上下文进行更多集成。

We are a relatively small team so far and so it restricts how many things we can do in parallel. We do think we're on to something with Scout. It wasn't obvious to us when we first put it out there. Like we all thought this makes sense and it's valuable. But when we put it out there it was unclear if people would get like this is not ChatGPT, right? This is not something that you go talk to and it's instantly real time talking back and forth. It's not Google search. And so it's this you're essentially monitoring for information that's going to happen in the future. And so like now when I describe it, I think it makes sense and everyone gets it. But at the time we weren't really sure if this is a thing that will land for users. And whether this is something that shows up in their lives frequently enough, is this valuable enough? And so yeah, but I do think with the users that we've seen, the retention that we've seen, conversions to the price plan that we've seen with all of that, it does feel like we are on to something here. But we don't want to sort of fall for that thing of like that local minima where this is good and let's not just keep optimizing in its local neighborhood. There's so much more to do in terms of getting closer and closer to like taking care of people's digital chores. And so yeah, there's a bunch of new capabilities that are coming at different points in time, including being able to monitor information that's sort of behind Owall's actual completion of these tasks on the web for you and sort of integration with other tools and your context more and more of your context over time.

个人用例 Personal Use Cases

Host

你个人使用这个工具有哪些方式?

What are some of the ways that you use the tool personally?

Devi Parikh

我用它做各种事情。比如一些很具体的事,比如当某位 AI 研究员宣布离职但还没公布下一站时,这是一种很好的方式,可以跟踪生态系统中人员流动的情况。

I use it for all kinds of things. So things like even sort of narrow things like whenever an AI researcher announces that they're leaving but they haven't announced where they're going next is sort of a nice way of just keeping up with what's happening in the ecosystem with people moving around.

Host

这真是个有趣的用例。就像一个活跃的书签,帮我盯着这个,当有新动态时通知我,而不是持续发生并给我所有新内容的摘要。

That's kind of an interesting use case. It's like an active bookmark, like keep a tab on this for me and let me know when something new happens around it, as opposed to this is a thing that's continually happening and give me summaries of all the new things.

Devi Parikh

是的。

Yeah.

Scout的用例 Use Cases for Scout

Host

这就像我刚才说的,当你想到 Google 快讯时,它非常基于关键词,比如你可以为某人的名字设置提醒。但如果是这种情况:每当一位 AI 研究员发帖说他们离开了当前工作,但没说下一步去哪,你怎么为这种事设置 Google 快讯呢?

And this is like what I just said is like when you think about Google alerts, right? It's very keyword based and so you could set up an alert for somebody's name for example. But if it's this kind of a thing that whenever an AI researcher changes like posts that they've left their current job and haven't said when they're going next, how would you set up a Google alert for something like that?

Devi Parikh

是的。我经常为一些时事设置提醒,比如几个月前我家乡发生了一起印度航空坠机事件,那几周我只想了解所有相关的最新动态,于是我为它设置了一个 Scout。我还设置了一个用于追踪旧金山引发热议的最新观点,它每天午饭后给我发一条通知。诸如此类。我最近迷上了风干粘土,以及可以用它做的各种手工,但我不打算花时间去工作室。就像陶瓷和陶艺,但用的是风干粘土。你不需要去工作室,在家就能做。所以我一直在寻找快速完成的手工点子,不需要太多准备,也不会花太多时间。于是我设置了一个 Scout,专门给我推送 Reddit 和其他地方的有趣项目创意。是的,很多人用它来找公寓,觉得非常有用。我们有招聘人员用它来寻找潜在候选人,销售团队也用它来挖掘潜在客户。所以我们可以看到它在个人用途上非常广泛,在工作场景中也有相当深入的应用。

Yeah. I often set some up for certain current events like there was this Air India crash that had happened in my hometown some number of months ago and so in those weeks I just wanted to know of all updates that are happening around that and so I set up a scout for that. I have one set up for just sort of the latest hot take that's causing a lot of debate in San Francisco. And it sends me a note every day after lunch. Things of that nature. I've recently gotten into air dry clay and like various projects that I can do with it but I'm not going to have the time to go to a studio. So like ceramics and pottery but with air dry clay. You don't need to go to a studio. It's just things that you can make at home. And so I'm always looking for ideas for quick things that I can do. Not a lot of overhead. It's not going to take me a lot of time. And so I have a scout set up for just sending me posts from Reddit and other places where there are neat project ideas for that. Yeah, a bunch of people have used it for apartment searching and they found that to be really useful. We have recruiters who've used it for lead generation. We have sales teams using it for lead generation. So we kind of see a pretty broad spectrum of usage across personal use cases and fairly deep usage even in the context of work.

摄取架构与工具 Ingestion Architecture and Tools

Host

你们是如何搭建数据摄取架构的?是爬取 Reddit,还是用它们的 API,或者 X(推特),你们有 X 信息流吗?这可能是领域问题的另一面:你们是预先配置好最热门的 X 个来源,然后让用户根据需求访问,还是让智能体自行判断哪些内容与特定用户查询相关,然后去查找,比如针对不同用例多次搜索 Google?

How have you approached setting up an ingestion architecture? Are you like crawling Reddit or using their APIs, or X, do you have an X feed? Are you doing that? And this is maybe kind of the other side of the domains question, but do you have the top X sources prefigured to come in and then you have that accessible for what different users are trying to do, or does the agent need to figure out what is relevant for a particular user's query and then go find that, and you're maybe searching Google 50 different times for different use cases or something?

Devi Parikh

是的。我们有 80 到 90 个工具可供编排层使用,智能体如果认为与查询相关就可以调用。所以我们有一个预定义的工具词汇表。但根据查询,它会决定使用哪些工具、使用多少次、哪些可以并行使用等等。在构建过程中,我们有时会给它一些工具,然后看到某些查询时,会想:等等,如果我们有这个工具,事情会简单得多。于是我们不断调整工具列表。是的。可能不太直观的是,当工具数量达到 80、90、100 个时,直接让编排层访问所有工具并不可靠。

Yeah. So we have 80 to 90 tools that are made available to this orchestration to the agents that they can use if they think it's relevant to the query. So we have this predefined vocabulary of tools that it has access to. But then based on the query, it's going to decide which one of these makes sense to use, how many times to use it, which ones can be used in parallel all at once, and things of that nature. And so there is a little bit of this flavor when we were building out that maybe we had given it access to some tools and then we saw certain queries and we were like wait if we had this other tool that would make it much easier. Then we sort of just kept adapting that tool list over time. Yeah. And what's maybe not so intuitive to people is that when you start getting to 80, 90, 100 tools, it's not very reliable to just give the orchestrator access to all of these tools at once.

Host

上下文会爆炸,它不知道该怎么办。

Context blows up and it doesn't know what to do.

Devi Parikh

完全正确。所以我们不得不以更层次化的方式构建,其中有一些子智能体,这些子智能体可以访问特定的工具。为了在工具数量规模下保证可靠性,我们做了这些工作。

Exactly. Exactly. And so we've had to build it out more hierarchically where there are certain sub agents and those sub agents have access to certain tools and so there are things that we've needed to do to make that reliable at the scale of number of tools.

多代理协作 Multi-Agent Collaboration

Host

你能再深入谈谈这方面吗?你之前提到了多智能体类型的协作和交互。这是你们使用多个智能体的主要场景,还是有其他方式?

Can you dig into that aspect of it a little bit more? You spoke earlier about multi-agent types of collaboration and interactions. Is that the primary place that you're using multiple agents, or are there other ways that you're using that?

Devi Parikh

它在每一步都发挥作用。所以你可以把 Scout 理解为带 cron 任务的智能体搜索,并且能访问它过去已经告诉你的历史记录。一种围绕智能体搜索的上下文 cron 任务。那么,就智能体搜索部分而言,它从一个待办事项列表和一个计划开始,计划如何响应用户的查询。它会根据这个计划并行调用多个工具。然后根据这些工具返回的结果,决定下一步该做什么。所以这个待办事项列表会根据第一步发现的结果进行调整。对吧?因此每一步都是多个智能体出去,带回它们获取的信息,然后编排层根据这些信息决定下一步做什么。所以没有两个查询会执行相同的工作流。而且每一步都在主动与网络交互,而不是仅仅使用网络索引,因为这是面向未来的。它们寻找新信息,包括像网球场预订这类网站的长尾内容。所以,这就是它的样子。

It's at every step along the way. So you can basically think of scouts as agentic search with a cron job around it with access to this history of what it has already told you in the past so far. A sort of contextual cron job around agentic search. And so if I think of the agentic search piece of it, it starts off with a bit of a to-do list and a plan of what it's going to do to address the query that has come in from the user. It will send out multiple tools in parallel based on this plan that it has come up with. And then based on what it gets back from these tools, it's going to decide what it should do next. So that to-do list is adapting based on the outcomes of what it found after that first step. Right? So each one of these steps is multiple agents going out coming back with the information that they have and based on that the orchestrator deciding what it's going to do next. And so no two queries have the same workflow that is being executed on. And each of these steps is actively engaging with the web as opposed to having an index of the web and using just that, because this is future facing. They're looking for new information including in sort of heavy tails of websites like the tennis court reservation or so on. And so yeah, that's what it looks like.

跨运行的冗余与学习 Redundancy and Learning Across Runs

Host

关于这些智能体和 cron 任务的后台运行,在让一切正常运转的过程中,有什么有趣的发现或经验吗?

In terms of the kind of the backgrounding of these agents and cron jobs, any interesting learnings or experiences in getting all that working?

Devi Parikh

我认为这方面有趣的是,如果你仔细想想,对于每个 Scout 查询,你都在反复执行相同的事情,但网络上的信息可能在上次执行和这次执行之间发生了变化,对吧?所以存在大量你可以利用的冗余。而我们目前只是刚刚触及表面。另一个相关的点是……

I think what's interesting in that aspect is that if you think about it, for each scout query, you're executing on that same thing over and over again, but the information on the web may have changed between the last time you did it and you did it this time, right? And so there is a good amount of redundancy that you could be taking advantage of. And we've only sort of scratched the surface of that so far. And the other relevant bit is what...

Host

冗余是指多个用户中,一个用户请求的内容对另一个用户也有用?

Redundancy in a set of multi one users what one user has requested being useful for another user?

Devi Parikh

冗余是指,即使对于同一个用户的同一个 Scout,每次执行搜索时,由于 cron 任务的特性,它会在该任务中反复执行相同的查询。所以对于同一个用户的同一个 Scout,在这些 Scout 运行之间,在我们作为 cron 任务一部分随时间进行的这些智能体搜索之间,存在冗余。因此,思考如何利用这一点很有趣。

Redundancy in the sense that even for the same user for a particular scout, every time it does the search, right? The cron job nature of it, within that cron job it's executing on the same query over and over again. So for the same user for the same scout, there is a redundancy across these scout runs across these agentic searches that we're doing over time as part of the cron job. And so it's interesting to think about how to exploit that.

代理工作流中平衡冗余与动态性 Balancing redundancy and dynamism in agentic workflows

Devi Parikh

你不想走向一个极端,即第一次运行时就把流程编码成确定性的工作流,然后一遍又一遍地执行,因为网页上的信息可能已经变了,对吧?所以同样的工作流可能不再有效。因此,你不想走到那个极端去利用冗余。但与此同时,完全忽略这里存在冗余似乎也不对。所以找到那个最佳平衡点很有意思。我认为我们只是刚刚触及了表面。与此相关的另一点是,有些信息变化很快,有些则变化很慢,对吧?或者有些信息只在白天出现。比如,如果你在关注 AI 发布,它不太可能在美国的夜间发生,对吧?更可能是在早上,也许是美国太平洋时间的早晨。所以,基于这些先验知识来判断何时值得再次检查,是另一个相关的事情。但同样,我们目前还没有在这方面做太多探索。

You don't want to go all the way to one extreme where the first time you run it, you encode that into a deterministic workflow and then you just execute that over and over again because the information on the web may have changed, right? And so that same workflow may no longer work, right? So you don't want to go to that extreme end of trying to take advantage of the redundancy. But at the same time, completely ignoring the fact that there is redundancy here also doesn't seem right. And so figuring out that sweet spot is interesting. And I think we've only scratched the surface of doing that. And the other bit that is relevant to this is some information is fast changing and some information is very slow changing, right? Or some information only tends to happen during the day. Like if you're looking for AI releases, it's unlikely to be happening at night time in the US, right? It's more likely to happen in the morning, maybe morning Pacific time in the US. And so having these priors of when is it worth checking again based on certain priors of when this kind of information is likely to change is another thing that would be relevant. But again, we haven't exploited that a whole lot so far.

Host

我们快到年底了。你对这个领域明年有什么预测?你认为它会如何发展?

We're coming close to the end of the year. Like what are your predictions for next year in this space? How do you think it will evolve?

Devi Parikh

我不知道。我总体上对预测持怀疑态度。是的,我想说的是,这不会很有见地,因为这个领域有很多事情在发生。所以,我很好奇它会如何随着浏览器中的 AI 功能而演变。我很想看看人们的反应。比如,我用过 Comet。我现在正在使用 Chad Gibby 的 Atlas,但我发现自己或多或少还是把它当作普通浏览器来用。我还没有找到让 AI 功能成为我工作流程中不可或缺部分的方法。我在与其他用户交流时也有这种感觉。我很好奇你是否在使用这些 AI 浏览器,以及你的体验如何。

I don't know. I'm a little skeptical of predictions in general. Yeah, I think what I'll say is it's not going to be very insightful like there is a lot happening in this space. And so yeah, I'm curious to see how that evolves with the AI features in browsers. I'm interested to see what the reaction to that is. Like I've used Comet. I'm using Chad Gibby's Atlas right now, but I find myself more or less using it as a regular browser. I haven't found ways for the AI features to be a very integral part of my workflows yet. That is a feeling that I get in talking to other users as well. I'm actually curious if you use any of these AI browsers right now and what your experience with them has been.

Host

我玩过它们,主要是想看看它们能做什么,但我也还没有找到所谓的杀手级应用。有一件我想做的事情,有点奇怪,但我在很多不同的平台上保存文章。X、LinkedIn、Hacker News 是主要的几个。访问 LinkedIn 的保存内容非常困难,因为我总是得去找那个东西在哪里。而且我没有找到任何可用的 API。所以有人提到用 Comet 来处理 LinkedIn。我觉得如果它能定期进去找到那些保存的文章,提取出来,格式化成其他形式,然后发给我或发布到别处,比如 AirTable,那会很有意思。这就是我觉得有用的东西。我不知道我们是否已经能处理旅行预订或人们常说的那些常见用例。至少对我来说,处理那个问题的方式比这些工具目前能做的要复杂得多。

I've played with them mostly to see what they can do and to play with them, but I also have not found a killer use case so to speak. The one thing that I've wanted to do is, it's kind of a weird one, but I save articles across lots of different platforms. X, LinkedIn, Hacker News, are the big ones actually. And accessing those LinkedIn saves is kind of super hard because like it's said, I always have to look for where that thing is. And there are no APIs that I have been able to find. And so someone was talking about using Comet I think for LinkedIn. And I thought it'd be kind of interesting if it could go in periodically and find those saved articles and pull them out and format them into some other way and then send them to me or post them someplace else, post them to AirTable. Like that's the kind of thing that I would find useful. I don't know that we're there yet for travel booking or those common use cases that people talk about. The problem for at least the way I approach that problem is so much more complex than any of these things are readily able to do.

Devi Parikh

我认为你所描述的是,我确实认为我们需要这些智能体在后台为你做有用的事情。对吧?比如现在有一种工作流,细节不重要,但有一种手动操作我每天要做多次,它非常适合让浏览器里的智能体来替我完成,但我经常不小心关掉标签页,然后它就没了,对吧?我需要重新发送。所以,如果我能设置成每次发生某事时,就去执行这个非常具体的手动操作,并在后台完成。无论我的笔记本电脑是开着还是关着都不重要。我的意思是,这很简单,但在某种意义上可能难以实现。比如,抛开当前技术的状态,如果你跟某人谈论代表你行动的智能体,你会想象它们生活在云端,做这些事情,并在有有趣的事情时通知你。但目前大部分情况并非如此。你需要去访问它们,向它们提问,与它们互动。这非常交互式,我还没有看到很多真正能展示其力量的例子,我不知道该怎么称呼它们,环境智能体系统之类的。我不知道我们是否已经创造了这个术语。

And I think what you're describing is like I do think we need these agents to be in the background doing useful things for you. Right? Like there is a certain workflow right now where the details are not important but there's a certain manual thing that I end up doing multiple times a day and it's so, like it would be a great fit for this agent in my browser, right, to do it for me, but I end up closing the tab accidentally every so often and then it's gone, right? And I need to sort of resend that. And so, if I could just set it up to be like every time blah happens, just go do this very specific manual thing and do it in the background. Whether or not my laptop is open or shut is not important. I mean it's amazing how simple that is and maybe elusive in a sense. Like when absent of the current state of the technology, if you talk to somebody about agents that act on your behalf, you would imagine them living in the cloud and doing these things and letting you know when they have something interesting for you. And for the most part it's not really like that right now. Like you're going to these things, you're asking them stuff, you're doing stuff with them there. It's very interactive and I've not seen a lot of really great examples that illustrate the power of I don't know what we want to call them, ambient agentic systems or whatever. I don't know that we've coined that term yet.

Host

是的。我的意思是,显然我有偏见,但我认为 Scouts 正是这样,对吧?你应该试试看。它就是这样。你设置一次,对吧?你告诉它,每当某事发生或有趣的事情发生时通知我,然后它就开始工作了,对吧?你不需要与它交互。它全天候为你监控网络。每当它有东西要报告时,它会给你发一封包含信息的邮件。所以经常有好几周我都没收到 scout 的消息,因为没发生相关的事情,然后它突然出现在我的收件箱里,因为它一直在外面寻找,这感觉很神奇。所以,是的,我在开始录音前就答应给你访问权限了。我会给你权限。你应该试试,然后告诉我你的想法。

Yeah. I mean, obviously I'm biased, but I think Scouts is exactly that, right? You should try it out. But it's very much this. You set it up once, right? You're telling it that like let me know whenever blah happens or whenever something interesting happens and then it's off, right? You're not interacting with it. It's off monitoring the web 24/7 for you. And whenever it has something to report, it's going to send you an email with that information. And so it's often the case that for several weeks I haven't heard from the scout because nothing relevant happened and then it shows up in my inbox because it was out there looking and that just feels quite magical. So yeah, I already promised to give you access before we started recording. So I'll get you access. You should try it out and then you should tell me what you think of it.

Devi Parikh

我一定会的。一定会的。那么,Devi,非常感谢你抽出时间来分享你最近在做的事情。非常酷。

I definitely will. I definitely will. Well, Devi, thanks so much for taking the time to jump on and share a bit about what you've been up to recently. Super cool stuff.

Host

谢谢你邀请我。谢谢你邀请我。

Thanks for having me. Thanks for having me.

互动版:逐字朗读 + 针对本期提问 →