构建物理 AI:Waymo 自动驾驶的经验教训

Building Physical AI: Lessons from Waymo's Autonomous Driving

德米特里·多尔戈夫 Dmitri Dolgov · Y Combinator · 2026-08-03 · 约 49 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

Waymo CEO 分享在物理世界中安全部署 AI 的七条技术经验,强调与数字 AI 的差异。

Waymo's CEO shares seven technical lessons for safely deploying AI in the physical world, emphasizing the differences from digital AI.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 23)

全文 · Full transcript(中英对照)

引言与概览 Introduction and Overview

Host

大家下午好。很高兴来到这里。我们经常谈论那些活在屏幕里、活在数字世界中的 AI。而今天我想和大家聊聊我们在 Waymo 一直在构建的一种不同的 AI——活在真实物理世界中的 AI。顺便问一下,你们当中有多少人坐过 Waymo?请举手。哇。好的,这很了不起。尤其是我知道你们很多人是从外地来的。那些来参观、还没机会体验 Waymo 的朋友,我希望你们在湾区的时候能试一试。既然这是创业学校,我把这次演讲组织成一系列的经验教训——我们在 Waymo 多年来学到的七条经验,关于构建并安全交付当今物理世界中最成熟的 AI 应用——Waymo 驾驶员。让我先看一段短视频。这是最近我和孩子们一起乘坐 Waymo 时的一段录像。正如你们看到的,我们正在向前行驶,通过一个路口时,几个人类司机突然插到我们前面,而 Waymo 驾驶员安全、平稳地做出了反应。事实上,孩子们在后座全神贯注,甚至都没注意到发生了什么。对我来说,这是一个非常有力的时刻。我从事这项技术和产品已经将近二十年了,而它刚刚做了一件相当重要的事:它安全地行动了,保护了我的孩子,保护了所有人,而没有人注意到。我认为这将成为物理 AI 的一个主题——最好的 AI 时刻看起来就像什么都没发生,只是任务安全、平稳地完成了。而像这样 Waymo 驾驶员保护所有人安全的时刻,每天都在我们的车队中发生。如今,Waymo 驾驶员每周提供约 500 次出行,每周在美国 15 个城市行驶超过 400 万英里的全自动驾驶里程。作为比较,这相当于每周超过 300 年的普通美国司机年均驾驶量。而 Waymo 驾驶员以超人的安全记录实现了这一点。那么,在物理世界中大规模构建和部署 AI 智能体需要什么呢?在硅谷,有一个常见的口号是“快速行动,打破常规”。然而,当你处理的是原子而非比特时,“打破”就不太行了。所以你必须做的是“快速行动,安全交付”。这是一件困难得多的事情。你必须从第一天起就构建稳健的系统。你必须构建 AI 模型,必须构建训练配方,让安全成为基础,而不是事后添加的东西。顺便说一句,物理 AI 的问题本身就和数字 AI 不同。如果你要为物理世界构建 AI,与数字世界相比,你必须应对四个主要差距。首先是错误成本差距。如果你有一个语言模型、聊天机器人或副驾驶,它犯了错误,通常只是让你重试一次。而在物理世界中,错误的代价可能以人的生命来衡量,而不是 token。根本没有撤销和重试按钮。其次是延迟差距。通常,当你运行一个 VLM 或数字助手时,它可能需要几秒甚至几分钟才能给你答案。而一辆以高速公路速度行驶的汽车,1 秒内移动约 100 英尺。所以毫秒真的很重要。你必须在车载算力上运行所有的推理、做出所有的决策,而这些算力要能放进汽车后备箱。接下来是数据差距。数字 AI 拥有互联网——这个我们积累的、巨大的、预先标记的人类知识和思想的缓存。而物理世界没有数字化的互联网版本。最后是验证差距。在数字 AI 中,你通常可以交付一个足够好的产品,让用户使用,他们发现边缘情况,这样你就可以在第一天几乎无限规模地部署,然后从那里迭代、爬山式地提升质量。而在物理 AI 中,情况不同。鉴于错误的高成本,你需要在第一天、在部署第一个机器人、在行驶第一英里自动驾驶里程之前,就达到非常高的安全水平和非常高的信心。同时,当你处理物理 AI 时,让你的智能体在真实世界中的实际体验是无价的、不可替代的。这些系统不是你在实验室里构建、做到完美,然后一夜之间全面部署的东西。所以,鉴于这两个因素,你真的需要极其清晰、极其明确地定义你的智能体的运行条件和部署参数,然后构建一个严谨的框架来指导你的部署,以便负责任地扩展。这绝对至关重要。这是你赢得客户、社区、监管者和你自己信任的方式。在 Waymo,我们当然是在自动驾驶汽车的背景下看待这些差距。但这些差距几乎会以某种形式出现在我们部署的任何非平凡的物理智能体中。而驾驶只是 AI 大规模跨越这四个差距、与公众互动的第一个领域。那么,让我们深入探讨我们在 Waymo 多年来从这个问题的解决中学到的经验,并谈谈我们如何应对这些差距。我这次演讲有七条经验。它们都是技术性的。构建公司和产品还有很多其他方面。但今天我只专注于构建物理世界 AI 的技术方面。我认为每一条经验本身都不会是惊天动地的。很多内容可能和你听过的其他地方重叠。但我希望这些经验在我们部署物理智能体并安全扩展的经历中的具体体现,以及我能补充的一些细微差别,对你们中许多在这个领域、正在构建产品和创业的人来说,会是有趣和有用的。那么,让我们开始吧。第一条经验与演示和真实产品之间巨大的、令人沮丧的、有时甚至令人崩溃的差异有关。一个能用的演示最多只是你必须做的工作的 1%。接下来的许多个九的性能、许多个九的可靠性,那才是真正的工作所在。如果你在座是创始人,很可能你正专注于让第一个原型、第一个演示落地。

Good afternoon everyone. It's great to be here. We talk a lot about AI that lives on your screen, lives in the digital world. And today I'd like to talk to you about a different kind of AI that we've been building at Waymo. AI that lives in the real physical world. How many of you by the way have been in a Waymo? Just raise your arms. Wow. Okay, that is impressive. Especially I understand many of you are out of town. The folks who are visiting and have not had a chance to check out Waymo, I hope while you're here in the Bay Area, give it a try. So this being a startup school, I structured this presentation as a sequence of lessons, seven lessons that we've learned over the years at Waymo around what it takes to build and safely ship today's most mature application of AI in the physical world, the Waymo driver. Let me start with a short video. This is a clip from a ride that I recently took in a Waymo with my kids. So as you see here, we're moving forward. We're proceeding through an intersection and a couple of human drivers just decide to cut in right in front of us and the Waymo driver reacted safely, reacted smoothly. In fact, so much so that the kids, my kids were preoccupied in the back seat. They didn't even notice that anything happened. And to me, this was a pretty powerful moment. I've been working on this technology and this product for close to two decades, and it just did something fairly important. It acted safely. It kept my kids safe. It kept everybody safe and nobody noticed. And that I think will be a bit of a theme in general when it comes to physical AI that the best AI moments will look like nothing happened. It's just the task got done safely and smoothly. And these sort of moments where the Waymo driver kept everyone safe are happening daily across our fleet. Today the Waymo driver is serving around 500 trips per week and driving over 4 million fully autonomous miles every week in 15 cities across the United States. Just for a comparison, that's over 300 years every week of an average American driver per year. And the Waymo driver is accomplishing that with a superhuman safety record. So what does it take to build and deploy an AI agent in the physical world at scale? Now in Silicon Valley there's a common mantra to move fast and break things. However, when you're dealing with atoms instead of bits, breaking things is not really okay. So the thing you have to do is to move fast and ship safely. And that's a much more difficult thing to do. You have to build systems that are robust from day one. You have to build AI models and you have to build training recipes where safety is the foundation and not an afterthought, not an add-on. And by the way, the problem itself of physical AI is different from digital AI. There are four main gaps that you have to contend with if you're building AI for the physical world versus the digital world. First, there is the cost of error gap. Have a language model or a chatbot or a co-pilot and it makes a mistake, usually it costs you a retry. In the physical world, the cost of a mistake can be measured in human lives, not tokens. There's simply not an undo and a retry button. Secondly, you have the latency gap. And typically when you're running a VLM or a digital assistant, it can take many seconds, sometimes minutes to come back with an answer to you. A car traveling at freeway speeds moves about 100 feet in 1 second. So milliseconds really matter. And you have to run all of your inference, make all of your decisions on board a compute that fits in a trunk of your car. Next, there's the data gap. Digital AI had the internet, this wonderful immense cache of pre-labeled human knowledge and human thought that we've ever assembled. There's no digitized version of the internet for the physical world. And lastly, there's the validation gap. In digital AI, often you can ship something that's good enough and you let your users use your product, they find the edge cases and that allows you to deploy on day one practically at unlimited scale and then you can just iterate and hill climb on quality from there. In physical AI, the situation is different. Given the high cost of errors, you need to have a very high level of safety and a very high level of confidence on day one before you deploy your first robot, before you drive your first autonomous mile. Now, at the same time, when you're dealing with physical AI, the actual experience of having your agent in the real world is invaluable and it's irreplaceable. These systems are not just something that you can build in the lab, get it perfect, and then deploy at full scale overnight. So given those two factors, you really need to super clearly and super crisply define the operating conditions and the deployment parameters of your agent and then build a rigorous framework to guide your deployment so that you can scale in a responsible manner. And this is absolutely critical. This is how you earn trust from your customers, from the communities, from the regulators, and yourself. So at Waymo, we see these gaps of course in the context of autonomous vehicles. But these gaps will show up in practically any sort of non-trivial physical agent that we will deploy in some shape or form. And driving is just simply the first domain where AI has crossed these four gaps at scale with the public interacting with our product. So let's dive into those lessons that we've learned over the years at Waymo from working on this problem and talk about how we address those gaps. I have seven lessons in this talk. They're all technical. There's a lot more that goes into building a company and building a product. But today I'll just focus on the technical aspects of building AI for the physical world. And each one of those lessons I think by itself will not be exactly earthshattering. A lot of it will overlap with likely things you've heard elsewhere. But I hope that the grounding of these lessons in our experience and some of the nuance that I can add about how they showed up in our experience of deploying a physical agent and scaling it safely will be interesting and useful for many of you who are in the space as you build your product, as you build your startup. So let's dive in. And the first lesson has to do with this massive, frustrating, sometimes soul-crushing difference between a demo and a real product. And a working demo is 1% at best of the work that you have to do. The many nines of performance, the many nines of reliability that follow, that's where the real work happens. And if you're a founder in the room, chances are you are focused on getting that first prototype, that first demo, off the ground.

演示与产品差距 The Demo vs. The Product Gap

Dmitri

当你做出第一个能用的系统版本,那最初的 90%,当演示真的跑通时,感觉妙不可言。你会觉得自己已经解决了问题,前途无量,开始向外推演。而在我们这一行,我们大约在 2010 年就达到了第一个里程碑,也就是那最初的 90%。所以当这个项目启动时,在我们开始搭建系统之前,我们给自己定下了几个相当宏大的目标。一个是让 100 辆自动驾驶汽车完成 10 万英里的自动驾驶里程。第二个目标是跑通 10 条路线,每条 100 英里,覆盖湾区各种路况。而且每条路线都必须从头到尾无人干预。当时我们团队大约有 12 名工程师,我们在大约一年半的时间里完成了这两个目标。要知道,这远在任何 AI 突破之前,在卷积神经网络、Transformer、大语言模型,以及我们今天谈论的所有这些东西之前。然而,我们还是做到了,而且按照演示的标准,自动驾驶在 2010 年就已经解决了,对吧?我们应对了所有情况:白天能开,晚上能开,能应对交通、行人、骑行者、红绿灯、施工区,在高速和地面道路都能开。所以我们可以说是“能力完备”。当时我们觉得自己站在世界之巅。但当我们开始打造产品时,很快就撞上了残酷的现实:做一次,或者跑通 10 条路线,和打造一个无人值守的可扩展服务,这两者之间有着天壤之别。我们又花了大约 10 年才开始提供服务,然后又花了 5 年才扩展到每周 50 万次行程。所以演示花了 18 个月,产品花了大约 15 年,但现在我们正在指数级扩张。到目前为止,我们已经完成了超过 2000 万次全自动驾驶行程,行驶里程超过 2 亿英里。我们在美国 15 个城市运营着只有乘客的车辆。我们正在指数级增长。我们花了 15 年才达到第一个 1 亿英里,而下一个 1 亿英里只用了大约 7 个月。从我们开始最初的只有乘客的运营,到在 4 个城市为乘客提供服务,我们花了大约 8 年。今年早些时候,我们一天之内就在 4 个城市上线。

And when you hit that first version of a system that works, that first 90%, when the demo actually works, it feels incredible. You feel like you solved it, the sky's the limit, you're extrapolating forward. And in our world, we hit that first milestone, that first 90% back around 2010. So when this project started, before we started building the system, we set a couple of pretty ambitious goals for ourselves. One was to drive 100 autonomous 100,000 miles in autonomous mode. The second goal was to drive 10 routes. Each one was 100 miles long, chosen to cover a variety of conditions across the Bay Area. And we had to do each one from beginning to end without a human intervention. We had at the time a team of about a dozen engineers and we accomplished both of these goals in about a year and a half. And keep in mind this was well before any of the AI breakthroughs, before CNNs, before transformers, before LLMs, before any of the stuff that we talk about today. And yet, you know, we got it done and kind of by demo standards, autonomous driving was solved in 2010, right? We handled everything: we could drive during the day, during the night, we handled traffic, pedestrians, cyclists, traffic lights, construction zones, on freeways, on surface streets. So we were, quote unquote, capability complete. And at the time we felt like we're on top of the world. But then we quickly ran, as we started building towards a product, we quickly ran into a brutal reality that there's a massive difference between doing something once or driving 10 routes once and building a scalable service with nobody behind the wheel. It took us about 10 more years to begin providing a service and then five more years to scale to half a million trips per week. So the demo took 18 months, the product took about 15 years, but now we're scaling exponentially. To date, we've served well over 20 million fully autonomous trips, and we've driven well over 200 million fully autonomous miles. And we have rider-only vehicles operating in 15 cities across the United States. And we're scaling exponentially. It took us 15 years to get to that first 100 million miles and about 7 months to drive the next 100 million. It took us about 8 years to go from the time when we started our initial rider-only operation to the time when we were serving riders in four cities. Earlier this year we launched four cities in just one day.

Dmitri

那么,为什么从演示到产品的跨越需要这么长时间?因为有一个残酷的工程现实,你无法真正作弊。可靠性和性能处在一个“9”的指数阶梯上。达到最初的 90% 或 99% 是容易的部分。但之后每增加一个 9,大约需要 10 倍的努力。所以你需要提前确切知道你的产品到底需要多少个 9。演示可能只需要一个 9。辅助产品或者副驾驶可能需要几个,但一个完全自主的 AI 智能体,我们要放到物理世界中,与公众互动,周围还有孩子跑来跑去,那需要一整套的 9。而在规模化的过程中,长尾就是问题空间,就是你整个问题的定义。当你每周行驶数百万英里时,一个百万英里才发生一次的罕见事件,就成了你的日常现实。而获得那些额外的 9 意味着每次都要做不同的事情。你不能通过把之前达到前两个 9 的方法做得更久来达到六个 9 的性能或可靠性。你必须做根本不同的事情,需要根本不同的方法。比如,以可靠性为例,你可以通过规范的工程和修复一些 bug 来达到前几个 9,但要达到接下来的几个,你需要投资于根本不同的方法。你需要构建完全冗余的系统,有分层回退架构等等。AI 模型的性能也是如此。所以这实际上意味着,在这个领域,起步非常容易,但要达到真正的产品可能极其困难,而且这种效应随着每一波技术突破而放大,自然会导致炒作周期。每一次 AI 突破,从深度学习到卷积神经网络到视觉语言模型,你能想到的,都让起步变得更容易。你的演示、原型,容易了 100 倍,但长尾,也就是难题所在,移动得少得多。它也在移动,但效果被削弱了。这就是为什么每个炒作周期都会产生一波惊艳的演示,但真正的产品却寥寥无几。而每个周期反复出现的错误是,把本该留给“9”的资源花在了演示上。

So why does bridging that gap from demo to product take so long? Well, because there's this harsh engineering reality that you can't really cheat. That reliability and performance lives on this exponential ladder of nines. So getting to that first 90% or 99%, that's the easy part. But then every next nine that you want to add, that takes about 10 times more effort. So you need to know upfront exactly how many nines your product actually needs. So a demo might need, you know, one nine. An assist product or, you know, co-pilot might need a few, but a fully autonomous AI agent that we're going to be putting out in the physical world that engages with the public, you know, with kids running around, that needs a whole stack of them. And at scale, the long tail is the problem space. It's your entire problem statement. When you drive millions of miles per week, a rare event that might happen once in a million miles, that just becomes your daily reality. And getting those next nines means doing something different every time. So you don't get to say six nines of performance or reliability by doing the same thing that you did to achieve the first two, but longer. You have to do fundamentally different things. It requires a fundamentally different approach. For example, we can take reliability. You can get to the first couple of nines by just doing proper engineering and doing some bug fixes, but to get to the next few, you need to invest in fundamentally different approaches. You need to build fully redundant systems, have tiered fallback architectures, and so forth and so on. And the same thing holds for the performance of AI models. So what that actually means is that in this space it's incredibly easy to get started but it can be excruciatingly difficult to get to the real product, and that effect is only amplified with every wave of technological breakthroughs, and that naturally leads to hype cycles. So every AI breakthrough, from deep learning to CNNs to VLMs, you name it, it makes it that much easier to get started. Your demos, your prototypes, they get a 100 times easier, but the tail, that's where the hard problems are, that moves much less. It moves, but the effect is muted. And that's why every hype cycle produces a wave of absolutely spectacular demos and very few real products. And the recurring mistake of every cycle is spending on the demo what you should be saving for the nines.

Dmitri

现在,你知道,这里是创业学校,我最不想做的就是给那些早期的魔力和兴奋泼太多冷水。这个时期绝对神奇,太棒了。珍惜它,利用它。但关键是要对自己正在构建的产品保持诚实,对产品所需的性能和可靠性的“9”的数量保持诚实,并且不要为了达到目标而偷工减料。否则,你以后可能会遭遇相当残酷的觉醒。所以,在数你的演示观看次数之前,先数数你的“9”。

Now, you know, this being a startup school, the last thing I want to do is throw too much cold water on the magic and the excitement of those early days. This time is absolutely magical. It's amazing. Cherish it, leverage it. But the key is to remain honest about the product that you're building, the number of nines in performance and reliability that that product demands, and not cutting corners to get there. Otherwise, you might be in for a pretty rude awakening later. So count your nines before you count your demo views.

架构由所需九数决定 Architecture Dictated by Required Nines

Dmitri

这就引出了第二个教训。一旦你知道你的产品实际需要多少个“9”,它就从根本上决定了你需要追求的系统架构和核心技术方法。现在,每项技术都有一条性能与努力程度的曲线,对吧?它们通常都是先陡峭上升,然后趋于平缓。而且,正如我刚才提到的,每增加一个“9”难度就会增加一个数量级。所以一个常见的失败模式是,选择了能让你最快早期上升的技术,沿着那条陡峭的曲线前进,感觉自己赢了,把那条陡峭的斜率投射到未来,觉得前途无量,然后撞上平台期,发现你选择的技术路径在远未达到产品所需性能之前就变平了。现在,你可能仍然会选择,至少在一段时间内,走那条陡峭的曲线,出于各种实际原因。你知道,也许你想做原型、做演示,或者为了学习而构建一些东西。但要对自己诚实,你是在为演示而建,为学习而建,还是为了真正的产品。让我们从我们领域的例子来看,自动驾驶和感知。关于自动驾驶到底需要什么样的传感器,一直存在争论。自然,更多的传感器意味着更高的性能,但也意味着更高的复杂性。所以人类当然可以只用眼睛开车。

And this brings us to the second lesson. Once you know how many nines your product actually needs, it fundamentally dictates the architecture and the core technical approach that you need to pursue. Now every technology has a performance versus effort curve, right? They all tend to start fairly steep and go up and then they flatten out. And you know, as I just mentioned, every other nine gets an order of magnitude more difficult. So a common failure mode is picking the tech that gives you the fastest early ramp, riding that steep curve, feeling like you're winning, projecting that steep slope into the future and feeling like the sky is the limit, and then hitting the plateau and discovering that the technology path that you picked actually flattens out way before the performance that is required by your product. Now, you might still choose to be, at least for a while, on that steep curve for a variety of practical reasons. You know, maybe you want to prototype something or demo something or build something in service of learning. But be honest with yourself whether you're building for the purpose of a demo, for the purpose of learning, or towards an actual product. So let's take an example from our domain, autonomous vehicles and sensing. There's been a long-standing debate about what kind of sensors you actually need for autonomous driving. Naturally, more sensors means higher performance, but also means higher complexity. So humans, of course, can drive with just eyes.

传感模态介绍 Introduction to Sensing Modalities

Dmitri

所以这就是存在的证明。现在,如果目标只是大致达到人类水平,或者构建一个辅助产品,那是一种非常合理的方式。然而,如果你的目标是完全自主,并且追求超人、强超人的性能,你会发现弱感知只会导致安全曲线过早趋于平缓。因此,在 Waymo,我们采用了一种使用多种感知模态的方法。我们使用摄像头、激光雷达和雷达,它们相互补充。摄像头提供高分辨率和色彩,但它们是被动式的,在黑暗和眩光下会退化。激光雷达直接测量周围世界的 3D 结构,而雷达非常擅长穿透环境条件和天气,如雾、雨或雪,并且可以直接使用多普勒测量速度。激光雷达和雷达是主动传感器,这意味着它们在漆黑一片或例如迎着刺眼夕阳行驶时也能看得一样清楚。这些不同的感知模态当然不是彼此的备份。在我们的技术栈中,每种模态都有一个编码器,来自所有这些传感器的信息被融合成一个对周围世界的单一视图,这个视图比任何单一传感器获得的都要精确得多,通常也优越得多。

So there's that proof of existence. Now, if the goal were to just approximately match human performance or to build an assist product, that's a very reasonable way to go. However, if you're targeting full autonomy and you're targeting superhuman, strongly superhuman performance, you find that weak sensing just leads to a safety curve that flattens out way too early. So, at Waymo, we've taken an approach where we use multiple sensing modalities. We use cameras, lighters, and radars, and they all complement each other. Cameras give you high resolution and color, but they're passive, and they degrade in darkness and glare. Lighter gives you a direct measurement of the 3D structure of the world around you, and radar is very good at punching through environmental conditions and weather like fog or rain or snow, and it directly can measure velocity using Doppler. Lighter and radar are active sensors. So that means they see just as well in pitch darkness or, for example, when driving into a blinding sunset. And these different sensing modalities, of course, they're not backups to each other. In our stack, each modality has an encoder, and the information from all of those sensors gets fused into a single view of the world around us that is much more precise and generally vastly superior to what you get with any one sensor.

多传感器融合实例 Examples of Multi-Sensor Fusion

Dmitri

让我给你们看几个例子。这是一个 Waymo 在凤凰城沙尘暴中行驶的场景。你们在这里看到的是我们相当先进的高分辨率、高动态范围摄像头所看到的场景。这非常接近人类在相同条件下会看到的情况,也就是几乎看不到什么。右边是激光雷达在完全相同的帧中看到的内容。你可以更清楚地看到路边站着一个行人。所以如果他们踏上马路,这种早期检测会对事态的发展和所有相关人员的安全产生非常大的影响。

So let me show you a few examples. Here's a scene where a Waymo is driving in a dust storm in Phoenix. What you see here is what the scene looks like to our fairly advanced high-resolution and high dynamic range camera. It's very close to what a human would see in the same conditions, which is not much. And here on the right is what the lighter sees for the exact same frame. And you can much more clearly see that there's a pedestrian standing on the side of the road. So if they were to step onto the road, that early detection can make a really big difference in how the situation plays out and the safety of everyone involved.

Dmitri

另一个例子。夜间行驶时,有几个行人正要翻过混凝土施工护栏跳到马路上。同样,在底部你可以看到摄像头,真的看不到太多东西。顶部是激光雷达的视图。再次,激光雷达对比摄像头。

Here's another example. At night driving along, and there are a couple of pedestrians who are about to jump onto the road over a concrete construction barrier. Again, at the bottom you see the camera, really can't see much. And the lighter view at the top. Again, lighter versus camera.

Dmitri

另一个例子。几只狗在追一头公牛,几个孩子在追狗。差别很大。这是摄像头看到的样子。这是激光雷达,对孩子的早期检测在侧面,那里没有车头灯,没有路灯,完全黑暗。所以这会产生很大的不同。

Here's another example. A couple of dogs chasing a bull, and a couple of kids chasing the dogs. And big difference. Here's what it looks like to the camera. Here's the lighter, and the early detection of the kids is off to the side, and there are no headlights. There are no lamps there. It's complete darkness. So it makes a big difference.

Dmitri

或者想想当有东西物理上遮挡传感器视野时会发生什么。如果感知没有冗余,一片叶子落在传感器上就可能让你的机器人完全停摆。所以你需要冗余。冗余当然不一定意味着多种感知模态,但既然你无论如何都需要冗余,不如在正常情况下从不同感知模态的互补物理特性中获益。

Or think about what happens when something physically obstructs the view of your sensors. If you don't have redundancy in sensing, you can have a single leaf land on your sensors and bring your robot to a full stop. So you need redundancy. Redundancy of course does not necessarily mean multiple sensing modalities, but if you need redundancy anyway, you might as well benefit from the complimentary physics of the different sensing modalities in the nominal case.

Dmitri

这里有一段视频,我们的一辆车挂上了一片叶子,实际上我想是一整根树枝,雨刷无法将其甩掉,汽车检测到了这一点,由于我们有感知冗余,它安全地返回了停车场进行适当清洁。

So here's a video of one of our cars that picked up a leaf, or actually I think a full branch of a tree, that our wipers were unable to shake, and the car detected that, and because we have sensing redundancy, it safely was able to get back to the depot for proper cleaning.

硬件成本与代际改进 Hardware Cost and Generational Improvements

Dmitri

所以具体到硬件,不要锚定今天的组件价格。我们目前是第六代 Waymo 驾驶系统,即 Waymo 硬件套件,每一代硬件不仅提供了惊人的能力,而且我们还能大幅简化并显著降低硬件成本。所以把你的公司、你的方法押在今天的硬件价格上,就是把公司押在一个保质期很短、即将过期的数字上。硬件会变化,许多组件会商品化并降价。所以为那个未来设计,并准备好升级。

So specifically when it comes to hardware, do not anchor to today's component prices. We are on the sixth generation of the Waymo driver, the Waymo hardware suite today, and with every generation the hardware not only delivered amazing capability but we're able to drastically simplify and radically reduce the cost of the hardware as well. So betting your company, betting your approach on today's hardware prices is just betting your company on a number that has a fairly short shelf life and is going to expire. So hardware will change. Many components will get commoditized and drop in price. So design for that future and be ready to upgrade.

驾驭技术浪潮 Riding Technology Waves

Dmitri

这就引出了下一个教训。教训三。技术发展得极其迅速,尤其是在当今。所以你需要准备好驾驭这些技术浪潮,并且反复这样做。当你这样做时,你不仅要考虑性能和能力的提升,还必须非常注意统一和简化。

And that brings us to the next lesson. Lesson number three. Technology moves incredibly fast, especially nowadays. So you need to be ready to ride those tech waves and do that repeatedly. And when you do, you have to not only think about the wins in performance and the wins in capability, you have to be very mindful about unification and simplification.

Dmitri

多年来,我们看到了许多重大技术突破,其中很多围绕 AI,每一波创新浪潮,我们几乎都围绕那波 AI 突破重建 Waymo 驾驶系统,并且我们经常自己推动这些领域的前沿。我们在 2013 年左右利用卷积网络进行计算机视觉和感知。然后当 Transformer 在 2017 年左右出现时,我们在感知以及行为预测、决策和规划任务上大力押注。

Over the years we've seen a number of major breakthroughs in technology, a lot of them around AI, and with every wave of innovation, we pretty much rebuild the Waymo driver around that major wave of AI breakthroughs, and we often push the state-of-the-art in those areas forward ourselves. We leveraged ConvNets around 2013 for computer vision and perception. Then when transformers came about around 2017, we bet big on them both for perception and for the task of behavior prediction and decision-making and planning.

Dmitri

事实证明,驾驶任务与语言建模任务并没有太大不同。由于驾驶的社会性,你有点像在与世界上的其他动态参与者进行对话,但你是通过你的智能体——你的车——的肢体语言来进行的,而不仅仅是文字语言。而且你是在序列中操作,局部连续性很重要,但全局上下文也很重要。

Turns out the task of driving is not that dissimilar from the task of modeling language. Because of the social aspects of driving, you're kind of having a conversation with other dynamic actors in the world, but you're doing that in the space of body language of your agent, your car, as opposed to just the language of words. And you operate in sequences and local continuity matters. But so does global context.

Dmitri

而今天,我们正在利用最新的视觉语言模型和前沿世界模型。现在,利用最新技术来获得能力和性能上的胜利,我不想说这很容易,但可以相当直接。孤立地进行应用研究,或者组建一个精英团队来原型化某项新技术,并不是最困难的部分。有许多公司、许多团队在这方面非常出色。

And today we're leveraging the latest in VLMs and frontier world models. Now using the latest tech for capability and performance wins, I don't want to say it's easy, but it can be reasonably straightforward. Doing applied research in isolation or starting a tiger team to prototype some new technology is not the most difficult part. There's many companies, many teams that are excellent in this.

Dmitri

更难建立的能力是将前沿研究带入生产,并在安全关键环境中部署,且不出现回退,同时不打断产品扩展的步伐。增加能力同样不是最困难的部分,但在增加能力的同时减少碎片化和降低复杂性,这才是真正重要的。

The much harder muscle to build is to carry that bleeding-edge research into production and deploy it in a safety-critical environment without regressions, and do it without breaking stride on the scaling of your product. And adding capability again is not the hardest part, but adding capability while at the same time reducing fragmentation and reducing complexity, that is really important.

Dmitri

最后,作为一家公司,更难建立的能力是能够通过多波技术创新和突破反复做到这一点。所以在这方面,我有两条建议。第一条:当一项新技术出现时,它可能非常令人兴奋,非常诱人,让你想启动一个新项目,组建一个精英团队去追求它。这很好,你绝对应该这样做。然而,当你这样做时,非常重要的一点是考虑在成功情景下你之后会怎么做。假设那个努力成功了,你应该非常清楚这项新创新对你公司、你整个产品、你整个系统的路径是什么。

And finally, the hard muscle to build as a company is to be able to do that repeatedly through multiple waves of technical innovation and technical breakthroughs. So on this front, I have two bits of advice. The first one: when a new technology shows up, it can be very exciting, very tempting to kick off a new effort, a tiger team to pursue it. And that's great. You should absolutely do that. However, when you do, it's very important that you consider what you would do after under a success scenario. Let's say that effort succeeds, you should be very clear on what the path of that new innovation is for your company, for your entire product, for your entire system.

技术采纳建议与Waymo基础模型 Advice on Tech Adoption and the Waymo Foundation Model

Dmitri

我经常看到一种失败模式:一个非常困难的技术项目成功了,然后却走进了死胡同。这可能会非常浪费,也完全令人泄气。我的第二条建议是,在追求新技术时,不要只问这项新技术在能力和性能上能给我带来什么,还要问它是否简化了我的技术栈,是导致了碎片化还是统一化。所以,把你的标准设得高一些,要求既要有突破性的性能,同时又要实现彻底的简化和统一。正是我们在 Waymo 多年积累的这种理念和能力,催生了我们最新的核心技术。而它的核心就是 Waymo 基础模型。

Often times I've seen a failure mode where a project, a very difficult technical project succeeds and then there's a dead end. So that can be very wasteful. That can be completely deflating. The second bit of advice I have here is when pursuing new tech, don't just ask what does this new tech give me in terms of capability and performance, also ask has it simplified my stack and has it led to fragmentation or unification. So set your launch bar to demand both breakthrough performance and at the same time radical simplification and unification. And this exact philosophy and this muscle that we've built at Waymo over the years is what produced our latest core technology. And the heart of it is the Waymo foundation model.

Dmitri

Waymo 基础模型是一个多模态世界动作语言模型。这有点拗口,让我拆解一下它的构成。它是多模态模型,因为它能够处理多模态传感器输入,包括摄像头、激光雷达和毫米波雷达。它是世界模型,因为它天生就理解世界如何运作,包括物理、动态,以及社会和语义层面。它是动作模型,因为我们不是被动地观察世界如何演变,我们是主动的参与者。所以模型需要理解我们智能体的动作对世界的影响,并能区分好坏。最后,它与语言对齐,这让我们能够从视觉语言模型中解锁通用世界知识,这在罕见语义情况的长尾分布中非常有用。

Now the Waymo foundation model is a multimodal world action language model. That's kind of a mouthful. So let me unpack the ingredients. It's a multimodal model because it is able to process these multimodal sensor inputs, cameras, lidars, and radar. It's a world model because it inherently understands how the world works, the physics, the dynamics as well as the social and semantic aspect of it. It's an action model because we are not just passively observing how the world evolves. We're an active participant. So the model needs to understand the effects of the actions of our agent on the world and be able to tell the good ones from bad ones. And finally, it's aligned with language and that allows us to unlock general world knowledge from visual language models and that's incredibly useful in the long tail of rare semantic situations.

Dmitri

更具体地说,架构是这样的。它有点像典型的编码器-解码器架构。编码器部分接收多模态感知数据,并将其压缩或编码成一种高效的表示,保留所有相关数据,所有对生成部分(即解码器)有用的信息。这是一个端到端模型,有几个很好的特性。它让我们能够有效地从我们真正关心的任务出发,将梯度一直反向传播到模型的早期层,并且让编码器能够学习到生成部分解决任务所需的丰富表示。它采用了系统一、系统二,即“思考快与慢”的架构,并利用视觉语言模型的通用世界知识来高效学习语义任务。

More specifically, this is what the architecture looks like. It's kind of your typical encoder-decoder architecture. The encoder part takes in the multimodal sensing and compresses it or encodes it into an efficient representation that retains all of the relevant data, all of the relevant information for the generative part or the decoder. It's an end-to-end model which has a couple of nice properties. It allows us to effectively backpropagate the gradient from the task that we actually care about all the way to the early layers of the model, and it allows the encoder to learn the right rich representations for what the generative part needs to solve the task. It uses a system one, system two, think fast, think slow architecture, and it leverages the general world knowledge of VLMs for efficient learning of semantic tasks.

Dmitri

让我们深入一点。首先是“思考快”的路径。这部分融合来自摄像头、激光雷达和毫米波雷达的原始数据,从而能够做出瞬间的安全关键决策。你可以把它想象成你的驾驶本能。比如,如果有行人冲进马路,或者附近的骑行者突然拐到你的车道上,这能让汽车瞬间刹车。这就像智能体的“蜥蜴脑”,处理大量几何任务,能在毫秒级做出反应。

Let's dive deeper. First, the think fast path. That part fuses the raw data from our cameras, our lidars, our radars, and that allows for split-second safety-critical decisions. So you can think of it as your driving instincts. This is what allows the car to brake instantly if, let's say, a pedestrian runs into the road or a cyclist that's nearby swerves into your path. This is like the lizard brain of your agent that deals with a lot of geometric tasks and can react in milliseconds.

Dmitri

其次是“思考慢”的路径。这部分负责更复杂的语义和场景级理解任务。这类任务通常不会在毫秒级变化,所以你可以承受一点延迟,用延迟换取更高的能力和更高水平的推理。例如,如果 Waymo 驾驶员遇到路边有一辆车着火了的情况,快速路径可能只会把它看作一个普通障碍物,并认为前方道路是畅通的。这时就需要慢速路径介入。这条路径可以利用深度语义推理来理解那个物体的语义,即汽车着火,以及更广泛的场景背景。这能让我们的驾驶员决定采取完全不同的行动,或者干脆换一条路线,即使从几何上看前方道路是畅通的。

Second is the slow path. That's the part that's responsible for the more complex semantic and scene-level understanding type tasks. And these sorts of tasks typically don't change in milliseconds. So there you can afford a bit more latency and you can trade that off for higher capability and higher levels of reasoning. For example, if the Waymo driver encounters a situation where there's a vehicle, let's say it's on fire on the side of the road, the fast path might just see it as a generic obstacle and reason that the path ahead of us is clear. And this is where the slow path comes in. That path can use deep semantic reasoning to understand the semantics of that object, the car being on fire, and the broader scene context. And that allows our driver to decide to take a very different action or a different route entirely, even if geometrically the path ahead of us is clear.

Dmitri

最后是生成组件,也就是解码器。这个组件理解并能够产生行为。它理解其他参与者如何行动,让我们能够做出预测并规划我们自己的驾驶决策。我们的 Waymo 基础模型为 Waymo 驾驶员提供动力,它运行在不同代际的硬件上,也运行在不同的车辆平台上。比如我们的第五代和第六代,捷豹路虎的 IPA、Ohigh 和现代 Ioniq。未来,它还将为不同的产品和商业应用提供动力,比如卡车运输和私人拥有的车辆。

Finally, there's the generate component. That's the decoder. That's the component that understands and can produce behavior. It understands how other actors behave and it allows us to make predictions and plan our own driving decisions. And our Waymo foundation model powers the Waymo driver that runs on different generations of hardware and runs on different vehicle platforms. You have our fifth generation and sixth generation, the JLR IPA, the Ohigh, and the Hyundai Ioniq. And in the future, it will power different products and different commercial applications like trucking and personally owned vehicles.

Dmitri

通过专注于车载模型的高容量基础这一策略,我们能够将大量复杂性转移到上游的大型共享基础上,这让我们能够使运行在车上的专业化层变得相当轻量,进而加快开发进程。所以这一课中最重要的能力是,你的公司不仅要利用当下的技术,还要有能力并锻炼出反复驾驭这些技术浪潮的肌肉,将创新的成果引入生产,而不出现回归,不打断部署和扩展的节奏,也不被复杂性淹没。

By leveraging the strategy of focusing on the high-capacity foundation of our onboard model, we're able to move a lot of complexity upstream to that large shared foundation, and that allows us to make that specialization layer that's running on the car pretty lightweight, and that in turn allows us to speed up the development process. So the most important muscle in this lesson is for your company to not just leverage the tech of the day, but have the ability and build that muscle to repeatedly ride those tech waves and pull in the results of that innovation into production without regression, without breaking stride in deployment and scaling, and without drowning in complexity.

Dmitri

我们进入下一课。AI 社区有一个众所周知的教训:利用大规模算力和大规模数据的通用方法,总是会击败依赖手工工程化人类知识的方法。这就是理查德·萨顿在 2019 年提出并阐述的所谓“苦涩的教训”。我们亲身经历过,也在每一次技术突破浪潮中都看到了这一点。每一次“苦涩的教训”都成立,那些在算力和数据上扩展得最好的方法总是胜出。顺便说一句,这也是我们押注于构建基础模型这一方法的原因之一。

Let's move to the next lesson. There is a well-known lesson in the AI community that general methods that leverage massive compute and massive data will always beat methods that rely on handcrafted engineered human knowledge. That's the so-called bitter lesson that Richard Sutton published and formulated in 2019. And we have lived this and we have seen this in every wave of technical breakthroughs. Each time the bitter lesson holds, methods that scale best with compute and data always win out. And by the way, this is one of the reasons why we bet on the approach of building the foundation model.

Dmitri

有一个众所周知的特性:如果你押注于高容量模型,并把你的数据和算力用在这上面,你会得到更好的缩放定律,然后你将其蒸馏成更小、更高效的模型,实时运行在你的智能体上。相比于直接专注于较小的模型,你会得到更好的缩放定律。所以这个教训体现的一个细微之处在于模型中结构的使用,而根据你如何使用结构,你最终可能站在“苦涩的教训”的任一边。本质上,与规模对抗的结构总是会输,而引导规模的结构总是会赢。这一点在关于端到端模型的讨论中尤为突出。正如我提到的,端到端模型有一些非常好的特性。

There is a well-known property that if you bet on a high-capacity model and you use your data and your compute on that, you just get better scaling laws, and then you distill into smaller more efficient models that are running on your agent in real time. You just get better scaling laws as opposed to just focusing on the smaller models directly. So one nuance area where this lesson shows up is the use of structure in your models, and depending on how you use your structure, you can end up on either side of the bitter lesson. Essentially, structure that fights scale will always lose, and structure that channels scale always wins. In particular, this comes up around the discussion of end-to-end models. As I mentioned, an end-to-end model has some very nice properties.

结构增强端到端 Structure-Augmented End-to-End

Dmitri

你要从最终任务一路把梯度正确传回整个模型,这让编码器和解码器之间的接口能够学会使用丰富的学习表示。你知道,这些模型是最容易构建和训练的。你可以从模仿学习开始,一个黑盒端到端模型会给你非常快的进展。你会沿着曲线最初那段陡峭的部分快速上升,对某些产品来说这就够了。但如果你需要在安全关键环境中达到完全自主智能体的超人水平,仅仅做那种基本的、朴素的端到端是不够的。这就是结构介入的地方。关键问题是:结构是助推 Scaling 还是与之对抗?它是限制和约束你的解空间,还是帮助你在不损失通用性的情况下实现 Scaling?

You back properly gradient from the final tasks all the way through the model, and it allows the API between the encoder and the decoder to learn to use rich learned representations. And you know, those are the easiest models to build and train. You can start with doing some imitation learning, and a kind of a black-box end-to-end model will give you very rapid progress. You will ride that initial steep part of the curve, and for some products that's enough. But if you need to reach superhuman levels of performance in a fully autonomous agent in a safety-critical environment, just doing that basic vanilla end-to-end is not enough. This is where structure comes in. The key question here is: does the structure boost scale or does it fight it? Does it limit and constrain your solution space, or does it help you scale without loss of generality?

Dmitri

让我用一个简单的思维实验和一个玩具问题来说明这一点。想象你在构建一个会下围棋的机器人,你希望它在物理世界中下棋。你有一个摄像头观察棋盘,还有一个执行器实际移动棋子。构建这样一个机器人的一种方式是采用端到端系统,直接从像素到执行动作,也许你通过给它一些人类下棋的视频来训练它。那可能是一个非常有趣的研究练习。然而,如果你的目标是构建世界上最好的下围棋机器人,那可能不是最有效的方式。原因是有一个非常简单的中间表示,它完全捕捉了游戏的状态——你试图解决的任务的状态。如你所知,19×19 的棋盘给你一个完全可观察、完整的世界状态,至少对于游戏对弈部分是这样。利用这种结构不会限制你的模型,不会约束你的解空间,但它给你一个非常有帮助的 Scaling 方式。

Let me illustrate this point with a simple thought exercise and a toy problem. Imagine you're building a robot that will play the game of Go, and you want it to play the game in the physical world. You have a camera observing the board, and you have an actuator that will actually move the pieces around. One way to build such a robot is to have an end-to-end system that goes directly from pixels to actuation, and maybe you train it by giving it some videos of how humans play the game. That could be a very interesting research exercise. However, if your goal was to build the world's best Go-playing robot, that's probably not the most efficient way to go. The reason is that there is a very simple intermediate representation that captures completely the state of the game—the state of the task you're trying to solve. As you know, a 19 by 19 board gives you a fully observable and complete state of the world that you care about, at least for the game-playing part. Leveraging that structure doesn't limit your model, doesn't constrain your solution space, but it gives you a very helpful way to scale.

Dmitri

当然,那只是一个玩具例子。任何你想在物理世界中部署的非平凡系统都不会具备那种特性。在物理世界中不存在如此简单、干净、工程化的表示,这正是我们需要端到端系统、学习表示和学习嵌入的全部原因。但在物理世界中,结构确实存在。你有物理定律,有交通规则,有以相当可预测方式行为的物体。你可以在学习表示之外利用这种结构来提升性能、简化验证,最终获得更好的缩放定律。

Now, that of course was a toy example. Anything that's not trivial that you're trying to deploy in the physical world will not have that property. The fact that such a simple, clean, engineered representation doesn't exist in the physical world is the whole reason why we need end-to-end systems and learned representations and learned embeddings. But in the physical world, structure does exist. You have laws of physics. You have rules of the road. You have objects that behave in reasonably predictable ways. You can use that structure in addition to the learned representations to boost your performance, simplify validation, and at the end of the day you just get better scaling laws.

Dmitri

这是我们在 Waymo 所采用的方法,我们称之为“结构增强的端到端”。我们通过用物化的结构表示来增强学习嵌入,从而超越基本的朴素端到端。这给了我们几个非常重要的优势。首先是推理时的验证。因为模型不仅仅是一个黑盒,传感器输入、执行命令输出,我们可以创建一个非常强大的正确性和安全验证层,当智能体部署在我们的车辆上时,你可以实时运行它。这对任何在物理世界中运行的智能体都非常重要。

This is the approach we are pursuing at Waymo, which we call structure-augmented end-to-end. We go beyond the basic vanilla end-to-end by augmenting the learned embeddings with materialized structure representations. That gives us a few very important advantages. First is validation at inference time. Because the model isn't just a black box where sensors go in and actuation commands go out, we can create a very powerful correctness and safety validation layer that you can run in real time when the agent is deployed on our vehicles. This is really important for any agent operating in the physical world.

Dmitri

其次,在模型生成部分(解码器)的大规模训练和评估方面,我们获得了巨大的效率提升。如果你只有一个黑盒端到端系统,你被迫在端到端设置中完成所有评估和训练,从传感器到决策再到执行。有了那个中间结构表示,你可以混合搭配。你可以在更大规模上进行一些训练,在那些紧凑结构表示的空间中进行一些评估,并在从传感器到决策的完整端到端空间中进行一些评估。

Secondly, we get great wins in efficiency when it comes to large-scale training and evaluation of the generative part of the model, the decoder. If all you have is a black-box end-to-end system, you are forced to do all of your evaluation and training in the end-to-end setup, all the way from sensors to decisions to actuation. Having that intermediate structured representation allows you to mix and match. You can do some training at larger scale and some evaluation in the space of those compact structure representations, and some in the full space of end-to-end from sensors to decisions.

Dmitri

最后,我们为评估和训练都获得了强大的可验证反馈信号,支持强化学习等。额外的物化结构为你提供了更强大的评估工具、指标工具,以及设计损失函数或强化学习方案的工具。这里的教训是,要押注一个最大化学习、最小化约束的系统,并有意识地利用结构来提升训练和评估中的性能和缩放定律。

Finally, we get strong verifiable feedback signals for both evaluation and training, supporting things like reinforcement learning. That additional materialized structure gives you much more powerful tools for evaluation, for metrics, as well as crafting your loss function or reinforcement learning recipes. The lesson here is to bet on a system that's maximally learned and minimally constrained, and leverage structure intentionally to boost performance and scaling laws both in training and in evaluation.

模拟器与世界模型 Simulators and World Models

Dmitri

这就引出了一个问题:你实际上如何训练和评估你的物理 AI 智能体?这把我们带到下一个教训。要在物理世界中构建并安全部署智能体,拥有一个良好的、大规模的、逼真的、高保真模拟器是绝对关键的。

Now that raises the question of how you actually train and evaluate your physical AI agent. That brings us to the next lesson. To build and safely deploy an agent in the physical world, it is absolutely critical to have a good, large-scale, realistic, high-fidelity simulator.

Dmitri

训练和评估有两种方式:开环和闭环。在开环中,你被动地观察输入输出对,你可以用它来进行评估和训练。模仿学习就是这样工作的。评估通常采取这样的形式:如果你发现自己处于这种情况,你会怎么做?然后你对此打分。这与闭环形成对比,在闭环中,你采取一个行动,你看到该行动对世界的影响,然后你通过传感器更新你对世界的看法,你再采取另一个行动,依此类推。你根据这些行动序列和世界演化序列进行评估和训练。

There are two ways you can do training and evaluation: open loop and closed loop. In open loop, you passively observe input-to-output pairs, and you can use that for evaluation and training. Imitation learning works like that. Evaluation usually takes the shape of: if you find yourself in this situation, what would you do? Then you score that. That's in contrast with closed loop, where you take an action, you see the effect that action has on the world, then you update your view of the world through your sensors, you take another action, and so on. You evaluate and train on those sequences of actions and sequences of world evolutions.

Dmitri

采取行动并评估那种反事实的能力,对于在物理世界中构建和部署安全关键型智能体绝对至关重要。真正的模拟器就是你实现这一目标的方式。而真正的模拟器不仅仅是放在你 AI 旁边的一些轻量级工具。它本身就是一个大型 AI 模型。构建一个良好的逼真模拟器的问题与构建智能体本身一样困难。

The ability to take an action and evaluate that counterfactual is absolutely vital for building and deploying safety-critical agents in the physical world. A real simulator is how you do that. And a real simulator isn't just some lightweight tooling that sits next to your AI. It is a big AI model in itself. The problem of building a good realistic simulator is just as hard as building the agent itself.

Dmitri

模拟器背后的 AI 确实需要理解世界如何运作——物理、语义、交通、天气等等。模拟器的质量必须足够高,不仅看起来好,而且足以高置信度地训练和评估你将要在安全关键环境中部署到世界的智能体。换句话说,你必须构建一个高度准确的生成式世界模型。

The AI behind the simulator really needs to understand how the world works—the physics, the semantics, the traffic, the weather, and so on. The quality of that simulator has to be high enough so that it doesn't only look good, but it's sufficient to train and evaluate with high confidence an agent that you're going to be putting in the world in a safety-critical environment. In other words, you have to build a highly accurate generative world model.

Dmitri

在 Waymo,多年来我们一直在构建我们所谓的“行为世界模型”。而且我们早在“世界模型”这个术语流行之前就在做这件事了。

At Waymo, for years we've been building what we call behavioral world models. And we were doing that way before the term world models even became popular.

端到端模型与感知世界模型 End-to-End Models and Sensing World Models

Dmitri

而在端到端模型时代,除了行为真实性,你还需要感知真实性。事实上,构建端到端模型已经相当容易有一段时间了。但在闭环中评估它,那才是问题的难点。所以我们转向构建感知世界模型。因为我们在模型中使用这种结构化增强表示,我们也可以在模拟中利用这种结构。我们的行为世界模型在结构化中间表示的空间中运行,而紧密耦合的传感器世界模型则产生逼真的传感器模拟。我们的世界模型利用了 Google DeepMind 在 Gen3 上的杰出工作,这使我们能够在行为和感知方面都生成可控且高度逼真的场景。这反过来又使我们不仅能在之前遇到过的情境中评估和训练新版本的智能体,还能在从未在现实世界中见过的纯合成罕见场景中进行训练和评估。

And now in the era of end-to-end models, you also need, on top of behavioral realism, sensing realism as well. In fact, building an end-to-end model has been fairly easy for quite a while now. But evaluating it in closed loop, that was the hard part of the problem. So we've moved on to building sensing world models. And because we're using that structured augmented representation in our models, we can also leverage that structure in our simulation. Our behavior world model operates in the space of structured intermediate representations, and the tightly coupled sensor world model then produces realistic sensor simulations. Our world model leverages the great work of Google DeepMind on Gen3, and that gives us the ability to produce controllable and highly realistic scenarios both in the behavioral as well as sensing aspects. And that in turn allows us to not just evaluate our agent and train new versions of our agent in situations that we've previously encountered, but it allows us to train and evaluate in purely synthetic rare scenarios that we've never seen in the real world.

闭环仿真示例 Closed-Loop Simulation Examples

Dmitri

所以你在这里看到的不仅仅是一个生成的视频。这是 Waymo 驾驶员在闭环中运行的完整生成式模拟。这里我们模拟的是,如果它在高速公路上遇到一辆停在车道上的汽车会发生什么。你还可以更进一步。这里有一架飞机在我们前方的高速公路上降落,你可以模拟一头大象在十字路口漫步、金门大桥上下雪,或者一只恐龙四处走动。所以这里的教训是,闭环模拟对于评估是绝对必要的,对于训练你的物理 AI 智能体也极其有价值。因此,你需要高度逼真的大规模模拟来进行训练和评估。

So what you're seeing here is not just a generated video. It's a full generative simulation of the Waymo driver operating in closed loop. Here we're simulating what would happen if it came across a car that was stopped in a lane on the freeway. And you can go further than that. Here's a plane that's landing on a freeway in front of us, where you can simulate an elephant on the loose walking through the intersection, snow on the Golden Gate Bridge, or a dinosaur walking around. So the lesson here is that closed-loop simulation is absolutely required for evaluation and is extremely valuable for training of your physical AI agents. So you need highly realistic large-scale simulation to train and evaluate.

第六课:构建生态与飞轮 Lesson Six: Build an Ecosystem and Flywheel

Dmitri

这就引出了第六课。当你处理如此复杂的问题时,你不能只是构建一个模型就完事了。你必须构建一个完整的生态系统。然后你还需要一个飞轮来驱动它。因为要让这在规模上运作,你不能只构建智能体。你需要构建三个。你为我们构建智能体,那是驾驶汽车的驾驶员。你还有模拟器,那是智能体在其中学习的虚拟游乐场。然后你还有评判者。评判者严格评估和判断智能体的表现,并告诉它如何改进。好消息是,这三者的基本推理和生成能力是共享的,这就是为什么在我们的案例中,它们基于同一个基础世界模型。

And this brings us to lesson number six. When you're dealing with a problem of that complexity, you can't just build a model and call it a day. You have to build an entire ecosystem. And then you also need a flywheel that powers it. Because to make this work at scale, you can't just build the agent. You and one AI, you need to build three. You're building the agent for us. That's the driver that drives the car. You also have the simulator, which is that virtual playground for the agent to learn in. And then you have the critic. And the critic is what rigorously evaluates and judges the performance of the agent and tells it how to improve. And the good news is that the fundamental reasoning and the generative capabilities of all three of those are shared, and that's why in our case they're based on the same foundation world model.

飞轮在行动 The Flywheel in Action

Dmitri

一旦你有了这三根支柱,你就可以创建一个极其强大的飞轮来加速你的进展。你的智能体在现实世界中的部署会产生数据。这些数据随后为模拟器提供依据,使其更加逼真。模拟器生成更难的边缘案例,供评判者评分和智能体学习。于是智能体变得更聪明,被部署到物理世界,产生更多数据,从而驱动飞轮并加速进展。但飞轮当然可以向任何方向旋转或原地打转。所以为了让它朝你想要的方向前进,你需要用指标来引导它。

Now once you have these three pillars, you can create an incredibly powerful flywheel to accelerate your progress. So a deployment of your agent in the real world generates data. That data then grounds the simulator and makes it more realistic. The simulator generates harder edge cases for the critic to score and for the agent to learn from. So the agent gets smarter, gets deployed in the physical world, generates more data, and that powers the flywheel and accelerates progress. But a flywheel of course will spin in any direction or in place. So in order to make it go in the direction you want, you need to guide it by metrics.

最终课:评估与指标作为战略模式 Final Lesson: Eval and Metrics as Strategic Mode

Dmitri

这就引出了最后一课:你的模型其实是入场券,但评估和指标才是你最重要的,那是你的战略模式。所以先构建你的评估,再构建你的技术。先构建你的评估和指标,再构建你的产品。如果你不能定量地定义“足够好”意味着什么,你就不是在真正构建产品。你只是在迭代你的演示。如今,最好的模型架构已经广为人知,新想法往往传播得很快。数据极其重要,但没有好的指标,你就是在盲目飞行。你没有利用最好的数据,也无法真正评估对其做出改变的 ROI。所以评估和指标才是你的基础,它们引导你的整个技术栈。

And that brings us to the final lesson: your model is really table stakes, but eval and metrics, that's your most important, that's your strategic mode. So build your eval before you build your technology. Build your eval and your metrics before you build your product. If you can't quantitatively define what good enough means, you're not really building a product. You're just iterating on your demo. So nowadays, the best model architectures are fairly well-known and new ideas tend to proliferate fairly quickly. Data is incredibly important, but without good metrics, you're just flying blind. You aren't leveraging the best data and you can't really evaluate the ROI on making changes to it. So really eval and metrics, that's your foundation and that steers your whole tech stack.

超越模型级评估 Beyond Model-Level Evaluation

Dmitri

但对于物理 AI 智能体,模型级评估是不够的。当你把 AI 智能体放入物理世界时,你的评估和验证需要更深入、更广泛。你需要评估和验证系统的每个组件,从物理层到在物理世界中运行的车载行为层,以及非车载组件和所有相关的运营流程。对我们来说,我们称之为“安全与就绪框架”,我们花了数年时间构建和完善它。这指导着我们的开发、部署和 Scaling,我认为它是最重要的资产之一。

But for physical AI agents, model-level evaluation is not enough. When you're putting an AI agent into the physical world, your eval and your validation needs to go much deeper and much broader. You need to evaluate and validate every component of your system, from the physical layer to the behavioral layer that's running on board in the physical world, as well as the offboard components and all of the operational processes around it. So for us, we call that the safety and readiness framework, and we spend years building and refining it. And that's what guides our development and our deployment and our scaling, and I consider that to be one of our most important assets.

信任与商业优势 Trust and Business Advantage

Dmitri

这再次说明了为什么它很重要:因为在物理世界中,信任就是一切,而评估和指标正是你赢得信任的方式。你不能仅仅通过谈论巧妙的技术解决方案或最先进的模型架构,或展示一个花哨的演示来赢得信任。你要通过在现场日复一日地不懈证明你的系统是安全的、有效的,来逐渐赢得信任。当然,你不能只在闭门造车中向自己证明这一点。这正是我们公开分享安全数据和持续安全研究的原因。于是,这种赢得的信任成为你最终的商业优势,对吧?你的模型可能被泄露,算法可能被复制,但数亿英里完全自主运营的真实世界数据,加上证据级评估和公开审计的证明,这要难复制得多。

And this again is the reason it's important: because in the physical world, trust is everything, and eval and metrics is how you go about earning that trust. You don't just win trust by talking about the clever technical solution or the clever state-of-the-art architecture of your models or doing a flashy demo. You earn it gradually day by day in the field by relentlessly proving that your system is safe and that your system works. And of course, you can't just prove that to yourself behind closed doors. And this is exactly why we openly publish our safety data and our ongoing safety research. So then that earned trust becomes your ultimate business advantage, right? Your models can be leaked, algorithms can be replicated, but hundreds of millions of miles of fully autonomous operations in the real world backed by evidence-grade evaluation and publicly audited proof, that is much, much more difficult to replicate.

整体剧本 The Playbook as a Whole

Dmitri

所以当你退一步看整个剧本,你会发现这些课程没有一个是独立起作用的。九项设定你的标准,确保你选择正确的技术和正确的技术方法,这样你就不会陷入局部最优。然后有意利用结构来提升 Scaling,以及驾驭技术创新浪潮的能力,帮助你达到正确的“九”的水平。而你的 AI 生态系统,包括智能体、模拟器和评判者,由评估和指标引导,这让你能够构建强大的飞轮,这就是所有这些效应如何复合的。正是我们多年来不断完善的这个剧本,让我们实现了 Waymo 驾驶员远超人类的安全表现。

So when you zoom out and look at this playbook as a whole, you realize that none of these lessons works alone. So the nine set your bar and ensure that you pick the right technology and the right technical approach so that you don't get stuck on the local minimum. Then intentional use of structure to boost scaling and the ability to ride technical waves of innovation helps you get to the right level of nines. And your AI ecosystem with the agent, the simulator, and the critic guided by eval and metrics, that's what allows you to build that powerful flywheel, and that's how all of these effects compound. And it's this playbook that we've been refining over the years that allows us to achieve the strongly superhuman safety performance of the Waymo driver.

安全数据快照 Safety Data Snapshot

Dmitri

这是我们发布的最新安全数据快照,基于超过 2.2 亿英里的完全自动驾驶里程。我们在那里看到,在我们运营的区域,Waymo 驾驶员在导致严重伤害的事故方面比人类驾驶员好约 17 倍。

This is a snapshot of the latest safety data we've released, based on over 220 million fully autonomous miles. And we're seeing there that in the areas where we operate, the Waymo driver is about 17 times better than human drivers when it comes to crashes that cause serious injury.

AI在物理世界中的安全影响 Safety Impact of AI in the Physical World

Dmitri

这一点非常重要,因为如今在世界某个地方,每 26 秒就有人在道路交通事故中丧生。按照目前的规模,这意味着 Waymo 每 8 天就能防止一次严重伤害。这不仅仅是仪表盘上的一个指标,而是意味着某人的亲人能够在一天结束时平安无事地走进家门。所以这些只是 AI 在物理世界中的早期安全效益,而且它们只会从这里继续增长。

And that really matters because today somewhere in the world every 26 seconds someone loses their life on a road to a crash event. And on the current scale, what that means is that Waymo is preventing a serious injury every eight days. And this isn't just a metric on a dashboard. That means that someone's loved one got to walk through the front door at the end of the day safe and unharmed. So these are just the early safety benefits of AI in the physical world, and they will only grow from there.

物理AI的机遇 Opportunity in Physical AI

Dmitri

如果你放眼更广阔的图景,这里的机会绝对是巨大的。当前的物理 AI 就像几年前的数字 AI,我们拥有所有正确的要素去追求它。我们有生成式世界模型,我们有架构,我们有负担得起的算力和传感技术,我们有经过验证的缩放定律,我们还有大规模运营的真实产品。过去十年的 AI 发生在数字世界,我认为下一个十年也将发生在物理世界。

If you look at the broader landscape, the opportunity here is absolutely massive. Physical AI right now is where digital AI was a few years ago, and we have all of the right ingredients to go after it. We have generative world models. We have the architectures. We have affordable compute and sensing. We have proven scaling laws. And we have a real product operating at scale. And the last decade of AI happened in the digital world. I think the next decade will also happen in the physical world.

给建设者的建议 Advice for Builders

Dmitri

对于那些决定在这个领域建设的人,祝你好运,玩得开心,并记住你为谁而建、你的使命和你的客户。这才是重要的。否则科技只是一个科学项目,而归根结底,无论科技多么令人兴奋和振奋,没有什么能比得上改变人们生活的喜悦。

And for those of you who decide to build in the space, good luck, have fun, and remember who you're building for, your mission, and your customers. That's what matters. Otherwise tech is just a science project, and at the end of the day, as exciting and exhilarating as the tech is, nothing really beats the joy of making a difference in people's lives.

首次Waymo乘车体验 First Waymo Ride Experience

Host

我们在做什么?

What are we doing?

Dmitri

我们在我们的第一辆 Waymo 里。

We're in our first ever Waymo.

Host

当我们在 Waymo 里意味着什么?

And what does it mean when we're in a Waymo?

Dmitri

这意味着没有人。

It means that there is nobody.

Host

没有人驾驶这个东西。这是一次完全自动驾驶的 Waymo 行程。

Nobody driving this thing. And this is a fully autonomous Waymo ride.

Dmitri

我简直不敢相信。这车比有人驾驶做得还好。

I cannot believe this. The car did a better job than if somebody was driving.

Host

卡车越过了黄线,所以 Waymo 刹车并靠边避让。

The truck was over the yellow line. So the Waymo braked and moved to the side.

Dmitri

它知道怎么念我的名字。

It knew how to pronounce my name.

Host

天哪,你看这个。

Oh my god. Look at this.

Dmitri

哦,这很好。我喜欢。这不是

Oh, it's nice. I love it. This is not

Host

这太酷了。

This is so cool.

Dmitri

我永远不会忘记这个。永远不会。抱歉。

I'll never forget this. Never. Sorry.

互动版:逐字朗读 + 针对本期提问 →