From Farm to Frontier: Alex Kendall on Wayve's End-to-End AI for Autonomous Driving
打开互动全文版(中英对照 + 朗读 + 问答)→Wayve 公司 CEO Alex Kendall 分享其端到端 AI 方法如何在 500 多个城市、多种车辆中实现零样本驾驶,目标是将无人驾驶技术推广至全球。
Alex Kendall, CEO of Wayve, shares how his company's end-to-end AI approach enables zero-shot driving in over 500 cities across diverse vehicles, aiming to scale driverless technology globally.
您正在收听的是《Gradient Descent》,一档探讨机器学习如何在现实世界中落地的节目。我是主持人 Lucas Bewald。好的,我们现在在 NVIDIA GTC 大会的 Core Weave House。我刚刚采访完 Wayve 的 CEO Alex Kendall,这家公司我仰慕已久,在自动驾驶领域做得非常出色。我想你其实能看到他就在我身后摆弄 Core Weave 的 F1 赛车。这是一次有趣的采访,希望你喜欢。Alex,很荣幸在 GTC 与你对谈。这是我第一次在自家地下室或办公室以外的地方录播客,所以挺兴奋的。我得先请你讲讲 Wayve 的故事,因为我觉得它特别有意思。
You're listening to Gradient Descent, a show about making machine learning work in the real world, and I'm your host, Lucas Bewald. All right. Here we are at the Core Weave House at NVIDIA GTC. I just got done interviewing Alex Kendall, who's the CEO of Wayve, a company that I've admired for a long time and has done phenomenally well in self-driving. And I think you can actually see him behind me messing with the Core Weave F1 car. So, it was a fun interview. I hope you enjoyed it. Well, honored to be here with you, Alex, at GTC. This is our first podcast I've recording anywhere except my basement or the office. So, this is pretty exciting. I have to start by asking you to tell me the story of Wayve because I think it's an especially interesting one.
嗯,其实是个有趣的故事。我刚刚在活动上做完演讲,在 Jensen 上台前 5 分钟,我的脸出现在了一个 5 万人体育场的大屏幕上。这让我想起来,我这辈子只在 5 万人面前讲过两次话。上一次是在 2018 年一个叫 Web Summit 的会议上,我们参加了一个种子公司路演比赛,我进了决赛,需要在人群面前做 3 分钟演讲。长话短说,观众投票我得了最后一名,但评委却把冠军给了我,因为他们喜欢我说的话。但这能让你明白,10 年前我创办公司时,我们做的事完全是反主流的。我们的理念是:自动驾驶本质上是用 AI 方法解决自动驾驶汽车问题。这是一个决策性的复杂推理挑战。我们采取的方法是,用端到端学习来攻克它。10 年前,我们开始构建一个单一模型,能够以尽可能大的规模进行推理、努力和泛化,而不依赖 AV 1.0(第一代方法)——那种依赖改装传感器、算力、高精地图和基础设施、且扩展性极差的方式。当时整个行业都嘲笑和否定我们。我们就这样默默耕耘。现在,不仅市场对自动驾驶汽车重新燃起热情,而且我们拥有一个 AI 模型——我们是第一家实现零样本驾驶超过 500 个城市的公司。这意味着这个模型已经跑遍了欧洲、亚洲和北美,还驾驶过超过 10 种不同的车辆,从电动车到货车到 SUV。它正在成为一个能够随时随地驾驶任何车辆的 AI 模型。未来一年,我们将全力冲刺,把它部署到无人出租车和你能买到的消费车辆上。
Yeah, actually it's a fun story. I just got done giving a speech here at the event and I put my face up on the big screen in front of a 50,000 person stadium before like 5 minutes before Jensen came on and it reminded me because the last time I've only spoken in front of like 50,000 people twice in my life. And the last time was at a conference called Web Summit back in 2018 where we had this seed company pitch competition that I got to the final and had to give a 3-minute pitch in front of the crowd. And long story short, there was a crowd vote that I ended up coming last in. Yet, the judges gave the competition to me because they liked what I said. But it just gives you the idea of 10 years ago when I started the company, what we were doing was completely contrarian. The idea that autonomous driving is all about looking at the AV problem with an AI approach. So, it's a decision-making complex reasoning challenge. And we took the approach of the best way to go tackle that is with end-to-end learning. And a decade ago, we started working on this building a single model that could reason and strive and generalize in the largest scale possible without the AV 1.0, the first generation approach which relied on retrofit sensors, compute, HD maps, infrastructure, and really limited scaling. And we started that a decade ago when this was really laughed at and dismissed by the industry. And we've sort of quietly building it away. And we're at the point now where not only is the market industry excited about AVs again, but we have an AI model that we're now the first company to have driven zero shot in over 500 cities. So, this means this model is driven throughout Europe, Asia, North America. It's also driven in over 10 different cars from electric vehicles to vans to SUVs. And so, it's now emerging as an AI model capable of driving any vehicle anywhere. And we've got a sprint ahead of us over the next year to get this deployed in both robo-taxis and consumer vehicles you can buy.
我记得在伦敦和你一起乘车时有过一次非常棒的经历,你让我决定模型往哪开。我在伦敦待的时间不长,但让我惊讶的是,那里的驾驶环境比旧金山更具挑战性——据我所知,旧金山的人更遵守交通规则。伦敦有更多意想不到的情况出现。我当时简直不敢相信你的模型表现那么好。那大概是两三年前的事吧。但在讨论模型质量之前,我想稍微回溯一下,问问你的个人经历。你是在新西兰的一个农场长大的,对吗?
You know, I remember having a really amazing experience driving with you in London where you let me decide where I wanted the model to go. And I hadn't spent much time in London, and I was astonished by how challenging it seemed compared to San Francisco where people I think obey the traffic laws as I understand it much more thoroughly than in London. And there's just much more surprising things coming at you. And I couldn't believe how well your model works at that time. And I think this was maybe two or three years ago. But before we get to the quality of your model, I kind of want to go back in time a little bit and ask you about your history, right? You grew up on a farm, I think, in New Zealand. Is that right?
是的,在南岛的基督城。
Yeah, in Christchurch in the South Island.
你觉得这如何影响了你的领导风格或对自动驾驶的看法?
And how do you think that's kind of affected your leadership or your thinking about autonomous vehicles?
嗯,回想起来,有几件事对我成长很重要。我一半的时间在造东西,不管是电子游戏、乐高、机器人还是其他小玩意儿。我花了很多时间捣鼓那些东西。另一半时间则在山里——爬山、冲浪、山地自行车等等。我把公司命名为 Wayve,是因为我希望开车的感觉能像冲浪一样好。但我真的认为,推动前沿技术本质上是一场冒险。你需要规划、风险管理、雄心勃勃、突破边界。我觉得我从那些经历和价值观中汲取了这些。这些至今仍能指引我们,因为我们所做的,在开始时看似不可能,但我相信并希望,未来构建这项技术会成为可能。我们很幸运,这么多因素汇聚在一起,让梦想成真,但这确实像是在为冒险而优化。我很幸运获得奖学金,从新西兰去了剑桥大学。那个地方就像中世纪的城堡,充满了最杰出的头脑,让我学习和探索。可以说,对冒险的追求在某种程度上引领我走到了今天。
Um, I think there are a few things that were really important to me growing up on reflection. I mean, I spent half my life building things, whatever it was video games or different Lego or robotics or other contraptions. I spent a lot of time messing around with that kind of stuff. And the other half was in the mountains climbing mountains, surfing, mountain biking, all this kind of stuff. And I called Wayve Wayve because I want riding a car to feel as good as surfing a wave. But I really think, you know, when it comes to pushing forward frontier technology, it's all about an adventure. It's all about doing something where you have to plan, you have to risk manage, you have to be ambitious, push boundaries. And I think there's a bit of that that I got from those experiences and those values. That we can take forward today because what we've been doing, it's one of those ones where when we started, this wasn't possible, but I thought and hoped and it's turned out that the future building this technology would become possible. And I think we've been lucky that so many things have come together that make that a reality, but it really did seem like optimizing for the adventure. I was fortunate to get a scholarship to take me from New Zealand to Cambridge University. Turning up to that place was like a castle from the Middle Ages full of the most amazing brilliant minds to just learn and explore. Again, like optimizing for adventures is somewhat led me to where we are today.
好了,你得跟我讲讲构建第一个端到端原型的事。我猜用你们那种方法解决问题会特别难,因为你没法作弊,对吧?如果你要做端到端机器学习的话。跟我说说那段经历,以及你们是怎么让它成功的。
Okay, you got to tell me about building the first end-to-end prototype because I would imagine that would be particularly hard with the way that you're approaching the problem. You can't really cheat, right? If you're trying to do end-to-end machine learning. Tell me about the experience of that and how you made it work.
2017 年,我刚完成博士论文,就能为不同的感知任务构建第一个深度学习模型:定位、立体视觉、语义分割、不确定性估计,所有这些系统都能将高维、数百万维的图像转化为压缩的世界表征,实现端到端学习。与此同时,DeepMind 的朋友们刚刚做出了 AlphaGo。这太惊人了。围棋的状态比宇宙中的原子还多,极其复杂。他们展示了通过自我对弈,你可以在模拟中学习,训练出一个能击败世界冠军的智能体。所以我看到这些后想:“好了,我们现在可以通过计算机视觉理解真实世界的复杂性。如果你通过模拟器拥有无限数据,在非常低维的状态下,你可以解决最难的围棋游戏。现在是时候把这些结合起来,带入物理世界了。”我一直对物理移动、冒险和机器人技术充满热情。所以我想我们可以去尝试。我们筹集了 150 万美元,召集了一些朋友,租了一栋房子,把车停在车库里,开始动手。
So, in 2017, I just finished my PhD thesis and I was able to build the first deep learning model for different perception tasks: localization, stereo vision, semantic segmentation, uncertainty estimation, all of these systems that could take high-dimensional, multi-million dimensional images and turn them into compressed representations of the world, learning end-to-end. At the same time, friends down the road at DeepMind had just produced AlphaGo. That's amazing. That's a system where there are more states of the game of Go than there are atoms in the universe. Extraordinarily complex. And they showed through self-play you could learn in simulation, you could learn an agent that could beat the world champion. So, I saw these things and I thought, "Okay, we can now understand the complexity of the real world through computer vision. If you have unlimited data through a simulator in a very low-dimensional state, you can solve the hardest game there is, Go. Now it's time to bring these together into the physical world," which I've always been excited about: physical mobility, adventure, and robotics. So I thought we could go try it. What we did is we raised 1.5 million dollars, got some friends together, rented a house, put our car in the garage, and started hacking away.
我们尝试的第一件事是在策略强化学习:把车放在路上,说:“好了,开吧,奖励是尽可能开远,无需任何人工干预。”奖励就是行驶距离。车子开始随机驾驶。经过几个月,我们开发了一个 Q-learning 算法,本质上允许你在它每次试图驶离道路时干预并抓住方向盘。重大突破是,仅仅 10 次干预,我们就让它学会了车道保持。我记得那是一个周末,我拼命努力,但什么都不管用。这类算法往往十次只有一次能成功。
The first thing we tried was on-policy reinforcement learning, where we put the car on a road and said, "Okay, go drive with the reward of drive as far as you can without any human intervention." Your reward is distance traveled. The car starts driving randomly. Over a period of months, we developed a Q-learning algorithm that essentially lets you intervene and grab the wheel every time it tries to drive off the road. The big breakthrough was the moment where with just 10 bits of intervention, we got it to learn how to lane follow. I remember it was on a weekend where I was pushing really hard, nothing was working. These algorithms tended to work one out of 10 times.
等等,让我确认一下我理解得对不对。所以你只是干预了,它对世界没有任何先验信息,只是看着摄像头。10 次干预之后,它就明白你想让它沿着车道行驶,并且能做到?
Wait, let me make sure I understand. So, you only intervened, it had no prior information about the world. It's just looking at a camera. And after 10 interventions, it understands that you wanted to follow lanes and it can follow lanes?
没错。输入是一个前置摄像头。车上有一块 GPU 做在策略优化。本质上就是试错,大概 10 次不同的试验,然后它就学会了策略。YouTube 上有个视频,我们静音了,因为音频里是我在它开始行驶时欢呼雀跃,但它确实在没有任何先验知识的情况下,仅凭 10 次反馈就相当稳健地学会了车道保持。
That's right. The input was a front camera. There was a GPU on board doing on-policy optimization. It was essentially trial and error, a few different 10 trials, and then it was able to learn a policy. There's a video on YouTube that we muted the audio because the audio is me whooping and cheering as it starts to drive, but it was essentially able to lane follow quite robustly after just 10 bits of feedback with no prior knowledge of the world.
你第一次上车时它没有任何先验信息,你害怕吗?比如你们真的在路上?
Were you afraid when you got in the car the first time and it had no prior information? Like were you actually on a road?
它开得很慢,所以你有时间在它驶离道路前干预。但这引出了一个好问题:那种方法显然无法规模化。你不能在全球范围内大规模进行在策略测试,那太不安全了。所以我们很快转向了模仿学习和离线强化学习范式,并开始扩展世界模型,这样我们就可以通过“做梦”来学习,通过经验来理解。后来在 2018 年,我们开始朝这个方向发展。这种增长把我们带到了今天,我们可以让任何车辆在世界任何地方行驶。
It was going slowly, so you had time to intervene before it drives off the road. But that brings up a good question: that method obviously doesn't scale. You can't do on-policy tests at scale around the world. That's pretty unsafe. So we quickly moved to an imitation and offline reinforcement learning paradigm and started scaling up world models so we could learn through dreaming and understanding through their experience. Later in 2018, that's where we started to take things. And growth there has brought us to this point today where we can drive any vehicle anywhere all around the world.
那么,在实现车道保持之后发生了什么?下一步是什么?
So, what happened after you got the lane following? What was kind of the next step?
下一步是开始规模化。我们尝试了很多不同的东西。我们尝试注入计算机视觉的先验知识。我们尝试构建世界模型,在“做梦”和想象状态中学习。我们尝试了模仿学习。我们尝试了 sim-to-real 学习,从模拟转移到现实世界。我们几乎什么都试了。我们写了所有这些博客文章,最终在所有这些不同领域都成为了行业首创。然后我们确定了一个似乎有效、似乎可扩展的方案,并开始大力推进。我们花了几年时间推进基础设施、车队建设,并真正把这项技术提升到可以在伦敦行驶的程度。回想起来,在伦敦学习很棒,因为它迫使我们构建能够应对普通伦敦街道混乱和复杂性的东西。任何时候你周围都有上百个行人和骑自行车的人。它迫使我们构建不依赖地图基础设施之类的东西,而是能够真正在一个有 2000 年历史的城市中行驶的系统。
Next step was starting to scale it up. We tried so many different things. We tried injecting priors around computer vision. We tried building world models and doing learning in dreaming and in imagination states. We tried imitation learning. We tried sim-to-real learning, simulation transfer to the real world. We just sort of tried everything. We wrote all these blog posts that ended up being industry first in all those different areas. Then we settled on a recipe that seemed to work, seemed to scale, and started really pushing. We had a couple of years of pushing infrastructure, fleet buildout, and actually bringing up this technology to a point where it could start to drive in London. Retrospectively, learning in London was great because it forced us to build something that could deal with the chaos and complexity of your average London street. You have like a hundred pedestrians and cyclists around you at any one time. It forced us to build something that didn't rely on mapping infrastructure or anything like that, but could actually drive in a city that's 2,000 years old.
我记得 10 年前你刚开始的时候,车辆领域非常拥挤,有很多不同的方法,大量风投资金涌入。现在似乎已经整合了,但似乎也确实在发挥作用。我一直在我的特斯拉上使用自动驾驶。我每天都坐 Waymo。我觉得它已经悄悄崛起,成为一项至少我们在旧金山一直在使用的技术。我很好奇你如何看待目前的竞争。它正在变成一种商品吗?Wayve 能做什么,或者你们的方法有什么不同?
I remember 10 years ago when you were getting started, the vehicle space was super crowded with lots of different approaches, lots of VC money flowing in. It seems like it's consolidated at this point, but also it seems like it's really working. I use autopilot in my Tesla all the time. I ride Waymos every day. I feel like it's kind of quietly snuck up and become a technology that at least here in San Francisco we use all the time. I'm curious how you think about the competition at this point. Is it becoming a commodity? Is there something different about what Wayve can do or how you approach this?
这是个好问题。我认为总的来说,我们在 2017 年起步时,那可能是炒作周期的顶峰。筹集 150 万美元去挑战大约 10 家刚刚宣布了数十亿美元雄心和交易的公司,有点不平衡。但可以说我们经历了幻灭的低谷。我认为很大程度上是因为巨大的开销,构建那种 AV 1.0 范式(地图、规则、基础设施)并使其成熟需要超过千亿美元。出现了很多整合,实际上只有一两家公司有资本意愿去推动并使其成功。与此同时,特斯拉 FSD 或我们自己成熟了 AV 2.0 方法,一个更实惠、可扩展的下一代技术栈,并且已经从 5 年前的演示发展到如今真正准备好成为全球汽车产品。我认为今天最令人兴奋的是,我们也看到了产品的巨大验证。你和其他许多人现在都在购买 FSD 并为之付费,而且它确实受到很多客户的喜爱。
That's a good question. I think broadly speaking, when we started in 2017, that was probably the peak of the hype cycle. Raising 1.5 million dollars to go tackle what was like 10 companies that just announced billion-dollar ambitions and deals was a bit lopsided. But it's fair to say we saw a trough of disillusionment. I think largely that was from the sheer expense, the hundred billion-plus dollars required to build that AV 1.0 paradigm of mapping, rules, infrastructure, and to get that to maturity. There was a lot of consolidation where really only one or two companies had the appetite of capital to push through and make that work. At the same time, Tesla FSD or ourselves matured an AV 2.0 approach, a next-generation stack that's more affordable, scalable, and has now gone from a demo we had 5 years ago to a point where it's actually ready for global automotive product. I think the most exciting thing today is we're also seeing huge validation of the product. You and many others are now buying FSD and paying for it, and it's actually loved by a lot of customers.
当你去上海和旧金山这样的地方,人们愿意付费,甚至比人类出租车体验支付更多费用来乘坐无人驾驶出租车,尽管等待时间更长、受地理围栏限制,但这是一个人们真正喜爱的产品。所以现在的问题是,如何将这项技术带到任何地方的任何车辆上?我认为这就是我们构建并能够提供的阶跃性变化。从昂贵的改装车辆转向大众市场车辆,这些车辆每辆售价 3 万、4 万、5 万美元,内置全球供应链中的硬件,不需要高精地图,因此可以在任何地方行驶,并且可以灵活地用于网约车车队或私人拥有。顺便说一句,这种雄心壮志和商业模式之所以可行,完全是因为我们构建了这种可泛化的 AI 驾驶员。但现在的问题是,如何将其推向全球规模?这就是我们希望引领行业实现的转变。我很高兴现在看到了市场条件:人们正在付费,这作为产品获得了真正的商业验证。监管已经到位,不仅在美国,我们还与英国、欧洲和日本密切合作。我们共同主持了联合国关于自动驾驶系统 D-CAST 法规的委员会,该法规将于今年晚些时候通过并实际合法化这些产品。当然,制造商现在正在以数百万辆的规模生产车辆,这些车辆拥有集中式计算、环绕传感器,以及通过无线更新和获取数据的能力,这使得车队学习产品成为可能。所有这些因素加在一起,让我对未来一年经历大规模商业拐点充满乐观。
When you go to Shanghai and San Francisco, places like this, people will pay, even pay more than a human taxi experience for a robotaxi, even though it has longer wait times, it's limited by a geofence, and it's a product that people really love. So now the question is, how do you bring this to any vehicle anywhere? I think that's the step change that we've built and can offer. Taking it from expensive retrofit vehicles to mass-market vehicles that you can buy or manufacture for 30, 40, 50,000 dollars each, that have inbuilt hardware in global supply chains, don't need an HD map, so can drive anywhere, and have the flexibility to be in a fleet for ride-hailing or in a car you own. By the way, this ambition, this business model is only available because we've built this generalizable AI driver. But now the question is, how do you bring it to global scale? That's the shift we want to take the industry through. I'm thrilled that we're now seeing the conditions in the market. People are paying, and there's real commercial validation of this as a product. Regulation is in place, not just in the US, but we work closely with the UK, Europe, and Japan. We co-chair the UN committee on D-CAST regulation for autonomous systems, which later this year will come through and actually legalize these products. And of course, manufacturers are now building vehicles at millions of scale volume that have centralized compute, surround sensors, and the ability to over-the-air update and get data off them, which makes a fleet learning product possible. All these factors together give me a lot of optimism for going through a massive commercial inflection point in the coming year.
有趣的是,你在那个列表中没有提到改进算法。而且当你坐在车里时,很难判断算法的质量或自动驾驶汽车的安全性。但我必须说,当我坐在你的车里时——那是大约两年前——它真的感觉已经准备好了。我觉得它比旧金山街上的 Cruise 汽车更安全,虽然我不知道是否应该相信自己的判断。但算法方面是否还有工作要做,才能让你对全球部署感到满意?还是说这不再是挑战,你已经转向物流和监管等问题了?
It's kind of interesting you didn't mention in that list improving the algorithm. And it's very hard to judge the quality of an algorithm or the safety of a self-driving car when you're in it. But I have to say when I was inside your car, and this is now I think 2 years ago, it really felt ready for primetime. I think it felt safer than the Cruise cars that were on the street in San Francisco, which I don't know if I should trust my own judgment or not. But are there still things to do with the algorithm to get it to the point where you feel good about using it globally, or is that no longer the challenge and you've kind of moved on to logistics and regulation and things like that?
当然,还有工作要做。我们今天还没有无人驾驶服务,但这不再是关键路径。我们看到,来自我们前沿具身 AI 科学团队的创新质量,加上我们拥有的数据和算力增长,正在推动复合规模的扩大。再加上与合作伙伴在车辆上的更深层次集成,你就能在车辆上获得感知和计算。这些因素相互叠加,性能不再是关键路径,但我们在这方面仍有工作要做。我们将继续推动其进入通用化状态。我认为一个大的挑战是评估和验证。如何向自己、监管机构和消费者证明性能水平足够?当然,还有商业部署和建立体系。过去几年对我们公司来说是增长的大年,我们建立了团队实力,与德国、日本和美国等主要市场的监管机构、汽车制造商和车队合作。这是为了推动集成。汽车行业有多年销售周期,现在将是这些系统的部署和启动。所以我认为这就是关键路径所在,但我们在性能方面不会放松油门,显然在未来的机会面前,性能仍有数量级的提升空间。
Absolutely, there are still things to do. We don't yet have a driverless service today, but this is no longer the critical path. We are seeing that the quality of innovation coming from our frontier embodied AI science team, plus the data and compute growth we have, is just driving compounding scale. Then you layer on deeper integration into vehicles with our partners, and you get sensing and compute on the vehicles. These things compound, and performance is no longer the critical path, but we still have work to do there. We'll keep pushing that into a generalized state. I think a large challenge is evaluation and validation. How do you prove the level of performance is sufficient to yourselves, to regulators, to consumers? And then of course, commercially deploying and setting up. The last couple of years for us has been a big year of growth for the company to build team strength, partner with regulators and automotive and fleets in major markets like Germany, Japan, and the US. It's to get the integration going. It's the multi-year sales cycle of automotive, and now it's going to be the deployment and bring-up of these systems. So I think that's where the critical path is, but we're not taking the foot off the accelerator in terms of performance, which clearly has orders of magnitude still to go in terms of opportunity ahead.
我看过的最喜欢的演示之一,是你团队中的某个人向我展示了将感知模型接入 LLM,这样你就能看到汽车在行驶时的想法,这看起来既令人惊讶又令人愉悦。那是一个你意识到不需要的玩具,还是你继续投资的东西?
One of my favorite demos I've ever seen was someone on your team who showed me plugging your perception model into an LLM so that you could actually see what the car was thinking as it was driving, which seemed surprising and delightful. Was that a toy thing that you realized you didn't need, or is that something that you've continued to invest in?
你知道吗?在 2021 年或 2022 年,远在 ChatGPT 之前,我团队中的一位成员 Vijay 来找我说:“嘿 Alex,我想开始研究驾驶中的语言。”我的第一反应是:“不行,老兄。我们要保持专注。我们还没发布过产品。拜托,专注于驾驶。”他阐述了论点:从文本数据中可以获得巨大的知识,并且可能产生新的有趣产品和交互模式。实际上,通过引入语言,我们可以改进表示、性能和产品体验。所以我们说:“好吧,让我们开始探索这个。”很快我们建立了一个原型,它并不是简单地接入 LLM,因为那时 LLM 还没有那么大规模。它更像是一个经过训练的视觉-语言-动作模型。我们训练了一个单一模型,可以观察世界、驾驶汽车并理解语言。它一开始是个噱头——你可以开着车,让它解释自己的驾驶行为,比如“我正在绕过一辆双排停放的公交车。我正在为红灯停车。”然后我们就解释它。你可以挑选出一些非常令人愉快的事情,但同样也有一些事情它只是在幻觉和出错。从那时起,我们开始训练它。我们实际上让它变得非常稳健。我看到了一个有趣的现象:如果你开车,对面来了一辆车,你想在对面车前面转弯进入一条小巷,如果那辆对面车停下来并朝你闪灯,这是让你先走的信号,即使你没有路权。我们的 AI 不会这样做。它会停下来让行,只是等着那辆车,因为它有路权。当我们开始将语言预训练引入 VLA 模型时,它实际上表现出了这种行为。当那辆车开始闪灯时,它现在会转弯了。我第一次看到这个。这是一个例子,说明通过将语言引入系统,我们实际上可以改进表示和推理能力。今天,它已成为我们训练的核心部分。它提升了性能,并开辟了新的可解释性能力。
Yeah, you know what? In 2021 or 2022, way before ChatGPT, one of my team members, Vijay, came to me and said, 'Hey Alex, I want to start working on language for driving.' My first reaction was, 'No way, man. We're going to stay focused. We've never shipped a product. Come on, stay focused on driving.' He laid out the argument that the knowledge you can get from text data is enormous and might produce new interesting products and interaction modalities. Actually, by bringing in language, we can improve the representation, the performance, and the product experience. So we said, 'Okay, let's start exploring this.' Very quickly we got up a prototype that wasn't plugging together an LLM because LLMs didn't really exist at that large scale yet. It was more of a trained vision-language-action model. We trained a single model that could see the world, drive a car, and understand language. It started off as a gimmick—you could drive along and have it explain its driving, like 'I'm going around a double-parked bus. I'm stopping for a red light.' And we'd just explain it. You could cherry-pick some really delightful things, but then again, there were also some things where it was just hallucinating and wrong. From that point, we started to train it up. We actually got it to a point where it would become very robust. One interesting thing I saw: if you drive and you have an oncoming car and you want to make a turn across that oncoming car to go into a side street, if that oncoming car stops and flashes its lights at you, it's a signal for you to go in front, even though you don't have right of way. Our AI wouldn't do that. It would stop and yield and just wait for that car because it had right of way. When we started bringing language pre-training into the VLA model, it actually demonstrated that behavior. When the car started flashing its lights, it would now turn. I saw that for the first time. That was an example of how we could actually improve the representation and reasoning capabilities by bringing language into the system. Today, it's a core part of our training. It gives us a boost in performance and is opening up new interpretability capabilities.
当然,这也开启了车辆个性化与交互的产品体验。现在你可以坐进我们的车,让它以不同的驾驶风格行驶。我觉得这会是一种很酷的体验,因为当你使用高级驾驶辅助系统或自动驾驶出租车时,如果车开得太保守,你会非常沮丧;如果开得太快,你又会吓到。不同的人对此有非常不同的期望。不同品牌也是如此——有些品牌想走运动风格,有些则想强调安全可靠。当然,不是说运动风格就不安全可靠,但你懂我的意思。所以,提供这种个性化是这项工作的一个很好的成果。
And of course, it opens up product experiences of personalization and interaction with the vehicle. You can now get in our car and prompt it to drive in different driving styles. I think that'll be a neat experience because when you get into an ADAS or a robotaxi, if that car is driving too conservatively, you're going to get really frustrated. Or if it's driving too fast for you, you're going to get freaked out. Different people have quite a wide variety of expectations there. Or different brands. You can have some brands that want to be sporty or some that want to be safe and reliable. And not saying the sporty won't be safe and reliable, but you know what I mean. So providing that personalization is quite a nice outcome of this work.
对于自动驾驶来说,现在是个有趣的时刻,至少在旧金山,突然到处都是 Waymo。我很喜欢,觉得这是一种很棒的体验,但我们也看到了各种疯狂的交通问题。而且,人们可能出于好玩或恶意,以你意想不到的方式去干扰它们。所以,这种能力似乎对于让自动驾驶汽车成为人们信任的系统至关重要。
It's a funny moment for self-driving, at least in San Francisco, where we suddenly have Waymos everywhere. I love it. I think it's an amazing experience, but we're also witnessing all kinds of crazy traffic issues being caused. And also, people, maybe for fun or maybe maliciously, messing with them in ways you wouldn't have expected. So it does seem like that capability might actually be completely critical to making a self-driving car a system people trust.
是的,不过有几点。首先,我们的车,随着时间的推移,你不会注意到任何不同,因为它们就是普通的车,只是内置了传感器。所以我希望人们不会对它们有不同于其他交通参与者的行为。其次,我们的系统设计成以非常像人类的方式驾驶。你看到它在繁忙的城市中像自动驾驶出租车一样接客送客,或者在人群中穿梭,做那些非常人性化的事情,而不是卡住不动。我认为这些真的会提高公众的接受度。我们设计系统是为了在我们生活的混乱城市中运行,而不需要基础设施行为改变。我认为这对于具身 AI 的采用非常重要。你需要使用大众市场的硬件,需要能够利用离策略数据的可用性,并且需要能够作为现有基础设施的即插即用组件运行,而不需要大量的资本交换。我认为如果你把这些做对了,那么具身 AI 就会是一种令人愉快的部署体验。
Yes, although a couple of things. Firstly, our cars, over time you won't notice anything different because they're just normal cars with built-in sensors. So I hope that people don't behave differently around them compared to other traffic. And secondly, our system is designed to drive in a very human-like way. The way you see it picking up and dropping off as a robotaxi experience in busy cities, or the way you see it nudging through crowds of pedestrians and doing things that are very human-like and not just being stuck. I think these kinds of things are really going to improve public acceptance. We design the system to operate in the messy cities we live in and not require infrastructural behavior change. I think that's really important for adoption of embodied AI. You need to use mass-market hardware. You need to be able to leverage availability of off-policy data, and you need to be able to operate as a drop-in to existing infrastructure, not requiring massive capital exchanges. I think if you get these things right, then embodied AI is a delightful deployment experience.
你是如何决定不自己造车,而是将技术授权给其他汽车公司的?
How did you make the decision to not build your own car and to license your technology to other car companies?
这一点从一开始我们就非常清楚。我认为我以及我们公司一直以来的心态是,自动驾驶是一个 AI 问题。我们想先解决最难的问题。所以我们开始研究 AI,并积极在云、汽车、基础设施、保险等方面建立合作,与周围最优秀的公司合作,这样我们就能专注于关键路径问题。通过始终先解决最难的问题,我们不会在前进道路上撞上某种天花板。所以我认为是这种心态。但回过头来看,这变得很重要,因为我认为专注于一种车辆形态是个错误,因为产品市场契合度究竟如何还不清楚。是像今天这样的普通 SUV?还是双向车辆,人们面对面坐着?还是两座电动车?或者还有其他真正起飞的应用?我认为这还不清楚,但归根结底,如果你只专注于一种形态,你也只能占据一小部分市场。所以我们试图专注于如何构建一个通用的驾驶系统,能够应对一切,然后跟随市场在任何阶段最先进的方向。我们从改装起步,当时还没有什么可以集成的东西。然后在疫情期间,杂货配送和车队利润开始飙升,所以我们开始与该领域合作。然后情况变了。但后来,那些从不与我们交谈、说这不安全、永远行不通的汽车公司,突然开始与我们对话,并构建软件定义的车辆以便我们集成。而今天,我们有了自动驾驶出租车的热潮。所以我们基本上是跟随市场,但对集成到任何东西都持开放态度。最终状态当然是消费车辆、配送、自动驾驶出租车、卡车运输、非汽车形式的机器人。我们希望授权我们的具身 AI,让所有这些都变得智能、安全且可行。
That was very clear to us from the beginning. I think the mindset I've always had, and we've always had as a company, is that self-driving is an AI problem. We want to work on the hardest problem first. So we started working on AI and aggressively partnering on cloud, on car, on infrastructure, on insurance, and working with the best companies around us so that we can focus on the critical path problem. By always focusing on the hardest problem first, we're not going to stumble into some glass ceiling down the road. So I think it was that mindset. But retrospectively, it's become important because I think it's a mistake to focus on one vehicle form factor because it's really unclear what will have product-market fit. Is it a normal SUV like today? Is it a bidirectional vehicle where you sit facing each other? Is it a two-seater electric vehicle? Or are there other applications that really take off? I think that's unclear, but at the end of the day, if you focus on one form factor, you're also going to have a small part of the market. So we've tried to focus on how to build a generalizable driver that can address everything and then follow where the market is most advanced at any one stage. We started off with retrofit when there wasn't really anything for us to integrate into. Then during COVID, grocery and fleet started skyrocketing in profits, so we started partnering with that sector. Then things changed. But then automotive, who would never speak to us, who said this is unsafe, it would never work, all of a sudden started talking to us and building software-defined vehicles so we can integrate with them. And now today, we've got the excitement of robotaxis. So we've kind of followed the market, but been open to integrate into anything. And the end state, of course, is consumer vehicles, delivery, robotaxi, trucking, non-automotive forms of robotics. We want to license our embodied AI and make all of them intelligent, safe, and possible.
作为 CEO,你现在对这些不同模式有没有一个优先级排序?哪个是你最关注的?
Where you sit right now as CEO, do you have like a stack ranking of those different modalities and which one's most top of mind for you?
是的,今天一切都围绕着消费车辆和自动驾驶出租车。重点是在城市和高速场景下,使用大众市场的消费车辆硬件进行点到点驾驶。这是一个与优秀合作伙伴一起商业化的对齐产品。这是我们今天真正的重点,能够产生非凡的业务。我的意思是,看看数据、收入、车辆数量、曝光度的规模。每年大约生产 1 亿辆汽车。人们常常忘记这一点。他们只想到自动驾驶出租车,但世界上只有不到 1 万辆自动驾驶出租车。所以,将这项技术原生集成到消费车辆中,打开了数千万辆汽车的机会,这将在短期和中期内远远超过自动驾驶出租车。当然,长期来看,每辆车都能实现无人驾驶操作,这显然是我们要达到的稳定状态,考虑到它能带来的安全性、便利性和产品优势。
Yeah, today it's all about consumer vehicles and robotaxis. It's about point-to-point driving in urban and highway situations with mass-market consumer vehicle hardware. So this is an aligned product that we can now commercialize with great partners. This is the real focus for us today that can generate extraordinary business. I mean, you look at the scale of data, revenue, number of vehicles, exposure. About 100 million cars produced each year. People often forget that. They only think of robotaxis, but there are less than 10,000 robotaxis in the world. So getting this natively integrated into consumer vehicles opens up tens of millions of vehicle opportunity that's going to far outscale robotaxis in the short and medium term. Before long term, of course, every vehicle is capable of driverless operation, which is clearly the steady state we're going to go to, given the safety, convenience, and product benefits it can bring.
你计划首先在哪些地区推出?
What geographies do you plan to launch in first?
我们确实希望系统是全球性的,但今天我们的重点是欧洲(包括英国)、日本和北美。我们在伦敦、斯图加特、东京和湾区设有车队和办公室。我们已经在欧洲和北美的大部分城市进行了测试,并在日本不断扩展。考虑到我们的客户,这些是我们起步的市场。
We're really looking to make sure the system is global, but we're focused today on Europe, including the UK, as well as Japan and North America. So we have fleets and offices in London, Stuttgart, Tokyo, and the Bay Area. We've driven around most cities across Europe and North America and are growing in Japan. Those are the markets we're starting with, given the customers we work with.
你感觉哪里会最先实现大规模市场?
Do you have a sense of where you will go mass market first?
就是这些市场。我认为总的来说,我们看到北美和日本进展更快,欧洲紧随其后。但当然,中国不幸对我们来说很难进入,反之亦然。我不想涉及地缘政治,但世界处于这种状态有点遗憾。不过在中国,这项技术发展非常迅速,创新速度令人印象深刻。
It's those markets. I think in general, we're seeing faster movement in North America and Japan, closely followed by Europe. But of course, China is unfortunately very hard for us to work in and vice versa. I won't get into geopolitics, but it's a bit of a shame that's where the world's at. But in China, this technology is moving so quickly and the rate of innovation is very impressive.
嗯,所以我认为从汽车技术的角度来看,这值得关注。另一个证据是,当这项技术被部署时,消费者真的很喜欢它。现在,它和价格、安全性一起,成为中国消费者购车的前三大理由。哇。这就是自动驾驶。那么,去年你在 AI 500 路演中取得了惊人的成果。我不知道你是否想谈谈,但你用一个单一的 AI 模型部署到了很多地方,而且其中许多地区之前没有任何训练数据。你当时对它的出色表现感到惊讶吗?
Um and so, I think it's worth looking at if you haven't from an automotive technology perspective. Another proof point that when this technology is deployed, consumers really love it. And it's now alongside price and safety, it's now a top three reason why people buy a car in China now. Wow. There's autonomy. So, you had amazing results at the AI 500 road show last year. I don't know if you want to talk about it, but you took one single AI model deployed all over the place with I think many of those geographies actually not having any prior training data on it. Were you surprised that it worked so well?
不惊讶,因为核心论点就是通过构建通用 AI 模型,我们可以开到任何地方。但实际看到它确实是一个巨大的验证。嗯,所以我们环游了世界。有一些很酷的轶事。比如,我们在巴黎的凯旋门附近开过。你开过那里吗?
Uh it's no, because that was the core thesis that by building a general-purpose AI model, we could drive anywhere. But to actually see it was huge validation. Um so yeah, we drove around the world. Some very cool anecdotes. I mean, we drove in places like around the Arc de Triomphe in Paris. Have you driven around there?
我去过,是的。嗯,不是开车,但我去过那里,是的。那里很混乱,对吧?
I have, yeah. Well, not driven, but I've gone there, yeah. It's chaos, right?
嗯,是的,我们去了芬兰和瑞典的北极圈以北,在 22 小时的黑暗、雪地等条件下测试驾驶。嗯,我们甚至在台风期间在东京,当时当地的火车和巴士都停运了。实际上,我们就是用这种方式进行了一堆媒体试驾。我们进行了一整天的记者试驾,没有脱离,全天完美自动驾驶,在我见过的最大的雨中。嗯,我们的摄像头就位于挡风玻璃后面,前置摄像头。雨刷不停地刮水。嗯,但它在东京市中心可靠地行驶,这非常了不起。那天路上的行人少了,但交通仍然繁忙。嗯,是的,我们看到了所有这些情况,这证明了我们是第一家做到这种规模的公司。它证明了我们确实可以泛化到任何地方的任何车辆,因为超过一半的城市我们之前没有训练数据。所以,这是一次真正的零样本泛化测试。
Uh but yeah, we went north of the Arctic Circle in Finland and Sweden and tested driving in 22-hour darkness, and snow, and things like that. Um we were even in Tokyo during a typhoon, where local trains and buses shut down. And actually, that's how we were giving a bunch of media drives. We had a day long of journalist drives, no disengagements, perfect autonomy the whole day, in literally the heaviest rain I've ever seen. Um and our cameras are situated just behind the windshield, the front camera. Windshield wiper was just constantly wiping water off it. Uh but it drove reliably around central Tokyo, which was quite remarkable. Less pedestrians on the road that day, but still busy traffic. But um yeah, we've seen all of these things in that, and it demonstrated as the first company to do something of that scale. And it demonstrated that we truly can generalize to any vehicle anywhere given that over half of these cities we had no prior training data in. And so, it was true zero-shot generalization test at scale.
嗯。这听起来几乎像是超人级别的表现。
Mhm. That almost sounds like superhuman-level performance.
是的,我会谨慎使用这个词。我的意思是,你和我,如果我们跳到一个新城市,你可以租一辆车开过去。但要声称超人,我的意思是,有一些很好的报告,比如我读过瑞士再保险公司的一份报告,他们测量了 Waymo 自动驾驶系统在其运营历史中的影响,发现它们的事故和安全事件比同等的人类驾驶员少得多。所以,我认为有一些有力的证据表明自动驾驶可以超越人类。嗯,我们还需要努力来证明这一点,但我们正在取得良好进展。但你也看到,不幸的是,超过 99% 的事故是由人为错误造成的。所以,即使只是让自动驾驶体验永远不会分心、醉酒、受损,能够同时看到 360°,每秒做出超过 10 次决策,我的意思是,有机会将事故率降至接近零。嗯,这就是我们想看到的,但我认为你会看到大规模部署来真正证明这一点。我们希望在接下来的几年里实现这一目标。
Yeah, I'd be careful with that word. I mean, you and I, if we jump in a new city, you can get a rental car and go drive there. But to claim superhuman, I mean, there's some great reports that I read one from Swiss Re, for example, with the insurer, where they measured the impact of an autonomous system of Waymo operating of their operational history and found that they had much fewer accidents and safety events than the equivalent human drivers. So, I think there's some robust evidence that's really showing autonomy can be superhuman. Um we have work to do to demonstrate that, but we're making good progress there. But then you also look at sadly, over 99% of accidents are caused due to human error. And so, even just making a self-driving experience that can never be distracted, drunk, impaired, that can see 360° all at once, make decisions over 10 times a second, I mean, there is an opportunity there to drive accidents down to near zero. Um and that's what we want to see, but I think you're going to see scale deployment to really show that. And we want to go do that in the coming years.
目前还有让你感到紧张的情况吗?什么样的配置会让你处于模型能力的边缘?
Are there situations that still make you nervous at this moment? What would be the kind of configuration that would put you on the edge of what the model can do?
我认为很难给出笼统的说法,因为当今模型往往在多种故障模式叠加时出现问题。如果是黑暗、高速交通、有人逆行或不礼让的对抗行为。嗯,也许是恶劣天气。你知道,就是当你有多个混杂因素时。然后可能还有传感器故障。当你遇到多个混杂因素时,你往往会发现挑战和故障模式。所以,我们需要通过更好的系统弹性或冗余来持续解决这个问题。我们正在通过一些下一代车辆架构来构建这一点。我们需要更多数据来改进泛化和风险评估。嗯,是的,确保我们负责任地部署在我们确信满足安全要求的领域。嗯,但随着我们的发展,我们将与合作伙伴负责任地做到这一点。我认为,从监督试验开始,部署到消费车辆和机器人出租车中的另一个好处是,我们可以与社区和监管机构逐步推进,最终实现全球范围内每辆车都自动驾驶的终极状态。
I think it's very hard to give general statements to that because where the models today tend to struggle is when you have a culmination of many failure modes. If it's dark, you have high-speed traffic, you have adversarial behavior from someone else driving on the wrong side of the road or not yielding. Um maybe it's bad weather. You know, it's when you have multiple confounding things. Then maybe have a sensor failure. When you get multiple confounding factors, that's when you tend to find challenges and failure modes. And so, we need to keep addressing that with better system resiliency or redundancy. We're building that with some next-generation vehicle architectures. We need more data to improve the generalization and risk assessment. And um yeah, making sure we deploy responsibly in areas that we're confident meet the required bar for safety. Um but we'll do that responsibly with our partners as we grow. And I think the other nice thing about being deployed in consumer vehicles and robotaxis, starting with supervised trials, is that we can incrementally grow this with communities and regulators to the end state where every vehicle is autonomous at a global scale.
我记得在 AI 伦理方面有一个时刻,也许是在自动驾驶火热的时候,人们在问一些问题,也许是玩具问题,比如汽车应该优先考虑乘客的安全还是行人的安全。你知道,端到端的方式,你似乎在回避这些问题,也许你没有介入。这是真的吗,还是你需要以某种方式推动数据来获得你想要的某种伦理行为?
I remember there was a moment in sort of ethics in AI, maybe when AI was like when self-driving was hot, that people were asking questions, maybe toy questions about should a car prioritize the safety of the passenger versus a pedestrian. And you know, and end to end, you're kind of sidestepping those questions and maybe you're not inserting yourself. Is that true, or do you need to push the data in certain ways to get certain types of ethical behavior that you want?
所以,这就是电车难题。我们有两个糟糕的决定,你必须选择糟糕的决定 A 或糟糕的决定 B。有很多种框架方式。实际现实是,作为一个行业,我认为我代表大多数自动驾驶公司发言,我们需要设计一个系统,使其始终有第三个好的选择。这被称为最小风险机动,即如果你处于糟糕的场景中,你尝试并看到风险,总有一个你可以采取的第三个选项,那就是最小化风险。这通常涉及降低动能、保持可预测的行动路线,以及停车或靠边停车等。我们总是需要为此做好准备,如果任何系统失效,需要有失效操作能力来做到这一点。所以我们设计这些系统,使其始终有第三个好的选择,而不是被困在两个糟糕的选择之间。第二点是,要进入那种两个糟糕选择的状态,必须有太多事情出错才能达到那个点。这类场景的概率非常小。即使你处于那个位置,你实际上能够准确感知那两个选项的概率,这使得这更像是一个理论上的思维练习,而不是行业需要做出的实际实际决策。嗯,所以我们专注于,行业也专注于构建这些最小风险机动,使我们远离那些情况。话虽如此,如果你确实进入了那种不太可能的情况,那么行为当然取决于你如何从训练数据和设计的系统中进行偏置。
So, this is the trolley problem. We have two bad decisions, and you have to choose bad decision A or bad decision B. There's many ways of framing it. The practical reality is as an industry, and I think I speak on behalf of most AV companies for this, is that we need to design a system so there's always a third good option. It's called a minimal risk maneuver, which is if you're in a bad scenario, you try to and you see risk, there's always a third option you can take, which is to minimize risk. And that typically involves reducing kinetic energy, maintaining a predictable course of action, and stopping or pulling over the vehicle or something like this. And we always need to prepare to do that if any systems fail, there needs to be a fail operational ability to do that. And so we design these systems to always have a third good choice, not to be stuck between two bad choices. The second point is that to get into that two bad choice state, so many things have to go wrong to get to that point. And the probability of those kind of scenarios is so small. And even if you are in that position, the probability that you can actually accurately sense in those scenarios that those two things, it makes this more of a theoretical thought exercise than actual practical decision the industry needs to take. Um so we focus on and the industry focuses on building these minimal risk maneuvers that take us away from those situations. Having said that, if you do get into that unlikely situation where you're in there, then of course the behavior will depend on how you bias it from your training data and the system that you design.
嗯,所以当然,我们需要谨慎对待我们提炼到系统中的原则和知识,以及我们生成的行为。我的意思是,有一些好的原则,比如管理风险、让系统可预测,当然还有最大化安全性。但我想,在你的端到端训练中,你并不会看到很多碰撞事故。所以如果有碰撞,你可能是在某种梦境数据中生成了它们。所以我认为,你输入系统的碰撞类型会赋予它一定的视角。
Um and so there, of course, we need to be thoughtful with the principles and the knowledge that we distill into the system and the behavior that we generate. I mean, there are some good principles there around managing risk, making the system predictable, and of course, maximizing safety. But I would imagine that in your end-to-end training, you're not seeing a lot of crashes. So if there are crashes, you're probably generating them in the dream data somehow. So I think that the types of crashes that you're feeding into your system would give it a certain amount of point of view.
实际上,我为我们的安全记录感到非常自豪。你知道,自 2018 年开始运营以来,我们没有发生过任何事故。所以全球范围内的安全记录非常出色。本质上,这里有几点很重要。首先,我们不仅仅从自己车辆行驶的数据中学习。我们有庞大的数据合作伙伴关系,为我们提供相当通用的知识。不幸的是,今天仍然有很多道路事故,我们与行车记录仪提供商、从消费者车辆获取数据的汽车公司合作。可悲的是,他们仍然经历了很多相当可怕的场景。好的一面是,我们可以利用这些数据来学习。其次,我们拥有世界模型 Gaia,它可以模拟那些在现实世界中过于罕见或不安全而无法看到的场景,并通过合成数据和模拟确保我们对这些场景具有鲁棒性。所以无论是这些策略,最终,我们需要建立对这些事件的鲁棒性。我们仍然在世界各地的道路上看到一些相当奇怪的事情。不幸的是,一些相当不安全的行为。我们可以重新模拟、增强或修改这些场景来建立对它们的韧性。这是一个持续的车队学习和数据迭代的游戏,以确保我们随着时间的推移不断提升性能。
I'm really proud of our safety record, actually. You know, we've had no incidents since we started operating in 2018. So really fantastic safety record around the world. And essentially, there are a few things that are important here. Firstly, we don't just learn from the data of our vehicles driving. We have enormous data partnerships that give us quite general purpose knowledge. Unfortunately, there are still a lot of road accidents today, and we partner with dashcam providers, car companies that get data from their consumer vehicles. And sadly, there are still a lot of quite frightening scenarios that they experience. The good thing is we can use that data to learn from it. Then secondly, we have our world model Gaia that can simulate these scenarios that are too rare or unsafe to see in the real world, and make sure that we're robust to them through synthetic data and simulation, too. And so whether it's those strategies, ultimately, we need to build robustness to those events. And we still see some pretty bizarre stuff on the roads around the world. Unfortunately, some pretty unsafe behavior. And we can resimulate or augment or modify these scenes to build resiliency to them. And it's a constant game of fleet learning and data iteration to make sure we're grinding up performance over time.
所以我记得当特斯拉推出时,我相信他们说的是端到端梦想全自动驾驶。我觉得当那个版本出来时,你真的能感受到,至少在我的特斯拉上是这样。其中一个非常显著的特点是它真的在旧金山的每个停车标志前都停下来,而实际上旧金山的司机并不这么做。所以如果你真的严格遵守那条规则,你会有点烦人。这让我想到,实际上每个人的平均驾驶可能并不是你想要的。你在端到端训练中如何处理这个问题?
So I remember when Tesla launched, I believe what they said was end-to-end dream full self-driving. You really felt it, I think, when that version came out, at least in my Tesla. And one of the things that was really notable about it was it really stopped at every stop sign in San Francisco, which actually San Francisco drivers don't really do. So it's kind of you're kind of annoying if you're really following that rule to a T. And it made me think, you know, that actually the average of everybody's driving is probably, you know, not what you want. How do you deal with that in your end-to-end training?
这是个好问题。我上周刚去了日本,那里每个人都超速约 20%。这引发了一个关于政策、法规、责任的有趣问题。你知道,如果是一个脱眼或无人驾驶系统,系统承担驾驶责任,那么它应该遵守道路规则。但如果你是消费者,当今许多驾驶员辅助产品允许你将设定速度设置得高于限速。所以这是一个非常有趣的问题。我希望自动驾驶能够提高道路安全,以至于我们实际上可以提高限速,因为足够安全。但我们可能还有一段距离才能进行这类争论。是的,所以今天首先要意识到的是,当你训练世界模型时,你希望看到尽可能多样化的数据。你想看到好的驾驶、坏的驾驶。你想真正理解世界中动态事件的完整谱系。然后,当涉及到实际如何驾驶的策略学习时,当然,与客户偏好、OEM 驾驶风格行为、运营所在国的规则和政策相匹配的数据分布变得非常重要。所以我们可以分离这两个数据分布。第一个,你希望尽可能多样化、尽可能大。第二个可以是高度精选的小数据集,无论是 RLHF 风格的反馈,还是专家演示,有很多方法可以确保你以非常安全且考虑周全的方式驾驶该应用。所以你可以分离这些数据分布。
It's a great question. I was just in Japan last week, and there everyone drives about 20% above the speed limit. And it opens up an interesting question around policy, regulation, liability. You know, if you're a eyes-off or driverless system where the system's taking liability for driving, it should meet the road rules. But if you're a consumer, a lot of products today in a driver assistance setting allow you to set the set speed above the speed limit. And so it's really interesting question. I hope that self-driving improves road safety to the point where we can actually increase speed limits because it's safe enough to do so. But we're probably some distance away from being able to have those kind of arguments. Yeah, so today the first thing to realize is when you're training a world model, you want to see as diverse data as possible. You want to see good driving, bad driving. You want to really understand the full spectrum of dynamic events in the world. Then when it comes to actual policy learning of how you drive, of course, a data distribution that matches whichever customer of ours preferences, the OEM driving style behavior, the rules and policy of the country they're operating in becomes really important. And so we can separate those two data distributions. The first one, you want to be as diverse as possible, as large as possible. The second one could be a highly curated small, whether it's RLHF style feedback, whether it's expert demonstrations, there's many ways to do it to make sure that you drive in a very safe and considered way for that application. So you can separate these data distributions.
实际上,作为同行创业者,我想问你一件事:你最初是 CTO,与另一位联合创始人一起,后来他离开了,你成为了 CEO。我想知道那是什么样的经历。我猜那是旅程的早期,但如果你愿意分享的话。
Actually, one thing I wanted to ask you somewhat as a kind of fellow entrepreneur is you started I think as CTO with another co-founder who then left, and you became CEO. I wonder what that experience was like. I guess it was earlier in the journey, but if you want to share about that.
是的,那是早期阶段,我相信你记得,每个人都做一点所有事情。所以当我们开始时,有那种感觉,但出于一些原因,我的联合创始人离开公司去追求其他事情,这不幸但合理。我们今天仍然保持良好联系。但那次过渡是我第一次真正经历公司的变革管理沟通。我想当时我们大约有 25 到 30 人。刚完成 A 轮融资。但这也符合我过去十年的经历,每年都是全新的挑战。比如,我启动了代码仓库,编写了很多最初的演示。然后我开始招聘团队、寻找商业合作伙伴、扩大投资。你知道,现在管理客户,并通过执行团队进行管理,我非常兴奋能与一些真正处于行业顶尖的人一起工作。他们是行业传奇。无论是在科学、工程、产品、金融、商业等各个方面,我们都拥有真正世界级的领导者。所以我认为有几点很重要:建立互补团队的重要性,与你能学习的人一起工作,以及习惯于一旦我擅长某项技能,很快就超越它并面临新的挑战。一旦我擅长某件事,它就不再是我需要从事的相关技能,因为我通常是在为组织填补空白,然后组织成长吸收它,接着就有新的挑战。不过我很享受这一点。我认为这是建立公司冒险的一部分。我认为接下来的一年对我来说是一个新阶段,我开始频繁飞行,与全球客户合作。这将是我们第一个十年,也将是我们的第一个产品发布。
Yeah, it was early on where kind of stage, I'm sure you remember, where everyone does a bit of everything. And so as we're getting started, there was a bit of that sense, but yeah, for a number of reasons that unfortunately made sense for my co-founder to leave the company and pursue other things. And we're still in good touch today. But the transition was my first real experience of change management communications of the company. I think we were about 25 or 30 people at the time. Just raised series A. But also, it's I guess matched my experience over the last decade where every year has been a completely new challenge. Like I started the code repo and was coding a lot of the first demos. Then I started to work on hiring the team, finding commercial partners, growing investments. You know, now managing customers and managing through an executive team where I'm so thrilled to be working with some people that are truly at the top of their game. They're industry legends. You know, whether it's in science, engineering, product, people in finance, commercial across the board, we've got truly world-class leaders. And so I think it's probably a few things: the importance of building complementary teams, working with people that you can learn from, and also getting used to as soon as I get good at a skill, pretty quickly I outgrow it and have a new challenge. As soon as I get good at something, it's no longer the relevant skill for me to be working on because it's usually me plugging a gap for the organization and then growing the organization to absorb it, and then there's a new challenge. I enjoy that though. I think that's part of the adventure of building a company. And I think the next year for me is now this phase is new one where I'm starting to live on a plane, work with global customers. It's going to be our first 10 years in, and it's going to be our first product launch.
如果你能给刚开始创业的自己发一条 30 秒的建议,你会说什么?
If you could send a 30-second message of advice back to yourself when you're just starting out, what would you put in there?
天哪,我觉得如果让我重新打造 Waymo,我不知道。如果重来一次,你觉得能快多少?我完全同意,大概能快两三倍吧?
Oh man, I think I feel like if I built Waymo again, I don't know. If you built things again, how much faster do you think you could build second time over? I would totally. It's like two or three X, right?
嗯。
Yeah.
我走了太多捷径,或者犹豫不决的事情,还有我们犯过的错误,我觉得可以直接拿战术手册来用。
I did so many shortcuts or things that I was hesitant to doing, or mistakes that we made that I think I could just take get the tactical playbook.
而且你得在更高层面提炼,因为只有 30 秒,不能只讲损失函数配方、关系和合作伙伴,得是真正战术性的东西。
And you had to distill it in the higher level because you only have a 30-second window to not just give the loss function recipe and relationships and partnerships, like just really tactical things.
我一直觉得一个有趣的问题是,如果我们带着今天的知识回到 200 年前,会做什么?因为我不知道你通过权重和偏差学到了什么,但我们是站在巨人的肩膀上,而 200 年前这些巨人根本不存在。所以时机很重要。从很多方面看,如果我的产品开发速度快三倍,我们可能领先于行业,但汽车行业还没有合适的基础设施或产品来接受我们的产品。所以我认为时机就是一切。在很多方面,自动驾驶就是要在保持公司生存和资本化的同时,推进下一代方法。但如果要提炼的话,我一直对这个方法有信心。我不会说“继续前进,Alex”,因为我对这方面一直很有韧性。我觉得我在构建这件事上已经足够大胆和勇敢了。老实说,最大的收获是两件事。一是关于 Wayve 的人才成长,我看到了很多优秀的人、杰出的人、聪明的人,以及公司专业知识的增长。从伦敦的博士起步,到今天能与这么多优秀的人合作,我觉得有巨大的机会,我本可以更充分地利用这一点。
I always thought the interesting question is I feel like we'd be a bit useless if we went back in time like 200 years with today's knowledge, what would you do? Because I don't know what you've learned through weights and biases, but you've built on top of stuff that just we're standing on the shoulders of giants that just did not exist 200 years ago. And so there's a bit about timing, you know, in many ways our product if I developed three times as fast, we may have been ahead of the industry where automotive did not have the right infrastructure or products to actually accept our product. So I think timing is everything. And in many ways, self-driving has been a question of like staying alive and keeping the company capitalized while making progress on our next generation approach. But yeah, if I was to distill it down, you know, I've always had conviction in this approach. It's not like I'd say, you know, keep going, Alex, cuz I've been kind of resilient towards that one. So I think yeah, I think I've been fine with how bold and brave I've been in building this. Honestly, the biggest thing has been I think yeah, I think two things actually genuinely come to mind. One is about talent growth at Wayve, and I think I've seen so many great things about great people in the fantastic people in the brilliant people and growth of the expertise we have in the company that I think moving faster and harder and starting out of a PhD growing from London compared to the scale of the amazing people we can work with today. I think there's a great opportunity there that I could have lent more into.
所以你的意思是尽早招到专家?
So, do you mean like try to get the experts earlier? Is that what
是的,我想是的。我们一开始是通才思维,试图自己摸索,很多东西都得自己造。但我觉得这可能是一种复杂的感觉。对我来说,更重要的一点是,我在算法创新上投入过多,而在基础设施上投入不足。我从深度学习中学到的大教训是:1% 是算法,99% 是基础设施。无论是数据的清洁度、学习循环的可靠性和迭代速度,我们多年来一直是 Weights & Biases 的忠实用户,它带来的加速效果非常显著。无论是内省、测量还是评估工具。我们在这方面走过一段路,但投入不足。我希望我们更早地在这方面下更大功夫,而不是追求光鲜亮丽的算法创新。事实是,我们多年来一直被一个稳健的平台所拖累。一旦我们有了一个稳健的机器人平台,就像给盲人 AI 戴上了眼镜,它突然就能以惊人的方式看到和驾驶了。所以,更早、更努力地关注这一点,建立成熟度,对提升性能至关重要,而我们在这方面做得太晚了。
Yeah, yeah, I think so. You know, we started off with a generalist mindset really trying to figure things out and had to build a lot ourselves, but I think maybe it's a mixed feeling on that one, but I think the better of the more bigger thing that comes to mind for me is I think I over-indexed on algorithmic innovation compared to infrastructure. And like the big learning I have with deep learning is that it's 1% algorithms, 99% infra. Whether it's cleanliness of your data, reliability and iteration speed of your learning loop, we're very happy adopters of Weights & Biases over the years, the acceleration you get with it. Whether it's introspection, measurement, evaluation tools. I think we've been on a journey there and under-invested on it. I think going harder and faster there rather than the shiny sexy algorithmic innovation is something that I wish we'd done earlier. The thing is is we were held back by a robust platform for many years. And as soon as we had a robust robot platform, it's like giving our AI being blind and giving it glasses. All of a sudden it could see and drive in a remarkable way. So, focusing on that earlier and harder and building that maturity, I think was quite key to bringing up performance, and we were too late in that.
在仿真和评估方面,我认为大语言模型和具身 AI 在评估模型性能上都是一个真正的挑战。这是一个开放性问题,今天还没有一个明确的指标。所以我们在开发这些技术的过程中,不断尝试不同的东西,从生成模型到程序化游戏引擎,到 NeRF 和高斯泼溅,再到现在的生成式世界模型。在这些过程中,我们从不害怕放弃沉没成本。我们花了几百万美元、几个月时间构建的东西,一旦发现更好的,就直接删除,继续前进。这种文化和历程我认为很重要。今天,我们正处于一个重大突破的边缘,这将真正改变行业。但这也是一个例子,说明更早地做对这件事是鸡生蛋蛋生鸡的问题。如果你知道如何衡量,你就可以针对那个目标进行优化。但做对这件事,我们本应更快行动。
In simulation and evaluation, I think both large language models and embodied AI is a real challenge in evaluating the performance of models. It's a open-ended problem. It's not like you can have a clear metric around that today. And so developing these techniques we've been on a journey. We've consistently tried different things from generative models to procedural game engines to NeRF and Gaussian splatting to now generative world models. And throughout those times, we haven't been afraid to forego sunk costs. And you know, we spent millions of dollars, months of time building something and then just delete it, move on cuz we found something better. And so that culture and that journey I think has been important. And today I think we have something on the cusp of a major breakthrough there that will really transform the industry. But that's another example where getting that right earlier, it's chicken and egg. If you figure out how to measure it, then you can just optimize against that target. But getting that right is something that we should have moved faster on.
好的,最后几个快问快答。当你带比尔·盖茨坐你的自动驾驶车去吃炸鱼薯条时,你紧张吗?你们聊了什么?
Okay, some quick questions to end with. When you took Bill Gates for fish and chips in your self-driving car, were you nervous that something would happen? And what did you talk about?
是的,那是我们系统第一次足够可靠。实际上,那次行程中有几次人工干预。但我觉得比尔当时还是很兴奋的。我们聊到,有趣的是,当年他在商业化微软和 Windows 时,需要赢得 PC 厂商的设计订单。类似地,我们也需要赢得全球最大汽车公司的设计订单。所以有一些有趣的相似之处,他对此很感兴趣。他在伦敦待了一天,我们去吃炸鱼薯条。他的一个团队成员进去买了炸鱼薯条,从车窗递进来,然后上车。比尔把食物放下,坐在那里,不打算吃。我说:“来吧,比尔,我们得……”他说:“好吧,我们不能不吃炸鱼薯条。”然后他说:“那好吧。”他拿起一块鱼,掰成两半,就在车后座和我一起吃起了那块裹着面糊的鱼。我很喜欢这一幕。他显然对油腻的炸鱼薯条情有独钟。
Yeah, that was one of the first times we've got the system reliable enough. And actually, I think there were a few interventions on that drive. So, I think Bill was still excited at the time. We talked about It's interesting how when he was commercializing Microsoft and Windows, he had to get a bunch of design wins with PC OEMs. And in a similar way, we've had to get design wins with the biggest car companies in the world. So, there's some interesting similar dynamics there that he was interested in. And then he was in London for a day when we went to get fish and chips. One of his team went in to get the fish and chips, handed it through the window, got in the car, and he just put down and sat there and wasn't going to eat it. I said, "Come on, Bill, we've got to" He said, "So, we've got to Come on, we can't we can't not eat the fish and chips." So, he goes, "All right, then." Grabs a piece of fish in his hand, brings it up, breaks in half, and just starts eating this battered piece of fish in the back seat of the car with me, which I love. He clearly has a soft spot for some greasy fish and chips.
好吧,那当你租的房子里造出第一个原型时,你说你办了一生中最大的派对之一。那是什么样的?
All right, what about when you built the first prototype in your rented house and you said you threw one of the biggest parties of your life. What was that like?
我们邀请了剑桥所有认识的人来家里。客厅就是我们工作的地方,摆满了电脑。
We invited everyone we knew in Cambridge to the house. The living room was where we worked with a bunch of computers.
那个小卧室就是我们的会议室。另一个卧室里塞满了服务器,因为 60 块 GPU 在跑,温度经常高达 40 摄氏度。然后我们把房子变成了一个派对屋。团队里有几个音乐人,我们一起即兴演奏。那种氛围真的非常非常有趣。是啊,那些日子非常特别,我很珍惜那段大家一起并肩作战的回忆。但那确实是个大转变。不知道你怎么想,但我总觉得我一直在追寻早期那种所有人一起拼命干的感觉。那种感觉真好。我上周刚去参观了我们在日本的办公室,大概有 20-30 人。真的很有那种感觉。太令人兴奋了。
The small bedroom was our board room. We had another bedroom full of servers and was like constantly 40°C because of the 60 GPUs that were running there. And then we just turned the house into a bit of a party house. We had a few musicians in the team. We were jamming. And that was a really, really fun vibe. Yeah, those days are very special and I really enjoyed the lasting memories of a team that we were really in the trenches together during that time. But that was quite a big one. I don't know about you, but I feel like I'm always chasing that feeling of the early days when everyone's cranking together. Such a good feeling. I just finished visiting our Japanese office last week, which is about 20-30 people. Really feels like that. That's super exciting.
然后我认为在不同规模下运营各有利弊。你知道,今天我们全球大约有 1000 人。但你在那个规模下设定的文化确实推动了我们走到今天。我认为在银级规模下运营确实需要不同的技能,但你能产生的影响是巨大的。但坚持颠覆性创新,不陷入创新者的困境,你知道,这些原则仍然是我们真正为之奋斗的东西,我认为在今天的 Wayve 仍然适用。
And then I think there are pros and cons to operating at a different level of scale. You know, today we're about 1,000 people globally. But the culture you set at that scale really drives through where we are today. And I think it does require different skills operating at the silver scale, but the impact you can generate is just tremendous. But holding on to disruptive innovation, not falling for the innovator's dilemma, you know, these kind of principles are still really things we really fight for and I think still hold true today at Wayve.
嗯,你们取得了一些真正的成功。看着你们成长真是一件乐事。
Well, you've got some really success. It's been a joy watching you guys grow.
Chris Sacca。谁说的?Chris Sacca。非常感谢收听本期 Greymatter Decent。请继续关注未来的节目。
Chris Sacca. Who said it? Chris Sacca. Thanks so much for listening to this episode of Greymatter Decent. Please stay tuned for future episodes.