构建空间智能:World Labs 与 Scenix 合并

Building Spatial Intelligence: World Labs and Scenix Merge

李飞飞 Fei-Fei Li · a16z 播客 · 2026-07-28 · 约 42 分钟 · 原视频 ↗

打开互动全文版(中英对照 + 朗读 + 问答)→

本期速览 · Overview

李飞飞与 Yinuo 讨论空间智能的愿景以及用于训练机器人的现实到模拟再到现实的管道。

Fei-Fei Li and Yinuo discuss their vision for spatial intelligence and the real-to-sim-to-real pipeline for training robots.

要点 · TL;DR

核心观点 · Key points

反共识 · Contrarian takes

本期章节 · Chapters(共 28)

全文 · Full transcript(中英对照)

介绍与World Labs Introduction and World Labs

Host

好的,很高兴你们两位都来了。Fei-Fei,对于可能不了解背景的听众,也许你可以介绍一下 World Labs 是做什么的。

All right. Well, it's great to have you both here. So, Fei-Fei, for the listeners that may not have the background, maybe you can give an overview of what World Labs does.

Fei-Fei

是的。World Labs 是一家成立两年的初创公司。我认为我们应该认识到它是一个前沿模型实验室。我们正在构建 AI 的下一个前沿,即我们所说的“空间智能”。空间智能是指创造能够生成、理解、推理并与之交互的 AI,无论是物理空间还是虚拟空间。实现空间智能的一种手段是构建大型世界模型,而这正是 World Labs 主要专注的领域。

Yeah. World Labs is a two-year-old startup. I think we should recognize it's a frontier model lab. We are building the next frontier of AI which is what we call spatial intelligence. Spatial intelligence is about creating AI that has the ability to generate, understand, reason with, and interact with spaces, whether physical or virtual. A means to an end towards spatial intelligence is building large world models, and that's what World Labs is mostly focused on.

空间智能与机器人收购 Spatial Intelligence and Robotics Acquisition

Host

你从一开始就一直在说机器感知、推理空间并对其行动的能力,但我一直以为对空间采取行动是某种遥远未来的事情。现在你收购了一家机器人公司。谈谈这其中的时机和意图吧。

So you've been saying this since the very beginning, the machine's ability to perceive and reason about spaces and act on spaces, but I always had the assumption that acting on spaces was some long-term future thing. Now you're acquiring a robotics company. Talk a little bit about the timeliness and intentions.

Fei-Fei

首先,在空间中行动或交互并不仅限于机器人。看看创意领域:VFX、游戏、设计。在许多用例中,你可以在虚拟空间中创造并行动。World Labs 的论点一直是,我们生活的世界可以是一个多元宇宙,我们创造技术让人们、建设者、开发者在不同空间中行动。话虽如此,在物理空间中行动的能力是未来 AI 最激动人心且至关重要的能力之一。机器人正是其中的一部分。World Labs 一直认为机器人是空间智能和世界建模的重要应用和用例。因此,与 Scenix 团队的联合是我们长期愿景和使命的一部分。我们一直致力于此。

First of all, it doesn't just take robotics to act within spaces or interact. Look at the creative field: VFX, gaming, design. In many use cases, you can create and act within virtual spaces. World Labs' thesis has always been that the world we live in can be a multiverse, and we create technology to allow people, builders, developers to act within different spaces. Having said that, the ability to act within the physical space is one of the most exciting and profoundly important capabilities of future AI. Robotics is very much that. World Labs has always believed that robotics is an important application and use case of spatial intelligence and world modeling. So joining forces with the Scenix team is part of our long-term vision and mission. We've always committed to that.

Yunu背景与Scenix Yunu's Background and Scenix

Host

太棒了。Yunu,你是 Scenix 的联合创始人。给大家简单介绍一下你的背景以及 Scenix 是做什么的。

Amazing. So, Yunu, you're the co-founder of Scenix. Provide everyone with a quick overview of your background and what Scenix does.

Yunu

好的。我是 Yunu,目前是 Scenix 的联合创始人,也是哥伦比亚大学的助理教授。我的研究始于麻省理工学院的博士阶段,之后在斯坦福大学与 Fei-Fei 一起做博士后。在我的职业生涯中,目标一直很简单:帮助机器人更好地感知并与物理世界交互。我是一个非常务实的人,我希望我的机器人在真实物理环境中能够工作。

Yeah. I'm Yunu, currently co-founder of Scenix and also assistant professor at Columbia University. My research started from my PhD at MIT and then postdoc with Fei-Fei at Stanford University. Throughout my career, my goal has been very simple: trying to help robots better perceive and interact with the physical world. I'm a very practical person. I want my robot to work in real physical environments.

真实-仿真-真实管线 Real-to-Sim-to-Real Pipeline

Yunu

对于 Scenix 来说,我们看到的独特机遇是,通用机器人的开发目前面临许多瓶颈,尤其是在训练和评估方面。因此,我们正在开发所谓的“真实到仿真到真实”管道。我们希望将真实环境映射到与真实环境高度对齐的数字世界中。这里的对齐是指:数字世界中发生的任何事情也会在真实环境中发生。这样,我们就可以用数字世界中可大规模生成的数据来替代真实环境中所需的所有数据和评估。这就是 Scenix 一切的起点。我们组建了一支非常强大的团队,涵盖机器人学、机器人学习、仿真和渲染,试图构建这个“真实到仿真到真实”堆栈,以解决一些关键瓶颈。

For Scenix, the unique opportunity we see is that there are a lot of bottlenecks currently faced in the development of general-purpose robots, especially around training and evaluations. So we are developing what we call a real-to-sim-to-real pipeline. We want to map real environments into a digital world that has the best alignment with the real environments. By alignment, we mean that whatever happens in the digital world also happens in the real environment. This allows us to replace all the data and evaluation we need in the real environment by using data that can be generated at scale in our digital world. That's how everything started at Scenix. We put together a very strong team around robotics, robot learning, simulation, and rendering, trying to build this real-to-sim-to-real stack to solve some of the key bottlenecks.

合作故事 Collaboration Story

Host

你们两位能合作真是太好了。

It's amazing that you two work together.

Fei-Fei

是的。这有个有趣的故事。你可能以为因为我们曾一起工作——他是我出色的博士后——我们讨论 World Labs 的整合已经很久了。但实际上并不是这样。他们是作为客户来到 World Labs 的。

Yeah. There's a funny story here. You would think because we worked together — he was my amazing postdoc — we've been talking about this World Labs integration for a long time. It's actually not true. They came to World Labs as a customer.

Host

真的吗?

Really?

Fei-Fei

去年冬天,大约十一月、十二月,我们发布了名为 Marble 的生成模型的第一个版本,Scenix 就签约成为了客户。

When we released the first version of our generative model called Marble last winter, around November, December, Scenix just signed up as a customer.

Host

不是吧?作为客户?

No kidding. As a customer?

Fei-Fei

是的。我当时甚至不知道那是什么。然后我意识到这是 Yunu 的公司。我打电话给他,说:“哇,这是你的公司。”然后我们发现彼此之间有很多协同效应。

YES. And I didn't even know what it was. Then I realized this is Yunu's company. I called him and said, 'Wow, this is your company.' And then we realized there's so much synergy.

Marble与数据稀缺 Marble and Data Scarcity

Host

也许你可以快速描述一下 Marble 是什么。

Maybe you could quickly describe what Marble is.

Fei-Fei

Marble 是 World Labs 一直在训练和迭代的基础模型的代号。目前 Marble 公开发布的基本功能是:根据一个提示(可以是一张图片、几张图片或一段文字),将其转化为一个几何上一致的世界,可以用 3D 几何表示,无论是高斯溅射还是网格。实际上,Scenix 团队正在尝试解决机器人领域一个极其困难的问题:数据匮乏——训练数据的匮乏和评估数据的匮乏。这与语言模型截然不同,后者的数据在互联网上非常丰富。

Marble is the code name for the base model that World Labs has been training and iterating on. The fundamental capability right now of Marble that is publicly released is to take a prompt — it can be an image, a few images, or text — and turn that into a geometrically consistent world that can be represented in 3D geometry, whether Gaussian splat or mesh. Really, what the Scenix team is doing is trying to solve an extremely difficult problem in robotics: the lack of data — the lack of data in training and in evaluation. This is very different from language models, where data is abundant on the internet.

规模与互补性 Scaling and Complimentarity

Host

我们知道,为了让机器人发挥作用,我们必须解锁缩放定律的力量。但缩放定律从何而来?这是机器人领域每个人都在努力解决的深刻问题。如果能聊聊这个就太好了。你们组建了一支非常有才华的团队。那么,在多大程度上存在重叠,又在多大程度上是扩展?请谈谈这一点。

We know that in order for robotics to work, we have to unlock the power of scaling laws. But where does that come from? This is a profound problem that everybody is battling with in robotics. It would actually be great to talk about this. You have put together a very talented team. To what extent is there overlap, and to what extent is this an extension? Talk a little bit about that.

Fei-Fei

简单来说,两者非常互补,且有着共同的使命。所以,(该公司)是三位技术联合创始人之一。

The TL;DR is that it's very complementary with a shared mission. So, (the company) is one of the three technical co-founders.

团队成员介绍 Introduction of Team Members

Fei-Fei

另外两位是哥伦比亚大学的教授 Changi Jan,他是一位世界级的模拟技术专家。他有视觉特效背景,曾在 Weta、腾讯工作,并且是一位企业家。还有 Sunonni,一位杰出的工程领导者,他曾在多年前被亚马逊收购的一家初创公司工作,在亚马逊从事计算机视觉领域的多个技术栈。当我们开始认真讨论时,我意识到 Phoenix 在人才方面有几个点与 World Labs 高度互补。一是他在机器人领域卓越的思想领导力和技术实力,从硬件到全栈机器人。即使他在斯坦福做我博士后时,已经拿到了教职,所以只待了一年。我本来希望他待更久,但他得去做一份真正的工作。他是一位从建模到硬件的全栈机器人研究员。当然,Phoenix 和他的学生们在 Synenix 组成了 World Labs 之前没有的人才库。而 Changi 这边,他拥有令人难以置信的模拟能力,是一位资深的模拟研究员和技术专家,而 World Labs 的工作与模拟紧密相连。因此,World Labs 所缺乏的是生成模型和计算机视觉 3D 重建方面的能力,而这正是我们的强项,也是 Synenix 需要的技术。这两方面结合起来,使团队更加完整。

The other two are Changi Jan, another Columbia professor who has been a world-class technologist in simulation. He has a background in VFX. He worked at Weta, Tencent, and has been an entrepreneur. Then there is Sunonni, a phenomenal engineering leader who was at a startup that was acquired by Amazon many years ago. He worked on many different tech stacks in computer vision at Amazon. When we started talking more seriously, I recognized that a couple of things that Phoenix has from a talent point of view are extremely complementary to World Labs. One is incredible thought leadership and technical prowess in robotics, from hardware to full-stack robotics. Even when he was my postdoc at Stanford, he had his faculty offer, so he stayed only one year. I wanted him for longer, but he had to take a real job. He was a full-stack researcher in robotics from modeling to hardware. And of course, Phoenix and his students at Synenix were that pool of talent World Labs hadn't had yet. On the Changi side, he has incredible simulation capability. He's such a senior researcher and technologist in simulation, and what World Labs is doing interfaces heavily with simulation. So, what they don't have is on the generative model side and computer vision 3D reconstruction, which we are very strong at in World Labs. That's a technology Synenix needs. Together, these two sides make a much more complete team.

加入World Labs动机 Motivation to Join World Labs

Host

Fei 这个想法的动机是作为进入机器人领域的延伸和补充。鉴于你曾有过决定何时出售公司的经历,我很想听听你对加入 World Labs 的看法,以及你为何做出这个决定。

Fei's motivation in this is as an extension and a complement to get into robotics. Having been in your situation, which is deciding when to sell a company, it would be great to hear from you on how you think about joining World Labs and the fit there, and why you made the decision to do it.

Fei-Fei

最开始,我们决定要继续独立发展。但和 Fei 聊过之后,看到了中间所有协同效应,双方联合起来就变得非常合理。从 Synenix 的角度看,我们一直在做真实到模拟再到真实的过程:通过捕捉环境的外观、几何和动态(即施加动作时环境如何变化)来重建环境。目前的张量重建还稍微有些重,而 World Labs 在做的事情涉及稀疏重建和生成方面的深厚能力。所以我们看到了很多机会,可以利用 World Labs 的 Marble 和其他能力来高效地重建和建模环境。

At the very beginning, we were deciding okay, we want to just keep going. But after chatting with Fei and seeing all the synergies happening in the middle, it just makes perfect sense for the forces to join each other. In any sense, as Synenix, what we have been doing is real-to-sim-to-real: we reconstruct the environment by capturing its appearance, geometry, and dynamics — how the environment changes when you apply actions. This tensor reconstruction right now is still a bit heavy, and what World Labs is doing involves profound capabilities around sparse reconstruction and generation. So we see a lot of opportunities to leverage Marble and other capabilities at World Labs for efficient reconstruction and modeling of environments.

机器人基础模型 Foundation Model for Robotics

Host

那么我们可以期待 World Labs 推出一个机器人基础模型吗?

So can we expect a foundation model for robotics from World Labs?

Fei-Fei

World Labs 正在构建一个基础模型。你知道的,Martin,我们在构建一个基座模型。随着技术发展,最令人兴奋的基座模型之一是全模态模型:它们接受多模态输入并产生多模态输出。什么是机器人基础模型?很可能涉及动作。除了世界状态之外,它很可能还要输出动作。我们绝对不排除这种可能。

World Labs is building a foundation model. As you know, Martin, we're building a base model. As the technology evolves, some of the most exciting base models are omnimodels: they take multimodal input and have multimodal output. What is a foundation model for robotics? It's very likely going to involve actions. It's very likely going to involve the output of actions in addition to the state of the world. We're definitely not ruling this out.

Host

是的,太好了。

Yeah. Great.

Fei-Fei

例如,一个基础模型本质上需要是一个多模态模型。它必须考虑帧、文本、图像、深度等不同模态,而动作是这些模态中非常重要的一部分。如果你把动作作为输入,那本质上就是一个前向模拟器,能够预测施加特定动作时环境会如何变化。当动作作为输出时,这本质上就是一个策略模型,根据给定目标预测在真实环境中应该采取什么动作才能更接近目标。这种全模态模型能够极大地帮助机器人社区理解如何建模环境以及如何在其中行动。它还可以作为一个骨干网络,通过微调用于特定的机器人应用,以确保客户所期望的可靠性和效率。

For example, a foundation model essentially needs to be a multimodal model. It has to take into account different modalities like frames, text, images, depth, and action is a very important part of those modalities. If you think about feeding actions as input, that is essentially a forward simulator that predicts how the environment will change when you apply a specific action. When the action is output, this is essentially a policy model that predicts, given a specific goal, what action you should take in the real environment to get closer to that goal. This kind of omnimodel can benefit the robotics community greatly in understanding how to model environments and how to act in them. It can also serve as a backbone that you can fine-tune for specific robotic applications to ensure the reliability and efficiency expected by clients.

视频模型vs.3D仿真 Video Model vs. 3D Simulation Approach

Host

如果你不介意一个外行投资者的问题:我看到很多机器人公司现在流行的做法是使用视频模型,这是主流方法。但你们这里用的是 3D 和模拟——一种非常不同的方法。你能对比一下这种只用视频的流行方法和你们的目标之间的区别吗?

If you don't mind a kind of a lay investor question, I see a lot of robotics companies and a very popular approach right now is to use a video model. That's the predominant method. But here you are using 3D and simulation — a very different approach. Could you contrast this popular approach of using video only versus what the ambition here is?

Fei-Fei

为了创建机器人可以学习的世界,正如我提到的,我们需要捕捉问题的本质结构。对这些世界的一个非常重要的要求是一致性。这也是我看到与 Marble 有极强协同的地方,因为我们正在构建一个一致的世界——在空间、时间、不同视角和不同交互类型上都一致。Marble 生成的世界也提供了整个世界的基础设施组件,我们认为这是机器人学习所必需的。想象一下,如果机器人向前推一个物体,物体却凭空消失了——这是许多现有视频预测模型存在的问题。那将无法为机器人提供足够好的信号来知道该做什么。当然,目前有很多研究在构建更好的视频模型。我们看到一种方式,我们构建的部分基础设施可以提供初始动力,并经历一个数据飞轮:从更基于模拟的模型到在真实环境中执行并收集新数据的机器人策略模型。收集的数据会反馈回来改进模型,这个模型不必是纯物理驱动或纯学习驱动,而是介于两者之间——既能捕捉问题的本质结构,又能随着数据积累而扩展和变得更好。

In order to create worlds where the robot can learn, as I mentioned, we need to capture the essential structure of the problem. One very important requirement for those worlds is consistency. That is where I see very strong synergies with Marble, because what we are building is a consistent world — consistent over space, over time, over different viewpoints, and over different types of interactions. Marble's generated worlds also provide an infrastructure component of that entire world that we believe is necessary for the robot to learn. Imagine if a robot pushes an object forward, and the object magically disappears — which has been a problem for many existing video prediction models. That won't provide a good enough signal for the robot to know what to do. But obviously there is a lot of investigation into building better video models. We see a way where some of the infrastructure we build can provide initial momentum and go through a data flywheel: from simulation-driven models to robot policy models that execute in the real environment, collecting new data. The data will come back to improve the model, which doesn't have to be purely physics-based or learning-based, but somewhere in between — able to capture the essential structure of the problem while scaling and becoming better as more data accumulates.

Synenix务实方法 Robot Aspirations and Tasks

Host

你知道,我和你已经密切合作了一段时间,你一直有这个北极星驱动着,你以各种方式阐述过它,比如 3D 以及许多其他方式。我只是想知道,对你来说,是否也有一个类似的哲学北极星,还是你更像我这样务实——构建系统,做实事。

You know, I've worked now fairly closely for a while and you've always had this north star which has driven this and you've articulated it variously as kind of 3D and in a number of other ways. I'm just wondering for you, is there also a similar philosophical north star or you're more the pragmatic like I am — like build the system, do the thing.

Fei-Fei

我的北极星是让机器人在真实环境中工作。我是一个非常务实的人。我希望机器人能工作。有一件有趣的事情实际上源于我在博士后期间与 Fei-Fei 的合作,我们正在构建这种基准测试。我们实际上向公众发送调查问卷,询问他们希望机器人为他们做什么。在收集到的上千个任务中,三分之一是关于清洁的。人们就是不喜欢做那些沉闷和肮脏的任务。而这些正是我们真正希望确保有机器人解决方案应对的场景。

My north star is to make robots work in the real environment. I'm a very practical person. I want the robot to work. One interesting thing that's actually coming from my collaborations with Fei-Fei during my postdoc, we are building this kind of benchmark. We actually send out surveys asking the general public what they want their robots to do for them. Among the thousand tasks we collected, one third of the tasks are about cleaning. People just don't like to do those dull and dirty tasks. And those are the scenarios we really want to make sure we have robotic solutions to deal with.

精度vs.创意与仿真保真度 Pragmatic Approach of Synenix

Host

关于 Synenix,马丁,我真正喜欢的一点——尤其是接着你的问题——有很多机器人公司在构建模型等等。我真正喜欢 Synenix 的一点是,Renu 和他的联合创始人有着非常务实的机器人方法。他们尤其是——他们来自学术界,对吧?Sunny 不是,但 Vindra 和 Chani 来自学术界,但他们的第一反应是与实际行业的设计合作伙伴和客户合作,无论是工业实验室的实验室还是仓库还是电子装配。这是一种非常令人耳目一新的机器人方法,这让我非常兴奋能与他们合作。

One thing I really like about Synenix, Martin, especially continuing your question, there are a lot of robotics companies building models and all that. One thing I truly like about Synenix is Renu and his co-founders have such an incredibly pragmatic approach to robotics. They especially — they come from academia, right? Sunny doesn't, but Vindra and Chani come from academia, but their first instinct is work with design partners and customers in real industry, whether it's labs at industrial labs or warehouses or electronics assembly. That is such a refreshing way of approaching robotics, and that really made me very excited to work with them.

仿真与反事实推理 Precision vs. Creativity and Simulation Fidelity

Host

也许这是问你的,但这只是我个人的好奇。在我看来,机器人技术必须非常精确。我的意思不是完美,但相当接近。但对于创造性的用例,World Labs 做了很多,你有点不需要这样,因为即使有时犯错也是风格化的或故意的等等。所以从技术角度来看,调和这两件事的挑战是什么,或者它们永远不会被调和——就像设计空间中总会有两个极点?

Maybe this is for you, but I'll just be this is personal curiosity. It seems to me that for robotics you have to be pretty exact. I mean not perfect but pretty close. But for the creative use cases which World Labs has done a lot of, you kind of don't need to because even sometimes like being wrong is stylistic or intentional or whatever. And so from a technical perspective, what is the challenge here for reconciling these two things or do they never get reconciled — like there will always be two points in the design space?

Fei-Fei

它们当然会在长期内得到调和。环境建模不必完美。在机器人技术中,模型不必完美。

They will be reconciled in the long term, of course. And modeling of the environments doesn't have to be perfect. The model doesn't have to be perfect in robotics.

Host

顺便问一下,有没有更正式的说法?比如不完美是什么意思?它必须相当接近。

And by the way, is there a bit more formal way to say that? Like what does that mean not to be perfect? It has to be pretty close.

Fei-Fei

让我这样说。例如,模型在各种机器人应用的发展中一直是一个非常重要的基石。如果你看看所有现有的机器人应用,如飞机、无人机、Roomba,甚至四足机器人、双足机器人,模型一直是它们实际工作并能够从模拟迁移到真实环境的方式。

Let me put it this way. For example, models over the development of all different kinds of robotic applications have been a very important cornerstone. If you look at all the existing robotic applications like planes, drones, Roomba, or even for quadruped robots, bipedal robots, model has been the way for them to actually work and be able to transfer from simulation to the real environment.

Host

我明白了。

I see.

Fei-Fei

但如果你看看那些运动机器人,比如四足机器人、双足机器人,它们可以在雪地上行走,可以在灌木丛中行走。但你不需要一个能够非常精确地模拟所有灌木和雪的模拟器。你需要一个能够捕捉问题本质结构并在数字环境中进行各种不同随机化的模拟。这就是我们的目标。所以基本上,通过 Synenix 与 World Labs 一起,我们试图研究我们需要什么保真度来建模除机器人之外的庞大世界,以便我们能够将在模拟环境数字世界中训练的机器人系统迁移回真实场景。

But if you look at those locomotion robots like quadruped robots, bipedal robots, they can walk on snow, they can walk on bushes. But you don't need to have a simulator that can simulate all the bushes and snow very precisely. You need to have a simulation that captures the essential structure of the problem and do a whole different kind of randomization inside the digital environment. So that is what we're aiming for. So basically with Synenix and together with World Labs, we're trying to investigate what is the level of fidelity we need to model the massive worlds besides the robots so that we will be able to transfer the robotic systems trained in simulated environment digital worlds back into the real scenarios.

Host

作为一名投资者,我听到其他研究人员,比如 Sergey Levine,说模拟最终总会偏离物理世界,真实世界的数据收集绝对是关键。那么也许谈谈这种方法(以模拟为基石)相对于其他方法的可行性。

As an investor, I've heard other researchers say like Sergey Levine say simulation will always eventually deviate from the physical world and real world data collection is absolutely critical. And so maybe talk a little bit about the viability of this approach where simulation is a cornerstone as opposed to some other approach.

Fei-Fei

它们并不矛盾。如果你思考模拟,模拟本质上是试图预测当你施加动作时环境会如何变化,这本质上是一个世界模型。它不一定必须是纯物理的。它可以是物理和学习的结合。我们正在收集真实世界数据。我们将使用这些真实世界数据。只是在这个数据飞轮的不同阶段。也许在最开始,我们更强调物理,以确保我们具有正确的一致性和正确的结构来学习世界,来训练机器人策略。但随着我们通过数据收集以及与客户的合作积累越来越多的数据,我们将拥有数据,这些数据将朝着更基于学习的环境建模发展。所以这种过渡和数据流确实是一个促成因素,能够同时获得物理、几何和一致性的最佳效果,以及数据和算力的所有力量和魔力。

They don't contradict with each other. If you think about the simulation, simulation is essentially trying to predict how the environment is going to change when you apply the actions, and this is essentially a model of the world. It doesn't necessarily have to be pure physics. It can be a combination between both physics and learning. We are collecting real world data. We will be using those real world data. It just at different stages of this data flywheel. Maybe at the very beginning we have stronger emphasis on physics to make sure we have the right consistency and right structure for us to learn the world, for us to train the robot policies. But as we accumulate more and more data both through data collection and also through the collaboration with our clients, we'll have the data that will be moving towards more learning-based modeling of the environments. So this kind of transition and data flow is really an enabling factor of both getting the best of both physics and geometry and consistency as well as all the power and magic from data and compute.

仿真优势:可靠与高效 Simulation and Counterfactual Reasoning

Host

我想补充一点,这里稍微哲学一些。模拟或没有模拟不是一个二选一的选择。所有这些共同作用才能使机器人工作。想想人类智能。我们在头脑中做很多模拟。为什么?模拟扮演着一个真实世界数据无法扮演的非常重要的角色,那就是反事实推理——你模拟那些尚未发生或不可能发生或你没有足够数据让它在真实世界发生的事件。当你模拟它时,你学会了如何在其中行动。人类一直在这样做。我知道你在世界杯现场。祝贺西班牙获胜。我相信在每场比赛的规划中都有模拟,无论是数字的还是白板上的等等。模拟扮演的角色就是反事实推理。这在机器人技术中非常重要,因为我们不可能有足够的真实世界数据来做到这一点。这里有一个现实生活中的例子:自动驾驶汽车行业。Waymo 已正式表示他们使用了数十亿小时的模拟,实际上 Waymo 更偏向模拟而不是真实世界数据。所以这些是真实的例子,而且你知道,马丁,汽车是最简单的机器人。所以很明显,模拟在机器人学习中扮演着巨大的角色。

I want to add to this and be slightly philosophical here. There isn't a binary choice between simulation or no simulation. All this comes together to make robotics work. Think about human intelligence. We do a lot of simulation in our head. Why? There is a very important role simulation plays that real world data doesn't play, which is counterfactual reasoning — you play out events that haven't happened or cannot happen or you don't have enough data to make it happen in real world. And while you play it out, you learn how to act in it. Humans do this all the time. I know you were at World Cup. Congratulations to Spain winning. I'm sure in the planning of every game there is simulation, whether it's digital or on the whiteboard or whatever. The role simulation plays is counterfactual reasoning. And that's really important in robotics because we just cannot possibly have enough real world data for that. Here's a real life example: the industry of self-driving cars. Waymo has officially said they use billions of hours of simulation, and actually Waymo is more simulation-heavy than just real world data heavy. So these are real examples, and as you know Martin, cars are the simplest kind of robots. So clearly simulation plays a huge role in robotic learning.

Fei-Fei

我也想说一点。

I also want to add to that.

用例:评估与训练 Simulation benefits: reliability and efficiency

Fei-Fei

更具体地说,仿真可以提供两个层面的好处。第一是可靠性,第二是效率。在可靠性方面,如果你想让机器人在真实环境中可靠运行,就需要数据来系统性地覆盖机器人可能遇到的所有状态空间和变化。这样你才能学会如何做到鲁棒。通过仿真,你可以系统地对光照、摩擦力、几何形状、物体类型以及各种物理参数进行随机化控制,确保对状态空间有充分的覆盖。这就是机器人系统可靠性的来源。第二是效率。目前很多人都在做遥操作,如果你观察遥操作设备,你会发现数据采集的速度实际上比人类亲自执行任务还要慢。

So, to be more specific, simulation provides two levels of benefits. The first is reliability, and the second is efficiency. For reliability, if you're thinking about a robotic system working reliably in real environments, you need data to provide systematic coverage of the entire state space and variations that robots might encounter. That's how you learn to be robust. With simulation, you can systematically randomize and control variations in lighting, friction, geometries, object types, and all kinds of physical parameters to ensure sufficient coverage of the state space. This is what gives robotic systems reliability. The second is efficiency. Right now, many people are doing teleoperation. If you look at teleoperation devices, you're collecting data at a speed slower than a human actually performing the task.

Host

但对于我们的很多客户来说,人类的速度还不够。他们想要比人类更快的速度。是的。

But for many of our clients, human speed is not good enough. They want faster than human speeds. Yeah.

Fei-Fei

所以想让机器人更快,并不只是简单地把机器人驱动得更快,因为重力不会改变。而在仿真中,你可以系统地加快机器人的行为速度来训练机器人,使其考虑到环境中所有的动态变化。这就是我们客户获得效率的方式。因此,可靠性和效率是仿真能够提供的独特价值。

So for the robot to move faster, it's not as simple as just driving the robot faster, because gravity doesn't change. But in simulation, you can systematically speed up the robot's behaviors to train the robot to account for all the dynamic changes in the environment. This is what gives our clients efficiency. So both reliability and efficiency are unique values that simulation can provide.

Host

你刚才谈到了技术和平台的功能。也许可以讲讲人们具体用它来做哪些用例?

You've talked about the technology and the platform, what it does. Maybe talk about the specific use cases people use it for.

平台与生态系统 Use cases: evaluation and training

Fei-Fei

基本上有两个具体的用例,主要围绕训练和评估。

There are essentially two specific use cases, especially around both training and evaluations.

Host

好。

Okay.

Fei-Fei

先从评估说起。评估在机器人领域往往被忽视,但如果你在调优机器人模型,你必须知道它表现如何。那是你迭代的唯一信息来源。

Starting from evaluations. Evaluation is something people often overlook in robotics, but if you're tuning robotic models, you have to know how well they work. That's the only source of information for iteration.

Host

是的。顺便说一句,每个 AI 从业者都很清楚什么是评估,并且一直在使用。非 AI 领域的人,这个词的含义可能略有不同。所以或许值得具体说明一下你所说的评估是什么意思。

Yeah. By the way, every AI person really understands what evals are and uses them all the time. Non-AI people, it often means something a little different. So maybe it's worth describing specifically what you mean by evaluation.

Fei-Fei

好的。我所说的评估是指,能够了解某个具体检查点的性能如何。例如,它在 95% 的时间里生效,还是 99.9%?行业里使用的关键标准是——需要多少墙上时间才能区分出 90% 的检查点和 92% 的检查点?如果你只在真实环境中做这件事,那要花很长时间才能区分。再想想今天在真实环境中进行的机器人评估,其迭代速度比语言模型的迭代要慢好几个数量级。

Okay. So what I mean by evaluation is being able to understand for a specific checkpoint how well it performs. Does it perform, for example, 95% of the time or 99.9% of the time? The key criteria people use in industry is how long does it take—how much wall-clock time does it take to distinguish between a checkpoint that is 90% from a checkpoint that is 92%? If you only do that in the real environment, it takes so long to make that distinction. If you think about robotic evaluations in real environments today, the iteration speed is multiple orders of magnitude slower than iterations of language models.

Host

是的。

Yeah.

Fei-Fei

所以机器人任务不仅非常多样化和分散……

So not only are robotic tasks very varied and diverse...

Host

哦对,因为你真的得去做这件事,对吧?原子必须在空间中移动。

Oh yeah, because you actually have to do the thing, right? Atoms have to move through space.

Fei-Fei

正是。但只有……

Exactly. But only there are...

Host

物理定律必须起作用。

The laws of physics have to do.

Fei-Fei

你看过那些机器人视频吗?每个视频都放到了 10 倍或 8 倍速,因为它动得太慢了。

Have you watched those robotics videos? Every video has like 10x or 8x speed because it moves so slowly.

Host

的确。

Exactly.

Fei-Fei

所以它不仅慢,而且危险、昂贵,同时速度还慢了好几个数量级。因此,我们的一些客户需要一个数字环境来评估他们的机器人系统。因为我们的数字环境已被证明与真实世界对齐,也就是说,仿真中发生的事很可能也会在真实环境中发生。如果某个检查点在仿真中表现更好,那么它在真实环境中也很可能表现更好,这在我们的一篇博客文章中已经讨论过。这给了客户很强的信心,让他们能够利用数字环境中的数据信号,进行可扩展、安全且更快的机器人系统评估。这是关于评估的。然后是训练:它关乎可控性。你需要控制状态、参数、光照、摩擦力、物理参数、甚至物体几何形状和物体类型的所有可能变化。你需要确保对所有不同场景有充分覆盖,从而为机器人生成能使其鲁棒的信息丰富的数据。这在真实环境中极难做到。如果做遥操作,数据采集速度慢,你还会受到机器人数量、遥操作设备数量的限制——围绕数据操作有一整套挑战。但在仿真中,一切都可以控制,一切都可以系统化,你能够确切地知道你覆盖了哪些分布,从而建立信心:在该分布内机器人将会正常工作。这种信心、效率和可扩展性,正是我们的客户在利用我们的数字世界训练机器人系统时所看重的。

So not only is it slow, it's dangerous, it's costly, but at the same time the speed is also multiple orders of magnitude slower. So some of our clients need this digital environment that can be used to evaluate their robotic systems. Because our digital environment has proven alignment with the real world, meaning whatever happens in the sim is also likely to happen in the real environment. If a checkpoint works better in simulation, it's highly likely to also work better in real environments, as we discussed in the blog post. That gives our clients strong confidence in using the data and signal from the digital environment for scalable, safe, and much faster evaluations of their robotic systems. So that's on evaluation. Then on training: it's about controllability. You want to control all possible variations of states, parameters, lighting, friction, physical parameters, even object geometry and object types. You want to ensure sufficient coverage of all different scenarios so you can generate informative data for robots to be robust. This is extremely hard to do in real environments. If you do teleoperation, the data collection speed is slow. You're limited by how many robots you have, how many teleoperation devices you have—a whole set of challenges around data operations. But in simulation, everything can be controlled, everything can be systematic, and you can understand exactly what distribution you've covered, developing confidence that within that distribution the robot will work. That kind of confidence, efficiency, and scalability is something our clients value when using our digital worlds for training robotic systems.

无关具体形态的平台 Platform and ecosystem

Host

有一件疯狂的事:甚至在 Synix 之前,我们就在谈论 Marble 的入站客户,他们已经在提出这类需求。我们当时无法服务这些客户。但我们已经接到很多机器人公司的电话——从早期开发模型的初创公司到下游非常务实的用例——我们看到了这些需求。当人们听说你要做机器人时,他们想象的是你拿出 3D 打印机,制作硬件,给机器人编写大脑,然后放进机器人里,你就得到了一个机器人。我不认为你们说的是这个。所以也许可以谈谈这在整个机器人创建生命周期中的位置,你们止步于哪里,而生态系统的其余部分又从哪里开始。

Here's the crazy thing: even before Synix, and we're talking our inbound customers for Marble were already seeing this kind of demand. We just couldn't serve these customers. But we were already getting a lot of phone calls from robotics companies—early-stage robotics companies developing their models all the way to downstream pragmatic use cases—and we were seeing these needs. When people hear you're going into robotics, what they envision is you pulling out a 3D printer, making hardware, programming the brain of a robot, sticking it in, and then you have a robot. I don't think that's what you guys are talking about here. So maybe talk about where this fits in the life cycle of creating a robot, where you will end, and where the rest of the ecosystem will begin.

Fei-Fei

所以我们可以把我们在构建的东西想象成一个基础设施,以及围绕这个基础设施的软件,让人们能够构建世界,让机器人可以学习和评估。而且这个基础设施天生就是模型无关的和形态无关的。

So what we have been building, you can imagine, is an infrastructure, with software around this infrastructure, for people to build worlds such that robots can learn and evaluate. And this infrastructure is naturally model-agnostic and embodiment-agnostic.

人形机器人预测与受限部署 Embodiment-agnostic platform

Host

所以我想说清楚一点——照你说的,这不是在造机器人,而是在搭建一个环境,让其他公司可以把它们的机器人大脑放进去导航和学习。

So I just want to be very clear. From what you said, that's not building a robot, it's building an environment which another company can place their robot brain to navigate and to learn.

Fei-Fei

没错。目前我们的客户有各种各样的机器人,有的用单机械臂,有的用固定臂,有的用移动操作平台,有的用夹爪,有的用更复杂的末端执行器。因此我们的平台天生就是具身无关的。我们可以非常容易地集成不同类型的机器人躯体,把它们放进我们生成的世界中,从而让每个机器人都能在真实环境里以恰当的可靠性和效率完成特定任务。

Exactly. For our customers right now, they have all different kinds of robots. Some are using single robot arms, some are using a fixed arm, some are using mobile manipulators, some use grippers, some use more elaborate end effectors. So our platform is naturally embodiment agnostic. We can very easily integrate different robotic embodiments and put them into the worlds we generate, so we can give those individual robots the capabilities to perform the right tasks at the right levels of reliability and efficiency in real environments.

Host

嗯。

Yeah.

Fei-Fei

而且我们也是模型无关的。我们可以利用世界生成的数据来训练不同的模型,要么从头训练,要么对现有的基础模型(比如视觉-语言-动作模型或世界-动作模型)进行后训练。所以对我们来说无所谓,我们只确保有基础设施和世界,让机器人能在真实环境中可靠工作。

And we are also model-agnostic. We can use the data generated by our worlds to train different models, either from scratch or doing post-training of existing foundation models like vision-language-action models or world-action models. So to us it doesn't matter; we just make sure we have the infrastructure and the worlds so that robots can work reliably in the real environment.

经济视角与能效比较 Predictions on humanoids and constrained rollouts

Host

你之前跟我说,你觉得很多人形机器人的预测有点过于激进了,我们更可能看到像仓库之类的受限部署。能谈谈这个吗?以及这对你们在 World Labs 的工作有什么影响?

You've told me that you think a lot of the predictions around humanoids were a little bit aggressive, and we're likely to see more constrained rollouts like warehouses or whatever. Can you talk a little bit about that and how that impacts what you're tackling here at World Labs?

Fei-Fei

这是个很好的问题。如果你看看机器人应用在真实环境中的演进,它总是遵循从完全结构化环境到半结构化环境,再到非结构化环境的趋势。完全结构化环境是指你对所有配置都有了解和控制,比如工厂或汽车制造,这些已经自动化了几十年。然后是半结构化环境,你对环境有一定控制来让机器人的任务更容易,比如亚马逊仓库、餐厅或酒店,但还有很多物体你无法控制,比如衣服。非结构化环境则是你的家或我的房子——那是终极挑战。鲁棒性来源于对机器人可能遇到场景的充分覆盖。所以现在更容易、更可行的是先聚焦半结构化环境,再进入完全非结构化环境。我们会朝那个方向走,但我们要采取更可持续、更现实的方法。

That's a very good question. If you look at the progression of robotic applications in real environments, it has always followed the trend from fully structured environments to semi-structured environments and then to unstructured environments. Fully structured environments are ones where you have knowledge and control over all configurations, like factories or car manufacturing, which have been automated for decades. Then you have semi-structured environments where you have some control over the environment to make tasks easier for robots, like Amazon warehouses, restaurants, or hotels. But there are many other objects you don't control, like clothes. Unstructured environments are like your home or my house—that is the grand challenge. Robustness comes from sufficient coverage of scenarios robots might encounter. So it's much easier and more approachable right now to focus on semi-structured environments before moving to fully unstructured ones. We will move in that direction, but we want to take a more sustainable and realistic approach.

Host

我觉得你的意思是,人形机器人模仿了人体,而进化已将人体优化用于非结构化环境。我们的手指和腿并不是完成某件特定事的最佳装置。如果人类作为一个物种的唯一目标是爬树,我们不会进化出这样的身体。但人类最终演化出了这种非常通用但并非每件事都最优的体形,是为了在非结构化环境中生存。从商业和务实的技术角度看,非结构化环境和通用躯体恰恰是最难解决的问题,甚至不是正确的解题思路——我们应该专精。所以要用更专门的躯体解决更窄的问题。对我们 World Labs 来说,挑战在于要更加躯体无关,让基础设施能够服务不同的躯体以及不同的半结构化环境。

I think your point here is that humanoids mimic the human body, and evolution has optimized the human body for unstructured environments. Our fingers and legs are not the best apparatus for one specific task. If our only goal as a species was to climb trees, we wouldn't have this body. But humans evolved into a body shape that is very general but not necessarily best at everything, for survival in unstructured environments. From a business and pragmatic technology point of view, this unstructured environment and a generalized body is actually the hardest problem to solve. It's not even the right way to solve the problem—we should specialize. So we take more specialized bodies to solve narrower problems. The challenge for us at World Labs is to be more body-agnostic so that our infrastructure can serve different bodies and different semi-structured environments.

机器人vs.软件性能 Economic lens and energy efficiency comparison

Host

你知道,审视这个问题的一个常用视角是经济视角。拿它跟生成式大语言模型比,大语言模型生成散文或代码比人类快一万倍,便宜得多。经济上的理由成立,因为我们的大脑在这方面效率不高。但我们在三维导航、在世界中移动或抓取物品方面,大脑和身体非常高效。这是个预测问题:你认为我们能否在可预见的未来造出机器人,在从事体力劳动(比如最低工资的工作)时达到人类的能量效率?还有多远?五年?还是需要很长很长时间?

You know, a common lens to look at this question is an economic lens. You compare it to generative LLMs, which can create prose or code 10,000 times faster than a human and much cheaper. The economic case makes sense because our brains aren't very efficient at that. However, our brains and bodies are very efficient at 3D navigation, moving through the world, or picking things up. This is a prediction question: do you believe we'll ever be able to build robots, at least in the foreseeable future, that have the power efficiency of a human being when it comes to menial tasks, like minimum wage jobs? How far away are we? Five years? Or will it take a very long time?

Fei-Fei

所以如果你认真想想真实环境中的机器人,它始终是一个系统。每一个在真实环境中工作的机器人都是一个系统。你需要非常用心地考虑这个系统如何组合——硬件、软件、大脑,甚至精细到手指的摩擦系数。有很多东西要考虑才能把它们变成现实,而且需要迭代。但让我兴奋的是,我一直处在机器人学习的前沿,并推动前沿发展。前沿总是比我预想的要快。我现在专注的东西跟我读博时已经截然不同。这说明整个生态系统进化有多快,所有移动部件开始拼凑在一起构建机器人系统。但我们也必须校准我们的预测。我们会看到很多进步,但达到人类级别的效率和能力需要更长时间。

So if you really think about a robot in the real environment, it will always be a system. Every working robot in the real environment is a system. You need to be very mindful and thoughtful about how the system comes together—the hardware, the software, the brain, even details like the friction coefficient of the fingers. There are many things to consider to make these a reality, and it will take iterations. But what I'm excited about is that I have always been at the state of the art of robot learning and pushing it forward. The state of the art is always moving faster than I expected. What I'm focusing on now is very different from when I started my PhD. That speaks to how fast the ecosystem has evolved and all the moving pieces are coming together to build robotic systems. But we also have to be calibrated about our predictions. We will see a lot of progress, but to achieve human-level efficiency and capabilities will take longer.

Host

我是说,就连大语言模型也不具备人脑的效率。

I mean, even LLMs do not have human brain efficiency.

Fei-Fei

人脑的功耗只有 30 瓦。没错,确实如此。所以我们离那个还很远。

Human brain operates on 30 watts. Yeah, that's true. So we are far from that.

对战略与仿真的影响 Robotics vs Software Performance

Host

那么,但从性能到功耗来看,可能很接近,对吧?在像生成图像或软件工程这样的狭窄任务中,是这样的吗?

So, but performance to power it may be close, right? In narrow tasks like generating an image or software engineering, it is right?

Fei-Fei

是的,我也这么认为。但在机器人领域,我们还远未接近。

Yeah, I think so. I don't think we're anywhere close when it comes to robotics.

与Synnex集成 Impact on Strategy and Simulation

Host

这是否改变了你对团队战略雄心的看法?我是说,它改变了,还是仍然符合你刚起步时的预期?

Does this change how you think about strategically the level of ambition that your team can go after? I mean, has it changed that or is it still very much in line with what you expected to do when you started?

Fei-Fei

它确实以非常深刻的方式改变了轨迹。我们看到,能够以更高效、更可扩展的方式完成整个环境建模过程,带来了很多解锁,尤其是与 World Labs 的合作。我还想补充一点,如果你考虑当前语言模型的状态,它们拥有不可思议的能力,但你仍然不会完全信任它们去预订机票或酒店,仍然需要有人阅读这些语言模型的输出。

It definitely changed the trajectories in a very profound manner. So we see a lot of unlock in being able to do this whole process of modeling the environments much more efficiently and in a much more scalable manner, especially in partnership with World Labs. And I also want to add, if you think about for example the current state of language models, those are models with incredible capabilities, but still you don't just blindly trust them to book your flight tickets or make your hotel reservations. You still have a person reading the output from those language models.

Host

嗯。

Yeah.

Fei-Fei

但这与人们使用机器人模型的方式截然不同,因为对于机器人模型,开箱即用,机器人必须在真实环境中可靠工作。我们甚至没有数据或必要的基础设施,让机器人能够在真实环境中立即可靠地工作。因此,能够创建这个可扩展的数字世界,让机器人在其中学习和评估,将释放巨大潜力,用从这些世界生成的数据取代真实环境中昂贵且不安全的数据,实现机器人的可扩展学习和评估。

But that is very different from how people will use robotic models, because for robotic models, out of the box, the robot has to work reliably in the real environment. And we don't even have the data or all the necessary infrastructure for the robots to just work reliably out of the box in real environments. So for that reason, being able to create this scalable digital world where the robot can learn and evaluate within it is going to unlock so much potential to replace costly and unsafe data in the real environment with data generated from the worlds for the robots to do scalable learning and evaluations.

地域扩展 Integration with Synnex

Host

你知道,我见过很多这类整合。它们在这一阶段实际上运作得很好,当有这么多对齐时,这很棒。但总有这个问题:你是现在就整合到当前事物中,还是保持分离,提供将在一年内实现的长期轨迹?你怎么看,Fei?这是立即整合的事情,还是独立的长期事务?

You know, I've seen many of these kinds of integrations. They actually work very well at this stage when they have this much alignment, which is great. But there's always this question: do you integrate now into what's happening or do you keep things separate and provide a long-term trajectory that will be realized in a year timeframe? How are you thinking about this, Fei? Is this something that integrates right away or is this a separate longer-term thing?

Fei-Fei

这是个好问题。目前,Changi、Sonny、Justin、Ben 和我一直在讨论这个问题。我们会慎重考虑。我们不会急于整合从代码库到团队的一切,因为 Synnex 拥有一个经过深思熟虑的、我不会说完全独立,但相当封闭的技术栈,以及他们的客户和正在构建的产品。我们会花时间。我们肯定会整合;我们已经开始在仿真侧和潜在的基础模型动作条件模型侧进行合作。我们开始对话,而且他们正在使用 Marvel 作为内部客户。所以我们会整合,但我们不会急于把团队像沙拉碗一样完全混合。

This is a great question. At this point, you know, Changi, Sonny, Justin, Ben and I have been talking about this. We're going to take it thoughtfully. We're not rushing to integrate everything from codebase to teams, because Synnex does have a very well-thought-out, I wouldn't call it standalone completely, but fairly contained tech stack, as well as their customers and the kind of products they're building. We're going to take time. We definitely will integrate; we already are on the simulation side and the potential base model action-conditioned model side. We're starting to talk, and also they are using Marvel as an internal customer. So we will be integrating, but we're not rushing to blend the team like a full salad bowl.

Host

嗯。

Yeah.

未来产品成功 Geographic Expansion

Host

你对这次搬迁的地理位置是怎么考虑的?它会留在同一个地方吗?

How are you thinking about geographies with this move? Is it going to stay in the same place?

Fei-Fei

Vundra 会搬。

Vundra is going to move.

Host

哦,那欢迎来这里。搬到旧金山。佛罗伦萨和文艺复兴。完美。

Oh well, welcome here. Moving to San Francisco. Florence and the Renaissance. Perfect.

Fei-Fei

我认为 World Labs 正在正式成为一家双岸公司,总部设在旧金山。我住在帕洛阿尔托,感觉像在不同的州。但我实际上很兴奋,我们将在纽约设立办公室,帮助我们吸引东海岸的人才。此外,我们一直在讨论确保在两个办公室都设置机器人,这样我们就可以测试和完善工程栈,以便远程与机器人合作,因为无论如何我们都必须为我们的客户这样做。

I think World Labs is officially becoming a bi-coastal company, with headquarters in San Francisco. I live in Palo Alto; I feel like I'm in a different state. But I'm actually excited that we're going to have an office in New York that can help us attract talent on the East Coast. Also, we've been talking about making sure that in both offices we set up the robots so that we can basically test out and mature our engineering stack so that we can work with robots remotely, because we have to do that for our customers anyway.

客户参与阶段 Future Product Success

Host

那么,也许说得具体一点,Fei,让我们勾勒一下两年内的完美成功案例是怎样的?你们有什么产品?谁在使用它?他们如何使用?就一个清晰的画面。

So maybe just to be very concrete, Fei, let's pencil out what is the perfect success case in two years? What product do you have? Who's engaging with it? How do they use it? Just the crisp picture.

Fei-Fei

我们很高兴 Synnex 团队和 World Labs 团队将在少数重要的垂直用例中获得经过验证的客户,在这些用例中,我们的系统和基础设施已被证明对他们的自动化需求真正有益。这些客户成为我们扩展业务的灯塔示例。

We're very happy that the Synnex team and World Labs team will have validated customers in a small number of important vertical use cases where our system and infrastructure have proven to be truly beneficial to their automation needs. And these customers become our lighthouse examples to scale our business.

何时联系World Labs Engagement Stage for Customers

Host

那么时间多早?假设听这期节目的某人经营一家机器人公司。他们在什么阶段与 World Labs 接触?是很早期,还是中间某个阶段?

And how early? Let's say someone listening to this is running a robotics company. At what stage do they engage with World Labs? Is it really early on? Is it somewhere in the middle?

Fei-Fei

现在,对于我们的客户,因为我们正在构建这种实到仿到实的管道,仿真本质上是我们提供的用于训练和评估的世界。一些客户只需要实到仿部分,他们希望数字化他们关心的任务,并能够对其机器人系统进行评估。一些客户需要整个管道,以便策略可以在他们的硬件上运行。因此,我们的平台根据客户需求灵活设计。同时,我们正在合作的客户实际上已经接近部署阶段。所以基本上,他们正在处理非常实际的任务,一旦我们有机器人解决方案,就能立即创造价值。他们至少有成百上千种这样的情况在考虑自动化。因此,与 World Labs 一起,我们将能够为这些场景开发可靠的解决方案,正如我们已经在博客文章中展示的那样。我们将能够进一步研究,看看它们如何真正解决实际部署面临的关键需求和约束。

So right now, for our customers, because we are building this real-to-sim-to-real pipeline where the simulation is essentially the worlds we provide for training and evaluation grounds. Some customers only need the real-to-sim part; they want to digitalize the task they care about and be able to do evaluations of their robotic systems. Some customers need the entire pipeline so they can have policies running on their hardware. So our platform is designed flexibly depending on what our clients need. At the same time, the clients we are working with are actually pretty close to the deployment stage. So basically, they are working on very practical tasks where, when we have a robotic solution, it can create value immediately. They have at least a ton or hundreds of these situations they are thinking about for automation. So together with World Labs, we'll be able to develop reliable solutions for those scenarios, as we have already shown in our blog post. We'll be able to further our investigation to see how they can actually solve the key requirements and constraints faced by real-world deployments.

When to Contact World Labs When to Contact World Labs

Host

太好了。我想非常具体地问这个问题:如果你是一家机器人公司,联系 World Labs 是太迟还是太早?

Great. I want to be very specific about this. Is it ever too late or too early to call World Labs if you're a robotics company?

Fei-Fei

不。我们希望每个人都联系我们。我们想了解你们的用例。

No. We want everybody to call us. We want to learn about your use case.

Host

太好了。如果你正在收听,并且你与机器人项目或机器人公司有任何关系,请关注 World Labs。

Wonderful. If you're listening to this and you're anywhere close to a robotics project or robotics company, please track World Labs.

Fei-Fei

是的。谢谢。

Yes. Thank you.

Host

当然在开展业务。

Definitely open for business.

Fei-Fei

我们在开展业务。好的。

We are open for business. All right.

Host

不会太早。

Not too early.

Host

好的。如果你在从事机器人领域,请致电 World Labs。非常感谢两位的到来。

All right. If you're doing robotics, call World Labs. Thank you both very much for coming.

Fei-Fei

谢谢。

Thank you.

互动版:逐字朗读 + 针对本期提问 →