Skild AI: Building a Universal Brain for Robots
打开互动全文版(中英对照 + 朗读 + 问答)→Skild AI 创始人探讨其使命:为任何机器人打造通用大脑,将机器人技术视为数据问题。
Skild AI founders discuss their mission to create a general-purpose brain for any robot, treating robotics as a data problem.
欢迎收听 NVIDIA AI Podcast。我是 Noah Kravitz。今天请到的是 Skild 公司的 Deepak Pathak 和 Abhinav Gupta。Skild 是一家机器人公司,正在打造 omni-brain——一个通用大脑,可以驱动任何形态的机器人去完成任何任务。这太棒了。我很高兴能听你们亲自讲讲。那我们开始吧。Deepak、Abhinav,欢迎你们。非常感谢你们来做客 AI Podcast。
Welcome to the NVIDIA AI Podcast. I'm Noah Kravitz. I'm here today with Deepak Pathak and Abhinav Gupta from Skild. Skild is a robotics company that's building the omni-brain, a universal brain that can power robots across any form factor to tackle any task. It's amazing stuff. I'm very excited to find out about it from the source. And so let's get into it. Deepak, Abhinav, welcome. Thank you so much for joining the AI Podcast.
非常感谢你们的邀请。
Thank you so much for having us.
那么 Deepak,也许你可以先讲讲公司的情况,讲讲 Skild,然后你们两位再谈谈各自的角色。
So Deepak, maybe you can start and tell us a little bit about the company, about Skild, and then you can both talk a little bit about your roles.
在 Skild,正如你所说,我们正在构建一个通用大脑。我们称之为“全具身智能”。任何机器人、任何任务、一个大脑。你可以把它想象成语言领域的 ChatGPT。我们正在为任何物理设备或任何类型的机器人构建一个通用大脑。这非常通用。你可以有一个人形机器人、一个狗形机器人、或者传送带上的机械臂,都由同一个共享大脑、共享智能在幕后控制。为什么我们要做得这么通用?原因是,机器人技术是一个数据问题。与语言或视觉不同,机器人领域没有太多数据。没有机器人数据的互联网。如果是这种情况,我们不能挑选使用哪些数据,所以我们要以最通用的方式去做。我们部署的每一个大脑实例,无论用于什么任务或什么形态,都有助于让大脑在未来场景中变得更好。这就是背后的主要目标。就我个人而言,我们之前都是教授,所以我们非常技术化。过去十多年里,我们一直参与推动机器人学习领域的技术发展。我们的角色既在技术方面,确保这些系统被构建出来,并且非常通用、可迁移,但我们也非常关注部署。我们不认为部署是事后才考虑的事情。比如,在 ChatGPT 或语言模型的情况下,人们研究了好几年,但一旦准备好,七天就有 100 万用户。也许一天,我不记得了。也许一个月 1 亿用户。对,增长最快的产品。物理 AI 不是这样的。部署需要时间。所以对我们来说,部署从第一天起就是我们的首要任务。
So at Skild, as you mentioned, we are building a general-purpose brain. So we call this omni-bodied intelligence. Any robot, any task, one brain. So think of like what ChatGPT is for language. We are building a general brain for any physical device or any kind of robot. So this is absurdly general. You can have a humanoid, or a dog-like robot, or a robotic arm on a conveyor belt, all being controlled by the same shared brain, shared intelligence behind the scene. So why do we go so general? And the reason is, robotics is a data problem. Unlike language or vision, there is not much data in robotics. There is no internet of robot data. So if that's the scenario, we cannot pick and choose which data we use, so we go in a most general fashion. Every single instance of our brain which we deploy for any kind of task or any form factor that contributes in making the brain better for the future scenarios. So this is the main goal behind this. And personally, in my role, like I have been, so we both have been professors before this. So we are extremely technical. We have been involved in bringing up these technologies in the robot learning area for the last decade and more. So our role is both on the technical side to make sure that these things get built and they are super general, transferable, but our focus is also a lot on deployments. Like we do not believe deployment to be a... It's not a hindsight scenario. Like for instance, in the case of ChatGPT or language models, folks did research for several years, but once it was ready, you have 1,000,000 users in seven days. Maybe one day, I don't remember. Maybe one hundred million users in one month. Right, fastest growing product. Physical AI is not like that. The things take time to deploy. So for us, deployment is our first priority from day one.
是的,有道理。你提到当过教授。你在卡内基梅隆大学?
Yeah, makes sense. And you mentioned being a professor. You're at Carnegie Mellon?
是的。
Yeah.
公司总部在匹兹堡?
And the company's based in Pittsburgh?
公司总部在匹兹堡,但我们在匹兹堡有办公室。现在我们在湾区也有办公室,在圣马特奥地区。
So the company has HQ in Pittsburgh, but we have offices in Pittsburgh. Now we are also in Bay Area, the San Mateo area.
哦,太好了。还有一个办公室在印度班加罗尔。
Oh, great. And one office in India, Bangalore.
太好了。
Fantastic.
那 Abhinav 呢?
And Abhinav?
是的,我想从我们为什么对此如此兴奋说起,因为我们几乎在重新思考传统机器人技术的做法。传统上,机器人技术是一个非常经典的、垂直导向的领域,对吧?我的意思是,在 AI 时代之前,你首先决定要把机器人放在哪个垂直领域。比如,我想造一个焊接机器人。然后你去制造硬件,专门针对焊接。你开始编写软件,专门针对焊接。这类部署的问题在于,很容易猜到前 80% 或 90% 的性能,但随后你会遇到一堵墙,叫做物理世界中的“边缘情况”。物理世界中有太多边缘情况,比如有人在你面前留下一个包裹,这就变成了一个边缘情况,等等。所以这就是为什么如果有边缘情况,因为你已经达到 90% 的性能,你仍然无法完全自动化。人类仍然需要在场,以确保边缘情况得到处理,等等。这就是为什么传统上,机器人技术并没有真正大规模主流化。然而,当 AI 出现时,事情发生了变化。比如,语言领域,在这一切出现之前,也非常垂直化。有一些不同的公司在构建聊天机器人。有不同的公司在构建搜索引擎。但一旦 LLM 出现,它们就成了横向平台。现在每个人都在这个横向 LLM 平台上构建。这正是我们现在对机器人技术的思考方式。我们正在构建这个横向的、通用的脑。这个通用脑然后可以被微调以适应不同的垂直领域。我们的论点是,如果一个垂直领域的边缘情况,在另一个垂直领域就变成了中心情况。所以现在数据来自四面八方。现在它将能够通过不同垂直领域的数据玩法来处理这些边缘情况。关于 Deepak 刚才说的,我们确实在形象上非常相似,因为我们都是教授。所以我们不会把工作分成“我做业务,你做这个”之类的。我们更像是……把彼此视为对方大脑的延伸,一起思考、制定战略,并且非常非常专注于部署。人类在某种意义上受到限制——我们无法进入彼此的大脑。我们正在以人类的方式融合全具身智能。
Yeah, I think one thing which I want to start from is like the reason we are actually so excited about this is because we are almost rethinking the way robotics is done traditionally. Traditionally, robotics has been a very classic, like a vertically oriented field, right? I mean, so what that means is if you think before this AI era, you first decide what vertical you want to place the robot in. So let's say I want to build a welding robot. Now you go and start making your hardware, which is very specific to welding. You start making your software, which is very specific to welding. Now, the problem with these kind of deployments has been is it's very easy to guess the first 80% or 90% of the performance, but then you hit this wall which is called the "corner cases" in the physical world. There are so many corner cases in the physical world, like someone might leave a package in front of you, and now it becomes a corner case, and so on. And so that is why if there is a corner case, now because you are at 90% performance, you will still not be able to get it completely automated. Human still needs to be around to make sure the corner cases are handled, and so on. And that is why it has not been traditionally, robotics has not really gone big mainstream, essentially. Now what, however, things have changed when AI came in. Like if you think, language also, before this whole came in, was very verticalized. There were some different companies building chatbots. There were different companies building search engines. But once LLMs came in, they became the horizontal platform. And now everyone is building on top of that horizontal LLM platform. That is exactly how we are now thinking about robotics. We are building this horizontal, general-purpose brain. And this general-purpose brain can then be fine-tuned for different verticals, essentially. And our thesis is that if there's a corner case of one vertical, it becomes the central case of the other vertical. So now the data is from everywhere. And so now it will be able to handle these corner cases through the data play with the different verticals. In terms of what Deepak was talking about, I mean, we are definitely like very similar in that profile, because both of us are professors. So we do not divide our work like, oh, I do business and you do this kind of stuff. We are more... think of it as extension of each other's brain and thinking about it, strategizing about it, and the whole... and really, really focusing on deployment. Like humans are limited in the sense—we cannot enter each other's brain. We are fusing the omni-bodied intelligence in the human way.
和你们聊了五分钟,我有一种感觉,你们可能比你们意识到的更接近融合大脑。我不知道。你们似乎在同一波长上。灵感是什么?我的意思是,你们讨论了,你知道,在某些方面,omni-brain 的灵感是构建那个横向平台。但你们是否看到了现有机器人基础模型中的缺陷或差距,或者真正的推动力是什么,让你们说,嘿,我们需要换一种方式来做这件事?
I have a feeling, from talking to you guys for five minutes, that you might be closer to fusing brains together than you realize. I don't know. You seem to be on the same wavelength. What was the inspiration? I mean, you discussed, you know, In some ways, the inspiration for omni-brain, building that horizontal platform. But were there deficiencies or gaps that you saw in existing robotics foundational models, or what was really the impetus to say, hey, we need to go do this a different way?
我认为如果你看看当前的系统,我想我已经暗示过了。在某种程度上,当机器人目前被部署时,它们更像机器,对吧?所以一切都是被测量的,一切,比如在工厂设置中,一切都是……例如,如果你看一条经典的自动化生产线,你会有一个机器人,但在机器人周围,会有一个大笼子,一切都被精确测量。整个设置的成本可能是机器人本身的几倍。然后如果有什么变化,你必须重新设计整个设置。然后人们谈论消费应用,那里事情会变化。
I think if you look at the current systems, I think I've already alluded to it. In a way, when the robots are currently deployed, they behave more like machines, right? So everything is measured, everything, like in factory setups, everything is... So for instance, if you look at a classical automation line, you will have a robot, but around the robot, you'll have a big cage, everything will be measured very precisely. The whole setup may cost several times more than the robot itself. Then if anything were to change, you have to redesign the whole setup. And then people talk about consumer applications, where things change.
比如说你的家,对吧?你不可能,无论你放多少传感器,你都无法把每一样东西都测量到 0.1 毫米的精度。
Let's say your home, right? You don't, you can't, no matter how many sensors you put, you cannot measure everything, single thing, to 0.1 millimeter accuracy.
当然。
Sure.
对吧?所以整个机器人学的范式……机器人学的主要转变已经发生,从编程行为转向学习行为。这意味着你从数据中学习。所以现在工程部分已经从“我的机器人该怎么动?可能会发生什么故障?”转变为思考“数据从哪里来?我怎样才能让它高质量?我怎样才能大规模获取?”这就是转变所在。
Right? So this whole paradigm of robotics has... the main shift in robotics has happened going from this programming in the behaviors to learning the behaviors. Which means you learn that from data. So now the engineering part has gone from, okay, how should my robot move? What failure may occur? To thinking, where the data will come from. Or how can I make it high quality? How can I get it at scale? And that's where the shift has come.
所以我们看到了学术界的转变。我们开始看到一个接一个的结果。比如我们今天得到一个结果,下周就能在会议上做现场演示。所以对我们来说,要么我们把它带给大众,要么我们最终以某种方式被它取代。所以对我们来说,这是机器人学的未来,这是毫无疑问的,而且……我认为这种认识同时也在整个领域发生。你可以在 GTC 上看到围绕物理 AI 的热情。我们正在与这个领域的几家主要参与者合作来推动这件事。所以这真的不是“哦,这发生了,所以这应该发生”。这是扩展的方式。如果你不这样做,几乎不可能像机器人领域过去那样扩展。
So we saw the shift in academia. Like we began seeing results, one after another. Like we could get a result today and demo, live demo, you know, in a conference the next week. So for us, it was like either we bring it to the masses, or we are the ones who just get eventually replaced by it in some way. So it was just a no-brainer for us that this is the future of robotics, and this is... I think this realization is also happening at the same time in the general field. You can see the excitement around physical AI at GTC. We are working with several major players in this space to bring this. So this is not really, oh, this happened, hence this should happen. This is the way to scale. If you do not do this, it is almost impossible to scale the way how things have been in the robotic space.
我注意到在你们的博客和网站上,我读到一篇关于用视频数据训练的文章。你能谈谈用视频数据训练的好处以及为什么你们要用视频数据训练吗?这是你们训练机器人的主要方式、唯一方式,还是你们也会从其他地方引入数据源?
I noticed on your blog, on the website, I was reading an article about training on video data. Can you talk a little bit about the benefits and why you're training on video data? And is that the primary way, the only way you're training your robots, or are you bringing data sources from other places as well?
所以我的意思是,在机器人学中,关于数据我们有多种选择。数据有三个主要来源。第一个数据来源是视频……或者也许我们先从机器人数据本身说起。现在你要做的是,你必须收集机器人执行任务的数据,而这些数据本身可以用来训练机器人。然而,这很难扩展,因为你是在用机器人收集数据。所以每一条数据点,你都需要一个机器人。你需要人类来控制机器人,因为目前机器人……我们称之为遥操作。所以你必须通过遥操作来收集数据。这种数据的好处是,它是最丰富的数据形式,因为机器人本身在执行任务。所以你可以读取所有的传感器值。你可以读取所有进入机器人的电机指令,等等。这种数据的问题是——它非常非常难以扩展。所以当它难以扩展时,就很难在此基础上学习大规模 AI 模型。
So I mean, when it comes to robotics, we have multiple choices when it comes to data. So there are three main sources of data. The first source of data is videos... or maybe let's start with the robot data itself. So now where you will do it is, you have to collect robot doing a task, and that data itself can be used to train the robot. However, this is very hard to scale, because you're collecting data with robots. So for every data point, you need a robot. You need humans to control the robot because currently robot... and we call this teleoperation. So you have to collect data with teleoperation. The good thing about this data is, it's the richest form of data, because robot itself is doing the task. So you can read all the sensor values. You can read all the motor commands that are going in the robot, and so on. The problem with this form of data—it's very, very hard to scale. And so when it becomes hard to scale, it's very hard to learn large-scale AI models on top of it.
第二种数据形式是类似视频的数据。在这种情况下,数据有巨大的多样性,因为我们在美国收集视频,人们在印度、中国,到处都在收集视频。所以你在各处都有巨大的动作多样性,等等。所以这是一种可扩展的数据形式,高度多样化。但这种数据的问题是它不够丰富。你不知道人们执行任务时具体做了什么动作,施加了多大的力。
The second form of data is something like videos. Now in this case, there's huge diversity of the data because we are collecting videos in U.S., people are collecting videos in India, China, everywhere. So you have huge diversity of the actions everywhere, and so on. So this is a scalable form of data, highly diverse. But the problem with this form of data is that it's not rich enough. You do not know what exact actions, what exact forces people are applying to do it.
然后还有第三种数据形式,即模拟数据。在这种情况下,它是高度可扩展的。模拟是尽可能可扩展的。例如,你可以在一天内收集数万亿个示例,等等。它也是……你可以在模拟器中测量所有的力,等等。但模拟器的问题是,总是存在人们所说的“模拟到现实差距”,比如模拟器不可能是真实世界的精确复制品。总是有一些差异。所以现在你必须通过算法或其他数据来弥合这个模拟到现实的差距,等等。
And then there's a third form of data, which is the simulation form of data. Now, in this case, it's highly scalable. Simulation is as scalable as it gets. You can collect trillions of examples in a day, for example, and so on. It is also... you can measure all the forces in a simulator, and so on. But the problem with simulator is there's always what people call sim-to-real gap, like simulator cannot be exact replica of the real world. There's always some difference. And so now you have to bridge this sim-to-real gap, either through algorithms or some other data, and so on.
所以在 Skild,我们实际上使用了所有三种不同形式的数据。我们相信……每一种数据形式都至关重要,因为每一种数据形式都与其他形式互补。比如,我的意思是,如果你认为视频是可扩展和多样化的,模拟是可扩展但不多样化的,那么第三种是机器人数据,它是最丰富的数据形式。所以每一种数据形式都有用。但有些数据有不同的指标。视频对于机器人训练的质量不如,例如,真实世界数据。所以我们所做的是使用视频数据来预训练我们的模型。这些数据已经有数十亿可用。所以我们可以预训练我们的模型来构建一个模型。
And so for at Skild, we use actually all three different forms of data. We believe... every form of data is critical, because every form of data is complementary to others. Like, I mean, if you think videos are scalable and diverse, simulation is scalable but not diverse, and then the third one is the robot data, which is the richest form of data. So every form of data is useful. But some data has different metrics. Videos is not as good quality for robot training as, for example, the real-world data. So what we do is we use the video data to pretrain our models. This is the data that is available in billions already. So we can pretrain our models to build a model.
然而,视频的问题是……如果我们能从视频中学到一切,Deepak 给我们举了一个很好的例子,如果我们能从视频中学习,我们所有人都会成为费德勒,因为我们会看费德勒,然后我们就会像费德勒一样打球,等等。所以这永远不会足够。仅仅观看视频是不够的。
However, the problem with videos is... If we can learn everything from videos, Deepak gave us a great example, that if we can learn from videos, all of us would be Federers, because we'll watch Federer, and we'll start playing like Federer, and so on. So that's never going to be sufficient. Just watching videos is not going to be sufficient.
如果这足够了,我就能扣篮了,但我不能。
If it was sufficient, I could dunk a basketball, but I can't.
没错。我们不能。所以这就是模拟发挥作用的地方。我们从视频中了解任务是什么、动作是什么,但然后我们在模拟中练习它。我们在模拟中使其更稳健。但同样,模拟仍然存在差距。记住,模拟到现实的差距仍然存在。现在,我们把这个模型,它已经在视频和模拟上进行了预训练,但在部署之前,我们在真实世界数据上对其进行后训练。在工厂或我们试图解决的任何任务中,我们能收集到的少量真实世界数据。这使它变得精确,并帮助它解决问题。
Exactly. We cannot. And so that is where, for us, simulation comes into play. We get the idea of what the task is, what the action is from video, but then we practice it in simulation. We robustify it in simulation. But again, simulation, there's still a gap. Remember, the sim-to-real gap still exists. And now, we take this model, which has been pretrained on videos and simulation, but before deployment, we post-train it on the real-world data. On the small amount of real-world data that we can collect in factories or whatever task we are trying to solve. And that makes it precise and help it solve.
所以你从预训练数据中获得了鲁棒性,比如边缘情况。记住我谈到的这些边缘情况。那些视频和模拟帮助你使其更稳健,而使其精确则是后训练数据的作用。
So you get the robustness from this pretraining data like the corner cases. Remember I was talking about these corner cases. Those videos and that simulation helps you to robustify, and to make it precise is where the post-training data comes in.
所以这……你也可以在语言中找到类比。我认为 AI 主要在语言数据的大规模上取得了成功,对吧,但同样的配方也在那里。就像你有这个,当你构建这个通用模型时,比如你先做通用模型,然后你转向专用模型,通用模型是在所有互联网数据上训练的,比如来自不同来源、不同文章,但然后假设你是 OpenAI,你构建了 ChatGPT,然后亚马逊过来说,哦,我将在我的 Amazon.com 网站上部署你的模型,那么你会拿那个模型进行微调。然后你部署它。所以仅来自 Amazon.com 的数据对亚马逊来说会非常高质量,但数量非常少。所以它用于后训练。互联网数据,也许质量较低,因为人们说不同的东西,也许在很多很多地方有垃圾文本。所以它质量低,但在预训练时规模巨大。所以预训练和后训练的分离是当前 AI 革命的管理方式。即使在 NVIDIA,对吧,你有用于推理的芯片,你有用于预训练的芯片。而这是我们正在构建到机器人学中的同样的分离。
And so this... You can also find analogies with language. I think AI has been mainly successful, right, at a massive scale for language data, right, but the same recipe is there. Like you have this, when you are building this general model, like you do a general first and then you go to a specialized model, the general model is training on all of internet data, like from different sources, different articles, but then let's say you are OpenAI, you build ChatGPT, and then Amazon comes and say, oh, I will deploy a robot... sorry... your model in my Amazon.com website, then you will take that model and you will fine-tune it. And then you deploy it. So then data from just Amazon.com will be very high quality for Amazon, but very low in amount. So it's used for post-training. Internet data, maybe it's low quality because people are saying different things and maybe there is junk text many, many, many places. So it's low quality, but at massive scale in pretraining time. So this separation of pretraining and post-training is how the current AI revolution is governed. Even at NVIDIA, right, you have chips for inference, you have chips for pretraining. And this is the same separation we are building to robotics.
这就是为什么我们能看到对各种应用的即时访问,否则你是无法做到的。你能……你之前已经谈过一点,但也许可以为观众和听众梳理一下脉络。你能谈谈构建、测试、部署以及将像 Omni-Brain 这样的东西推向市场需要什么过程吗?
And which is why we are seeing this immediate access to a variety of applications, which you would not have otherwise. Can you... you've talked about this a little bit, but maybe kind of to put a narrative around it for the viewers and listeners. Can you talk about kind of what it takes, the process of building, testing, and deploying, bringing to market something like the omni-brain?
是的,这是一个非常复杂的问题,因为它确实取决于具体场景,对吧?比如在语言方面,这很容易,因为你可以提问,提示词就能搞定一切。所以我们正在走向的通用方案是,幕后的“大脑”是共享的,明白吗?所以你的每一个动作都会改进这个大脑。那么,我们如何编排这个大脑的部署呢?思路是这样的:假设你有一个任务,如果我们以前见过这个任务,比如移动、行走或跳跃障碍,我们已经能做得很好。在这种情况下,你只需把大脑装上机器人,它就能开箱即用。然后你可以在上面构建应用,比如“我想用机器人自拍或做安全检查”,这是第二部分,对吧?但假设你现在遇到一个不同的任务,比如机器人在传送带上组装 GPU。这跟人们通常做的任务非常不同,连人类都需要训练。所以在这种情况下,我们会在那台机器人上收集几天的数据,或者如果你已经有资产,我们就在仿真中收集数据,两种方式都行。然后我们用这些数据对模型进行后训练,之后模型接管并直接控制机器人。这样,你就通过添加实际任务的数据,弥合了之前所见与全新任务之间的差距,这叫做领域特定数据。
Yeah. So it's a very complex question because it really depends on the scenario, right? Like in language, it's very easy because you can ask a question, it's just prompt does everything. So the general recipe which we are going towards is that the behind-the-scenes brain is shared, okay? So any single action you will take will improve the brain. Now, how do we orchestrate the deployment of this brain? So the idea is, let's say if you have some task, if we have seen that task before, let's say if it's a task of moving around or walking or jumping over things, we can do that already very well. So in that case, you can just take the brain, put on the robot, and we'll just work off the shelf. Then you can build applications on top, like, okay, I want to use the robot for taking a selfie or security inspection. That's the second part, right? But let's say now you go to a different task, where the robot is, I don't know, assembling a GPU on a conveyor belt. Now, it's a super different task compared to what people generally do. Even humans need training. So in that scenario, what we do is, on that robot, we may collect data for a few days. Either do that, or if you already have the assets, then we'll collect data in simulation. Either way. Then we use the data and we post-train the model, and then that model takes over and it turns on the robot directly. So in this case, now what you have done, you have bridged the gap between what you saw before to a very different task by adding data from the actual task. So it's called domain-specific data.
对。
Right.
现在,随着你部署越来越多的机器人,想象你得到了一支专家舰队,而它们都来自一个通才,对吧?这很像高中时你学很多科目。对,对。我读了博士,现在几乎不懂化学、物理了,对吧?对。但我需要那些知识才能获得现在的知识。所以当你有了这个专家,数据可以从所有专家那里回流到幕后的同一个大脑,这在人类身上不会发生,但我们可以在计算机里做到。现在,当这种情况发生时,当你面对下一个任务时,你需要的下一个任务的数据就会更少。这就是我们所说的数据飞轮。你可能听说过自动驾驶里的这个术语,比如人类开车。现在,我们跨垂直领域编排这个数据飞轮。所以从工厂开始,它们作为半结构化场景(如医院、杂货店、酒店)的数据飞轮。那里的数据飞轮帮助你达到终极挑战,比如家庭、消费机器人。这基本上就是我们如何编排每个开发中的自持续数据飞轮循环。这也是为什么你现在能理解我们为什么有 Omni-Brain,因为你希望利用每一个数据点,并将其用于下一个复杂任务。
Now, as you deploy more and more of these robots, imagine you are getting a fleet of specialists, which all came from a generalist, right? So it's very much like, you know, when you're in high school, you know many subjects. Right, right. I did Ph.D. I barely know any chemistry, physics at this point, right? Right. But I needed that to get the knowledge I have now. So then when you have this specialist, Then the data can pull back from all of them and come to the same brain behind the scene, which is not how what happens in humans, but we can do it in a computer. And now this happens, now when you have a next task to go to, you may need, you will need less data for the next task. Now this act as a... this is what we call, in other words, a data flywheel. Like you may have heard this term for self-driving, like humans drive cars. So this data flywheel, now we orchestrate this across verticals. So you start with factories. They act as a data flywheel for semi-structured scenarios like hospitals, grocery stores, I don't know, like hotels. Data flywheel from there helps you get to the ultimate challenge, which is like homes, consumer robots. So this is basically how we are orchestrating the self-sustaining data flywheel loop from every development. And this is why you can probably understand now why do we have an omni-bodied brain? Because you want to take benefit of every single data point and use it for the next complex task.
同样的概念适用于不同的形态吗?
And does the same concept apply to different form factors?
是的,我的意思是,在工厂里是机械臂;在家里可能是人形机器人或其他形态;在安全和巡检方面是狗形机器人;在配送方面又是不同的形态。所以跨越所有形态。
Yeah, I mean, on factory, it's a robotic arm. In home, probably some humanoid or some other form factor. For security and inspection, be a dog-like robot. In delivery, a different form factor. So across all factors.
所以我想问你们一下,你们如何使用 NVIDIA 技术,特别是围绕合成数据和仿真,正如你提到的,但基本上就是开放式问题。你们在用哪些 NVIDIA 的东西,它们如何融入?
So I want to ask you guys a little bit about how you're using NVIDIA technology, and specifically around synthetic data and simulation, as you mentioned, but really just kind of open-ended. How are you, what NVIDIA stuff are you using, and how does it fit in?
我的意思是,我们公司成立两年半了,但我个人从 2018 年就开始与 NVIDIA 合作,不是在那里工作,而是与他们合作。所以有整个仿真套件,比如 Isaac Sim,早前还有 PhysX 和 Isaac Gym。我们使用其中的物理组件来创建海量场景,让我们可以尝试和练习,就像 Abhinav 描述的练习和学习。所以我们是资深用户,现在我们也在与 NVIDIA 合作 Newton。事实上,我们正在共同开发更好的物理求解器。哦,太好了,是的。可能我们会一起开源它们。这是仿真方面的合作。第二方面是视频模型,比如 Cosmos 和其他模型。我们用它们做数据增强,比如每个数据点,你可以得到它,并用这些生成式 AI 模型创建多个变体。所以我们利用并在这方面合作。我认为最重要的是整个计算平台。因为机器人是下一代设备,对吧?那种适用于 LLM 的服务器大 GPU 解决方案,对机器人来说会非常不同。因为机器人如果摔倒,没有时间连接服务器,它必须立即反应。所以在设备端边缘计算方面,我们也在合作。
I mean, so our company is two-and-a-half years old, but I have been working personally with NVIDIA, I think, since 2018, like not at NVIDIA, working with them. Like, so there is this whole, this suite of simulation, like Isaac Sim, back in the day, there was PhysX and Isaac Gym. So we use that, the physics component of that. to really create these gazillion scenarios on which we can try and practice, like what Abhinav was describing, practicing and learning. So that's, we are basically the OG user, and we are now working with NVIDIA on like Newton as well. And in fact, we are co-developing better physics solvers. Oh, great, yeah. Probably we'll open source them together. That's one collaboration on simulation side. Second side is the video models, like the Cosmos and other models. So we use them to data augmentation, like every data point, you can get that, and you can create multiple variations with these generative AI models. So we leverage, we partner on that front. And I think the biggest of all is the whole compute platform. Because robots are the next-generation device, right? This solution that worked for LLMs of big GPUs in servers, it will look very different for a robot. Because a robot doesn't have time to connect to a server if it's falling. It has to react immediately. So on-device edge compute, this is where we are partnering as well.
太好了。那么当你们测试 Omni-Brain 时,也许是在与新伙伴合作或开发新功能时,你们有没有一个首选的测试案例或场景?或者带我们了解一下,在准备部署之前测试东西是什么样的?
Excellent. So when you're testing omni-brain, maybe when you're using it with a new partner or developing a new feature, do you have kind of a go-to test case, a go-to scenario that you put it through? Or walk us through what that's like, kind of testing something before you're ready to deploy it.
是的,我觉得这是个好问题。我的意思是,虽然这也很难,因为那是一个通用的问题,对吧?
Yeah, I think that's a great question. I mean, although this is also very hard, because that's a problem is something general purpose, right?
对,对。
Yeah, yeah.
这正是 Deepak 在谈论的通用大脑。现在,如果你针对某个专门任务进行微调,比如我带来一个专用大脑,它是否应该忘记通用部分?通用部分重要吗?可能不重要,但如果出现边缘情况,那就重要了。所以这些才是关键。这就是为什么我们一直在尝试制定非常具体的测试策略。首先,当然,我们必须在任务本身上测试。假设我们正在做,以我们与 NVIDIA 合作的 GPU 为例,比如在服务器上的 GPU 机架上安装母线。现在有两个要求:第一,必须正确安装,这是精度部分;第二,安装需要多长时间?如果安装一条母线需要一天,那对任何部署来说都不够好,等等。所以我们的测试有这些 KPI,我们首先在这些 KPI 上测试。这些是我们试图匹配的任务驱动型 KPI,等等。
And that's what Deepak was talking, about a general-purpose brain. Now, if you're fine-tuning it for something specialized, like I'm bringing a special brain, should it forget the general part of it is, does it matter, general part of it, or not? It probably does not matter, but then it matters if there was a corner case that was coming in and so on. So those are the kind of things that matter. So this is why we have been trying to develop a very specific strategy of testing these out. So the first thing, of course, we have to test out is on the task itself. Let's say we are putting, let's take the example of the GPU that we have been working with NVIDIA as well as a partner as well. Like putting a busbar on a GPU rack on a server. Now, there are two requirements. First, it has to be put properly. So that's the accuracy part of it. And then how much time does it take you to put? If it takes you one day to put one busbar, that's not good enough for any deployment, and so on. So our testing has these KPIs that we first test on. These are the task-driven KPIs that we are trying to match, and so on.
但仅仅做 KPI 是不够的,因为这就是为什么有 90% 通过 KPI 完成、或者 95% 通过 KPI 完成,但剩下的 5% 也很重要,而这正是我们去测试泛化能力的地方。我们说,好吧,如果有人在这里留下一个箱子,或者如果灯光完全熄灭,或者我们改变这些条件,我们有一系列想要测试的条件,即使这些情况发生,机器人要么继续工作,但仍然要安全。安全也是第三个方面,在所有这些条件下,我们必须确保机器人是安全的,不会做出任何意外行为,等等。所以我们基本上有整个流程,首先从任务指标开始,然后是泛化指标,比如如果出了问题。我的意思是,这是你意想不到的事情,但你仍然希望你的机器人对这些事情具有鲁棒性。我们在部署之前制定了一整套清单。好的,这些是我们想要在泛化方面测试的内容。最后是安全,在任何情况下都不应该违反安全规定,等等。所以我们在部署之前也设置了所谓的安全护栏。这确保,比如说,不知怎么的……有人弄断了电线,切断了摄像头线,因为现在机器人失明了,它什么都看不见。所以这是一个安全指标,我们需要确保护栏现在介入并说,好吧,如果我也看不到摄像头,我应该停下来……或者至少我不应该越过那些东西给我的边界。所以这些都是你必须测试的内容。再说一次,物理世界的问题是,它不像一夜成名,你把它放在网页上,现在每个人都可以访问它,等等。我们必须经过非常严格的测试,才能把任何东西上线部署。
But just doing KPIs is not sufficient, because that is where the whole idea that 90% is done through KPIs, or 95% is done through KPIs, but the rest of the 5% is also what matters, and that's where we go and test for generalization. We say okay, what if someone left a box here, or what if somehow the lights were completely off or like we change these conditions, and we have these set of conditions that we want to test in, like, even if these things happen, the robot will either continue to work. But still be safe. Safety is the third aspect of it as well, like in all these conditions, we have to ensure that the robot is safe and it's not doing any unexpected behavior, and so on. So we basically have this whole pipeline where we first start from task metrics. Then generalization metrics, like if things go wrong. I mean, this is something which you're not expecting, but you still want your robot to be robust to those kind of things. And we have like a whole list that we develop before we deploy that. OK, these are the things that we want to test on when it comes to generalization. And last is the safety. That in no scenarios that you should break the safety violations, and so on. So we put something called safety guardrails also before the deployments. That ensures that, let's say, somehow... somehow someone broke the wire and cut the camera wire, because now the robot is blind, it doesn't see anything. So that's a safety metric that we need to make sure that now the guardrails come in and say okay, if I'm not seeing a camera either, I should stop... or at least I should not cross the boundaries that I have been given by those things. So these are all the things that you have to test for. Again, the problem with the physical world is that It's not like an overnight sensation that you can become, you put it on a web page, and now everyone can access it, and so on. We have to go through very rigorous tests before we can put anything online for deployment.
完全同意。这是我在收尾时最喜欢问的问题之一。你认为机器人的未来会是什么样子?而且,你知道,我们试着设定一个时间框架,明年、未来两年。现在事情发展得太快了。尤其是,正如你谈到的物理 AI,你知道,AI 的具身化真的,你知道,尤其是今年,我觉得我们看到了更多这样的东西。但你认为机器人技术在未来几年、五年,或者任何合适的时间框架内会如何发展?
Absolutely. So this is one of my favorite questions I always ask as we start to wrap up. What do you think the future of robotics looks like? And, you know, we try to put a timeframe when next year, next two years. Things are moving so quickly these days. And particularly, as you were talking about with physical AI, you know, the embodiment of AI is really, you know, this year in particular, and I think we're seeing so much more of it. But how do you see robotics developing in, you know, the next few years, five years, whatever the right timeframe is?
我认为在更长的时间线上,我们将能够自动化人类在物理世界中能做的每一个动作。对吧?因为我们遵循的方法与自然界中事物实际发生的方式非常相似。现在的时间线……在某种意义上,你走得越远,就越意识到这是实现通用智能的途径。目前,我们迄今为止所拥有的是语言模型、视觉模型的结果。这都是人们所说的数字智能。但数字世界,如果你仔细想想,不超过 50 年的历史。
I think in the longer timeline, we will be able to automate every single action that humans can take in the physical world. Right? Because we are following the approach which is very similar to how this actually things happen in nature. Now the timeline... and in some sense the longer you go, the more you realize that this is the way to achieve general intelligence. Currently, what we have so far are the results in language models, vision models. It is all what people call digital intelligence. But digital world, if you think about this, is not more than 50 years old.
这是个好观点,是的。
That's a good point, yeah.
在那之前人类就不聪明了吗,对吧?所以这是长期愿景,对吧?现在,这如何编排?嗯,在我们看来,你将在很短的时间内开始看到这类模型使事情自动化,但……首先是高复杂度、可重复、可能变化较少的场景。所以这就像我们所说的非结构化、半结构化,比如工业任务仓库。它们充当垫脚石,我之前说过,通往更非结构化或半结构化的场景,更多半结构化的场景。这是一个谱系。结构化就像一切都映射好了,比如微波炉。在微波炉内部,你并不真正关心,你在它运行时不会把手伸进去,它是一个完全独立的系统,对吧?另一端是家庭,完全非结构化,这是一个谱系。所以就在今年,我们将开始看到在工厂、仓库、人群周围的部署,这为下一个领域提供动力,比如医院、酒店、服务业,这又为最终的消费级机器人提供动力。很难预测最终家用机器人的时间线,但你肯定会开始看到机器人。而且你已经在今年或未来几年看到这种情况发生了。
Were humans not intelligent before that, right? So this is the longer-term vision, right? Now, how does this orchestrate? Well, in our opinion, you will start to see already things getting automated with these kind of models in a very short horizon, but... high-complex, repeatable, maybe less variable scenarios first. So it's like what we call unstruc-, semi-structured, like industrial task warehouses. They act as a stepping stone, I was saying earlier, to get to more unstructured or semi-structured scenarios, more semi-structured scenarios. This is a spectrum. Structured is like everything is mapped, like a microwave. Inside microwave, you don't really care, you don't put your hand in when it's running, it's a completely separate system, right? Other part is home, which is completely unstructured, it's a spectrum. So in this year itself, we'll start to see deployments in like factory, warehouse, around people, that bootstraps the next one, like hospitals, hotels, service industry, that bootstraps the ultimate—consumer robots. It's very hard to predict the timeline for the ultimate home robots, but you will start to see robots for sure. And you're already seeing that happening in this year or in the next couple of years.
我认为从长远来看,我们都同意机器人将无处不在,做每一项任务。我想每个人都同意。所以短期内,我们,至少在公司内部,都同意今年我们将让工厂和仓库等结构化场所越来越自动化。比如渗透率将在今年年底开始上升,越来越高的渗透率。而中间部分是不清楚的,这就是为什么我们公司内部总是有赌注池,比如冰淇淋赌注之类的,我们一直在打赌这些什么时候会实现?每个人都有不同的看法,比如有些人认为家用机器人可能还需要两三年,但有些人则认为两三年仍然很难,我的意思是我们必须诚实,我们必须说,好吧,现实世界中可能发生的不确定性非常高。而且,虽然你在人形机器人领域看到这么多硬件,但这些硬件今天甚至能可靠地放进家庭吗?没有人这样做过。因为安全,再说一次,是个大问题。比如当你把它们放在家里时,如果它摔倒了,旁边有个孩子,之类的事情,对吧?所以我们公司内部有所有这些赌注在进行,等等。我认为我们俩在短期和长期上都有共识,但中间部分没人知道,我们只是在摸索,好吧。我们边走边看。
I think in the long run, we all agree that robots are going to be everywhere, doing every task. And I think everyone agrees. And so, shorter term also, we are, like at least in the company, we are all in agreement that this year, we are going to have like the structured places like factories and warehouses being more and more automated. Like the penetration will start to happen by the end of this year, more and more penetration. and it's a middle which is unclear, and that's where we always have a betting pool inside a company also like gelato bets and all these kind of bets that we keep going on then when when will these things come into play? Everyone has a different view, like some people believe that home robots might still come in two, three years, but then some people are arguing that two, three years is still very hard, I mean we have to be honest and we have to say, okay, like the kind of uncertainty that can happen in the real world is very, very high. And while you're seeing so much hardware in humanoid space also, are these hardware reliable to be even put in homes today? No one has put them. Because safety, again, is a big issue. Like when you are putting them in home, what if it falls and there's a child around and something like that, right? So we have all these kind of, within the company, all these pools going on, and so on. And I think both of us are kind of like agree on the short term and the long term, but it's middle where no one knows, and we are just figuring it out, okay. We are playing it as along.
有趣的是,事情的发展非常令人惊讶。我的意思是……
The interesting part is, it's very surprising how it's playing out. I mean...
怎么说?
How so?
我的意思是,因为从 AI 的角度来看,对吧,当我在 2008 年攻读博士学位时,我绝对猜不到我们在 AI 领域会走到今天这一步。而且它实际上继续带来越来越多的惊喜,比如……如果你三年前问我我们今天会在哪里,那也是非常令人惊讶的。
I mean, because I mean from the AI perspective, right, when I was doing my Ph.D. in 2008, would have never guessed this where we are in AI. And it actually continues to surprise even more and more, like... if you asked me three years ago where we would be today, that also is very surprising.
是的。
Yeah.
所以算力的进步和硬件成本的下降让这一切变得如此令人惊讶,以至于我想说,即使是像我们这样在这个领域工作了 20 年的专家,也不敢在网上说什么。可能你知道这个,对吧,这是一句引语,我不记得是谁说的,但可能是比尔·盖茨在某个地方提到过——人类在短期内极其乐观,在长期内却悲观。我认为这很适用,这就像一个现实世界的悖论。
And so the progress of compute and the hardware costs coming down has just made this all so surprising that I would say even the experts like us, who have been working in this for 20 years, are scared to say anything online. Probably you know this thing, right, this is a quote, I'm not remembering from whom, but probably Bill Gates mentioned it somewhere—Humans are extremely optimistic in the short term and pessimistic in the long term. I think this applies, this is like a real-world paradox.
所以我的百万美元问题是,我什么时候能有一个能叠衣服的机器人?这就是我想要的。
So my million-dollar question is, when am I going to have a robot that can fold my laundry? That's the task I want.
嗯,问题是,你今年就能拥有那个机器人,但如果它只是待在角落里做那件事,你还得给它送衣服,还得给它送……你真的会想要它吗?
Well, the thing is, you can have that robot this year, but if it does just that in a corner, you have to bring it clothes, you have to bring it, like, would you really want it?
不,说得对……我觉得这正是关键所在。
No, fair... That's the whole point, I think.
是啊,完全有道理。
Yeah, no, absolutely fair point.
但如果你能实现同样的事情,而且它是在工厂里执行更复杂的任务,需要每天无人值守运行,那你会想要它吗?当然会。人们都排着队想要呢。所以这其实是同一件事,只是视角不同。
But if you can do the same thing, and it's doing something maybe more complex in a factory where you have to run lights out every day, then would you want it? Of course. People are in line for that. So it's just the same thing, but different perspective.
没错,绝对如此。那么 Skild 接下来有什么计划?你们现在在做什么?在技术方面有没有探索新领域,或者正在开拓的新行业、新业务方向?公司的路线图是什么样的?
No, absolutely. And so what's next for Skild? What are you guys working on now? Are there new areas you're exploring on the technical side, new industries or business avenues that you're breaking into? What's the company roadmap look like?
有一件事,根据发布时间,可能就在这几个月里,我们一直高度专注于如何将这个通用模型转化为可快速大规模部署的专用系统。比如,通过少量微调,在几天内让新系统上线运行,并利用这一策略扩展到尽可能多的场景。背后的原因是真正启动这个通用数据飞轮。飞轮需要时间建立,需要时间积累动力,如果我们希望这些事按计划的时间表发生,就必须现在开始,这是我们的主要焦点之一。并不是说技术上我们已经到位,所有问题都解决了,但机器人领域的大规模部署本身就是一个技术挑战。不像语言或其他领域,你构建了东西,它就会被部署,因为人们会使用它或想办法使用它。但在这里,部署本身就是巨大的技术挑战。如何大规模协调?以前从未有人做过。这就是我们重点投入的方向。
One thing like in this... depending on when it gets released, in these couple of months, we have been ultra-focused on how do we take this general model and convert it into specialized systems which can be deployed at scale very quickly. Like get a new system up and running in a couple of days with a small amount of fine-tuning and use that strategy to scale to as many scenarios as possible. And the reason behind that is to really get started on this general data flywheel. Flywheel takes time to set up, takes time to get momentum, and if these things are to happen in the timeline we want them to happen, we have to start now, and this is one of our main focus. Not saying that technologically we are there, like everything is solved, but this is a big deployment in robotics is a technical challenge. Unlike language or other areas, where if you build the thing, it will get deployed, because people will use it or figure out how to use it. But here, deployment, in itself, is a big technical challenge. And how do you orchestrate that at scale? It has not been done before. So this is what we are focusing on a lot.
这太棒了,用你的话来说,它不会减速,只会越来越惊人,至少短期内是这样,对吧?谁知道长期会带来什么?但真是令人着迷。祝你们两位好运,再次感谢 Deepak 和 Abhinav 抽出时间参加播客。
It's amazing stuff, and to sort of paraphrase you, it's not gonna slow down, it's only gonna get more and more amazing, at least in the short term, right? So who knows what the long term has to bring? But just fascinating stuff. Best of luck to both of you, and again, Deepak and Abhinav, thank you so much for taking the time to join the podcast.
非常感谢你们的邀请。
Thank you so much for having us.