Exploring how to build general-purpose robots that can perform any task in any environment, drawing lessons from language models and large-scale data collection.
要点 · TL;DR
多样化的真实机器人数据,而非仅仅是规模,是通用机器人的关键。 Diverse real robot data, not just scale, is key for generalist robots.
在所有数据上预训练,然后在精选数据上微调,能解锁复杂任务。 Pre-training on all data then fine-tuning on curated data unlocks complex tasks.
来自语言模型的合成数据帮助机器人遵循开放式指令。 Synthetic data from language models helps robots follow open-ended prompts.
核心观点 · Key points
规模是通用机器人模型的必要条件,但非充分条件;数据的多样性才是关键。 Scale is necessary but not sufficient for generalist robot models; diverse data is key.
先在所有数据上预训练,再在精选数据上微调,能解锁叠衣服等复杂任务。 Pre-training on all data then fine-tuning on curated data unlocks complex tasks like laundry folding.
多样化数据使机器人能在未见环境中成功,缩小泛化差距。 Diverse data enables robots to succeed in unseen environments, closing the generalization gap.
来自语言模型的合成数据帮助机器人遵循开放式指令和插话。 Synthetic data from language models helps robots follow open-ended prompts and interjections.
强化学习将在后训练中发挥重要作用,以提高成功率和速度。 Reinforcement learning will play a large role in post-training for higher success rates and speed.
反共识 · Contrarian takes
来自工业自动化或 YouTube 的大规模数据缺乏通用机器人所需的多样性。 Massive scale from industrial automation or YouTube lacks diversity needed for generalist robots.
在所有数据上训练对复杂任务无效;在精选数据上预训练和后训练效果更好。 Training on all data fails for complex tasks; pre-training and post-training on curated data works better.
前沿模型因缺乏物理世界数据,在机器人视觉理解上表现不佳。 Frontier models struggle with visual understanding for robotics due to lack of physical world data.
机器人领域的合成数据并非模拟的类比,而更接近强化学习。 Synthetic data in robotics is not analogous to simulation but closer to reinforcement learning.
更多资源可能导致算力浪费;约束反而能促进更严谨的研究。 More resources can lead to wasteful compute; constraints can foster more careful research.
本期章节 · Chapters(共 25)
引言:机器人技术的问题Introduction: The Problem with Robotics
规模与数据源的作用The Role of Scale and Data Sources
物理智能方法:真实机器人数据Physical Intelligence's Approach: Real Robot Data
案例研究:叠衣机器人Case Study: Laundry Folding Robot
初始叠衣测试Initial Laundry Folding Tests
叠衣结果Folding Laundry Results
预训练与后训练的定量评估Quantitative Evaluation of Pre-training and Post-training
泛化到其他任务Generalization to Other Tasks
机器人基础模型Takeaways and Limitations
新环境中的训练与测试Foundation Models for Robotics
定量结果与数据多样性Training and Testing in Novel Environments
失败模式与未来工作Quantitative Results and Data Diversity
结论与未来方向Failure Modes and Future Work
分层视觉-语言-动作模型用于开放式机器人控制Conclusion and Future Directions
结束语与问答Hierarchical Vision-Language-Action Models for Open-Ended Robot Control
机器人强化学习Closing Remarks and Q&A
资金与应用Reinforcement learning for robots
世界模型与 VLA 集成Funding and applications
模型规模与检索World models and VLA integration
引言与致谢Model size and retrieval
物理智能领域建设者的机遇Introduction and appreciation
机器人合成数据Opportunities for builders in physical intelligence
学术界与工业界的机器人研究Synthetic data for robotics
架构限制与分词Academia vs industry for robotics research
Architecture limits and tokenizationArchitecture limits and tokenization