Ego-Pi: 自己中心視点の人間・ロボットデータによるVLAファインチューニング
Ego-Pi: VLA Fine-Tuning for Ego-Centric Human and Robot Data
ロボット操作のデータ不足を解決するため、自己中心視点の人間データを活用してVLAモデルをファインチューニングし、ロボットが新しいタスクの意味を学習し、既存スキルを組み合わせて新行動を生み出せることを示した論文。
著者: Ji Woong Kim, Ke Wang, Zipeng Fu, Sirui Chen, Cong Zhao, Jeff Lai, Chelsea Finn
分類: cs.RO, cs.AI
原文アブストラクト
Robotics faces a fundamental challenge of data scarcity. Unlike language or vision research, there is no internet-scale dataset for robotic manipulation. A promising path forward is to leverage egocentric human data, which can be collected more easily, with greater breadth, and at a larger scale. Towards this end, we investigate key design choices for learning across human and humanoid embodiments equipped with dexterous five-finger hands, using the $π_{0.5}$ model as a foundation. Our results show that human data enables robots to learn new task semantics and compose existing skills into novel behaviors without corresponding robot data. The paper website is here: https://egopipaper.github.io/