WholeBodyWAM:スケーラブルな動作事前分布を用いた全身ワールドアクションモデルの学習
WholeBodyWAM: Learning Whole-Body World Action Models with Scalable Motion Priors
人間やヒューマノイドの大規模動作データを事前学習に活用し、全身動作を予測するMotion Expertを組み込んだワールドアクションモデルを提案。実機ヒューマノイドの全身マニピュレーション性能を向上させた。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Bowei Zhang, Qiyao Zhang, Shuanghao Bai, Xinhua Wang, Meng Li, Yilei Wang, Leiwang Zhang, Jian Tang, Lu Zhou, Lei Sun, Zhengping Che
分類: cs.RO
原文アブストラクト
Humanoid whole-body manipulation requires coordinated whole-body dynamics, yet large-scale trajectories from a target robot are expensive to collect and difficult to scale. In contrast, whole-body motion from human and humanoid sources is abundantly available, although such data cannot be directly used as embodiment-specific robot actions. This work asks whether these scalable motion resources can instead provide a transferable predictive prior for humanoid world-action modeling. We introduce WholeBodyWAM, a humanoid world-action model that learns whole-body dynamics from large-scale heterogeneous motion before target-robot training. We curate UniMotion-4K, a motion corpus spanning more than 4K hours from human videos, native 3D motion datasets, and heterogeneous humanoid platforms, and canonicalize these diverse sources into a unified motion space. A language-conditioned Motion Expert is then pretrained to predict future whole-body motion without target-robot action supervision. During robot post-training, the pretrained Motion Expert is integrated with Video and Action Experts through asymmetric Mixture-of-Transformers (MoT) attention, enabling predictive scene dynamics and whole-body motion to jointly inform embodiment-specific action generation. Experiments show that WholeBodyWAM consistently benefits from increased motion-pretraining scale, improves future-motion prediction and downstream task performance, and transfers effectively to real-world humanoid manipulation. Moreover, the pretrained motion prior substantially improves data efficiency under limited target-robot demonstrations.
関連論文
- LYRIC: 言語駆動の物理ベース全身接触リッチ物体インタラクション制御全身マニピュレーション
- InterPrior: 物理ベースの人間-物体インタラクションのための生成的制御のスケーリング全身マニピュレーション
- AdaptManip: オンライン再帰的状態推定による適応的な全身物体持ち上げ・運搬の学習全身マニピュレーション
- HumanoidExo: ウェアラブル外骨格によるスケーラブルな全身ヒューマノイド操作全身マニピュレーション
- 人間型ロボットによる大型物体の抱え込み:強化学習を用いた全身マニピュレーション全身マニピュレーション
- SimGenHOI: 生成モデルと強化学習による物理的にリアルな全身ヒューマノイド-物体インタラクション全身マニピュレーション