日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
全身マニピュレーションarXiv:2609.18197

WholeBodyWAM:スケーラブルな動作事前分布を用いた全身ワールドアクションモデルの学習

WholeBodyWAM: Learning Whole-Body World Action Models with Scalable Motion Priors

シェア:XThreadsFacebookLINEはてブBluesky

人間やヒューマノイドの大規模動作データを事前学習に活用し、全身動作を予測するMotion Expertを組み込んだワールドアクションモデルを提案。実機ヒューマノイドの全身マニピュレーション性能を向上させた。

詳しい要約

1. どんなもの?

- ヒューマノイドの全身マニピュレーションのための World-Action Model である WholeBodyWAM を提案。 - 大規模で異種な全身運動データを予測事前分布として活用し、対象ロボットの行動生成に転移する。 - 4K時間超の UniMotion-4K を構築し、言語条件付き Motion Expert を事前学習。 - ロボット後学習では Video/Action Expert と非対称 MoT 注意で統合。

2. 先行研究と比べてどこがすごい?

- 対象ロボットの大規模軌道収集は高コストでスケール困難という課題に対し、人間・ヒューマノイドの豊富な運動データを予測事前分布として利用。 - 異種運動を embodiment 固有の行動に直接使うのではなく、転移可能な予測事前分布として扱う点が新しい。 - 運動事前学習の規模拡大に伴い性能が一貫して向上することを示す。 - 限られた実機デモでのデータ効率を大幅に改善。

3. 技術・手法の肝は?

- UniMotion-4K を構築:人間動画、ネイティブ3D運動データセット、異種ヒューマノイド平台から4K時間超を収集。 - 多様な運動源を統一運動空間へ正規化(canonicalize)。 - 言語条件付き Motion Expert を対象ロボット行動監督なしで事前学習し、将来の全身運動を予測。 - 後学習で Motion Expert を Video/Action Expert と非対称 Mixture-of-Transformers (MoT) 注意で統合。 - 予測シーン動力学と全身運動を embodiment 固有の行動生成に共同で活用。

4. どうやって有効だと検証した?

- 運動事前学習の規模増加に伴い WholeBodyWAM が一貫して恩恵を受けることを実験で示す。 - 将来運動予測と下流タスク性能が向上。 - 実世界ヒューマノイドマニピュレーションへ効果的に転移。 - 事前学習済み運動事前分布が限られた対象ロボットデモ下でのデータ効率を大幅改善。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として World-Action Model、Mixture-of-Transformers (MoT)、Motion Expert、Video Expert、Action Expert が挙げられる。 - 同分野の定番として humanoid whole-body manipulation、motion prior、world model に関する研究が次に読むべき候補。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Bowei Zhang, Qiyao Zhang, Shuanghao Bai, Xinhua Wang, Meng Li, Yilei Wang, Leiwang Zhang, Jian Tang, Lu Zhou, Lei Sun, Zhengping Che

分類: cs.RO

原文アブストラクト

Humanoid whole-body manipulation requires coordinated whole-body dynamics, yet large-scale trajectories from a target robot are expensive to collect and difficult to scale. In contrast, whole-body motion from human and humanoid sources is abundantly available, although such data cannot be directly used as embodiment-specific robot actions. This work asks whether these scalable motion resources can instead provide a transferable predictive prior for humanoid world-action modeling. We introduce WholeBodyWAM, a humanoid world-action model that learns whole-body dynamics from large-scale heterogeneous motion before target-robot training. We curate UniMotion-4K, a motion corpus spanning more than 4K hours from human videos, native 3D motion datasets, and heterogeneous humanoid platforms, and canonicalize these diverse sources into a unified motion space. A language-conditioned Motion Expert is then pretrained to predict future whole-body motion without target-robot action supervision. During robot post-training, the pretrained Motion Expert is integrated with Video and Action Experts through asymmetric Mixture-of-Transformers (MoT) attention, enabling predictive scene dynamics and whole-body motion to jointly inform embodiment-specific action generation. Experiments show that WholeBodyWAM consistently benefits from increased motion-pretraining scale, improves future-motion prediction and downstream task performance, and transfers effectively to real-world humanoid manipulation. Moreover, the pretrained motion prior substantially improves data efficiency under limited target-robot demonstrations.

関連論文

PR本紙発行元 EmplifAI