日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.16644

WholeBodyWAM:世界行動事前分布を全身制御に接地させたヒューマノイド移動操作

WholeBodyWAM: Generalizing Pre-trained World-Action Priors to Humanoid Loco-Manipulation via WBC-Grounded Coordination

シェア:XThreadsFacebookLINEはてブBluesky

視覚予測と行動を同時に学習する世界行動モデルを、全身制御器の意味づけと協調に接地させることで、ヒューマノイドの移動を伴う操作へ汎化させる手法を提案した。

詳しい要約

1. どんなもの?

- ヒューマノイドの loco-manipulation 向けに、visual dynamics・manipulation actions・whole-body control intents を同時予測する WholeBodyWAM を提案 - World Action Models (WAMs) の pre-trained world-action priors を活用 - テーブルトップ/arm-centric 中心だった WAM 研究を humanoid 全身協調へ拡張 - 全体で simulation 91.9% のタスク成功率を報告

2. 先行研究と比べてどこがすごい?

- 従来の WAM は tabletop や arm-centric manipulation が中心で humanoid loco-manipulation は未開拓 - 既存の humanoid 全身行動はゼロから再学習する傾向 - 本研究は pre-trained world-action priors を保持しつつ WBC 意味論を接地 - 再学習ではなく構造的な WBC grounding と coordination で拡張する点が新しい

3. 技術・手法の肝は?

- 未来の visual dynamics・manipulation actions・whole-body control intents を joint に予測 - pre-trained world-action priors を維持 - 異種の whole-body controller (WBC) semantics を grounding - whole-body behavior を coordination する枠組み - 詳細なアーキテクチャや学習手順は要旨からは不明

4. どうやって有効だと検証した?

- 広範な実験を実施 - simulation で overall task success rate 91.9% を達成 - real-world out-of-distribution task progress が 0.23 改善 - WBC 間の success-rate variance を 70% 削減 - 各 baseline との比較で有効性を検証

5. 議論はある?

- pre-trained world-action priors を WBC grounding と coordination で拡張する道筋を示唆 - scalable な humanoid whole-body intelligence への可能性を議論 - 限界・失敗事例・計算コスト・安全性などの詳細は要旨からは不明

6. 次に読むべき論文は?

- World Action Models (WAMs) の代表的研究 - humanoid loco-manipulation の既存研究 - whole-body controller (WBC) 関連研究 - 具体的な参照論文名は要旨からは不明

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhuo Li, Yiming Yao, Jim Tan, Mengjie Jing, Zhipeng Dong, Fei Chen

分類: cs.RO

原文アブストラクト

World Action Models (WAMs) offer a promising approach to general-purpose robot manipulation by jointly modeling visual dynamics and actions. However, most WAM studies focus on tabletop or arm-centric manipulation, while humanoid loco-manipulation remains less explored. To address this gap, we introduce WholeBodyWAM, which jointly predicts future visual dynamics, manipulation actions, and whole-body control intents for generalizable humanoid loco-manipulation. It preserves pre-trained world-action priors while grounding heterogeneous whole-body controller (WBC) semantics and coordinating whole-body behavior. Extensive experiments show that WholeBodyWAM achieves an overall simulation task success rate of 91.9%, with a 0.23 improvement in real-world out-of-distribution task progress and a 70% reduction in success-rate variance across WBCs relative to the respective baselines. These results suggest a path toward scalable humanoid whole-body intelligence by extending pre-trained world-action priors through structured WBC grounding and coordination, rather than relearning whole-body behavior from scratch. Project page: https://wholebodywam.github.io/.

関連論文