日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
モバイルマニピュレーションarXiv:2609.35652

MM-ABC: 見て、協調し、想像する汎用モバイルマニピュレーション

MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining

シェア:XThreadsFacebookLINEはてブBluesky

移動とマニピュレーションを別々のストリームとして扱い、マスク付きジョイントアテンションと未来予測の追加監督で協調させる基盤モデルを提案。

詳しい要約

1. どんなもの?

- モバイルマニピュレーションのための基盤モデル MM-ABC を提案。 - Seeing, Coordinating, Imagining の3要素で Arm-Base Collaboration を実現。 - 疎な multi-level VLM 特徴で空間知覚、world imagination と geometric intent を訓練時のみの future branch で補助監督。 - MM-APT が manipulation と mobility の別ストリームを masked joint attention と clean-action x-prediction で協調。 - 5,000+時間、400K+エピソード、12データセット、17 embodiments で事前学習。

2. 先行研究と比べてどこがすごい?

- 従来は明示的3D表現や予測的 world model で幾何を強化し、mobility と manipulation を別アクションストリームに分離。 - 本手法は分離だけでなく、ストリーム間協調を支える表現を重視。 - 訓練時のみの future branch で world imagination と geometric intent を追加監督し、知覚と manipulation-intent 予測を強化。 - MM-APT により異種アクションの協調を実現。

3. 技術・手法の肝は?

- 疎な multi-level VLM 特徴を空間知覚に利用。 - 訓練時のみの future branch が world imagination と geometric intent を補助監督として使用。 - MM-APT が manipulation と mobility の別ストリームを masked joint attention で協調。 - clean-action x-prediction を採用。 - 大規模異種ロボットデータで事前学習。

4. どうやって有効だと検証した?

- 制御された ablation を実施。 - clean-action prediction を velocity prediction に置換すると RoboCasa365 composite-seen で成功率が32.8%から29.2%に低下。 - future supervision や multilevel conditioning の除去でさらに大きな低下。 - EBench, RoboCasa365, ManiSkill-HAB, LIBERO, LIBERO-Plus, 実世界モバイルマニピュレーションで評価。 - EBench 44.71%、RoboCasa365 61.2%、LIBERO 99.1%、LIBERO-Plus 82.8%、実世界5タスク平均83%。

5. 議論はある?

- 有効なモバイルマニピュレーションには分離だけでなく、ストリーム間協調を支える表現が必要と主張。 - future supervision と multilevel conditioning の重要性を ablation で示す。 - その他の限界や議論は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として explicit 3D representations, predictive world models, decoupled mobility and manipulation が挙げられる。 - 同分野の定番として mobile manipulation, VLM-based robot learning, world models の論文を読むと良い。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Qiwei Liang, Guangyu Chen, Shaolong Zhu, Zikuan Xiao, Jinxuan Lu, Yifan Xie, Renjing Xu, Wenbo Ding, Tianxing Chen

分類: cs.RO

原文アブストラクト

Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base actions. Existing approaches strengthen geometry through explicit 3D representations or predictive world models, and often decouple mobility and manipulation into separate action streams. We argue that effective mobile manipulation requires not only decoupling, but also representations that support efficient cross-stream collaboration. We present MM-ABC, a foundation model built around Seeing, Coordinating, and Imagining Arm-Base Collaboration. MM-ABC combines sparse multi-level VLM features for spatial perception; a training-only future branch that uses world imagination and geometric intent as extra supervision, strengthening perception and manipulation-intent prediction and improving the overall learning signal; and MM-APT, which coordinates separate manipulation and mobility streams through masked joint attention and clean-action x-prediction. In controlled ablations, replacing clean-action prediction with velocity prediction lowers success on RoboCasa365 composite-seen tasks from 32.8% to 29.2%, and removing future supervision or multilevel conditioning causes larger drops. We pretrain MM-ABC on 5,000+ hours of heterogeneous robot data spanning 400K+ episodes, 12 datasets, and 17 embodiments. Experiments cover EBench, RoboCasa365, ManiSkill-HAB, LIBERO, LIBERO-Plus, and real-world mobile manipulation. MM-ABC achieves 44.71% success on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, 82.8% on LIBERO-Plus without perturbation training, and 83% mean success on five real-world tasks.

関連論文

PR本紙発行元 EmplifAI