日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
移動マニピュレーションarXiv:2610.07511

MobileVISTA: 移動マニピュレーションにおける姿勢汎化のための生成データ拡張

MobileVISTA: Generative Data Augmentation for Pose Generalization in Mobile Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

単一姿勢の実演データから、視覚観測の拡張と動作の再ターゲティングを組み合わせて姿勢摂動に頑健な訓練データを生成し、ヒューマノイドや双腕ロボットの移動マニピュレーションの汎化性能を向上させるフレームワーク。

詳しい要約

1. どんなもの?

- MobileVISTAは、mobile manipulator(例:humanoid)の模倣学習におけるpose generalizationを改善するdata generation framework。 - 単一のcanonical poseで収集されたdemonstrationを、多様なpose-perturbed training dataへ変換する。 - 対象はegocentric platformで、cameraがactuated chain上にあり、robot自身がframe内を占める設定。 - simulated tasks(humanoid, bimanual embodiments)とreal Galaxea R1 Proで検討。

2. 先行研究と比べてどこがすごい?

- 従来法はcameraがactuated chainからrigidly mountedされている、またはarticulated robot geometryがframe外という仮定を置く。 - MobileVISTAは、cameraがkinematic chainの影響を受け、かつchainを観測するegocentric platform(humanoid等)との互換性を狙う。 - 追加のdemonstration収集やtrained generative modelを必要とせず、test時のout-of-distribution poseへのrobustnessを改善。 - 特にcameraがactuated chainに乗り、robotがframeの多くを占めるhumanoidで効果が最大と報告。

3. 技術・手法の肝は?

- demonstrationsをcanonical poseからdiverseなpose-perturbed training dataへ変換するframework。 - 二つをjointlyに行う:(1) egocentric visual observationsのaugmentation、(2) base pose変化を補償するaction retargeting。 - これによりego-centric observationsとend-effector trajectoriesの分布ずれを緩和する。 - 詳細なアルゴリズムや実装は要旨からは不明。

4. どうやって有効だと検証した?

- simulated tasksでhumanoidおよびbimanual embodimentsを対象に検討。 - real Galaxea R1 Proでも検討。 - MobileVISTA-augmented dataで訓練したpolicyが、test時に遭遇する以前はout-of-distributionだったposeに対しrobustness向上を示す。 - 追加のdemonstration収集やtrained generative modelなしでこの改善が得られると報告。

5. 議論はある?

- MobileVISTAの利点は、cameraがactuated chainに乗りrobotがframeの多くを占めるhumanoidで最大となる。 - 単一poseのdemonstrationに依存するend-to-end manipulation policyの脆さ(cmスケールのposeずれで性能急落)が動機。 - 限界や失敗事例、計算コスト、他embodimentへの一般性の詳細は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている具体的な先行研究名は明示されていない。 - 関連手法として、imitation learning、data augmentation、action retargeting、egocentric vision、mobile manipulation、humanoid manipulationの定番研究を挙げる。 - プロジェクトページ(https://mobilevista.github.io)のvideos/appendixも参照候補。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Suzannah Wistreich, Stephen Tian, Isabella Huang, Vitor Campagnolo Guizilini, Sergey Zakharov, Katherine Liu, Jiajun Wu

分類: cs.RO, cs.CV, cs.LG

原文アブストラクト

Mobile manipulators such as humanoid robots are increasingly deployed in dynamic, unstructured environments to perform dexterous manipulation tasks. However, end-to-end manipulation policies trained to imitate demonstration data collected from a single robot pose are brittle: even centimeter-scale deviations in robot pose at deployment can drive ego-centric observations and end-effector trajectories out of the training distribution, leading to sharp drops in performance. We introduce MobileVISTA, a data generation framework that transforms demonstrations captured at canonical poses into diverse, pose-perturbed training data by jointly (1) augmenting egocentric visual observations and (2) retargeting actions to compensate for base pose changes. Unlike prior methods, which assume a camera rigidly mounted off the actuated chain or non-trivial articulated robot geometry largely out of frame, MobileVISTA targets compatibility with egocentric platforms (e.g., humanoids) where the camera is both influenced by and must observe the robot's kinematic chain as it moves. We study MobileVISTA in simulated tasks spanning humanoid and bimanual embodiments, and on a real Galaxea R1 Pro. We find policies trained on MobileVISTA-augmented data demonstrate improved robustness to previously out-of-distribution poses encountered at test time, without additional demonstration collection or a trained generative model. Additionally, we find MobileVISTA's benefit is largest on tested humanoids, where the camera rides the actuated chain and the robot fills much of the frame. Additional videos and appendix can be found on our website: https://mobilevista.github.io

関連論文

PR本紙発行元 EmplifAI