日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.37793

MVG-WAM: ロボットマニピュレーションのための多視点幾何認識型ワールドアクションモデル

MVG-WAM: Multiple View Geometry-Aware World-Action Modeling for Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

多視点カメラの幾何学的関係を明示的に組み込んだワールドアクションモデルを提案し、LIBEROやRoboTwin 2.0で高い成功率を達成した。

詳しい要約

1. どんなもの?

- ロボット操作のための **MVG-WAM**(Multi-View Geometry-Aware World-Action Model)を提案。 - World-Action Models (WAMs) は視覚動態と行動予測を結合し、事前学習済み video model の prior を活用する。 - 従来の多視点インターフェースは画像をタイル化または token 連結するため、同期カメラ間の幾何関係が暗黙的。 - 本手法は多視点観測を「一枚のキャンバス上の別画像」ではなく「同一物理世界の関連する投影」として組織化。 - epipolar 制約付き global state と view-indexed geometric states を統合し、action prediction の表現を明示的に構造化する。

2. 先行研究と比べてどこがすごい?

- 従来の WAMs は多視点入力を単純に tile または token 連結し、カメラ間の幾何関係を暗黙的にしか扱わない。 - その結果、global scene context と interaction に必要な local geometry の接続が困難だった。 - MVG-WAM は epipolar 制約と view-indexed geometric states により、多視点を一つの物理世界の投影として明示的に構造化。 - camera-aware routing で各 video region に幾何 context と共有 global state を供給。 - これにより global と local の幾何情報を明示的に結びつけ、action prediction の表現を改善。

3. 技術・手法の肝は?

- 同期観測から epipolar-constrained global state と view-indexed geometric states を jointly 推論。 - camera-aware routing により、各 video region に対応する geometric context と共有 global state を供給。 - これにより action prediction に用いる表現を明示的に構造化。 - multi-horizon future-depth supervision により geometry-aware 表現を metric scale に接地。 - action rollout 時に depth decoding を必要としない点が特徴。

4. どうやって有効だと検証した?

- LIBERO で平均成功率 99.1% を達成。 - RoboTwin 2.0 で平均成功率 92.07% を達成。 - 両ベンチマークで競争力のある性能を示す。 - 実世界実験は Cobot Magic 上で 3 つの manipulation task、150 trials にわたり 91.3% の成功率を達成。

5. 議論はある?

- 要旨からは、限界や失敗事例、計算コスト、スケーラビリティに関する議論は明記されていない。 - 多視点幾何の明示的構造化と metric scale の接地が性能に寄与する可能性が示唆されるが、詳細な ablation や議論は要旨からは不明。

6. 次に読むべき論文は?

- World-Action Models (WAMs) の代表的研究(要旨で具体的名称は挙げられていない)。 - LIBERO および RoboTwin 2.0 ベンチマークに関する研究。 - 多視点ロボット操作における epipolar geometry や camera-aware routing を用いた手法。 - 事前学習済み video model をロボット操作に応用する研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Wenbo Chen, Tianfu Li, Haoxuan Xu, Zhihao Cao, Zhenghan Chen, Zhengming Zhu, Zizhou Luo, Guosheng Yang, Yuan Liu, Lujia Wang, Wen Chen, Haoang Li

分類: cs.RO

原文アブストラクト

World-Action Models (WAMs) couple visual dynamics with action prediction, bringing the rich priors of pretrained video models to robotic manipulation. However, their multi-view interfaces typically tile images or concatenate tokens, leaving the geometric relationships among synchronized cameras implicit. This makes it harder to connect global scene context with the local geometry required for interaction. We introduce the Multi-View Geometry-Aware World-Action Model (MVG-WAM), which organizes these observations as related projections of one physical world rather than separate images on a canvas. Our model combines an epipolar-constrained global state with view-indexed geometric states jointly inferred from synchronized observations. Camera-aware routing supplies each video region with its corresponding geometric context and the shared global state, explicitly structuring the representation used for action prediction. We further ground the geometry-aware representation in metric scale through multi-horizon future-depth supervision, without requiring depth decoding during action rollout. MVG-WAM achieves average success rates of 99.1% on LIBERO and 92.07% on RoboTwin 2.0, demonstrating competitive performance across both benchmarks. Real-world experiments on Cobot Magic further demonstrate a 91.3% success rate across 150 trials spanning three manipulation tasks.

関連論文

PR本紙発行元 EmplifAI