日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.18243

メートル単位で行動する:精密ロボット操作のための計量的相互作用の学習

Acting in Meters: Learning Metric Interactions for Precise Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

物体とシーンに対する動作の計量的な関係を明示的にモデル化し、VLAやWAMモデルの操作精度を向上させるフレームワークを提案した。

詳しい要約

1. どんなもの?

- Vision-Language-Action models (VLA) と World-Action Models (WAM) のための **metric interaction framework** を提案。 - 物体レベルとシーンレベルの相互作用を物理的な Cartesian 空間で同一の metric scale でモデル化。 - 物体レベルでは **Interaction-Centric Tokens (ICTs)** が end-effector pose trajectories を操作対象との相対で表現し、actions と jointly denoise。 - シーンレベルでは **Metric Action Interaction Field (MAIF)** が action と ICT queries を用いて metric scene point-cloud features に attend し、geometry-conditioned action corrections を学習。 - 2段階適応により、多様な VLA/WAM ベースラインを少数の追加…

2. 先行研究と比べてどこがすごい?

- 従来の VLA/WAM は actions, objects, scene geometry 間の metric relations を暗黙的にしか扱わなかった。 - 本研究は人間の操作が意味理解と空間フィードバックを組み合わせることに着想を得て、物理的 Cartesian 空間で明示的に metric interaction をモデル化。 - 物体レベルとシーンレベルの両方で metric scale を共有し、幾何条件付きの action 補正を学習する点が新しい。 - 既存ベースラインに対し、少数の追加パラメータと訓練ステップで大幅な性能向上を実現。

3. 技術・手法の肝は?

- **Interaction-Centric Tokens (ICTs)**: end-effector pose trajectories を操作対象との相対で表現し、actions と jointly denoise することで物理的に根拠のある interaction supervision を提供。 - **Metric Action Interaction Field (MAIF)**: action と ICT queries を用いて metric scene point-cloud features に attention し、geometry-conditioned action corrections を学習。 - **2段階適応**: 既存の VLA/WAM ベースラインに少数の追加パラメータと訓練ステップで組み込む。

4. どうやって有効だと検証した?

- **LIBERO** で平均成功率 **0.80 percentage points** 向上。 - **RoboTwin 2.0** で平均成功率 **3.59 percentage points** 向上。 - 実世界タスクで **6.80 percentage points** 向上。 - その out-of-distribution variants で **7.45 percentage points** 向上。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- **Vision-Language-Action models (VLA)** - **World-Action Models (WAM)** - **LIBERO** - **RoboTwin 2.0**

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Lijie Wang, Zheng Lu, Yiming Wang, Heyang Yu, Kenghou Hoi, Bowen Hu, Di Cui, Tianyu Xin, Haoran Liao, Wanqi Zhong, Xingjie Fan, Yizhao Xu, Ziliang Wang, Fei Gao, Yiming Li

分類: cs.RO, cs.LG

原文アブストラクト

Vision-Language-Action models and World-Action Models have advanced language-conditioned robotic manipulation, yet often leave metric relations among actions, objects, and scene geometry implicit. Human manipulation combines semantic understanding of task-relevant objects with spatial feedback that guides hand motion relative to objects and their surroundings. Inspired by this, we introduce a metric interaction framework that models object-level and scene-level interactions in physical Cartesian space at a shared metric scale. At the object level, Interaction-Centric Tokens (ICTs) explicitly represent end-effector pose trajectories relative to manipulated objects and are jointly denoised with actions, providing physically grounded interaction supervision. At the scene level, the Metric Action Interaction Field (MAIF) uses action and ICT queries to attend to metric scene point-cloud features and learns geometry-conditioned action corrections. Through two-stage adaptation, our framework improves diverse VLA and WAM baselines with a small number of additional parameters and training steps. Experiments demonstrate average success-rate gains of 0.80 and 3.59 percentage points on LIBERO and RoboTwin~2.0, respectively, alongside gains of 6.80 percentage points on real-world tasks and 7.45 percentage points on their out-of-distribution variants.

関連論文

PR本紙発行元 EmplifAI