日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.31207

相互作用中心モデリングによる2本指グリッパ操作の統一的クロスドメイン表現の実現

Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling

シェア:XThreadsFacebookLINEはてブBluesky

2本指グリッパの共通構造をパラメータ化した抽象化で捉え、VLMとSAMで相互作用 triplet を推定し、Flow-Matching Transformerで7自由度動作を生成することで、異なるロボット間や視点間でのゼロショットsim-to-real転移を可能にした模倣学習手法。

詳しい要約

1. どんなもの?

- 二本指グリッパの操作のための統一的なクロスドメイン表現を可能にするフレームワーク。 - 言語とRGB-D観測からVLMがサブタスクを推論し、interaction triplet (gripper, held, target)をグラウンディング。 - SAM 2.1でマスクを追跡し、VLMクエリを削減。 - ターゲット/衝突人工ポテンシャル場とセグメント化されたグリッパフレーム点群を組み合わせたハイブリッド特徴を設計。 - Flow-Matching Transformerで滑らかな7-DoFアクションチャンクを予測。

2. 先行研究と比べてどこがすごい?

- 従来の模倣学習はタスク意味論とハードウェア固有の視覚幾何を不可分に絡めていた表現の欠陥があった。 - 本手法はinteraction-centricフレームワークにより、二本指グリッパの共有構造をパラメータ化されたuniversal gripper abstractionで活用し、canonical gripper-frame表現を生成。 - これにより、異なるロボットプラットフォームへの極端なクロスエンボディメント/クロスビューポイントゼロショットsim-to-real転移を同時に達成した初の模倣学習アプローチ。

3. 技術・手法の肝は?

- パラメータ化されたuniversal gripper abstractionを用いて二本指グリッパの共有構造を活用。 - VLMが言語とRGB-Dからサブタスクを推論し、interaction tripletをグラウンディング。 - SAM 2.1がマスクを追跡し、VLMクエリを削減。 - ターゲット/衝突人工ポテンシャル場によるグローバルガイダンスと、セグメント化されたグリッパフレーム点群によるローカル幾何を組み合わせたハイブリッド特徴。 - Flow-Matching Transformerで滑らかな7-DoFアクションチャンクを予測。

4. どうやって有効だと検証した?

- シミュレーションと実世界タスクでの実験を実施。 - 競争力のあるベンチマークスコアと、完全に異なる異種ロボットプラットフォームへの極端なクロスエンボディメント/クロスビューポイントゼロショットsim-to-real転移を同時に達成することを示した。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。関連手法として、VLM、SAM 2.1、Flow-Matching Transformer、模倣学習、sim-to-real転移、クロスエンボディメント一般が挙げられる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Guanlin Li, Shifeng Bao, Yihan Zhao, Haitao Shen, Haoyang Li, Chen Zhao, Tong Yang, Jie Tang, Jing Zhang

分類: cs.RO, cs.CV

原文アブストラクト

Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverages the shared structure of two-finger grippers via a parameterized universal gripper abstraction, yielding a canonical gripper-frame representation. Given language and RGB-D observations, a VLM infers the subtask and grounds an interaction triplet (gripper, held, target), while SAM~2.1 tracks masks to reduce VLM queries. We design concise hybrid features that combine target/collision artificial potential fields for global guidance with segmented gripper-frame point clouds for local geometry, and use a Flow-Matching Transformer to predict smooth 7-DoF action chunks. Experiments in simulation and real-world tasks demonstrate that ours is the first imitation learning approach to simultaneously achieve competitive benchmark scores and extreme cross-embodiment/cross-viewpoint zero-shot sim-to-real transfer to completely distinct, heterogeneous robot platforms.

関連論文

PR本紙発行元 EmplifAI