相互作用中心モデリングによる2本指グリッパ操作の統一的クロスドメイン表現の実現
Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling
2本指グリッパの共通構造をパラメータ化した抽象化で捉え、VLMとSAMで相互作用 triplet を推定し、Flow-Matching Transformerで7自由度動作を生成することで、異なるロボット間や視点間でのゼロショットsim-to-real転移を可能にした模倣学習手法。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Guanlin Li, Shifeng Bao, Yihan Zhao, Haitao Shen, Haoyang Li, Chen Zhao, Tong Yang, Jie Tang, Jing Zhang
分類: cs.RO, cs.CV
原文アブストラクト
Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverages the shared structure of two-finger grippers via a parameterized universal gripper abstraction, yielding a canonical gripper-frame representation. Given language and RGB-D observations, a VLM infers the subtask and grounds an interaction triplet (gripper, held, target), while SAM~2.1 tracks masks to reduce VLM queries. We design concise hybrid features that combine target/collision artificial potential fields for global guidance with segmented gripper-frame point clouds for local geometry, and use a Flow-Matching Transformer to predict smooth 7-DoF action chunks. Experiments in simulation and real-world tasks demonstrate that ours is the first imitation learning approach to simultaneously achieve competitive benchmark scores and extreme cross-embodiment/cross-viewpoint zero-shot sim-to-real transfer to completely distinct, heterogeneous robot platforms.
関連論文
- シミュレータ非依存の布操作のための簡易グリッパインタフェースマニピュレーション
- 全手把持のためのリアルタイム力制御フレームワークマニピュレーション
- WRAP: 治具不要の力覚考慮型マルチロボット組立計画マニピュレーション
- 受動的実行から能動的探索へ:実環境におけるエージェント型身体性マニピュレーションマニピュレーション
- CALM: 電流整合型リンクマニピュレーションによる単腕での大型物体持ち上げマニピュレーション
- Streaming-WAM: 非同期ロボットマニピュレーションのための行動条件付きワールドアクションモデルマニピュレーション