GeoProp: ロボット状態を視覚に接地する汎用操作のための手法
GeoProp: Grounding Robot State in Vision for Generalist Manipulation
ロボットの状態(プロプリオセプション)と視覚特徴を明示的に幾何学的に接地する軽量アダプタを提案し、操作ポリシーの性能を向上させた。
著者: Guoyang Zhao, Quanhao Qian, Gongjie Zhang, Wenhao Li, Jiuniu Wang, Xiaowei Lu, Deli Zhao, Ran Xu
分類: cs.RO, cs.AI
原文アブストラクト
Proprioception is fundamental to robotic manipulation, yet standard fusion methods often treat it as an isolated vector lacking explicit alignment with visual tokens. Without a direct correspondence between 3D kinematics and 2D feature maps, manipulation policies struggle to ground the robot's state within the scene, frequently underperforming even vision-only baselines. To address this, we introduce GeoProp, a lightweight, plug-and-play adapter that aligns proprioception with vision through explicit geometric grounding and spatial feature sampling. GeoProp projects the robot state onto the image plane to sample localized visual features, constructing a grounded state token. It then injects state-derived spatial priors into the corresponding visual features via FiLM modulation. To capture motion intent, GeoProp further samples features at a short-horizon predicted coordinate derived from recent kinematics, providing look-ahead visual context. Across 67 tasks, GeoProp improves Diffusion Policy by 8.7% on 63 simulation tasks and pi_0 by 4.0% on the RoboTwin subset, and yields a 10.6% average gain across both policy families in the real world, while adding only 2-3% to the parameter count. These results demonstrate that GeoProp is a simple yet high-impact inductive bias for generalist embodied policies. Project page: https://alibaba-damo-academy.github.io/GeoProp/.
関連論文
- 否定制約付き器用把持のためのポテンシャル誘導粒子ステアリングマニピュレーション
- Facet-0: 接触を伴う精密操作のためのロボット基盤モデルマニピュレーション
- Peg-in-Bench: 高精度ロボット挿入のためのモジュール式ベンチマークマニピュレーション
- Zeva: 文脈内因果学習による汎用身体操作の実現マニピュレーション
- Motus2: 巧みな操作のための自己進化型汎用世界モデルマニピュレーション
- SUN: 言語に基づく制御から学習、実機への永続的プログラムマニピュレーション