GIFT: 行動指向の構造的監督による誘導中間特徴学習を用いたロボット操作
GIFT: Guided Intermediate Feature Training via Action-Oriented Structural Supervision for Robotic Manipulation
視覚と言語の事前学習で得られる特徴と制御に必要な情報の乖離(行動十分性ギャップ)を埋めるため、中間特徴に幾何・アフォーダンス・目標領域の3つの制御関連構造を学習させるフレームワークGIFTを提案し、複数のポリシーで性能向上を確認した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Yupeng Zheng, Xiang Li, Songen Gu, Yuhang Zheng, Shuai Tian, Weize Li, Linbo Wang, Chaoyue Li, Qichao Zhang, Haoran Li, Zhongpu Xia, Ya-Qin Zhang, Shuicheng Yan, Dongbin Zhao
分類: cs.RO
原文アブストラクト
Vision-language pre-training and predictive world modeling provide robot policies with rich semantic and dynamic visual features, but their native action and visual-prediction objectives may omit critical physical and task structure while retaining control-irrelevant visual redundancy. We call this mismatch between visual richness and control utility the action-sufficiency gap. We investigate whether this gap can be bridged by guiding intermediate features to preserve three control-relevant structure in robotic manipulation: geometry governing motion feasibility, affordance encoding instruction-relevant entities, and goals grounding instructions in task-relevant regions. To this end, we present GIFT (Guided Intermediate Feature Training), an architecture-flexible framework for learning intermediate features that translates these structures into training-time constraints through geometry alignment, affordance prediction, and goal-region reconstruction. We instantiate GIFT in a Vision-Language-Action (VLA) policy, a direct-action World-Action Model (WAM), and an inverse-dynamics WAM while retaining each model's action formulation. Under zero-shot transfer to LIBERO-Plus, GIFT-VLA, GIFT-WAM-Fast, and GIFT-WAM-IDM outperform StarVLA-OFT, Fast-WAM, and Fast-WAM-IDM by 4.6, 12.6, and 5.2 points, reaching 79.6%, 72.6%, and 87.8%, respectively. On RoboCasa, the three GIFT variants reach 61.4%, 83.6%, and 82.3%, outperforming their counterparts by 12.6, 9.0, and 8.4 points, respectively. Together, these results establish learning functionally structured intermediate features as a reusable principle across model-specific action formulations, with especially large gains on articulated-object tasks and high-precision real-world manipulation under unseen visual and spatial perturbations. Project page: https://openphoenix-team.github.io/GIFT-pages.
関連論文
- 適応的視覚言語把持:構成可能な基盤事前知識と汎化可能な把持合成による実現マニピュレーション
- MS-MEM: 不確実性・外乱を考慮した行動選択によるマルチスキル操作強化マッピングマニピュレーション
- HINT: 長期的ロボット操作のための人間意図の注入マニピュレーション
- リアルタイム動力学に基づくトルクサンプリングMPPIによるコンプライアントで力認識のマニピュレーションマニピュレーション
- Facet-0: 接触を伴う精密操作のためのロボット基盤モデルマニピュレーション
- 否定制約付き器用把持のためのポテンシャル誘導粒子ステアリングマニピュレーション