RoboMP-DINOv2: フィルタではなくプロンプトでロバストなロボットマニピュレーション
RoboMP-DINOv2: Prompts, Not Filters for Robust Robot Manipulation
マスクを視覚フィルタではなく空間プロンプトとして扱うDINOv2ベースの全シーン視覚エンコーダを提案し、外観変化や clutter 下での操作成功率を向上させた。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
分類: cs.RO
原文アブストラクト
Robot manipulation policies must generalize across visual shifts while preserving scene context relevant to action. General-purpose vision encoders are not tailored to visuomotor control, while object-centric approaches often use segmentation masks as hard filters that discard potentially useful context. We propose RoboMP-DINOv2 (Robotics Mask-Prompted DINOv2), a full-scene vision encoder that treats masks as spatial prompts rather than visibility filters. It extracts dense DINOv2 features from the full observation, injects learned region-specific embeddings at masked locations, and jointly contextualizes prompted and unprompted tokens for action prediction. We further introduce masked-region color randomization (MCR) to improve appearance robustness, yielding RoboMP-DINOv2-MCR. Across seven simulated manipulation settings, RoboMP-DINOv2 achieves 60.7% success under spatial shifts and 59.7% under scene clutter, compared with 50.7% and 41.0% for a DINOv2-based Diffusion Policy. Under unseen object colors, RoboMP-DINOv2-MCR achieves 72.5% success versus 35.1% for the strongest color-randomized baseline. Additional experiments and representation analyses show improved robustness while preserving behaviorally relevant scene information. Code is available at https://github.com/han20192019/RoboMP_DINOv2.
関連論文
- エンドタスク成功を超えて:ロボティクスにおける視覚経験検索の監査手法マニピュレーション
- 複数把持点における非無視可能な物理応答を伴う線形変形物体の安定性保証付きマニピュレーションマニピュレーション
- PAKT: 強化学習のための物理的整合性を備えたキネステティック教示マニピュレーション
- 不完全データを活用した高精度ロボットマニピュレーションマニピュレーション
- SafeLoop: 視覚言語行動マニピュレーションのためのリスク認識ロールバックマニピュレーション
- ローカルコーディングエージェントによるマニピュレーションスキルの汎化マニピュレーション