日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.25506

RoboMP-DINOv2: フィルタではなくプロンプトでロバストなロボットマニピュレーション

RoboMP-DINOv2: Prompts, Not Filters for Robust Robot Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

マスクを視覚フィルタではなく空間プロンプトとして扱うDINOv2ベースの全シーン視覚エンコーダを提案し、外観変化や clutter 下での操作成功率を向上させた。

詳しい要約

1. どんなもの?

- ロボットマニピュレーション向けの視覚エンコーダ - 名称は RoboMP-DINOv2 (Robotics Mask-Prompted DINOv2) - マスクを可視性フィルタではなく空間プロンプトとして扱う - 全シーン観測から密な DINOv2 特徴を抽出 - マスク位置に学習済み領域埋め込みを注入 - プロンプト付き/無しトークンを jointly 文脈化し行動予測 - 外観頑健性のため masked-region color randomization (MCR) を導入 - 派生版 RoboMP-DINOv2-MCR も提案

2. 先行研究と比べてどこがすごい?

- 汎用視覚エンコーダは visuomotor control に特化していない - object-centric 手法は segmentation masks を hard filters として使い文脈を捨てる - 提案はマスクを spatial prompts として扱い文脈を保持 - 7つの simulated manipulation 設定で検証 - spatial shifts で 60.7% 成功、DINOv2-based Diffusion Policy は 50.7% - scene clutter で 59.7%、比較は 41.0% - 未見物体色で RoboMP-DINOv2-MCR は 72.5%、最強 color-randomized baseline は 35.1%

3. 技術・手法の肝は?

- 全観測から密な DINOv2 特徴を抽出 - マスク位置に学習済み領域特異埋め込みを注入 - プロンプト付き/無しトークンを jointly 文脈化 - その表現から行動予測 - masked-region color randomization (MCR) で外観頑健性を向上 - MCR 適用版が RoboMP-DINOv2-MCR

4. どうやって有効だと検証した?

- 7つの simulated manipulation 設定で評価 - spatial shifts で 60.7% 成功 - scene clutter で 59.7% 成功 - DINOv2-based Diffusion Policy は 50.7% / 41.0% - 未見物体色で RoboMP-DINOv2-MCR は 72.5% - 最強 color-randomized baseline は 35.1% - 追加実験と representation analyses で頑健性と行動関連文脈保持を確認

5. 議論はある?

- マスクを hard filter でなく spatial prompt として使う設計の妥当性 - 全シーン文脈保持が行動予測に寄与する点 - MCR による外観頑健性向上 - representation analyses で behaviorally relevant scene information の保持を示す - 限界や失敗事例の詳細は要旨からは不明

6. 次に読むべき論文は?

- DINOv2-based Diffusion Policy - Diffusion Policy - DINOv2 - object-centric な segmentation masks を用いる手法 - masked-region color randomization (MCR) 関連の color randomization baseline

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Han Qi, Heng Yang

分類: cs.RO

原文アブストラクト

Robot manipulation policies must generalize across visual shifts while preserving scene context relevant to action. General-purpose vision encoders are not tailored to visuomotor control, while object-centric approaches often use segmentation masks as hard filters that discard potentially useful context. We propose RoboMP-DINOv2 (Robotics Mask-Prompted DINOv2), a full-scene vision encoder that treats masks as spatial prompts rather than visibility filters. It extracts dense DINOv2 features from the full observation, injects learned region-specific embeddings at masked locations, and jointly contextualizes prompted and unprompted tokens for action prediction. We further introduce masked-region color randomization (MCR) to improve appearance robustness, yielding RoboMP-DINOv2-MCR. Across seven simulated manipulation settings, RoboMP-DINOv2 achieves 60.7% success under spatial shifts and 59.7% under scene clutter, compared with 50.7% and 41.0% for a DINOv2-based Diffusion Policy. Under unseen object colors, RoboMP-DINOv2-MCR achieves 72.5% success versus 35.1% for the strongest color-randomized baseline. Additional experiments and representation analyses show improved robustness while preserving behaviorally relevant scene information. Code is available at https://github.com/han20192019/RoboMP_DINOv2.

関連論文

PR本紙発行元 EmplifAI