DeicticVLA: 言語と指示ジェスチャに基づく指示モードを単一のVLAで統合する
DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA
言語指示、視覚言語指示、視覚指示の3つの指示モードをテキストプロンプトと指示マスクに正規化し、単一の事前学習済みVLAで扱えるようにした。シミュレーションと実世界タスクで、未見の表現や物体に対する汎化性能を評価し、視覚指示が言語指示より優れていることを示した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Kango Yanagida, Tatsuya Aoki, Yuichiro Yoshikawa, Takato Horii
分類: cs.RO, cs.CV
原文アブストラクト
Vision-Language-Action models (VLAs) allow users to specify manipulation tasks in natural language, but distinguishing a target or placement goal among objects of the same category or similar appearance requires detailed expressions that VLAs may not use reliably. We propose DeicticVLA, which canonicalizes Language Instruction (LI), Vision-Language Instruction (VLI), and Visual Instruction (VI) into a text prompt and deictic masks through text-prompt completion and deictic gesture grounding, enabling a single pretrained VLA to handle all three instruction modes. With a shared backbone, demonstrations, and matched training steps, we compare two RGB visual prompting methods, two separate-channel mask prompting methods, and three training strategies in simulation. Under two-stage training, the four prompting methods achieve high in-distribution success but differ in their ability to use deictic masks in unseen layouts. Across methods, training-strategy ablations show that two-stage training improves such use, while retaining second-stage LI data mitigates forgetting without reducing VLI and VI performance. In three real-world tasks, one policy supports all modes. VLI and VI outperform LI under unseen expressions, appearance changes, and novel objects. For unseen categories, both achieve 100% success, compared with 16.7% for jointly trained LI. These results demonstrate the unified three-mode interface and guide DeicticVLA design.