日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2608.24603

グリッパー認識型視覚言語行動モデル

Gripper-aware Vision Language Action Models

シェア:XThreadsFacebookLINEはてブBluesky

異なるグリッパー種別に応じた戦略を学習できるよう、5種類のグリッパーを含む大規模データセットと、専用トークナイザーとポリシールーティングを備えたモデルを提案した。

詳しい要約

1. どんなもの?

本論文は、ロボットの把持操作におけるグリッパー種別の影響を考慮したVision Language Action Models (VLAs)を提案する。従来のVLAはグリッパー不変性を暗黙に仮定していたが、parallel-jawとsuctionなど異なるグリッパーでは同じタスクでも異なる戦略が必要である。そこで、5種類のグリッパー、複数ロボット、103,000デモを含むマルチグリッパーデータセットMiGAを構築し、新しいマルチグリッパートークナイザとアダプタベースのポリシールーティングを組み合わせたGVLAを提案する。GVLAはグリッパー条件付き表現を学習し、シミュレーションと実機で評価され、ベースラインを上回る性能を示す。

2. 先行研究と比べてどこがすごい?

先行研究のVLAはグリッパー不変性を仮定し、主にparallel-jawグリッパーのデータセットで学習されていたため、グリッパー固有の戦略を学習できなかった。本研究は、複数グリッパーを含む大規模データセットMiGAを導入し、グリッパー条件付きの表現学習を可能にした点が新しい。また、グリッパーエンコーディングとポリシールーティングの組み合わせにより、パラメータ共有と戦略分化のバランスを取る点が優れている。

3. 技術・手法の肝は?

手法の核は、新しいmulti-gripper tokenizerとadapter-based policy routingを組み合わせたGVLAである。グリッパーエンコーディングは構造化された埋め込み情報を誘導し、パラメータ共有と戦略分化のバランスを実現する。また、層ごとのプロービングにより、グリッパー条件付き表現がVLAにとって意味があることを確認している。

4. どうやって有効だと検証した?

シミュレーションと実機ロボットの両方で、複数の設定においてベースラインと比較してGVLAが優れた性能を示すことを検証した。さらに、新しい物体や未見タスクに対するzero-shot汎化やfew-shot適応、およびグリッパー適応の効率性も評価している。

5. 議論はある?

要旨からは、グリッパー条件付き表現の学習が有効であることが示されたが、データセットのバイアスや他のグリッパータイプへの拡張性、実世界での複雑なタスクへの適用限界などについては議論されていない。また、グリッパーエンコーディングの設計選択の詳細や、ポリシールーティングのメカニズムの理論的裏付けについては要旨からは不明である。

6. 次に読むべき論文は?

要旨で参照されている関連研究として、既存のVLAモデル(例:RT-2、OpenVLAなど)や、グリッパーに依存しない操作学習の研究が挙げられる。また、マルチエンボディメント学習やアダプタベースの転移学習に関する論文も関連する。具体的な論文名は要旨に明記されていないため、同分野の定番であるVLAの基盤モデルや、ロボット操作のデータセット構築に関する研究を読むことが推奨される。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Hanyi Zhang, Zihong Luo, Tianyu Li, Khang Nguyen, Basu Hela, Shreyas Kumar, Ngoc Duy Tran, Feng Dai, Charith Munasinghe, Jorge Peña Queralta, Giovanni Toffetti, Khoa Vo, Ngan Le, Ravi Prakash, Quan Vuong, Tung D. Ta, Long Hu, Anh Nguyen, Baoru Huang

分類: cs.RO

原文アブストラクト

Vision language action models (VLAs) have advanced general purpose robotic grasping and manipulation by enabling robots to interpret visual observations and natural language instructions to generate executable action sequences. However, existing VLAs often implicitly assume gripper invariance, despite grasping strategies being inherently embodiment-dependent. Different gripper types, such as parallel-jaw and suction, usually require distinct interaction strategies to achieve the same grasping objective. Moreover, current datasets for VLAs predominantly rely on parallel-jaw grippers, limiting gripper-aware learning. To address this gap, we introduce MiGA, a multi-gripper-aware dataset spanning five distinct gripper types across multiple robots with 103,000 demonstrations, explicitly capturing strategy divergence under shared task objectives. We further propose GVLA, which combines a new multi-gripper tokenizer with adapter-based policy routing. Our new gripper encoding induces structured embedding information that balances parameter sharing and strategy differentiation, while layer-wise probing confirms meaningful gripper-conditioned representations for VLAs. Intensive experiments in both simulation and real-world robots show that our GVLA outperforms the current baselines across evaluated settings. Our method also improves zero-shot generalization or few-shot adaptation to new objects or unseen tasks, and enable more efficient gripper adaptation.

関連論文