日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.11508

WARP-VLA: 手首カメラ適応による視点に頑健なVLAポリシー実行

WARP-VLA: Wrist-Camera Adaptation for View-Robust Policy Execution in Vision-Language-Action Models

シェア:XThreadsFacebookLINEはてブBluesky

手首カメラの取り付け位置ずれに頑健なVLAを提案。視点ごとの特徴変換を専門家が学習するMoE構造で、カメラ外部パラメータなしに多様な配置へ適応する。

詳しい要約

1. どんなもの?

- Vision-Language-Action models (VLAs) のロボットマニピュレーションにおいて、カメラ設定の変化に頑健な WARP-VLA を提案。 - 特に wrist camera の多様な配置に対応するため、Mixture-of-Experts (MoE) アーキテクチャを採用。 - 各 expert が view-specific な特徴変換を学習し、router が暗黙の view 情報に基づいて統合。 - カメラ外部パラメータを追加入力として必要とせずにポリシーを展開可能。 - LIBERO benchmark と実機実験で有効性を検証。

2. 先行研究と比べてどこがすごい?

- 従来の VLAs はカメラ設定の変化に敏感で、cross-setup 展開が困難。 - 訓練時のカメラポーズを正確に再現することはほぼ不可能。 - 固定外部視点と異なり、wrist view はカメラがロボットと共に動くため、小さな取り付け変動でも細かい幾何学的手がかりが変化する。 - WARP-VLA はこの問題に対処し、wrist-view 摂動下で pi-0.5 の平均成功率を 39.2% から 78.3% に改善。 - 実機実験でシミュレーションで学習した特徴レベルの適応が多様な展開設定に転移することを示した。

3. 技術・手法の肝は?

- Mixture-of-Experts (MoE) アーキテクチャを採用。 - 個々の expert が view-specific な特徴変換を学習。 - router が暗黙の view 情報に基づいて expert を組み合わせる。 - カメラ外部パラメータを追加入力として要求しない。 - これにより、多様な wrist camera 設定にロバストなポリシー実行を実現。

4. どうやって有効だと検証した?

- LIBERO benchmark での実験を実施。 - wrist-view 摂動下で pi-0.5 の平均成功率を 39.2% から 78.3% に向上。 - 実機実験を実施し、シミュレーションで学習した特徴レベルの適応が多様な展開設定に転移することを確認。 - 再現性と将来研究のために、wrist viewpoint robustness benchmark と plug-and-play 実装を公開。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- pi-0.5 (ベースラインとして比較) - LIBERO benchmark (評価に使用) - Mixture-of-Experts (MoE) アーキテクチャ (関連手法) - Vision-Language-Action models (VLAs) (関連分野)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Junmyeong Lee, Dongmin Shin, Min-Gyu Park, Wooseok Jeon, Inho Chang, Hae-Gon Jeon

分類: cs.RO, cs.CV

原文アブストラクト

Despite recent advances in Vision-Language-Action models (VLAs) for robotic manipulation, their performance remains sensitive to changes in camera configuration. The problem becomes more evident in cross-setup deployment, as reproducing the exact camera pose used for training is nearly impossible. Unlike fixed external views, wrist views are more challenging because the camera moves with the robot, causing even small mounting variations to alter fine-grained geometric cues. To address this, we propose WARP-VLA, a camera-view robust VLA for diverse wrist camera configurations. WARP-VLA adopts a Mixture-of-Experts (MoE) architecture where individual experts learn view-specific feature transformations, and a router combines them based on implicit view information. This allows the policy to be deployed without requiring camera extrinsic parameters as additional input. Through experiments on the LIBERO benchmark, WARP-VLA improves the average success rate of pi-0.5 from 39.2% to 78.3% under wrist-view perturbations. The real-robot experiments further show that the feature-level adaptation learned in simulation successfully transfers to diverse deployment settings. To facilitate reproducibility and future research, we release our wrist viewpoint robustness benchmark and a plug-and-play implementation.

関連論文

PR本紙発行元 EmplifAI