日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.24350

LIBERO-VPro:ロボット基盤モデルの閉ループ視覚ロバスト性ベンチマーク

LIBERO-VPro: Benchmarking Closed-Loop Visual Robustness of Robotic Foundation Models

シェア:XThreadsFacebookLINEはてブBluesky

実行中の視覚情報を意図的に乱すことで、ロボット基盤モデルの閉ループ視覚ロバスト性を体系的に評価するベンチマークLIBERO-VProを提案し、VLAとWAMのロバスト性プロファイルの違いを明らかにした。

詳しい要約

1. どんなもの?

- ロボティック基盤モデルの閉ループ視覚ロバスト性を評価するベンチマーク LIBERO-VPro を提案。 - 実行中の視覚証拠を摂動させ、4次元(Visual Evidence Degradation, Camera Staleness, Visual Source Consistency, Task-Relevant Scene Variation)を網羅。 - 12 challenge categories、96 experimental settings、3,296 task-condition cases を含む。 - 3つの vision-language-action models と3つの world-action models を約196,000 simulated episodes で評価。 - さらに Franka Research 3 での200 real-world rollouts を実施。

2. 先行研究と比べてどこがすごい?

- 従来の操作ベンチマークは、実行中に clean, timely, consistent な視覚観測を仮定。 - LIBERO-VPro は閉ループ実行中の視覚証拠を系統的に摂動し、視覚ロバスト性を多次元的に診断。 - 名目性能が高くても視覚的 grounding や適応に大きな弱点が隠れることを明らかに。 - VLAs と WAMs のロバスト性プロファイルが異なることを示し、アーキテクチャ依存性を提示。

3. 技術・手法の肝は?

- 視覚証拠の摂動を4次元で設計:Visual Evidence Degradation, Camera Staleness, Visual Source Consistency, Task-Relevant Scene Variation。 - 12 challenge categories、96 experimental settings、3,296 task-condition cases を構築。 - 3つの VLA と3つの WAM を約196,000 simulated episodes で評価。 - Franka Research 3 を用いた200 real-world rollouts で補完。 - 閉ループ実行中の視覚入力の変化に対する成功率や適応を測定。

4. どうやって有効だと検証した?

- 約196,000 simulated episodes と200 real-world rollouts で評価。 - 3つの VLA と3つの WAM を対象に、4次元・12カテゴリ・96設定・3,296ケースで検証。 - 結果として、強い名目性能が視覚的 grounding と適応の弱点を隠すことを確認。 - 物体レベルの重度オクルージョンでは成功する一方、局所的な相互作用手がかりの破壊や空間事前の違反で急激に性能低下。 - 古い/欠落した観測に敏感で、タスク前提の変化への行動適応が困難。

5. 議論はある?

- 視覚ロバスト性は多次元であり、アーキテクチャ依存であると議論。 - VLAs と WAMs は異なるロバスト性プロファイルを示す。 - 名目性能だけでは視覚的 grounding や適応の弱点を評価できない。 - 課題視覚条件下で信頼性高く行動を grounding・適応させる基盤モデル開発のための診断フレームワークを提供。 - 具体的な限界や今後の課題は要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として vision-language-action models (VLAs) と world-action models (WAMs) が挙げられる。 - 同分野の定番として LIBERO などの操作ベンチマークや、ロボティック基盤モデルの評価研究を読むべき。 - 具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Huiqiong Li, Zhiting Mei, Anirudha Majumdar, Jingjing Chen, Yu-Gang Jiang, Bin Zhu

分類: cs.RO, cs.CV

原文アブストラクト

Robotic foundation models achieve impressive performance on standard manipulation benchmarks, yet these evaluations typically assume clean, timely, and consistent visual observations throughout execution. We introduce LIBERO-VPro, a benchmark for systematically evaluating the closed-loop visual robustness of robotic foundation models by perturbing the visual evidence available during execution. LIBERO-VPro covers four complementary dimensions, including Visual Evidence Degradation, Camera Staleness, Visual Source Consistency, and Task-Relevant Scene Variation, spanning 12 challenge categories, 96 experimental settings, and 3,296 task-condition cases. We evaluate three vision-language-action models and three world-action models over approximately 196,000 simulated episodes, complemented by 200 real-world rollouts on a Franka Research 3. Our results reveal that strong nominal performance can mask substantial weaknesses in visual grounding and adaptation. Models often remain successful despite severe object-level occlusion, yet degrade sharply when local interaction cues are disrupted or familiar spatial priors are violated. They are also highly sensitive to stale or missing observations and struggle when changed task preconditions require behavioral adaptation. Finally, VLAs and WAMs exhibit distinct robustness profiles, showing that visual robustness is multi-dimensional and architecture-dependent. LIBERO-VPro provides a systematic diagnostic framework for developing robotic foundation models that can more reliably ground and adapt their actions under challenging visual conditions.

関連論文

PR本紙発行元 EmplifAI