日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.25562

IndustrialVLA-Bench:オープンロボットポリシーモデルの追跡可能な多軸評価

IndustrialVLA-Bench: A Traceable Multi-Axis Evaluation of Open Robot Policy Models

シェア:XThreadsFacebookLINEはてブBluesky

6つのVLA/WAMモデルを統一プロトコルで評価し、クリーン性能・ロバスト性・言語感度・実行コストを比較したベンチマークを提案。

詳しい要約

1. どんなもの?

- オープンなロボットポリシーモデルを統一的に評価するベンチマーク - 対象は VLA (vision-language-action models) と WAM (world-action models) の6システム - 評価軸は clean capability (LIBERO)、non-language robustness (LIBERO-Plus)、instruction sensitivity (LIBERO-Para)、observed execution cost - 各タスクスコアは固定 checkpoint・推論設定で異なる random seed の3回完全評価を集約 - 各システムに inference latency、peak memory、runtime mode、evidence status を付与 - protocol-faithful、near-reproduction、pending-verification を区別し、厳密比較は protocol-faithful のみ

2. 先行研究と比べてどこがすごい?

- 従来は VLA と WAM が異なる評価プロトコルで報告され、能力・頑健性・言語感度・配備コストのトレードオフが不明瞭だった - 本研究は同一の reporting schema で両パラダイムを比較 - clean LIBERO の平均差は6システムで1.58点のみだが、robustness と paraphrase の要約は14.62点・31.08点に広がる - protocol-faithful な3システムに限定しても効果は保持 (1.36、14.62、23.10点) - 弱い evidence tier に依存せず診断的分離が得られる点が新しい

3. 技術・手法の肝は?

- 統一 reporting schema による evidence-aware 評価 - 4つの評価軸を分離: clean capability (LIBERO)、non-language robustness (LIBERO-Plus)、instruction sensitivity (LIBERO-Para)、observed execution cost - 固定 checkpoint と推論設定の下、異なる random seed で3回の完全評価を集約 - 各システムに inference latency、peak memory、runtime mode、evidence status を記録 - protocol-faithful、near-reproduction、pending-verification を可視的に分離し、厳密比較の可否を明示

4. どうやって有効だと検証した?

- 6つの公開 VLA/WAM システムを対象に評価を実施 - clean LIBERO 平均差1.58点、robustness 14.62点、paraphrase 31.08点という結果 - protocol-faithful な3システムに限定しても同様の傾向 (1.36、14.62、23.10点) - これにより診断的分離が弱い evidence tier に依存しないことを確認 - コードと評価記録を GitHub で公開

5. 議論はある?

- どちらかのパラダイムの普遍的な優位性を主張するものではない - 共有された実用的基準でリリース済みロボットポリシーを比較するための追跡可能な証拠を提供 - protocol-faithful なエントリのみが厳密比較を支持する - near-reproduction や pending-verification は可視的に分離される - 具体的な議論の詳細は要旨からは不明

6. 次に読むべき論文は?

- VLA (vision-language-action models) に関する研究 - WAM (world-action models) に関する研究 - LIBERO、LIBERO-Plus、LIBERO-Para の各ベンチマーク - 要旨で参照/比較されている個別の6システムの論文 - 同分野の定番として Open X-Embodiment や RT-1/RT-2 などのロボットポリシー研究

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yiqi Wang, Zhifeng Rao, Jiaqi Zhang, Xiaoyang Li, Zhangkai Wu, Yiqun Duan, Mingkai Zheng, Fei Wang, Shan You, Taotao Cai

分類: cs.RO, cs.AI

原文アブストラクト

Open robot policies increasingly follow two paradigms: vision-language-action models (VLAs) directly map observations and instructions to actions, whereas world-action models (WAMs) incorporate learned video or world dynamics into policy learning or action generation. Although both target the same manipulation tasks and represent alternative design choices, they are commonly reported under different evaluation protocols, leaving their capability, robustness, language sensitivity, and deployment-cost trade-offs unclear. We present IndustrialVLA-Bench, an evidence-aware evaluation of six released VLA and WAM systems under a unified reporting schema. It separately evaluates clean capability on LIBERO, non-language robustness on LIBERO-Plus, instruction sensitivity on LIBERO-Para, and observed execution cost. Reported task scores aggregate three complete evaluations with distinct random seeds under a fixed checkpoint and inference configuration. Across all six systems, clean LIBERO averages differ by only 1.58 points, whereas robustness and paraphrase summaries span 14.62 and 31.08 points. Restricting every comparison to the three protocol-faithful systems preserves the effect (1.36, 14.62 and 23.10 points), so the diagnostic separation reported here does not depend on the weaker evidence tiers. We additionally report observed inference latency, peak memory, runtime mode, and an evidence status for every system. Protocol-faithful, near-reproduction, and pending-verification entries remain visibly separated; only protocol-faithful entries support strict comparisons. Rather than claiming universal superiority of either paradigm, IndustrialVLA-Bench provides traceable evidence for comparing released robot policies on shared practical criteria. Code and evaluation records are available at https://github.com/xiaoqi-7/IndustrialVLA-Bench.

関連論文

PR本紙発行元 EmplifAI