日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2610.01351

成功だけでは不十分?卓上マニピュレーションにおけるVLAの挙動に対する入力摂動の影響調査

Is Success All You Need? Investigating the Impact of Input Perturbations on VLA Behaviour in Tabletop Manipulation Tasks

シェア:XThreadsFacebookLINEはてブBluesky

VLAモデルのロバスト性をタスク成功率だけでなく、成功軌道の動きの滑らかさや効率、グリッパ挙動といった振る舞いの観点から評価するフレームワークを提案し、LIBERO系ベンチマークで摂動が成功時の挙動を変化させることを示した。

詳しい要約

1. どんなもの?

- Vision-Language-Action (VLA) モデルのロボットマニピュレーションにおけるロバスト性評価に関する研究。 - 従来は Task Success Rate (TSR) のみでロバスト性を測っていたが、成功軌道の振る舞いの変化を評価する benchmark-agnostic な枠組みを提案。 - LIBERO と LIBERO-Plus を拡張し、3つの state-of-the-art VLA モデル、4つの LIBERO タスクスイート、7つの摂動条件で評価。 - 運動の滑らかさ、効率、グリッパー動作などの指標を用いて、成功軌道の典型性と変動性を分析。

2. 先行研究と比べてどこがすごい?

- 先行研究は VLA のロバスト性を主に TSR で評価していた。 - 本研究は TSR だけでは捉えられない成功軌道の振る舞いの変化に着目。 - 同一摂動条件下で TSR が同等でも、成功軌道の振る舞いが大きく異なるケースを特定。 - タスク完了の有無だけでなく、完了中のロボットの振る舞いも評価すべきと主張。

3. 技術・手法の肝は?

- benchmark-agnostic な評価フレームワークを提案。 - LIBERO と LIBERO-Plus を拡張し、摂動下での成功軌道の実行を特徴づける。 - 運動の滑らかさ、効率、グリッパー動作などの指標を導入。 - 成功軌道の典型性と変動性を評価。

4. どうやって有効だと検証した?

- 3つの state-of-the-art VLA モデル、4つの LIBERO タスクスイート、7つの摂動条件で評価。 - 成功軌道の振る舞いの変化と変動性を測定。 - TSR だけでは推測できない現象を発見。 - 同一摂動条件下で TSR が同等でも振る舞いが大きく異なるケースを特定。

5. 議論はある?

- 摂動が成功軌道の振る舞いを変えうることを発見。 - TSR だけではロバスト性を十分に評価できないと議論。 - タスク完了の有無だけでなく、完了中の振る舞いも捉える指標が必要と主張。 - TSR を補完する行動評価指標の重要性を強調。

6. 次に読むべき論文は?

- LIBERO - LIBERO-Plus - その他の VLA モデル評価に関する研究(要旨からは具体的な論文名は不明)

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Sophie Higham, Riccardo Andrea Izzo, Matteo Matteucci, Alessandro Suglia

分類: cs.RO

原文アブストラクト

Vision-Language-Action (VLA) models have achieved high task success rates on robot manipulation task benchmarks. More recently, there has been an emphasis on evaluating the robustness of VLA models to perturbations. However, this robustness is still predominantly measured through Task Success Rate (TSR). In this work, we propose a benchmark-agnostic evaluation framework to measure the behavioural robustness of models by characterising how successful trajectories are executed under perturbation. We implement this methodology by extending the widely-used LIBERO and LIBERO-Plus benchmarks. Across three state-of-the-art VLA models, four LIBERO task suites and seven perturbation conditions, we evaluate changes in both typical successful behaviour and its variability, including metrics of motion smoothness, efficiency and gripper behaviour. We find that perturbations can alter the behaviour of successful trajectories, a phenomenon which cannot necessarily be inferred from TSR alone. Across LIBERO suites, we identify cases where state-of-the-art VLA models achieve comparable TSR under the same perturbation condition, yet behaviour on successful trajectories diverges substantially. Therefore, to have a more robust assessment of task performance, we argue that suitable measures of robustness should capture not only whether a task is completed, but also how the robot behaves while completing it. When evaluating the robustness of VLA models, TSR may be complemented by behavioural evaluation metrics that characterise the nature and variability of successful task execution by robots.

関連論文

PR本紙発行元 EmplifAI