日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
VLAarXiv:2609.36518

LIBERO-MAX:世界が変化したときロボット政策は適応するか?

LIBERO-MAX: Do Robot Policies Adapt When the World Changes?

シェア:XThreadsFacebookLINEはてブBluesky

タスク実行中に環境変化が起きた際のロボット政策の頑健性を評価する8,000ペアのベンチマークLIBERO-MAXを提案し、14のVLA政策が11.0〜25.7ポイントの成功率低下を示すことを明らかにした。

詳しい要約

1. どんなもの?

- LIBERO-MAXは、ロボット政策が実行途中の環境変化に適応できるかを評価するベンチマーク。 - 8,000のペアケースを含み、geometry、observations、appearance、clutter、pathsの8種類の変化をカバー。 - 各ペアは、タスク実行中にイベントがある場合とない場合を比較し、タスク、初期状態、政策seed、イベント前の行動系列を固定。 - この制御された比較により、イベントに関連する性能低下と、変化なしでも生じる失敗を区別。

2. 先行研究と比べてどこがすごい?

- 多くのシミュレーションロバスト性ベンチマークは、リセット時に外部条件を固定するため、時間的な課題が十分に検討されていない。 - LIBERO-MAXは、実行途中の変化に焦点を当て、ペア設計と時間的診断により、イベント関連の回帰を分離。 - 14の現在のVLA、hybrid、world-action政策を評価し、イベントが成功率を11.0-25.7ポイント低下させることを示した。

3. 技術・手法の肝は?

- 8,000のペアケースを構築し、各ペアでタスク、初期状態、政策seed、イベント前の行動系列を固定。 - イベントあり/なしの実行を比較し、イベントに関連する性能低下を測定。 - イベントプロファイルにより、geometryとobservationの変化に対する共通の脆弱性を明らかに。 - カメラ制御を用いて、ロバスト性が変化条件下での能力と、その条件に遭遇する軌跡の両方を反映することを示す。 - クエリ頻度を変えてもギャップは解消されない。

4. どうやって有効だと検証した?

- 14のVLA、hybrid、world-action政策をLIBERO-MAXで評価。 - イベントにより成功率が11.0-25.7ポイント低下することを定量的に示した。 - イベントプロファイルが共通の脆弱性を明らかにし、政策ファミリー間のランキングが交錯することを示した。 - カメラ制御とクエリ頻度の変化による実験で、ロバスト性の要因を分析。

5. 議論はある?

- イベントプロファイルは、geometryとobservationの変化に対する共有された脆弱性を明らかに。 - 政策ファミリーのランキングは交錯し、単一の指標では捉えきれない。 - ロバスト性は、変化条件下での能力と、その条件に遭遇する軌跡の両方を反映。 - クエリ頻度を変えてもギャップは解消されず、根本的な課題が残る。 - 要旨からは、具体的な議論の詳細や限界については不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究は明示されていない。 - 関連手法として、VLA、hybrid、world-action policiesが挙げられる。 - 同分野の定番として、LIBEROベンチマークやロボット政策のロバスト性評価に関する研究が考えられるが、要旨からは特定できない。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yunbei Zhang, Zijian Jin, Yuanzhe Liu, Janet Wang, Xilun Zhang, Yuyou Zhang, Zhenyu Zhang, Daoan Zhang, Shuaicheng Niu, Gen Li, Jianfei Yang, Jihun Hamm, Ismini Lourentzou, Weirui Ye, Bo Liu, Peter Stone, Marco Pavone

分類: cs.RO, cs.AI

原文アブストラクト

Robots must often continue a task after a target moves, the viewpoint shifts, or an obstacle appears, even though their earlier observations and committed actions reflect the previous scene. Many simulation robustness benchmarks fix external conditions at reset, leaving this temporal challenge underexamined. We introduce LIBERO-MAX, a benchmark of 8,000 paired cases spanning eight types of changes to geometry, observations, appearance, clutter, and paths. Each pair compares task execution with and without a mid-task event, holding the task, initial state, policy seed, and pre-event action sequence fixed. This controlled comparison distinguishes event-associated regressions from failures already present without the change. Across fourteen current VLA, hybrid, and world-action policies, events reduce success by 11.0-25.7 percentage points. Event profiles reveal shared vulnerabilities to geometry and observation changes, while policy-family rankings interleave. Camera controls show that robustness reflects both competence under the changed conditions and the trajectory from which they are encountered; varying query cadence does not eliminate the gap. Together, the paired protocol and temporal diagnostics establish LIBERO-MAX as a reproducible testbed for diagnosing failures under mid-execution changes and measuring progress toward robot policies that remain effective as the world changes.

関連論文

PR本紙発行元 EmplifAI