日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ベンチマークarXiv:2609.28952

RoboRecover: 実行逸脱下におけるロボット方策の回復力ベンチマーク

RoboRecover: Benchmarking Robot Policy Recovery under Execution Deviations

シェア:XThreadsFacebookLINEはてブBluesky

実行中の逸脱状態から元のタスクを回復する能力を評価するベンチマークを提案し、初期状態の性能が回復性能を決めないことを示した。

詳しい要約

1. どんなもの?

- ロボット政策の回復性能を評価するベンチマーク「RoboRecover」を提案。 - 実行中の逸脱状態から元のタスクを継続できるかを評価。 - RoboTwinとLIBEROで各1,000シナリオ、計2,000シナリオ。 - 各プラットフォームで800/200のtrain/test分割。 - 初期状態性能と回復性能は別次元であることを示す。

2. 先行研究と比べてどこがすごい?

- 従来のベンチマークは多様なタスクやOOD条件を扱うが、定義済み初期状態からの完全軌道評価が中心。 - 初期シーンと最終結果に注目し、動的相互作用過程を軽視。 - RoboRecoverは実行中に生じるoff-nominal中間状態からの回復を評価。 - 回復をロボット政策評価の独立した次元として確立。

3. 技術・手法の肝は?

- 軌道から逸脱状態を選択し、action prefixesをreplayして状態を再構築。 - 再構築した状態から元のタスクで政策を評価。 - 訓練分割を用いて回復介入の研究も可能。 - 評価は初期状態ではなく、実行誘発中間状態に焦点。

4. どうやって有効だと検証した?

- RoboTwinとLIBERO上で2,000シナリオを構築。 - 各プラットフォーム1,000シナリオ、800/200のtrain/test分割。 - 初期状態性能が回復性能を決定しないことを実験で示す。 - 政策ごとにシナリオ間で回復強度が異なることを確認。

5. 議論はある?

- 初期状態性能と回復性能の乖離を指摘。 - 政策の回復強度がシナリオ依存であることを議論。 - 訓練分割により回復介入研究の基盤を提供。 - 実行誘発中間状態からの回復を評価軸として提案。

6. 次に読むべき論文は?

- RoboTwin - LIBERO - 要旨で参照/比較されている個別の先行研究は明記されていないため、同分野の定番ベンチマークとして上記を挙げる。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Yang Li, Chen Zhao, Zhuoran Wang, Jiankang Wang, Chao Shao, Yihan Lin, Haitao Shen, Jing Zhang

分類: cs.RO

原文アブストラクト

Robot-policy benchmarks increasingly cover diverse tasks and preset out-of-distribution conditions, but typically evaluate complete trajectories from predefined initial states. These evaluations often focus on the initialized scene and the final outcome, while paying less attention to the dynamic interaction process. During closed-loop execution, actions and contacts can alter object relations and task progress, producing off-nominal intermediate states that need recovery. Recovery requires a policy to infer how task progress has changed, correct the relevant relations, and continue the original goal. We introduce RoboRecover, a benchmark for robot policy recovery under execution deviations. RoboRecover selects deviation states from trajectories, reconstructs them by replaying action prefixes, and evaluates policies on the original task. RoboRecover contains 2,000 scenarios across RoboTwin and LIBERO, with 1,000 scenarios and a fixed 800/200 train/test split on each platform. Results show that initial-state performance does not determine recovery performance and policies exhibit different recovery strengths across scenarios. Using its training split, RoboRecover further supports study on recovery interventions. RoboRecover establishes recovery from execution-induced intermediate states as a distinct dimension of robot policy evaluation.

関連論文

PR本紙発行元 EmplifAI