日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.01178

Recova: 自律ロボットマニピュレーションのためのエージェント誘導型失敗回復

Recova: Agent-Guided Failure Recovery for Autonomous Robotic Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

デジタルツイン上でエージェントが失敗を診断し回復行動を学習、実機での検証と人間デモを組み合わせて回復能力を拡張する枠組みを提案し、成功率と人間介入の削減を実証した。

詳しい要約

1. どんなもの?

Recovaは、自律ロボットマニピュレーションにおける失敗からの回復をエージェントが導くフレームワーク。再構成されたdigital twin内でタスク実行と回復を同時に開発し、実世界経験で検証・洗練する。twin内でエージェントが失敗を診断し、修正プログラムを試し、成功したタスクと回復のrolloutを別々のpolicy用に収集。展開時には進捗を監視し、学習済みまたはプログラムによる回復を呼び出し、シーン復元を検証して実行を再開する。適切な回復がない場合、人間のdemonstrationが失敗を解決し学習ループに入り、回復能力を拡張する。物理rolloutと人間demonstrationは対応するpolicyにDAgger訓練用に振り分けられる。

2. 先行研究と比べてどこがすごい?

従来の失敗回復はスケーラブルな失敗探索と物理的groundingが課題だった。Recovaはagent-guided frameworkにより、digital twinでの失敗診断・修正プログラムテスト・rollout収集と、実世界での検証・洗練を統合。6つのLIBERO-Pro設定と4つのMolmoSpacesカテゴリで平均成功率78.8%と64.9%を達成し、最強baselineの71.7%と38.0%を上回る。実ロボット4ワークステーション並列収集でDAgger fine-tuningにより平均成功率23.8%から77.5%へ、回復スキルで87.5%へ向上。1タスク4ラウンドで人間介入が87.5%から0%に減少。

3. 技術・手法の肝は?

肝は、digital twin内でエージェントが失敗を診断し修正プログラムをテスト、成功したタスクと回復のrolloutを別々のpolicy用に収集する点。展開時は進捗監視、学習済み/プログラム回復の呼び出し、シーン復元検証、実行再開。回復不能時は人間demonstrationが学習ループに入る。物理rolloutと人間demonstrationを対応policyへDAgger訓練用に振り分け。これにより失敗を再利用可能な能力に変換。

4. どうやって有効だと検証した?

6つのLIBERO-Pro設定と4つのMolmoSpacesカテゴリで評価。Recovaは平均成功率78.8%と64.9%、最強baselineは71.7%と38.0%。実ロボット4ワークステーション並列収集でDAgger fine-tuningが平均成功率23.8%から77.5%へ、回復スキル追加で87.5%へ。1タスク4収集ラウンドで人間介入が87.5%から0%に減少。

5. 議論はある?

結果はagent-guided recoveryが失敗を再利用可能な能力に変え、堅牢性を改善しつつ人間介入を漸減させることを示す。ただし要旨からは、digital twinの再構成精度やsim-to-realギャップ、回復スキルの汎化性、人間demonstration依存の限界など詳細な議論は不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究としてLIBERO-Pro、MolmoSpaces、DAggerが挙げられる。関連手法としてdigital twin、sim-to-real transfer、failure recovery、imitation learning、reinforcement learningなどが次に読むべき候補。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Isabella Liu, An-Chieh Cheng, Johan Bjorck, Zhiding Yu, Hongxu Yin, Jan Kautz, Linxi Fan, Yuke Zhu, Sifei Liu

分類: cs.RO

原文アブストラクト

Manipulation failures can leave scenes in states from which a task policy cannot recover. Learning corrective behaviors requires scalable failure exploration and physical grounding. We present Recova, an agent-guided framework that jointly develops task execution and recovery in a reconstructed digital twin, then verifies and refines both through real-world experience. In the twin, the agent diagnoses failures, tests corrective programs, and collects successful task and recovery rollouts for separate policies. During deployment, it monitors progress, invokes a learned or programmatic recovery, verifies scene restoration, and resumes execution. When no suitable recovery is available, a human demonstration resolves the failure and enters the learning loop, allowing the system to expand its recovery capabilities. Physical rollouts and human demonstrations are routed to the corresponding policy for DAgger training. Across six LIBERO-Pro settings and four MolmoSpaces categories, Recova achieves 78.8% and 64.9% mean success, compared with 71.7% and 38.0% for the strongest baselines. With parallel collection across four real-robot workstations, DAgger fine-tuning raises mean success from 23.8% to 77.5%, and recovery skills further raise it to 87.5%. Over four collection rounds on one task, observed human intervention falls from 87.5% to 0%. Together, these results show how agent-guided recovery turns failures into reusable capabilities, improving robustness while progressively reducing human intervention. Project page: https://www.liuisabella.com/Recova

関連論文

PR本紙発行元 EmplifAI