日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.35575

F4R: 失敗駆動型の認識・再構築・洗練・再展開によるロボットの継続的自己改善

F4R: Failure-Driven Recognition, Reconstruction, Refinement, and Redeployment for Continual Robot Self-Improvement

シェア:XThreadsFacebookLINEはてブBluesky

ロボットの失敗を自動で診断し、シミュレーション環境で再構築して強化学習で改善する閉ループ学習フレームワークを提案。実世界での追加デモ収集なしに、4つのマニピュレーションタスクで高い成功率を達成した。

詳しい要約

1. どんなもの?

本論文は、vision-language-action models (VLA) の実世界性能を制約する要因(expert demonstrations の限界と物理的相互作用の理解不足)に対処するため、Failure for Rising (F4R) を提案する。F4R は failure-driven な real-to-sim-to-real 閉ループ学習フレームワークであり、実世界の失敗を対象とした policy 改善に変換する。具体的には、agent が rollout から失敗を自動識別・診断し、各失敗を task-relevant な空間的・物理的条件を保持した interactive で object-centric な table-top 環境として再構築する。その後、failure-conditioned sim-real co-training と再構築環境での targeted reinforcement learning により policy を洗練し、改善された policy を再展開する。新たに観測された失敗は次の再構築・学習サイクルに継続的にフィードバッ…

2. 先行研究と比べてどこがすごい?

従来の一般的な対処法は、新たに遭遇した失敗の実世界 demonstrations を追加収集することだが、これは高コスト・非効率・潜在的に危険・スケール困難である。F4R は追加の実世界 corrective demonstrations を収集することなく、失敗を real-to-sim-to-real 閉ループで活用して policy を改善する点が優れている。実世界評価では、4つの manipulation tasks において In-Distribution で 93.75%、Out-of-Distribution (OOD) で 90.0% の成功率を達成し、budget-matched な Targeted BC baseline を OOD 条件下で 18.75 パーセンテージポイント上回った。

3. 技術・手法の肝は?

F4R の肝は以下の通り。 - agent による rollout からの失敗の自動識別と診断。 - 各失敗を task-relevant な空間的・物理的条件を保持した interactive で object-centric な table-top 環境として再構築。 - failure-conditioned sim-real co-training と再構築環境での targeted reinforcement learning による policy 洗練。 - 改善された policy の再展開と、新たに観測された失敗の次の再構築・学習サイクルへの継続的フィードバック。 これにより real-to-sim-to-real の閉ループ学習を実現する。

4. どうやって有効だと検証した?

実世界評価を4つの manipulation tasks で実施した。結果として、In-Distribution で 93.75%、Out-of-Distribution (OOD) で 90.0% の成功率を達成した。また、budget-matched な Targeted BC baseline と比較し、OOD 条件下で 18.75 パーセンテージポイント上回った。追加の実世界 corrective demonstrations を収集することなくこの性能を達成した。

5. 議論はある?

要旨からは不明。

6. 次に読むべき論文は?

要旨で参照・比較されている研究として Targeted BC baseline が挙げられる。関連手法として vision-language-action models (VLA)、real-to-sim-to-real 学習、failure-driven learning、reinforcement learning、sim-real co-training などが考えられるが、具体的な論文名は要旨からは不明。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Zhuoyuan Yu, Jiacheng Wang, Tianle Liu, Yihua Ren, Peng Yu, Chen Bai, Ziheng Zhang, Yufei Jia, Jindou Jia, Yuhang Zhang, Xinrui Zhang, Shang Yujing, Yuxiang Chen, Chuhao Zhou, Tiancai Wang, Jianfei Yang

分類: cs.RO, cs.AI

原文アブストラクト

The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrations and their insufficient understanding of physical interactions. A common remedy is to collect additional real-world demonstrations of newly encountered failures. However, this process is costly, inefficient, potentially unsafe, and difficult to scale. To address this challenge, we propose Failure for Rising (F4R), a failure-driven real-to-sim-to-real closed-loop learning framework that converts real-world failures into targeted policy improvement. F4R first uses an agent to automatically identify and diagnose failures from rollouts. It reconstructs each failure as an interactive, object-centric table-top environment that preserves the task-relevant spatial and physical conditions. The policy is then refined through failure-conditioned sim-real co-training followed by targeted reinforcement learning in the reconstructed environments. The improved policy is subsequently redeployed, while newly observed failures are continuously fed back into the next reconstruction and learning cycle. Real-world evaluations on four manipulation tasks show that F4R achieves 93.75% In-Distribution and 90.0% Out-of-Distribution (OOD) success, outperforming the budget-matched Targeted BC baseline by 18.75 percentage points under OOD conditions without collecting additional real-world corrective demonstrations.

関連論文

PR本紙発行元 EmplifAI