FLARE: 視覚言語ロボット操作における自動修正と回復のための障害認識フレームワーク
FLARE: A Failure-Aware Framework for Autonomous Correction and Recovery in Visual-Language Robotic Manipulation
視覚言語行動モデル(VLA)が実行中に起こす失敗(掴み損ね、落下、衝突など)を自動で修正・回復できるようにするフレームワークを提案。デモに摂動や橋渡しセグメントを注入して「リトライ」能力を、MLLMによる失敗分析で「リセット」スキルを獲得し、オンライン監視で切り替える。
著者: Ganlong Zhao, Zijia Tang, Xingping Chen, Zhanghui Kuang, Ye Tian, Guanbin Li
分類: cs.RO
原文アブストラクト
Vision-Language-Action Models~(VLAs) have demonstrated significant promise in generalizing to complex, long-horizon robotic manipulation tasks. However, their performance remains brittle, as they are typically trained on trajectory-monotonic, failure-free demonstrations. This reliance on ``perfect" data leaves them unable to recover from common execution errors, such as a missed grasp, a dropped object, or an unexpected collision. In this paper, we propose FLARE, a novel framework that endows VLAs with robust error recovery capabilities through a ``Retry" and ``Reset" paradigm. First, we introduce a ``Retry" mechanism by injecting perturbation and bridging segments that decouple robot pose from environment state into demonstrations, enabling the policy to autonomously handle execution deviations. Second, to address critical, state-breaking (OOD) failures, we introduce a ``Reset" pipeline. We leverage an MLLM for offline failure analysis to automatically identify OOD states from execution videos. This analysis enables the efficient, targeted collection of a small library of object-centric ``Reset" skills, which are trained to restore the environment to a task-valid state. Our full framework integrates these learned policies. At inference, an online MLLM monitor arbitrates between task execution and ``Reset" skills. Experiments on challenging, contact-rich manipulation tasks show our approach significantly improves task success and robustness.
関連論文
- PredVLA: サブミリオンパラメータの予測符号化ポリシーによるロボット操作マニピュレーション
- リー群制約付きMeanFlowによる高速生成的把持マニピュレーション
- VISTA: 視覚から推定する空間接触注意による高密度接触操作マニピュレーション
- PRISM: 投影統合型サンプリングベースMPCとベイズコスト調整による双腕マニピュレーションマニピュレーション
- ロボット操作のための軌道レベル連続行動表現マニピュレーション
- ブラウザ制御と混合ステッパードライバ構成による再現可能な視覚誘導6自由度ロボットアームマニピュレーション