オンライン対話はリカバリに必要か?摂動によるロバスト計画のミニマリストアプローチ
Is Online Interaction Necessary for Recovery? A Minimalist Approach to Robust Planning via Perturbation
既存のデモのみを用いて、予測軌道の後半部分を元のデモに固定しつつ前半を自由にすることで、摂動状態からのリカバリを学習するPARSを提案。追加の環境対話や専門家なしで行動クローンのロバスト性を向上。
著者: Bumgeun Park, Donghwan Lee
分類: cs.RO
原文アブストラクト
Behavior cloning (BC) is vulnerable to covariate shift during closed-loop execution, where small prediction or execution errors can drive the robot toward states poorly covered by the demonstration data. We focus on action-sequence planning, where a policy predicts a finite-horizon sequence of actions as a reference trajectory for robot execution. Existing approaches to covariate shift often rely on collecting additional corrective demonstrations, requiring further environment interaction and access to an expert or reference policy. We propose Perturbation-Augmented Recovery Supervision (PARS), a simple training approach for improving recovery using only existing demonstrations. PARS perturbs the robot's proprioceptive state and anchors only the later portion of the predicted trajectory to the original demonstration, leaving the earlier portion free to generate corrective motion and recover toward the demonstrated behavior within the planning horizon. Unlike conventional input-noise augmentation, which preserves the original supervision target over the entire prediction horizon, PARS explicitly provides trajectory-level supervision for recovery from perturbed states. PARS requires neither additional environment interaction nor expert queries and can be instantiated across different action-sequence policy classes with only minor modifications to their BC objectives. Experiments on 51 RLBench manipulation tasks with flow-based, transformer-based, and diffusion-based policies demonstrate that PARS improves robustness to covariate shift during closed-loop execution.