REDIRECT: ロボットの悪い癖を1%の修正で直す
REDIRECT: A 1% Fix for Bad Robot Habits
混合品質のテレオペデータから、問題のある局所的な行動分岐だけを特定し、1%の学習予算で修正する手法を提案。
著者: Yu Zhang, Jiazhuo Li, Yancong Wei, Kangkang Dong, Xiaojun Zhu, Houde Liu
分類: cs.RO
原文アブストラクト
Robots can acquire bad habits from a few defective moments in otherwise useful teleoperation. In mixed-quality robot data, normal and problematic demonstrations share most task behavior and differ only at a local action continuation. Full retraining is costly, while fine-tuning on clean data alone offers limited recovery under a small update budget. We ask whether the local difference can instead be fixed with only 1% of the full-training sample-backward budget. Using only episode-level retained/problematic labels, REDIRECT localizes the branch, assigns a coherent retained continuation to the original observations around it, and anchors shared behavior, without frame-level annotations or additional interaction. Across three randomized ManiSkill tasks and three seeds, REDIRECT raises the mean success rate from 68.2% to 90.3%, recovering 88.8% of the clean-retraining gap versus 22.1% for matched-compute fine-tuning. On a PiPER arm, it recovers 86.7-92.9% of the clean-only gap across cup insertion and towel folding. Local robot habits can therefore be repaired by spending the update budget on the behavioral difference rather than relearning shared behavior.