日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
模倣学習arXiv:2610.03997

REDIRECT: ロボットの悪い癖を1%の修正で直す

REDIRECT: A 1% Fix for Bad Robot Habits

シェア:XThreadsFacebookLINEはてブBluesky

混合品質のテレオペデータから、問題のある局所的な行動分岐だけを特定し、1%の学習予算で修正する手法を提案。

著者: Yu Zhang, Jiazhuo Li, Yancong Wei, Kangkang Dong, Xiaojun Zhu, Houde Liu

分類: cs.RO

原文アブストラクト

Robots can acquire bad habits from a few defective moments in otherwise useful teleoperation. In mixed-quality robot data, normal and problematic demonstrations share most task behavior and differ only at a local action continuation. Full retraining is costly, while fine-tuning on clean data alone offers limited recovery under a small update budget. We ask whether the local difference can instead be fixed with only 1% of the full-training sample-backward budget. Using only episode-level retained/problematic labels, REDIRECT localizes the branch, assigns a coherent retained continuation to the original observations around it, and anchors shared behavior, without frame-level annotations or additional interaction. Across three randomized ManiSkill tasks and three seeds, REDIRECT raises the mean success rate from 68.2% to 90.3%, recovering 88.8% of the clean-retraining gap versus 22.1% for matched-compute fine-tuning. On a PiPER arm, it recovers 86.7-92.9% of the clean-only gap across cup insertion and towel folding. Local robot habits can therefore be repaired by spending the update budget on the behavioral difference rather than relearning shared behavior.

関連論文

PR本紙発行元 EmplifAI