Res-HIL: 人間誘導型残差強化学習によるサンプル効率の良い器用なマニピュレーション
Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation
模倣学習で得た方策を凍結し、その上に人間の介入から学ぶ残差方策を重ねることで、短時間のオンライン学習で接触の多い器用な操作タスクを高精度に実現する手法を提案。
著者: Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv
分類: cs.RO
原文アブストラクト
Imitation learning enables robots to acquire manipulation skills from demonstrations, but the resulting policies can fail outside the training data, while collecting more demonstrations requires substantial human effort. Human-in-the-loop reinforcement learning uses corrective feedback during online training, but typically learns the complete task policy rather than refining a pretrained imitation policy. We introduce Res-HIL, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy. Each human intervention provides two complementary learning signals: direct supervision of the residual policy and reward shaping of preceding autonomous behavior. Res-HIL combines these signals with zero initialization of the residual policy to stabilize and accelerate online learning. We evaluate Res-HIL on five contact-rich manipulation tasks spanning high-precision and long-horizon behaviors. With only 20 initial demonstrations, Res-HIL outperforms state-of-the-art full-policy human-in-the-loop reinforcement learning and residual fine-tuning without human guidance on every task after ten minutes of online training. Res-HIL improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations. An ablation study shows that direct residual supervision is critical to performance, while intervention-aware reward shaping substantially improves training efficiency.
関連論文
- Streaming-WAM: 非同期ロボットマニピュレーションのための行動条件付きワールドアクションモデルマニピュレーション
- シミュレータ非依存の布操作のための簡易グリッパインタフェースマニピュレーション
- WRAP: 治具不要の力覚考慮型マルチロボット組立計画マニピュレーション
- 受動的実行から能動的探索へ:実環境におけるエージェント型身体性マニピュレーションマニピュレーション
- 全手把持のためのリアルタイム力制御フレームワークマニピュレーション
- 剛体-空気圧ハイブリッドマニピュレータの連成状態空間モデリング・制御・方策蒸留マニピュレーション