日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2609.30023

Res-HIL: 人間誘導型残差強化学習によるサンプル効率の良い器用なマニピュレーション

Res-HIL: Human-Guided Residual Reinforcement Learning for Sample-Efficient Dexterous Manipulation

シェア:XThreadsFacebookLINEはてブBluesky

模倣学習で得た方策を凍結し、その上に人間の介入から学ぶ残差方策を重ねることで、短時間のオンライン学習で接触の多い器用な操作タスクを高精度に実現する手法を提案。

著者: Mariia Iavorskaia, Christian Dietz, Sebastian Albrecht, Majid Khadiv

分類: cs.RO

原文アブストラクト

Imitation learning enables robots to acquire manipulation skills from demonstrations, but the resulting policies can fail outside the training data, while collecting more demonstrations requires substantial human effort. Human-in-the-loop reinforcement learning uses corrective feedback during online training, but typically learns the complete task policy rather than refining a pretrained imitation policy. We introduce Res-HIL, a human-in-the-loop residual reinforcement learning framework that learns corrective actions on top of a frozen imitation policy. Each human intervention provides two complementary learning signals: direct supervision of the residual policy and reward shaping of preceding autonomous behavior. Res-HIL combines these signals with zero initialization of the residual policy to stabilize and accelerate online learning. We evaluate Res-HIL on five contact-rich manipulation tasks spanning high-precision and long-horizon behaviors. With only 20 initial demonstrations, Res-HIL outperforms state-of-the-art full-policy human-in-the-loop reinforcement learning and residual fine-tuning without human guidance on every task after ten minutes of online training. Res-HIL improves its pretrained base policies and outperforms imitation policies trained with five times more demonstrations. An ablation study shows that direct residual supervision is critical to performance, while intervention-aware reward shaping substantially improves training efficiency.

関連論文

PR本紙発行元 EmplifAI