日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
マニピュレーションarXiv:2610.07527

タスク空間模倣ガイダンスによる効率的強化学習

Task-Space Imitation Guidance for Efficient Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

模倣方策をタスク空間の進捗推定器として活用し、疎な報酬の卓上マニピュレーションにおいて密な進捗報酬と事前学習を組み合わせることで、サンプル効率と安全性を向上させる手法を提案。

詳しい要約

1. どんなもの?

- スパース報酬の卓上ロボットマニピュレーション向けに、報酬設計と事前学習を組み合わせたフレームワーク TIGER を提案。 - action-chunked imitation policy を実行用コントローラや action prior としてではなく、局所的な task-space の進捗推定器として扱う。 - 予測された action chunk を controller-aware action-to-motion mapping で短ホライズンの end-effector reference に変換し、RL エージェントに密な進捗報酬を与える。 - スパースな環境報酬は依然として支配的な目的として維持される。

2. 先行研究と比べてどこがすごい?

- 従来の RL および IL-RL ベースラインと比較して、評価タスクで初期の sample efficiency を改善。 - 測定された safety violations を削減。 - 最終 success rate はベースラインと同等以上。 - imitation policy を action prior や実行コントローラとして使うのではなく、task-space の進捗推定に用いる点が特徴的。

3. 技術・手法の肝は?

- action-chunked imitation policy の予測 action chunk を、controller-aware action-to-motion mapping により短ホライズンの end-effector reference へ変換。 - その reference への密な progress reward を RL エージェントに与える。 - 事前学習では imitation-guided look-ahead signals を用い、task-space の進捗が見込まれる行動に対する保守的な value penalty を緩和。 - これにより初期オンライン RL における off-manifold exploration を低減。

4. どうやって有効だと検証した?

- シミュレーションと実ロボット実験の両方で評価。 - 評価タスクにおいて、先行 RL および IL-RL ベースラインと比較。 - 初期 sample efficiency の改善、safety violations の削減、最終 success rate の同等以上を確認。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- 要旨で参照/比較されている prior RL および IL-RL baselines に関する研究。 - action-chunked imitation policy を用いた関連手法。 - スパース報酬のロボットマニピュレーションにおける RL 研究。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Salar Asayesh, Hossein Darani, Todd Cao, Evgeny Andriash, Mani Ranjbar

分類: cs.RO

原文アブストラクト

We introduce Task-Space Imitation Guidance for Efficient Reinforcement Learning (TIGER), a reward-construction and pretraining framework for sparse-reward tabletop robotic manipulation. TIGER treats an action-chunked imitation policy not as an executable controller or action prior, but as a local task-space progress estimator: predicted action chunks are converted, using controller-aware action-to-motion mapping, into short-horizon end-effector references, and the RL agent receives dense progress rewards toward these references while the sparse environment reward remains the dominant objective. During pretraining, TIGER uses imitation-guided look-ahead signals to relax conservative value penalties for actions predicted to make task-space progress, reducing off-manifold exploration during early online RL. Across simulation and real-robot experiments, TIGER improves early sample efficiency and reduces measured safety violations while matching or improving final success rates relative to prior RL and IL-RL baselines on the evaluated tasks.

関連論文

PR本紙発行元 EmplifAI