高次行動監督による強力な方策クラス
Higher-Order Action Supervision Makes A Strong Policy Class
模倣学習や強化学習において、0次(行動ラベル)だけでなく1次(行動の時間変化)も同時に監督する損失を提案し、既存の方策モデルに構造変更なしで組み込めるようにした。連続制御タスクで性能とロバスト性が大幅に向上。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Peng Cheng, Yunxian Hou, Zhi Zhou, Qian Zhang, Chang Huang, Xianyuan Zhan
分類: cs.RO, cs.AI
原文アブストラクト
Modern data-driven decision-making methods, such as imitation learning (IL) and reinforcement learning (RL), have achieved great success in solving many complex tasks. However, these methods often suffer from serious control instability and robustness issues when applied in real-world applications such as robotics and autonomous driving, posing notable challenges for their practical deployment. We argue that this instability issue stems largely from their limitations in solely supervising and optimizing zeroth-order actions (i.e., the action labels), failing to account for higher-order action dynamics and temporal consistency. In this paper, we show that simultaneously supervising both zeroth- and first-order actions can dramatically enhance policies' performance and control robustness. To achieve this, we introduce a novel and elegant loss scheme supported by formal theoretical guarantees that can equip any off-the-shelf policy model (e.g., deterministic, stochastic, or flow policies) with the capability for higher-order action supervision, without requiring any structural modifications. Moreover, our proposed method can serve as a lightweight plug-and-play module that seamlessly integrates with a broad spectrum of existing offline RL frameworks. Extensive evaluations on OGBench and D4RL demonstrate that our approach yields substantial performance and robustness improvements across a wide range of continuous control environments. Notably, our method can also enhance policies' out-of-distribution (OOD) generalization capability in the challenging low-data regime, making it an ideal tool in tackling many real-world control problems.
関連論文
- ConTrack: 適応的トレードオフ制御による拘束付き手の動作追跡模倣学習/強化学習
- 不完全なデモンストレーションからの学習:時間行動木ガイドによる軌道修復模倣学習/強化学習
- IN-RIL: 模倣学習と強化学習を交互に組み合わせた方策ファインチューニング模倣学習/強化学習
- DFM: 表現豊かなダンス動作学習のための深層フーリエ模倣模倣学習/強化学習
- 強化学習によるリザバーダイナミクス変調を用いた効率的なロボットスキル合成模倣学習/強化学習
- BMP:Bスプラインと運動プリミティブの橋渡し模倣学習/強化学習