計画と学習のループを閉じる:学習済み世界モデルによるロボット制御
Beyond Policy Alignment: Closing the Planning-Learning Loop for Robot Control with Learned World Models
学習済み世界モデルを用いたモデル予測制御(MPC)において、批評家の監督と計画の終端価値推定を改良し、計画と学習のフィードバックループを強化するPL-MPCを提案。HumanoidBenchで性能向上と実機へのゼロショット転移を実証。
著者: Kowndinya Boyalakuntla, Yuhan Liu, Abdeslam Boularias
分類: cs.RO
原文アブストラクト
Planning with learned world models combines online trajectory optimization with learned value and policy functions for high-dimensional control. Because the planner determines the experience used for learning, while the learned critic and actor in turn score and propose future plans, planning and learning form a closed feedback loop. TD-MPC is a prominent instance of this design. Recent policy-constrained variants strengthen one part of the loop by aligning the learned policy with planner behavior. We introduce PL-MPC (Planning-Learning MPC), which additionally modifies critic supervision and planner terminal-value estimation. Hybrid multi-step TD targets expose critic updates to more realized rewards before bootstrapping; disagreement-aware terminal estimates reduce the influence of uncertain critic values during MPPI planning; and return-weighted actor distillation emphasizes planner-executed actions from high-return episodes. The world-model architecture and MPPI optimizer are otherwise unchanged. On HumanoidBench, the largest gains occur on \texttt{balance-hard}, where Total Average Return (TAR) increases from $98\pm18$ to $387\pm255$, and \texttt{hurdle}, from $199\pm13$ to $466\pm200$; performance across the broader benchmark remains task dependent, and PL-MPC remains competitive on DMControl. Controlled ablations show different component interactions across the two tasks. We further demonstrate zero-shot sim-to-real transfer on wrench-nut alignment with a 7-DoF KUKA IIWA14, obtaining higher observed success than TD-M(PC)^2 on the training object size and two unseen sizes. Code and data will be available at: https://pl-mpc-humanoid.github.io.
関連論文
- CEMにおける世界モデルは提案メカニズムでもあるモデルベース強化学習
- 速度と精度の両立:油圧ショベル制御のためのサンプル効率の高いオンラインモデルベース強化学習モデルベース強化学習
- 表現世界モデル:表現空間における状態・遷移・実行可能計画の学習モデルベース強化学習
- CAST: 交互状態価値目標と拡張方策勾配によるモデルベース強化学習モデルベース強化学習
- ニューロシンボリック世界モデルによるゼロショットタスク転送に向けてモデルベース強化学習
- BRICKS-WM: インターフェース合成力学による構造化世界モデルの再利用性構築モデルベース強化学習