未モデル状態依存性下での学習クリティックによる最適制御
Optimal Control with Learned Critics under Unmodeled State Dependencies
解析モデルに残差動力学ネットワークを加え、学習した行動価値クリティックをiLQRに統合することで、接触などモデル化困難な状態依存性に対応する学習ベースMPCを提案。GPU加速バッチiLQRで効率的に学習する。
著者: Philipp Schoch, Markus Ryll
分類: cs.RO, cs.AI
原文アブストラクト
Model Predictive Control (MPC) provides a structured and constraint-aware mechanism for decision-making, but its reliance on optimization-friendly analytical dynamics models limits its use in tasks with contacts and other hard-to-model state dependencies. Model-free reinforcement learning avoids explicit modeling assumptions but typically requires large amounts of interaction data. We present a learning-based MPC framework that combines the data efficiency and structure of local model-based planning with learned components that compensate for incomplete dynamics and finite-horizon myopia. The method augments a nominal analytical model with a residual dynamics network that learns missing state-dependent effects from data and combines the resulting planner with a learned action-value critic that injects long-horizon MDP structure into the local iLQR optimization. To make this practical at reinforcement-learning scale, we develop a GPU-accelerated batched iLQR solver that evaluates learned dynamics and critic networks inside the optimal-control loop and solves thousands of trajectory-optimization problems in parallel. The complete system is integrated into a robotics simulator, enabling scalable model-based reinforcement learning under incomplete dynamics. Experiments on biased and incompletely modeled control tasks show that the approach improves closed-loop control performance while preserving the model-based structure needed for efficient constrained trajectory optimization.
関連論文
- VIGOR: モデルベース強化学習における潜在空間一貫性によるゼロショット視覚汎化モデルベース強化学習
- 行動なし時系列からの動的埋め込みによる転移可能な方策学習モデルベース強化学習
- CEMにおける世界モデルは提案メカニズムでもあるモデルベース強化学習
- 計画と学習のループを閉じる:学習済み世界モデルによるロボット制御モデルベース強化学習
- 速度と精度の両立:油圧ショベル制御のためのサンプル効率の高いオンラインモデルベース強化学習モデルベース強化学習
- 表現世界モデル:表現空間における状態・遷移・実行可能計画の学習モデルベース強化学習