すべてのモデルは誤り、どこが誤りかを知ることが有用:強化学習におけるモデル不確実性について
All Models are Wrong, Knowing Where is Useful: On Model Uncertainty in Reinforcement Learning
モデルベース強化学習におけるモデルの不正確さを扱う枠組みを提案し、不確実性を適切に扱うことでモデルの悪用を防ぎ、ハードウェア上での直接学習や安全な探索の成功を示した。
著者: Bernd Frauenknecht, Devdutt Subhasish, Artur Eisele, Friedrich Solowjow, Sebastian Trimpe
分類: cs.LG, eess.SY
原文アブストラクト
Model-based reinforcement learning (MBRL) infers information about the environment from a learned dynamics model and bears the potential to address open problems such as data efficient and safe learning in robotics. However, inaccuracies of the learned dynamics model are typically exploited by the agent, substantially hampering the capabilities of MBRL methods. We present a framework for dealing with inaccuracies of probabilistic models through targeted handling of uncertainty that effectively mitigates model exploitation. We present recent successes in learning directly on hardware and safe exploration, and discuss future directions for uncertainty-aware MBRL.
関連論文
- ニューロシンボリック世界モデルによるゼロショットタスク転送に向けてモデルベース強化学習
- BRICKS-WM: インターフェース合成力学による構造化世界モデルの再利用性構築モデルベース強化学習
- PRISM: ワールドモデルにおける事前知識誘導型想像サンプリングモデルベース強化学習
- 勾配ペナルティ付き潜在ダイナミクスによる滑らかでサンプル効率的な夢の学習モデルベース強化学習