DreamerV3-XP:不確実性推定による探索の最適化
DreamerV3-XP: Optimizing exploration through uncertainty estimation
世界モデル強化学習DreamerV3に優先経験再生とアンサンブル不確実性に基づく内的報酬を追加し、疎な報酬環境での探索と学習効率を改善した。
著者: Lukas Bierling, Davide Pasero, Jan-Henrik Bertrand, Kiki Van Gerwen
分類: cs.LG, cs.AI
原文アブストラクト
We introduce DreamerV3-XP, an extension of DreamerV3 that improves exploration and learning efficiency. This includes (i) a prioritized replay buffer, scoring trajectories by return, reconstruction loss, and value error and (ii) an intrinsic reward based on disagreement over predicted environment rewards from an ensemble of world models. DreamerV3-XP is evaluated on a subset of Atari100k and DeepMind Control Visual Benchmark tasks, confirming the original DreamerV3 results and showing that our extensions lead to faster learning and lower dynamics model loss, particularly in sparse-reward settings.
関連論文
- 効率的な一次強化学習のための局所・大域世界モデルの結合強化学習/世界モデル
- 因果関係を考慮した強化学習のためのオブジェクト中心の世界モデル強化学習/世界モデル
- 注目を学ぶ:部分観測強化学習における構造的注意機構による情報履歴の優先強化学習/世界モデル