日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
モデルベース強化学習arXiv:2603.07083

Dreamer-CDP: 連続決定論的表現予測による再構築不要ワールドモデルの改善

Dreamer-CDP: Improving Reconstruction-free World Models Via Continuous Deterministic Representation Prediction

シェア:XThreadsFacebookLINEはてブBluesky

Dreamerなどのモデルベース強化学習エージェントで、観測再構築を使わずにJEPAスタイルの連続決定論的表現予測を導入し、Crafter環境で再構築ベース手法と同等の性能を達成した。

著者: Michael Hauri, Friedemann Zenke

分類: cs.LG

原文アブストラクト

Model-based reinforcement learning (MBRL) agents operating in high-dimensional observation spaces, such as Dreamer, rely on learning abstract representations for effective planning and control. Existing approaches typically employ reconstruction-based objectives in the observation space, which can render representations sensitive to task-irrelevant details. Recent alternatives trade reconstruction for auxiliary action prediction heads or view augmentation strategies, but perform worse in the Crafter environment than reconstruction-based methods. We close this gap between Dreamer and reconstruction-free models by introducing a JEPA-style predictor defined on continuous, deterministic representations. Our method matches Dreamer's performance on Crafter, demonstrating effective world model learning on this benchmark without reconstruction objectives.

関連論文