目標条件付き強化学習のための世界モデル政策アービター
World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning
複数の凍結済み目標条件付き政策をポートフォリオとして活用し、学習した世界モデルで各政策の将来を想像評価して最適な政策を選択するテスト時フレームワークを提案。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Junwei Quan, Evgenii Opryshko, Nicholas Rhinehart, Igor Gilitschenski
分類: cs.LG
原文アブストラクト
Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task. Rather than deploying only the best-performing policy, we ask whether a set of frozen goal-conditioned policies can be used collectively as a portfolio, deciding at every state which policy should act. Choosing a policy at each state is not straightforward. The policies' own value functions cannot be compared directly: they may use different scales, and some policies have no value function. We need to judge each policy by the states it is likely to reach, even though we can execute only one policy at a time. We also need to avoid switching so often that control becomes unstable. To address these challenges, we introduce World-Model Policy Arbiter (WMPA), a test-time framework that, given a set of frozen policies as input, rolls out each frozen policy in a learned state-space world model, evaluates the imagined futures with a shared goal-conditioned value function, and executes the highest-scoring policy for a short commitment interval before the next round of arbitration (policy selection). WMPA assumes access to a bank of frozen goal-conditioned policies and requires neither policy retraining nor privileged task-specific knowledge. Under the official OGBench evaluation protocol on 18 state-based datasets spanning maze navigation as well as cube, scene, and puzzle manipulation, WMPA improves the macro-average success rate from the 44% achieved by the best policy selected per dataset to 58%, with statistically significant gains on 12 datasets. These gains include +33 percentage points on cube-double-play and +36 percentage points on scene-play.
関連論文
- アイコナール制約付き階層的準距離強化学習による目標到達目標条件付き強化学習
- スケーラブルな目標条件付き強化学習のためのマルチステップ準距離学習目標条件付き強化学習
- 目標条件付き強化学習のためのヌル反事実的因子相互作用目標条件付き強化学習
- 制約なし目標ナビゲーションのための世界モデル学習目標条件付き強化学習
- もつれ表現と到達可能性計画を組み合わせた目標条件付き強化学習目標条件付き強化学習