LASER: 潜在空間随伴マッチングによるサポート制約付きエントロピー正則化オフライン強化学習
LASER: Latent Space Adjoint Matching for Support-Constrained Entropy-Regularized Offline RL
フローマッチングで学習した振る舞いクローンポリシーの潜在空間内で、エントロピー正則化付きオフラインRLを行うLASERを提案し、OGBenchの40タスクで最先端性能を達成した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Songyuan Zhang, Oswin So, Eric Yang Yu, Matthew Cleaveland, Peter Crowley-Dolen, Chuchu Fan
分類: cs.LG, cs.AI, cs.RO, math.OC, stat.ML
原文アブストラクト
While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performing RL within its constrained latent space. However, naively optimizing the latent policy can easily cause the policy to collapse into a brittle mode or exploit sharp artifacts of the learned critic. In this work, we find that entropy regularization is essential in latent-space RL for addressing these challenges. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized latent-space RL with expressive flow policies while avoiding backpropagation through time. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses fixed method-specific hyperparameters across all tasks and outperforms the evaluated baselines, including those with task- and dataset-specific tuning, which highlights the robust applicability of LASER. Project website: https://mit-realm.github.io/laser/.
関連論文
- 対話制約付きオフライン強化学習による自動運転オフライン強化学習
- 役割適応型方策最適化によるオフライン強化学習オフライン強化学習
- VGFM: フローマッチングにおける密な価値誘導による表現力豊かなロボット方策オフライン強化学習
- オフライン強化学習における拡散ポリシーのためのノイズ空間ポリシー勾配オフライン強化学習
- オフライン方策改善に1ステップで十分か?オフライン強化学習
- CoDrift: オフライン強化学習のための合成的ドリフトオフライン強化学習