日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
オフライン強化学習arXiv:2610.08989

LASER: 潜在空間随伴マッチングによるサポート制約付きエントロピー正則化オフライン強化学習

LASER: Latent Space Adjoint Matching for Support-Constrained Entropy-Regularized Offline RL

シェア:XThreadsFacebookLINEはてブBluesky

フローマッチングで学習した振る舞いクローンポリシーの潜在空間内で、エントロピー正則化付きオフラインRLを行うLASERを提案し、OGBenchの40タスクで最先端性能を達成した。

詳しい要約

1. どんなもの?

- オフライン強化学習(offline RL)の新アルゴリズム LASER を提案。 - 静的なデータセットから方策を最適化するが、out-of-distribution(OOD)行動のリスクが課題。 - 近年の手法は flow matching で behavior-cloning 方策を学習し、その潜在空間内で RL を行う。 - しかし潜在方策の素朴な最適化は、脆いモードへの崩壊や critic の鋭いアーティファクトの悪用を招く。 - 本研究では潜在空間 RL において entropy regularization が不可欠であることを見出し、LASER を導入。 - LASER は latent-space adjoint matching を適用し、表現力豊かな flow 方策で entropy 正則化付き潜在空間 RL を実現。 - backpropagation through time を回避する。

2. 先行研究と比べてどこがすごい?

- 従来のオフライン RL は OOD 行動のリスクに悩まされる。 - 最近の手法は flow matching による behavior-cloning 方策の潜在空間で RL を行うが、素朴な最適化はモード崩壊や critic のアーティファクト悪用を招く。 - LASER は entropy regularization を潜在空間 RL に導入し、これらの問題に対処。 - 40 の挑戦的な OGBench タスクで、データセット品質が異なる設定において state-of-the-art 性能を達成。 - 特筆すべきは、LASER が全タスクで固定の method-specific hyperparameters を使用し、タスク・データセット固有のチューニングを行ったベースラインを上回ること。 - これにより LASER の堅牢な適用性が示される。

3. 技術・手法の肝は?

- 潜在空間における adjoint matching を適用し、entropy 正則化付き潜在空間 RL を実現。 - 表現力豊かな flow 方策を用いる。 - backpropagation through time を回避する。 - entropy regularization が潜在空間 RL の課題(モード崩壊、critic のアーティファクト悪用)に対処する鍵。 - 具体的なアルゴリズムの詳細は要旨からは不明。

4. どうやって有効だと検証した?

- 40 の挑戦的な OGBench タスクで、データセット品質を変化させて包括的な実験を実施。 - LASER が state-of-the-art 性能を達成することを示した。 - 全タスクで固定の method-specific hyperparameters を使用し、タスク・データセット固有のチューニングを行ったベースラインを含む評価対象を上回った。 - これにより LASER の堅牢な適用性を検証。

5. 議論はある?

- 要旨からは不明。 - ただし、entropy regularization が潜在空間 RL において不可欠であるという知見が示されている。 - 固定ハイパーパラメータでの堅牢性が強調されている。

6. 次に読むべき論文は?

- 要旨で参照/比較されている研究:flow matching を用いた behavior-cloning 方策と潜在空間 RL の最近のアプローチ。 - 評価に用いられた OGBench タスク。 - 関連手法として、offline RL の一般的な手法(例:CQL, IQL)や flow matching ベースの手法(例:FQL)が挙げられるが、要旨では具体的な論文名は明示されていない。 - 同分野の定番として、offline RL のサーベイや OGBench の原論文を読むことが推奨される。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Songyuan Zhang, Oswin So, Eric Yang Yu, Matthew Cleaveland, Peter Crowley-Dolen, Chuchu Fan

分類: cs.LG, cs.AI, cs.RO, math.OC, stat.ML

原文アブストラクト

While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performing RL within its constrained latent space. However, naively optimizing the latent policy can easily cause the policy to collapse into a brittle mode or exploit sharp artifacts of the learned critic. In this work, we find that entropy regularization is essential in latent-space RL for addressing these challenges. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regularized latent-space RL with expressive flow policies while avoiding backpropagation through time. Through comprehensive experiments on 40 challenging OGBench tasks with varying dataset qualities, we show that LASER achieves state-of-the-art performance. Notably, LASER uses fixed method-specific hyperparameters across all tasks and outperforms the evaluated baselines, including those with task- and dataset-specific tuning, which highlights the robust applicability of LASER. Project website: https://mit-realm.github.io/laser/.

関連論文

PR本紙発行元 EmplifAI