日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
強化学習arXiv:2610.02198

FERPO: 前方エントロピー正則化方策最適化

FERPO: Forward Entropy-Regularized Policy Optimization

シェア:XThreadsFacebookLINEはてブBluesky

クリティックの行動微分を使わず、前方KL目的と自己正規化重要度サンプリングで方策改善を行う最大エントロピー強化学習アルゴリズムを提案し、連続制御タスクで性能とサンプル効率を向上させた。

詳しい要約

1. どんなもの?

- 連続制御のオンライン強化学習アルゴリズム。 - 名前は Forward Entropy-Regularized Policy Optimization (FERPO)。 - on-policy の maximum entropy RL。 - critic の action 微分を使わず、critic の value のみで policy improvement を行う。 - entropy と KL divergence で正則化した policy-improvement 目的から最適な target action distribution を導出。 - actor は forward-KL 目的でこの target に fit する。

2. 先行研究と比べてどこがすごい?

- 従来の SOTA は learned critic の action gradients で policy を改善するが、critic は return 予測用で value 予測が正確でも action 微分が正確とは限らず、policy update が不安定になりうる。 - FERPO は critic を action で微分せず、value のみを使うため、この問題を回避。 - reverse-KL 目的は target 分布の一部の mode を好むが、forward-KL は複数の high-value mode の coverage を促し exploration を促進。 - REPPO より actor update が高速。

3. 技術・手法の肝は?

- entropy と KL divergence で正則化した policy-improvement 目的を解き、最適な target action distribution を導出。 - actor を forward-KL 目的の最小化で target に fit。 - forward-KL は rollout policy から得た action を用いた self-normalized importance sampling (SNIS) で推定。 - KL 正則化により target 分布の rollout policy からの逸脱を制限し、importance weights を安定化。 - forward-KL は複数の high-value mode の coverage を促す。

4. どうやって有効だと検証した?

- MuJoCo Playground と ManiSkill で実験。 - 競争力のある性能と sample-efficiency の向上を示す。 - ablation を実施。 - computational benchmark で REPPO より actor update が高速であることを示す。

5. 議論はある?

- 要旨からは不明。

6. 次に読むべき論文は?

- Relative Entropy Pathwise Policy Optimization (REPPO) - maximum entropy reinforcement learning - self-normalized importance sampling (SNIS) - MuJoCo Playground - ManiSkill

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv

分類: cs.LG, cs.AI, cs.RO, stat.ML

原文アブストラクト

Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the actor to this target by minimizing a forward-KL objective, estimated using self-normalized importance sampling (SNIS) with actions drawn from the rollout policy. By limiting the target distribution's deviation from the rollout policy, the KL regularization helps keep these importance weights well behaved. In contrast to reverse-KL objectives, which can favor a subset of the target distribution's modes, the forward-KL objective encourages coverage of multiple high-value modes and thereby promotes exploration. Experiments and ablations on MuJoCo Playground and ManiSkill show competitive performance and sample-efficiency gains. Computational benchmarks also demonstrate faster actor updates than Relative Entropy Pathwise Policy Optimization (REPPO).

関連論文

PR本紙発行元 EmplifAI