日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
安全強化学習arXiv:2609.32952

制約付きフローポリシー更新:一般化シュレーディンガー橋からの視点

Constrained Flow Policy Updates: A Generalized Schrödinger Bridge View

シェア:XThreadsFacebookLINEはてブBluesky

安全制約付き強化学習において、フローポリシーの生成経路を通じて拡張ラグランジュ目的を直接微分し、スコア推定を不要にした新しいオフポリシーアクター・クリティック法を提案。

著者: Boyang Li, Matthew Kim, Sylvia Herbert

分類: cs.LG, cs.RO

原文アブストラクト

Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. Reward and safety can induce multimodal action distributions, challenging the prevailing primal-dual methods: Gaussian actors may collapse onto a single suboptimal mode, and optimization over the nonconvex Lagrangian landscape can be unstable. Diffusion and flow policies can represent such distributions, but recent work with a diffusion actor relies on estimating and matching the score of an augmented-Lagrangian target policy. Instead, we differentiate the augmented objective directly through the generation path of a flow policy, so no score needs to be estimated. Because a flow policy lacks a readily available action log-density for entropy regularization, we build on the density-free kinetic-energy regularizer of FLAC, a recent reward-only method, and propose Reparameterized Augmented-Lagrangian Flow Actor with Least Energy (RAFALE), an off-policy actor-critic method for safe RL. We formulate its update as a constrained one-ended generalized Schrödinger bridge and show that, for each source draw, this path-space problem is exactly an entropy-regularized problem in action space. At positive noise, its solution reweights the reward-only action distribution only where the estimated cost exceeds a threshold set by the Lagrange multiplier. As the noise vanishes, the optimal value converges to that of a least-energy map objective that the flow policy optimizes directly. Across seven Safety-Gymnasium tasks, RAFALE achieves competitive reward with mean final cost within budget on every task, whereas strong baselines trade one for the other; ablations support the necessity of both its augmented objective and its flow actor.

関連論文

PR本紙発行元 EmplifAI