提案条件付きリファインメントフローによる拡散方策の改善
Diffusion Policy Improvement with Proposal-Conditioned Refinement Flows
拡散・フロー方策のKL正則化の限界を克服するため、批評家に基づく提案選択と条件付きリファインメントフローを組み合わせたPReFlowを提案し、オフラインRLとオンライン微調整で高い性能を達成した。
著者: Junhyun Ha, Juho Lee, Byungwoo Park
分類: cs.LG, cs.AI, cs.RO
原文アブストラクト
Diffusion and flow policies can model complex behaviors in offline reinforcement learning (RL). However, penalizing their KL divergence from the behavior policy can discourage actions having high critic values with low behavior density. Directly refining behavior proposals may be an alternative, yet Gaussian or deterministic editors limit expressiveness to represent multiple separated modes for the same proposal. In this work, we introduce Proposal-Conditioned Refinement Flows (PReFlow), a policy extraction method combining critic-based proposal selection with a conditional refinement flow. To optimize proposal selection and refinement together, we formulate a KL-regularized objective whose optimum induces a Gibbs policy over final actions under a Gaussian-smoothed behavior prior. The refinement flow can represent multiple high value modes, while a proposal-centered Gaussian reference regulates large action changes. This Gaussian reference further enables us to make use of simulation-free, closed form adjoint matching targets from sampled endpoints and critic gradients, yielding a single velocity regression loss without a backward adjoint solve. On 50 OGBench tasks, PReFlow achieves competitive offline performance and the highest aggregate score among the compared methods after online fine-tuning, reaching 91\% after 500K environment steps.