リプシッツ正則化クリティックによる遷移ダイナミクス不確実性に対する方策ロバスト性
Lipschitz-Regularized Critics Lead to Policy Robustness Against Transition Dynamics Uncertainty
PPOに敵対的状態計算(PGD)とリプシッツ正則化クリティックを組み合わせ、遷移ダイナミクスの不確実性下でもロバストな方策を学習する手法を提案し、実機ロボット歩行を含む実験で有効性を示した。
著者: Xulin Chen, Ruipeng Liu, Zhenyu Gan, Garrett E. Katz
分類: cs.LG
原文アブストラクト
Uncertainties in transition dynamics pose a critical challenge in reinforcement learning (RL), often resulting in performance degradation of trained policies when deployed on hardware. Many robust RL approaches follow two strategies: enforcing smoothness in actor or actor-critic modules with Lipschitz regularization, or learning robust Bellman operators. However, the first strategy does not investigate the impact of critic-only Lipschitz regularization on policy robustness, while the second lacks comprehensive validation in real-world scenarios. Building on this gap and prior work, we propose PPO-PGDLC, an algorithm based on Proximal Policy Optimization (PPO) that integrates Projected Gradient Descent (PGD) with a Lipschitz-regularized critic (LC). The PGD component calculates the adversarial state within an uncertainty set to approximate the robust Bellman operator, and the Lipschitz-regularized critic further improves the smoothness of learned policies. Experimental results on two classic control tasks and one real-world robotic locomotion task demonstrate that, compared to several baseline algorithms, PPO-PGDLC achieves better performance and predicts smoother actions under environmental perturbations.