日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
ロバスト強化学習arXiv:2506.16753

敵対的観測に対するロバスト強化学習のためのオフポリシーアクター・クリティック:対称的政策評価による仮想代替訓練

Off-Policy Actor-Critic for Adversarial Observation Robustness: Virtual Alternative Training via Symmetric Policy Evaluation

シェア:XThreadsFacebookLINEはてブBluesky

敵対的観測に強い強化学習を目指し、エージェントと敵対者の政策評価の対称性を利用して追加の環境相互作用なしに学習するオフポリシー手法を提案した。

著者: Kosuke Nakanishi, Akihiro Kubo, Yuji Yasui, Shin Ishii

分類: cs.LG, cs.AI, cs.RO

原文アブストラクト

Recently, robust reinforcement learning (RL) methods designed to handle adversarial input observations have received significant attention, motivated by RL's inherent vulnerabilities. While existing approaches have demonstrated reasonable success, addressing worst-case scenarios over long time horizons requires both minimizing the agent's cumulative rewards for adversaries and training agents to counteract them through alternating learning. However, this process introduces mutual dependencies between the agent and the adversary, making interactions with the environment inefficient and hindering the development of off-policy methods. In this work, we propose a novel off-policy method that eliminates the need for additional environmental interactions by reformulating adversarial learning as a soft-constrained optimization problem. Our approach is theoretically supported by the symmetric property of policy evaluation between the agent and the adversary. The implementation is available at https://github.com/nakanakakosuke/VALT_SAC.

関連論文

PR本紙発行元 EmplifAI