日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
オフライン強化学習arXiv:2609.40149

役割適応型方策最適化によるオフライン強化学習

Role-Adaptive Policy Optimization for Offline Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

オフライン強化学習において、価値学習と実行での方策の役割に応じて更新係数を適応させるRAPOを提案し、TD3+BCやIQLの性能を向上させた。

著者: Seonvin Cho, Soohyun Choi, Songnam Hong

分類: cs.LG

原文アブストラクト

Policy regularization in offline reinforcement learning balances policy improvement against reliance on uncertain value estimates. This balance can differ between selecting actions for execution and supplying actions for critic bootstrapping, yet methods such as TD3+BC couple these roles through a shared policy. We propose Role-Adaptive Policy Optimization (RAPO), which adapts policy-update coefficients according to their roles in value learning and execution. RAPO learns these coefficients by differentiating through candidate policy updates formed using the base algorithm's actor objective. For TD3+BC, RAPO separates bootstrap and execution actors and adapts their coefficients independently: the bootstrap objective penalizes policy-induced changes in target values, while the execution objective evaluates a local policy-improvement surrogate. For IQL, whose value learning is already independent of the execution actor, RAPO preserves the original value updates and adapts only the inverse temperature in advantage-weighted policy extraction. Experiments on D4RL locomotion and AntMaze tasks show improvements over both base algorithms, with larger gains for TD3+BC, whose RAPO instantiation outperforms baselines on average.

関連論文

PR本紙発行元 EmplifAI