日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
オフライン強化学習arXiv:2609.03842

オフライン方策改善に1ステップで十分か?

Is One Step Enough for Offline Policy Improvement?

シェア:XThreadsFacebookLINEはてブBluesky

オフライン強化学習における方策改善を多段近接方策改善(MPI)として定式化し、固定全ホライズンでの分割と局所ホライズンでの追加精緻化の効果を分析した。TD3+BCとIQLで実験し、分割が有用な全ホライズンの範囲を広げ、追加精緻化がリターンを改善することを示した。

著者: Soohyun Choi, Seonvin Cho, Songnam Hong

分類: cs.LG

原文アブストラクト

Behavior regularization in offline reinforcement learning limits the exploitation of critic errors, but strong anchoring can also restrict policy improvement. We study how policy improvement is composed through multi-step proximal policy improvement (MPI), which re-centers each proximal objective on the preceding policy. We parameterize the procedure by a nominal total horizon $T$ and $K$ stages with local horizon $T/K$, distinguishing subdivision at a fixed total horizon from additional refinement at a common local horizon. Our analysis shows that sequential re-centering can reach endpoints unavailable to any single proximal step and characterizes how subdivision reduces the leading local discretization error of ideal updates under a fixed critic. We consider TD3+BC and IQL-based policy extraction to examine how improvement composition interacts with actor objectives and policy geometry. TD3+BC experiments on D4RL locomotion suggest that subdivision can broaden the range of useful total horizons, while adding refinement stages at a fixed small local horizon can improve return. The results identify improvement composition as a design choice alongside regularization strength, with distinct effects from horizon subdivision and additional policy extraction.

関連論文

PR本紙発行元 EmplifAI