日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
オフライン強化学習arXiv:2509.06782

オフライン目標条件付き強化学習のための物理情報付き価値学習

Physics-informed Value Learner for Offline Goal-Conditioned Reinforcement Learning

シェア:XThreadsFacebookLINEはてブBluesky

アイコナール方程式に基づく物理情報正則化を価値関数学習に導入し、オフライン目標条件付き強化学習の性能と汎化を改善した。

著者: Vittorio Giammarino, Ruiqi Ni, Ahmed H. Qureshi

分類: cs.LG

原文アブストラクト

Offline Goal-Conditioned Reinforcement Learning (GCRL) holds great promise for domains such as autonomous navigation and locomotion, where collecting interactive data is costly and unsafe. However, it remains challenging in practice due to the need to learn from datasets with limited coverage of the state-action space and to generalize across long-horizon tasks. To improve on these challenges, we propose a \emph{Physics-informed (Pi)} regularized loss for value learning, derived from the Eikonal Partial Differential Equation (PDE) and which induces a geometric inductive bias in the learned value function. Unlike generic gradient penalties that are primarily used to stabilize training, our formulation is grounded in continuous-time optimal control and encourages value functions to align with cost-to-go structures. The proposed regularizer is broadly compatible with temporal-difference-based value learning and can be integrated into existing Offline GCRL algorithms. When combined with Hierarchical Implicit Q-Learning (HIQL), the resulting method, Eikonal-regularized HIQL (Eik-HIQL), yields significant improvements in both performance and generalization, with pronounced gains in stitching regimes and large-scale navigation tasks.

関連論文