動的計画法を超えて:スコアライフプログラミングによる強化学習
Beyond dynamic programming
無限ホライズンの行動列を有界区間の実数に写像することで、方策関数を必要とせず最適な行動列を直接計算する新理論を提案し、非線形最適制御問題で有効性を示した。
分類: cs.LG, cs.AI, cs.RO
原文アブストラクト
In this paper, we present Score-life programming, a novel theoretical approach for solving reinforcement learning problems. In contrast with classical dynamic programming-based methods, our method can search over non-stationary policy functions, and can directly compute optimal infinite horizon action sequences from a given state. The central idea in our method is the construction of a mapping between infinite horizon action sequences and real numbers in a bounded interval. This construction enables us to formulate an optimization problem for directly computing optimal infinite horizon action sequences, without requiring a policy function. We demonstrate the effectiveness of our approach by applying it to nonlinear optimal control problems. Overall, our contributions provide a novel theoretical framework for formulating and solving reinforcement learning problems.