隠れ状態は価値勾配である:リカレント方策のポントリャーギン構造
Hidden States as Value Gradients: The Pontryagin Structure of Recurrent Policies
リカレント方策の隠れ状態がポントリャーギンの最小原理における随伴状態(価値関数の勾配)として解釈できることを示し、それを活用した随伴損失で歩行タスクの方策を改善した。
著者: David Leeftink, Max Hinne, Marcel van Gerven
分類: cs.LG
原文アブストラクト
A key capability of intelligent agents is to act effectively under incomplete state observations. Recurrent policies address this by compressing observation histories into a hidden state. In this work, we show that the hidden state of a recurrent policy admits a control-theoretic interpretation: it plays the role of the co-state in Pontryagin's minimum principle, and the readout that maps it to actions implements control-Hamiltonian minimization. The hidden state thus tracks the gradient of the value function, encoding the optimality structure of the underlying control problem. We formalize this correspondence through a class of policies we refer to as co-state policies (CPs) and show that several modern recurrent cells implicitly realize this structure. The correspondence also allows for a co-state loss for actor-critic training, in which the critic's gradient serves as a target for the actor's hidden state. Empirically, we find that hidden states trained with the co-state loss encode co-state information beyond what is linearly decodable from the environment state alone, and that the loss improves policies both within and outside this class on challenging locomotion tasks, including the H1 and Berkeley humanoids. By connecting the minimum principle to recurrent memory, we provide a control-theoretic account of what hidden states compute in continuous control and a mechanism for shaping them toward optimality.