行動なし時系列からの動的埋め込みによる転移可能な方策学習
Learning Transferable Policies from Action-free Time Series Through Dynamical Embeddings
行動が観測されない時系列データから、関連システム間で共有される動的構造を低次元埋め込みで捉え、その埋め込みを再利用して制御方策を学習する階層的モデルベース強化学習を提案した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Niklas Emonds, Georgia Koppe
分類: cs.LG
原文アブストラクト
Learning control from action-free recordings is challenging because intervention effects are unobserved and policies may exploit errors in reconstructed dynamics. We present a hierarchical model-based reinforcement learning framework that uses shared structure across related systems to learn system-specific control policies from action-free recordings. A hierarchical dynamical system reconstruction model captures shared dynamics and individual variation through low-dimensional embeddings. These embeddings are then reused to parameterize shared policy and value networks, linking differences in reconstructed dynamics to differences in control. Policies are trained entirely via simulation under an explicit intervention model with additive latent perturbations. Piecewise-linear recurrent neural networks enable mechanistic analyses of the controlled dynamics, while decoder-based constraints make the immediate effects of interventions interpretable in observation space and permit interventions on one modality while protecting another from direct manipulation. On Lorenz-63 and double-pendulum systems, hierarchical policies improve transfer over independently trained policies. On Lorenz-63, they also achieve a higher mean reward than repeated planning with the same reconstructed models, perform comparably to methods trained with controlled interactions, and generalize to systems absent from policy training after embedding inference alone. Applications to neural-behavioral recordings demonstrate suppression of predicted movement under constrained neural perturbations. Together, these findings show how shared dynamical representations support transferable control and mechanistic hypothesis generation from action-free recordings.
関連論文
- VIGOR: モデルベース強化学習における潜在空間一貫性によるゼロショット視覚汎化モデルベース強化学習
- CEMにおける世界モデルは提案メカニズムでもあるモデルベース強化学習
- 計画と学習のループを閉じる:学習済み世界モデルによるロボット制御モデルベース強化学習
- 速度と精度の両立:油圧ショベル制御のためのサンプル効率の高いオンラインモデルベース強化学習モデルベース強化学習
- 表現世界モデル:表現空間における状態・遷移・実行可能計画の学習モデルベース強化学習
- CAST: 交互状態価値目標と拡張方策勾配によるモデルベース強化学習モデルベース強化学習