日本フィジカルAI新聞

世界のフィジカルAIを、日本語で。

週刊ニュースレター購読
モデルベース強化学習arXiv:2610.03065

行動なし時系列からの動的埋め込みによる転移可能な方策学習

Learning Transferable Policies from Action-free Time Series Through Dynamical Embeddings

シェア:XThreadsFacebookLINEはてブBluesky

行動が観測されない時系列データから、関連システム間で共有される動的構造を低次元埋め込みで捉え、その埋め込みを再利用して制御方策を学習する階層的モデルベース強化学習を提案した。

詳しい要約

1. どんなもの?

アクションなしの時系列データから、関連するシステム間で共有される構造を利用して、システム固有の制御ポリシーを学習する階層的モデルベース強化学習フレームワーク。階層的動力学システム再構成モデルが低次元埋め込みを通じて共有動力学と個体差を捉え、その埋め込みを共有ポリシー・価値ネットワークのパラメータ化に再利用する。ポリシーは明示的な介入モデルと加法的潜在摂動の下でのシミュレーションのみで訓練される。Lorenz-63と二重振り子システムで評価され、神経行動記録への応用も示す。

2. 先行研究と比べてどこがすごい?

アクションなし記録からの制御学習は介入効果が未観測で、再構成動力学の誤りをポリシーが利用しうる点が課題。本研究は関連システム間の共有構造を階層的埋め込みで活用し、独立訓練ポリシーより転移が改善。Lorenz-63では同じ再構成モデルでの繰り返し計画より高い平均報酬、制御付き相互作用で訓練した手法と同等の性能、埋め込み推論のみでポリシー訓練にないシステムへ汎化。

3. 技術・手法の肝は?

階層的動力学システム再構成モデルが低次元埋め込みで共有動力学と個体差を捕捉。埋め込みを共有ポリシー・価値ネットワークのパラメータ化に再利用し、再構成動力学の差異と制御の差異を結びつける。ポリシーは明示的介入モデルと加法的潜在摂動下のシミュレーションで訓練。区分線形リカレントニューラルネットワークで制御動力学の機械論的分析を可能にし、デコーダベース制約で介入の即時効果を観測空間で解釈可能にし、一モダリティへの介入と他モダリティの直接操作保護を実現。

4. どうやって有効だと検証した?

Lorenz-63と二重振り子システムで、階層的ポリシーが独立訓練ポリシーより転移を改善。Lorenz-63では同じ再構成モデルでの繰り返し計画より高い平均報酬、制御付き相互作用で訓練した手法と同等、埋め込み推論のみでポリシー訓練にないシステムへ汎化。神経行動記録への応用で、制約付き神経摂動下で予測運動の抑制を実証。

5. 議論はある?

アクションなし記録からの制御学習は介入効果が未観測で、再構成動力学の誤りをポリシーが利用しうる点が課題として挙げられる。共有動力学表現が転移可能な制御と機械論的仮説生成を支えることが示唆される。その他の限界や議論は要旨からは不明。

6. 次に読むべき論文は?

要旨で参照/比較されている研究は明示されていない。関連手法として、model-based reinforcement learning、hierarchical dynamical system reconstruction、piecewise-linear recurrent neural networks、Lorenz-63、double-pendulum、neural-behavioral recordings が挙げられる。同分野の定番として、Dreamer、PlaNet、World Models などが次に読むべき候補。

※ AIが要旨から生成した要約です。正確性は原文をご確認ください。

著者: Niklas Emonds, Georgia Koppe

分類: cs.LG

原文アブストラクト

Learning control from action-free recordings is challenging because intervention effects are unobserved and policies may exploit errors in reconstructed dynamics. We present a hierarchical model-based reinforcement learning framework that uses shared structure across related systems to learn system-specific control policies from action-free recordings. A hierarchical dynamical system reconstruction model captures shared dynamics and individual variation through low-dimensional embeddings. These embeddings are then reused to parameterize shared policy and value networks, linking differences in reconstructed dynamics to differences in control. Policies are trained entirely via simulation under an explicit intervention model with additive latent perturbations. Piecewise-linear recurrent neural networks enable mechanistic analyses of the controlled dynamics, while decoder-based constraints make the immediate effects of interventions interpretable in observation space and permit interventions on one modality while protecting another from direct manipulation. On Lorenz-63 and double-pendulum systems, hierarchical policies improve transfer over independently trained policies. On Lorenz-63, they also achieve a higher mean reward than repeated planning with the same reconstructed models, perform comparably to methods trained with controlled interactions, and generalize to systems absent from policy training after embedding inference alone. Applications to neural-behavioral recordings demonstrate suppression of predicted movement under constrained neural perturbations. Together, these findings show how shared dynamical representations support transferable control and mechanistic hypothesis generation from action-free recordings.

関連論文

PR本紙発行元 EmplifAI