ニューロシンボリック世界モデルによるゼロショットタスク転送に向けて
Towards Zero-Shot Task Transfer with Neurosymbolic World Models
報酬予測を潜在状態のシンボリックな部分に依存させる新しい世界モデルを提案し、同じシンボリック状態空間上の新しい報酬関数へ環境相互作用なしで適応できることを示した。
詳しい要約
1. どんなもの?
2. 先行研究と比べてどこがすごい?
3. 技術・手法の肝は?
4. どうやって有効だと検証した?
5. 議論はある?
6. 次に読むべき論文は?
※ AIが要旨から生成した要約です。正確性は原文をご確認ください。
著者: Isidoro Tamassia, Lennert De Smet, Giuseppe Marra
分類: cs.AI, cs.LG
原文アブストラクト
State-of-the-art model-based reinforcement learning methods learn neural world models that allow policy improvement by planning in a latent space, without assumptions on the structure of the underlying environment. While expressive, these models are generally task-dependent: they learn uninterpretable latent representations that are tied to the training task and thus hard to generalize to new tasks. In this work, we present a novel world model formulation where the reward prediction only depends on a subset of structured, symbolic components of the whole latent state. Decoupling observation reconstruction and reward prediction allows us to learn world models that can adapt zero-shot, i.e. without further environment interactions, to new reward functions defined over the same symbolic state space. We discuss the main advantages and challenges of learning these neurosymbolic world models and demonstrate the strong generalisation properties of our approach over purely neural methods.