MA-JEPA: マルチエージェント強化学習のための結合埋め込み世界モデル
MA-JEPA: Joint-Embedding World Models for Multi-Agent Reinforcement Learning
観測再構成の代わりに自己教師あり結合埋め込み予測(JEPA)を用いた確率的な世界モデルを提案し、集中学習・分散実行型のマルチエージェント強化学習を実現した。SMACの8マップ中4つで最強の比較手法と同等以上の勝率を達成。
著者: Brandon Gary Kaplowitz, Osaze James Obahor, Christian Schroeder de Witt
分類: cs.LG, cs.AI, cs.MA
原文アブストラクト
World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (JEPA) can provide this learning signal for multi-agent reinforcement learning. We introduce MA-JEPA, a stochastic world model that replaces observation reconstruction with prediction of target representations, enabling model-based multi-agent reinforcement learning with centralized training and decentralized execution. A categorical latent state and a causal Transformer are trained with posterior and action-conditioned dynamics prediction objectives and are then used for actor-critic learning from latent imagination. A training-only joint predictor conditions on all agents' local states and actions to predict each agent's next local observation embedding. These predictions are passed through the same local posterior used during real interaction with a centralized critic that is used only for value learning, with execution remaining decentralized. Our experiments show that this architecture performs strongly on SMAC, matching or exceeding the strongest reported comparator mean win rate on four of eight evaluated maps.
関連論文
- MATES: 凍結した単一エージェント方策の観測変換によるマルチエージェント相互作用の学習マルチエージェント強化学習
- 完全ビザンチン耐性マルチエージェント強化学習マルチエージェント強化学習
- 山火事対応における自律UAV探査のためのマルチエージェント強化学習マルチエージェント強化学習
- 予測シールディングによる分散型安全マルチエージェント強化学習マルチエージェント強化学習
- MARS-RA: マルチモーダル比較による具現化マルチエージェント協調におけるクレジット割り当てのためのランク集約マルチエージェント強化学習
- Dreamer-CPC: 世界モデルを用いたメッセージ学習による分散型マルチエージェント強化学習マルチエージェント強化学習