他者の夢を見る:マルチエージェント強化学習における世界モデル内の潜在チームメイトモデリング
Dreaming Of Others: Latent Teammate Modeling In World Models For Multi-Agent Reinforcement Learning
協調型マルチエージェント強化学習において、チームメイトの内部方策や意図を世界モデル内で潜在変数としてモデル化し、心の理論ヘッドで推論することで、多様な協調相手への適応を可能にするアーキテクチャを提案した。
分類: cs.MA, cs.AI, cs.LG
原文アブストラクト
In cooperative multi-agent reinforcement learning (MARL), agents must coordinate with partners whose internal policies and intentions are not directly observable. While world models such as Dreamer have demonstrated strong generalization and sample efficiency in single-agent settings, their application to MARL remains limited by an inability to handle teammate-induced uncertainty. We propose a new perspective: treat teammates as structured, learnable components within the agent's world model. We introduce an architecture that factorizes the latent state of a Dreamer-style recurrent state-space model (RSSM) into environment and teammate components, and learns an auxiliary Theory-of-Mind (ToM) head to infer latent embeddings of partner behavior such as character, intent, and predicted actions from partial trajectories. These teammate latents condition the actor and critic, enabling the agent to imagine and adapt to diverse collaborators. We outline how this approach can support zero-shot and few-shot coordination in partially observable settings and propose a set of benchmarks and evaluation protocols to assess its impact. This work positions world models as not only predictors of environmental dynamics, but as simulators of social behavior, opening new directions for generalizable, human-compatible AI.
関連論文
- MARS-RA: マルチモーダル比較による具現化マルチエージェント協調におけるクレジット割り当てのためのランク集約マルチエージェント強化学習
- Dreamer-CPC: 世界モデルを用いたメッセージ学習による分散型マルチエージェント強化学習マルチエージェント強化学習
- マルチエージェントデモから暗黙の因果世界モデルを学習するマルチエージェント強化学習
- 通信喪失下でのロバストなマルチエージェント協調のための価値認識予測マルチエージェント強化学習
- 遅延を考慮した能動三角測量:不確実性駆動型マルチエージェント強化学習による対UAS応用マルチエージェント強化学習
- マルチエージェント強化学習による協調的な長縄跳びマルチエージェント強化学習