微分可能な信念に基づく対戦相手形成
Differentiable Belief-based Opponent Shaping
対戦相手の信念を状態として捉え、その信念更新を微分可能な形でモデル化することで、報酬構造から最適な戦略を自動的に学習する新しい対戦相手形成手法を提案した。
著者: Aarav G Sane, Karthik Sivachandran, Rohan Paleja
分類: cs.AI, cs.LG
原文アブストラクト
Human coordination often relies on the ability to influence the beliefs of others through strategic action. In multi-agent reinforcement learning, opponent shaping attempts to replicate this influence, though existing methods typically operate within an opponent's parameter, policy, or value space. Meanwhile, belief-manipulation techniques in hidden-role games often rely on hard-coded objectives, such as deception or belief saturation. We propose Differentiable Belief-based Opponent Shaping (D-BOS), a first-order method that treats each observer's belief as the shaped opponent state and differentiates through $k$-step softmax-Bayes belief dynamics. Rather than explicitly rewarding deceptive or cooperative behavior, our method treats the belief state as the target for shaping. This allows the optimal strategy to emerge naturally from the environment's reward structure. This belief-space formulation provides an opponent-shaping signal by differentiating through opponent belief updates, and naturally extends to multiple observers by aggregating gradients over their individual inferred belief trajectories. Empirically, D-BOS outperforms PPO and BBM in hidden-role games, with the largest gains in mixed-motive settings.
関連論文
- MARS-RA: マルチモーダル比較による具現化マルチエージェント協調におけるクレジット割り当てのためのランク集約マルチエージェント強化学習
- Dreamer-CPC: 世界モデルを用いたメッセージ学習による分散型マルチエージェント強化学習マルチエージェント強化学習
- マルチエージェントデモから暗黙の因果世界モデルを学習するマルチエージェント強化学習
- 通信喪失下でのロバストなマルチエージェント協調のための価値認識予測マルチエージェント強化学習
- 遅延を考慮した能動三角測量:不確実性駆動型マルチエージェント強化学習による対UAS応用マルチエージェント強化学習
- マルチエージェント強化学習による協調的な長縄跳びマルチエージェント強化学習